Unified model construction method for open world and out-of-distribution generalization
By building a unified model for field adaptation and using meta-learning and field adaptive technologies, the open world and external distribution generalization problems are solved, and the model's generalization ability under unknown categories and environment changes is improved, and the accuracy rate is achieved.
Patent Information
- Application Number
- CN202510310059.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology is difficult to effectively solve the problems of open world and external distribution generalization, especially when faced with unknown categories and environmental changes, the generalization ability and robustness of the model are insufficient.
Using meta-learning and domain adaptation methods, by building a unified model of domain adaptation, using ResNet50 autoencoder and binary knowledge evaluator, we learn shared features and knowledge from multiple tasks or fields, and combine triple loss and mean square error loss functions to improve the performance of the model in unknown situations.
On a typical data set, the Rank-1 accuracy improvement of 31.94% in the off-distributed identification task was achieved, and the average accuracy rate reached 67.24% in open world scenarios, significantly improving the generalization ability and robustness of the model.
Smart Images

Figure CN120279172A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of out-of-distribution generalization, and particularly relates to a method for constructing a unified model for open-world and out-of-distribution generalization problems. Background Art
[0002] The out-of-distribution (OOD) generalization problem describes the generalization ability of a machine learning model in a new environment or scenario outside the training data distribution. The open-world problem describes the ability of a machine learning system to make reasonable judgments when facing various unknown situations in the real world. The two can be distinguished by whether the new and old environments / worlds have different semantics. For example: OOD generalization focuses on the test performance on specific tasks such as person re-identification data, while the open world involves the perception ability in a completely unknown environment, such as training on person re-identification data and testing on face recognition data.
[0003] The out-of-distribution (OOD) generalization problem involves the problem of model performance degradation caused by the distribution change between training and test data. Both the OOD generalization problem and the open-world problem aim to improve the model's performance when facing unknown situations. In a classification task, both the open-world and OOD generalization models need to have the ability to distinguish unknown classes. The difference between the two is that the OOD problem deals with test data with a distribution bias from the training data, while the open-world problem requires the model to be able to correctly handle unknown classes or examples and remain robust. The test datasets for OOD generalization and open-world problems are both composed of open-set instances, and the test dataset for the open-world problem has significant differences in visual features and semantics from the training dataset.
[0004] Out-of-distribution (OOD) recognition has become a major challenge in the field of machine learning because models often struggle to generalize samples from unknown distributions. Despite extensive research efforts in solving the OOD problem, the more general open-world recognition problem remains an important and unsolved issue, which involves handling unknown classes and adapting to new domains. OOD generalization plays an important role in addressing ethical issues in artificial intelligence.
[0005] The open-world problem is an important challenge in the field of artificial intelligence, referring to the difficulties faced by a system when dealing with unseen classes or examples. Under the traditional closed-world assumption, a model usually only needs to handle known classes and data. However, in real-world applications, data and the environment are often dynamically changing, and new classes and situations keep emerging. In this case, the capabilities of existing models are limited and they cannot effectively perform reasoning and decision-making.
[0006] The complexity of open-world problems lies in the fact that they require the system to not only be able to recognize known categories but also make reasonable responses when faced with unknown information. This involves how to flexibly adapt to new data distributions, how to effectively expand knowledge, and how to learn effectively in the absence of labeled data. Solving open-world problems is crucial for enhancing the generality and robustness of artificial intelligence systems and is a key direction for promoting the development of intelligent systems to a higher level.
[0007] Open-world problems and out-of-distribution (OOD) generalization problems are important challenges in the field of machine learning. Solving these problems requires the model to possess capabilities such as unsupervised learning and domain adaptation. Supervised learning is based on the assumption of independent and identically distributed (i.i.d.) training and test data, which is systematically violated in the context of open-world and OOD generalization problems.
[0008] Although some methods can be tried to improve the robustness of the model to unknown samples to solve the OOD problem, a more comprehensive and integrated approach may be needed to fully solve open-world problems. This requires considering the generalization ability of the model, the ability to identify and handle unknown categories, and its adaptability to new domains and tasks.
[0009] Traditional machine learning models usually assume that the training data and test data follow the same distribution. However, in practical applications, the test data may have a distribution shift, that is, it is quite different from the training data. This will lead to a decline in the generalization performance of the model. Techniques such as domain adaptation, unsupervised representation learning, and data augmentation aim to improve the generalization ability of the model in the case of distribution shift. Among them, domain adaptation uses the method of transfer learning to adapt the model from the source domain to the target domain to alleviate the distribution shift problem. Techniques include aligning feature distributions, adversarial training, generative adversarial networks, etc. Unsupervised representation learning learns invariant feature representations of data to improve the generalization ability of the model in the case of distribution shift. Techniques include autoencoders, generative adversarial networks, self-supervised learning, etc. Data augmentation artificially generates diverse training data through methods such as data transformation to improve the generalization ability of the model. Techniques include image transformation, mixed data, adversarial sample generation, etc.
[0010] There are a large number of unknown things and concepts in the real world, and traditional closed machine learning models cannot handle these unknown things. Open-world problems require the model to be open and able to identify and handle unknown categories or things.
[0011] To address the open-world generalization problem, researchers have proposed many methods and techniques. A common approach is to use anomaly detection or outlier detection techniques to identify unknown classes or abnormal samples. These methods attempt to detect samples that do not match by modeling the distribution of known classes. Another approach is to use generative models such as generative adversarial networks (GANs) or variational autoencoders (VAEs) to learn the distribution of data and distinguish between known and unknown classes by comparing samples of these classes.
[0012] These methods aim to enhance the ability of classification models to handle unknown classes and adapt to the challenges of open-world classification tasks. These methods are theoretically somewhat effective, but there are some limitations in practical applications, which prevent them from well solving out-of-distribution generalization and open-world problems. The main reasons include:
[0013] 1. Data diversity: The data distribution in an open-world environment may be extremely diverse. Anomaly detection methods usually rely on predefined models or thresholds and are difficult to adapt to newly emerging classes or features.
[0014] 2. Sample imbalance: In many cases, the number of samples of unknown classes is much lower than that of known classes, making it difficult for the model to effectively learn how to identify these classes.
[0015] 3. Limitations of generative models: Although GANs and VAEs can generate samples similar to the training data, they may not be able to capture the complex features of unknown classes, resulting in significant differences between the generated samples and real unknown samples.
[0016] 4. Dependence on the training set: These methods usually rely on the quality and representativeness of the training data. If the training data fails to cover potential unknown classes or abnormal samples, the generalization ability of the model will be limited.
[0017] 5. Environmental changes: In open-world tasks, environmental conditions and data collection methods may change, resulting in poor performance of the model under different conditions. Existing anomaly detection and generative models may not be able to effectively cope with these environmental changes.
[0018] In summary, although these methods may be effective in some cases, in the face of the complexity of open-world and out-of-distribution generalization problems, they often fail to provide comprehensive solutions. More comprehensive strategies are needed to improve the adaptability and generalization ability of the model.
[0019] In an open-world environment, due to changes in environmental conditions and camera settings, the performance of a model trained on one dataset may be very poor on another dataset. Domain adaptation techniques aim to bridge this gap and improve the generalization ability of the model, but they solve the problem of limited samples rather than the open-world problem. Summary of the Invention
[0020] Objective of the Invention: Aiming at the above problems, the present invention proposes a unified model construction method for open-world and out-of-distribution generalization. Through meta-learning and domain adaptation, the model learns shared features and knowledge from multiple tasks or domains, thereby improving its performance in unknown situations.
[0021] Technical Solution: To achieve the objective of the present invention, the technical solution adopted by the present invention is: A unified model construction method for open-world and out-of-distribution generalization, comprising the following steps:
[0022] Collect an input image data set and divide it into a source domain training sample data set and a target domain training sample data set;
[0023] Construct a domain adaptation unified model: Define a binary classification knowledge evaluator, apply a meta-learning algorithm to the binary classification knowledge evaluator, construct a meta-feature logistic regression model, and determine the classification of samples;
[0024] The binary classification knowledge evaluator extracts shared features from the source domain and the target domain, and learns the intrinsic features of all data, and trains the binary classification knowledge evaluator through the feature distance and its related meta-features.
[0025] Furthermore, the domain adaptation unified model adopts an open-world learning framework and consists of a main model and two knowledge evaluators; the two knowledge evaluators are a source domain knowledge evaluator and a target domain knowledge evaluator respectively;
[0026] Adopt an autoencoder based on ResNet50 as the main model. The main model learns the distance features of the input samples, and the loss function of the main model is updated separately;
[0027] Obtain the meta-features of the source domain sample features and the target domain sample features output by the main model, and use them as the inputs of the source domain knowledge evaluator and the target domain knowledge evaluator respectively; the meta-features include statistical indicators in the feature space, the correlation between attributes, and the data distribution;
[0028] Both the source domain knowledge evaluator and the target domain knowledge evaluator are logistic regression models based on meta-features. Train the two knowledge evaluators to perform binary regression on the input meta-features, that is, perform binary classification on the meta-features;
[0029] The binary classification knowledge evaluator extracts shared features from different domains, and learns the intrinsic features of all data, and searches for the meta-features of the maximum intra-domain distance and the minimum inter-domain distance;
[0030] The output result of the knowledge evaluator indicates that the sample belongs to the nearest neighbor domain class, or the sample does not belong to the nearest neighbor domain class.
[0031] Furthermore, the loss update in the unified model training stage for the said domain includes two parts:
[0032] One is the weighted aggregation of the mean squared error losses of the source data and the target data in the autoencoder, and the other is the weighted aggregation of the mean squared error loss of the source data and the triplet loss;
[0033] The triplet loss is used to learn the compact representation of the source training samples in the feature space, and train the model to distinguish samples of different classes. A sample is randomly selected from the source domain training sample dataset as the anchor sample for comparison with other samples. Samples belonging to the same domain as the anchor sample are defined as positive samples, and samples from different domains from the anchor sample are defined as negative samples; The triplet loss is expressed as:
[0034] TripletLoss = max(d(a, p) - d(a, n) + α, 0)
[0035] where TripletLoss is the triplet loss value, d(a, p) represents the distance metric between the anchor sample and the positive sample, d(a, n) represents the distance metric between the anchor sample and the negative sample, and α is a margin or threshold that determines the minimum separation distance required between the positive sample and the negative sample;
[0036] During the training process, the model learns to minimize the distance between the positive sample and the anchor sample, while maximizing the distance between the negative sample and the anchor sample;
[0037] The goal of the autoencoder is to minimize the reconstruction error and make the reconstructed output x' equal to the input x. This goal is achieved by minimizing the mean squared error loss:
[0038]
[0039] where x' represents the reconstructed output, x and y represent the input image data and the features output by the encoder respectively, represents the predicted value of the i-th sample, represents the true value of the i-th sample, and n is the total number of samples.
[0040] Furthermore, the tuple of the said meta-features is defined as: [dist_ap, dist_an, avg_dist_ap, avg_dist_an, cv_ap, cv_an, mycv], and each parameter of the tuple represents [the maximum intra-class distance, the minimum inter-class distance, the average maximum intra-class distance of the whole sample, the minimum inter-class distance of the whole sample, the coefficient of variation of the maximum intra-class distance, the coefficient of variation of the minimum inter-class distance, the deformation measurement parameter based on the coefficient of variation];
[0041] where the calculation method of the coefficient of variation is as follows:
[0042] Coefficient of variation of the distance of the anchored positive samples: Among them, σ represents the standard deviation, μ is the mean, and AP represents the feature distance between the input image to be recognized and the positive samples of the candidate category;
[0043] Coefficient of variation of the distance of the anchored negative samples: Among them, σ represents the standard deviation, μ is the mean, and AN represents the feature distance between the input image to be recognized and the negative samples of the candidate category.
[0044] Furthermore, the calculation method of the deformation measurement parameter mycv based on the coefficient of variation is as follows:
[0045] mycv = (CV_AP - CV_AN) / (CV_AP + CV_AN),
[0046] Among them, mycv is the deformation measurement parameter, which normalizes the difference to the range of -1 to 1, normalizes the coefficient of variation between the target domain sample data and the source domain sample data with known labels, and measures the symmetric features in the source domain dataset and the target domain dataset.
[0047] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:
[0048] The method of the present invention proposes a domain adaptation unified model, which solves the problems of out-of-distribution generalization and open-world two different generalization scenarios through unsupervised learning and meta-learning knowledge evaluator. The problem is transformed into a binary classification problem, that is, to judge whether a sample belongs to the nearest candidate category. The binary knowledge evaluator is trained through the feature distance and its related meta-features. Extensive experiments conducted on typical datasets show that the proposed method achieves a 31.94% improvement in Rank-1 accuracy in the out-of-distribution recognition task compared with the existing state-of-the-art technology. In addition, the proposed method also shows good performance in dealing with open-world scenarios, with an average accuracy rate of 67.24%. Description of the drawings
[0049] Figure 1 is the system model diagram of the method of the present invention.
[0050] Figure 2 is the schematic diagram of the triplet in binary classification.
[0051] Figure 3 is the schematic diagram of the loss update in the training stage.
[0052] Figure 4 is the flowchart of the knowledge evaluator query stage.
[0053] Figure 5 is the schematic diagram of meta-learning in the training and testing stages. Detailed implementation manners
[0054] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0055] According to the classification of the OOD generalization learning method, especially in terms of the model design strategy, the present invention uses domain adaptation as an unsupervised domain generalization method to enhance the model's ability to understand and represent data. Solving the sample bias in domain transfer / adaptation remains an active research direction.
[0056] In addition, the present invention further adopts an autoencoder based on ResNet50 as the main model, as Figure 1 shown, to transform the input data and improve the effectiveness of data representation, with the focus on improving the generalization ability by learning the shared features and knowledge of multiple domains. When solving the open-world problem, more attention is paid to how to identify and process unknown classes, and how to adapt to new, unseen data. The above two problems can both be transformed into binary classification problems on triplets, as Figure 2 shown. The triplet loss is a metric learning loss function used to train the model to distinguish samples of different classes. Anchor is a sample randomly selected from the training dataset. It is the core of the triplet and is used to compare with other samples. Positive is a sample belonging to the same class as the Anchor. During the training process, the model will learn to minimize the distance between the Positive and the Anchor to ensure that the feature representations of samples of the same class are closer. Negative is a sample different from the Anchor. The model will learn to maximize the distance between the Negative and the Anchor to ensure that the feature representations of samples of different classes are far from each other. The objective function of the triplet loss usually includes two parts: the distance between the Positive and the Anchor: it is desired that this distance is as small as possible. The distance between the Negative and the Anchor: it is desired that this distance is as large as possible. The rule for identification is: if the rule is satisfied, the new instance belongs to the nearest candidate class; if the rule is not satisfied, it does not belong to the candidate class and forms a new class. To find the above rules, the present invention needs to adopt a meta-learning method with adaptability.
[0057] ID-Discriminative Embedding is a technique for deep learning, mainly used to improve the discrimination ability of the model in specific tasks (such as image recognition, object tracking, etc.). Due to the limitations of the label space and classification granularity in the open world, the present invention abandons the traditional IDE (ID-discriminative embedding) network and defines a knowledge evaluator, as Figure 1As shown, to determine the classification of the sample. When faced with data from unknown domains, traditional logistic regression models may not generalize well. Therefore, the present invention applies a meta - learning algorithm (learning shared features and patterns in completely new data) to a binary - classification knowledge evaluator, which extracts shared features from different domains and learns the intrinsic features of all data, including completely new data.
[0058] The classification problem of the method of the present invention is essentially a problem of inter - class distance. The meta - features of the knowledge evaluator are related to the minimum inter - domain distance and the maximum intra - domain distance. Meta - feature logistic regression uses the concept of multi - task learning to achieve meta - learning. In multi - task learning, shared meta - features are used to capture common features and patterns between tasks. In this way, the knowledge and features learned from one task can be transferred to other tasks. Therefore, this method not only promotes OOD generalization, but also enables the knowledge evaluator to apply the knowledge and experience learned on the target training set to the test set of new categories, thereby enhancing the learning ability of new data in the open world. Both the OOD generalization and classification problems in the open world can be transformed into binary - classification problems related to the class distance of the above - mentioned meta - logistic regression model.
[0059] Both the OOD and open - world generalization problems study how the model copes with unknown challenges: in the old world and the new world, the features of objects have different units and dimensions, and there is no generally accepted standard to measure the general features of objects for identification. There are significant differences in data distributions and feature representations between different domains. When the model learns new information, the lack of feedback on new knowledge leads to unbalanced data samples. When using an open - world model for incremental identification, due to challenges brought by the infinite label space, such as the imbalance between old and new training samples, computational complexity, and the risk of forgetting knowledge.
[0060] The present invention proposes an open - world learning framework, which consists of a main model and two knowledge evaluators. The main model learns distance features, while the knowledge evaluators perform binary classification based on the meta - features of the distance features learned by the main model. As Figure 3 shown, multiple loss functions of the main model are updated separately, thus achieving faster convergence and preventing partial overfitting. The present invention introduces meta - features based on the maximum intra - class distance (i.e., the maximum intra - domain distance) and the minimum inter - class distance (i.e., the minimum inter - domain distance) as a representation of the basic understanding of the knowledge evaluator.
[0061] The domain adaptation method of the model is reflected in the introduction of unsupervised learning, which allows semi - supervised training on the same model, thus providing the possibility of comparing and integrating new and old knowledge. Considering the iterative comparison of new and old knowledge in the evaluator, the present invention proposes knowledge evaluators from the source domain and the target domain, which are optimized synchronously and communicate with each other.
[0062] Human learning begins with known knowledge categories, understanding the basic characteristics of these categories, and generalizing to new knowledge systems. The essence of the classification problem lies in determining whether a sample belongs to a candidate category, which is a research question related to feature distance. Existing open-world models, such as the NCM (Neighborhood Contrastive Model) and NNO (Neighborhood Network Optimization), use pre-trained fixed parameters for category measurement. The metric learning in NCM and NNO is limited to a closed set of known object categories and cannot effectively explore the subsequent unknown world. As Figure 4 shown, the knowledge evaluator is a system for knowledge evaluation and discrimination. In the model of the present invention, two discriminators (knowledge evaluators) are used in the feature spaces of the source training domain and the target training domain to evaluate knowledge, maintaining synchronous communication for understanding the essence of things.
[0063] Main model training phase: The update process of the source domain loss in the main model consists of two parts: calculating the triplet loss on the source domain data using the encoder, and calculating the mean squared error (MSE) loss on the source domain data using the autoencoder. The target training set is an optional step in which the mean squared error loss on the target training set (if available) is calculated and updated accordingly. The main model learns the distance features and searches for the candidate category closest to the test image.
[0064] Knowledge evaluator training phase: The knowledge evaluator first learns from the distance-related meta-features of the source training set and the target training set. This enables the knowledge evaluator to make better judgment decisions by combining knowledge from the target training set (if available) and the source training set.
[0065] Testing phase: Logistic regression predicts the classification based on the category distance features and related meta-features of the test image. It is determined whether the test image and the nearest candidate category belong to the same category.
[0066] The present invention defines the new world / new knowledge system as the target data set, while the old world / old knowledge system refers to the training set with known labels. Figure 5The meta - feature data shown refers to the features that describe other features. It can include statistical metrics in the feature space, the correlation between attributes, and the data distribution. The present invention defines statistical metrics such as the mean and variance, measures the correlation coefficient between features, such as variance. The tuple of meta - features is defined as: [dist_ap, dist_an, avg_dist_ap, avg_dist_an, cv_ap, cv_an, mycv], and each parameter of the tuple represents [the maximum intra - class distance, the minimum inter - class distance, the average maximum intra - class distance of the whole sample, the average minimum inter - class distance of the whole sample, the coefficient of variation of the maximum intra - class distance (cv_ap), the coefficient of variation of the minimum inter - class distance (cv_an), the deformation measurement parameter based on the coefficient of variation (the deformation of cv_an and cv_ap)].
[0067] Specifically, the present invention introduces a standardized soft metric called the coefficient of variation (CV).
[0068] (1) Coefficient of variation of the distance anchored to positive samples CVAP: Where σ represents the standard deviation, μ is the mean, and AP represents the feature distance between the input image to be recognized and the positive samples of the candidate class.
[0069] (2) Coefficient of variation of the distance anchored to negative samples CVAN: Where σ represents the standard deviation, μ is the mean, and AN represents the feature distance between the input image to be recognized and the negative samples of the candidate class.
[0070] The new knowledge system and the old knowledge system may have different units or dimensions. The coefficient of variation can eliminate the uneven distribution of feature sizes by normalizing the mean. It can adapt to features with different class distances in the new and old worlds. The present invention also defines a deformation measurement parameter mycv based on the coefficient of variation:
[0071] mycv = (CV_AP - CV_AN) / (CV_AP + CV_AN)
[0072] This measurement method normalizes the difference to the range of - 1 to 1 without considering the absolute value. This makes the comparison of differences between different meta - feature data more fair and reliable. The formula (a - b) / (a + b) is used for the coefficient of variation in the new and old knowledge systems, normalizes the coefficient of variation between the two fields, and calculates the relative difference between the coefficients of variation. In addition, due to its symmetry, this formula is very suitable for capturing symmetric features in the new and old knowledge systems, effectively transferring the learning experience in the old knowledge system to the new knowledge system. With its high robustness, this flexible parameter can adapt to the measurement requirements of different fields, making the comparison and enhancement of similarities and differences between the new and old worlds more flexible.
[0073] Such asFigure 4 As shown in Figure 4 , during the query phase of the source-target knowledge evaluator (i.e., meta-feature based logistic regression):
[0074] Traction: The traction of the knowledge evaluator for understanding the target domain training set (i.e., the new world / new knowledge system / ) comes from the communication between the metric systems R_m and R'_m (i.e., the target domain knowledge evaluator R_m and the source domain knowledge evaluator R'_m) in the meta-feature space of the source domain training set with known labels (i.e., the old world / old knowledge system). The coefficient of variation (CV) eliminates the unbalanced distribution between the target domain feature space R_E and the source domain feature space R'_E. Through the above communication, the source domain training set provides an empirical reference for the cognitive understanding of the target domain training set (the old knowledge system for the new knowledge system), thus guiding and accelerating the learning of the target domain training set (the new knowledge system).
[0075] Enhancement: The target domain knowledge evaluator R_m and the source domain knowledge evaluator R'_m receive the label feedback $\hat{y}^t$ and $\hat{y}^s$ from the target domain and source domain training sets respectively. Through statistical indicators such as mean and variance extracted from the meta-features, the evaluation criteria of the knowledge evaluator in the new world are enhanced from multiple perspectives.
[0076] Facilitation: Since the deformation measurement parameter mycv can measure the symmetric features in the old knowledge system and the new knowledge system, assuming that by comparing the measured feature spaces R_E and R'_E, the target domain meta-feature R_m and the source domain meta-feature R'_m obtain an asymmetric difference $\delta R_n$, which can be mapped to the meta-feature space $\delta R_m$ through a kernel function. Through sufficient communication between the knowledge evaluator and the old world, and combined with their respective state transitions, the evaluation criteria are improved, facilitating the transfer of experience between the old world and the new world.
[0077] Table 1 shows the main layers and parameter settings of the ResNet-50 autoencoder. The input size is (256, 128, 3), where the height is 256 pixels, the width is 128 pixels, and the number of channels is 3 (RGB image). Max pooling is a commonly used downsampling operation for extracting important feature information. In the table, each row represents a layer of the network. The columns include the layer level (encoder or decoder), type (convolutional layer, pooling layer, transposed convolutional layer, etc.), input size, output size, as well as the convolutional kernel size, stride, and padding. The fully connected layer in ResNet-50 is used for label classification, but it is not required in the model design of the present invention. The decoder structure gradually restores the features to (1, 1, 64) through a series of transposed convolutional layers and upsampling layers. The convolutional kernel size of the transposed convolutional layer is 10x10, the stride is 2, and the padding is 4. This configuration helps with upsampling and restoring the size of the feature map.
[0078] Table 1 ResNet-50 Autoencoder Layers and Parameter Settings
[0079]
[0080] The multi-loss function of the model is updated separately to achieve faster convergence and prevent partial overfitting. The loss update is divided into two parts: one is the weighted loss aggregation of the source data and the target data in the autoencoder, as Figure 3 shown; the other is the weighted aggregation of the mean squared error (MSE) loss and the triplet loss of the source data, as Figure 3 shown. The triplet loss is a loss function used to train the embedding model, aiming to learn a compact representation of samples in the feature space.
[0081] TripletLoss = max(d(a,p) - d(a,n) + α, 0)
[0082] where TripletLoss is the triplet loss value, d(a,p) represents the distance metric between the anchor sample and the positive sample, d(a,n) represents the distance metric between the anchor sample and the negative sample, and α is a margin or threshold that determines the minimum separation distance required between the positive sample and the negative sample.
[0083] The goal of the autoencoder is to minimize the reconstruction error so that the reconstructed output x′ is as close as possible to the input x. This goal is achieved by minimizing the mean squared error (MSE) loss.
[0084]
[0085] where x′ represents the reconstructed output, x and y represent the input image data and the features output by the encoder respectively, represents the predicted value of the i-th sample, represents the true value of the i-th sample, and n is the total number of samples.
[0086] In the development of artificial intelligence, significant progress has been made in learning about the known world, especially in cognitive and classification tasks of known categories, assuming a closed-world environment. However, exploring and dealing with the unknown world is the direction for the further development of artificial intelligence. Between the known world and the unknown world, there are significant obstacles in terms of knowledge categories and data distribution, hindering the progress of artificial intelligence exploration. The out-of-distribution generalization problem refers to the significant decline in the performance of the model when facing data in unknown domains that are significantly different from the training data. The open-world problem refers to the important and yet unsolved problem faced by the model when it encounters unseen categories or examples. The method of the present invention proposes a unified domain adaptation model to solve the problems of two different generalization scenarios, namely out-of-distribution generalization and open world, through unsupervised learning and a meta-learning knowledge evaluator. The problem is transformed into a binary classification problem, that is, to judge whether a sample belongs to the nearest candidate category. A binary knowledge evaluator is trained through the feature distance and its related meta-features. Extensive experiments conducted on typical datasets show that the proposed method achieves a 31.94% improvement in Rank-1 accuracy compared to the existing state-of-the-art technology in the out-of-distribution recognition task. In addition, the proposed method also shows good performance in dealing with the open-world scenario, with an average accuracy rate reaching 67.24%.
Claims
1. A unified model construction method for open-world and out-of-distribution generalization, characterized in that It includes the following steps: Collect an input image dataset and divide it into a source domain training sample dataset and a target domain training sample dataset; Construct a domain adaptation unified model: Define a binary classification knowledge evaluator, apply a meta-learning algorithm to the binary classification knowledge evaluator, construct a meta-feature logistic regression model, and determine the classification of samples; The binary classification knowledge evaluator extracts shared features from the source domain and the target domain, learns the intrinsic features of all data, and trains the binary classification knowledge evaluator through the feature distance and its related meta-features.
2. The unified model construction method according to claim 1, wherein The domain adaptation unified model adopts an open-world learning framework and consists of a main model and two knowledge evaluators; the two knowledge evaluators are a source domain knowledge evaluator and a target domain knowledge evaluator respectively; Use an autoencoder based on ResNet50 as the main model. The main model learns the distance features of the input samples, and the loss function of the main model is updated separately; Obtain the meta-features of the source domain sample features and the target domain sample features output by the main model, and use them as the inputs of the source domain knowledge evaluator and the target domain knowledge evaluator respectively; the meta-features include statistical metrics in the feature space, the correlation between attributes, and the data distribution; Both the source domain knowledge evaluator and the target domain knowledge evaluator are logistic regression models based on meta-features. Train the two knowledge evaluators to perform binary regression on the input meta-features, that is, perform binary classification on the meta-features; The binary classification knowledge evaluator extracts shared features from different domains, learns the intrinsic features of all data, and finds the meta-features with the largest intra-domain distance and the smallest inter-domain distance; The output result of the knowledge evaluator indicates that the sample belongs to the nearest neighbor domain class or the sample does not belong to the nearest neighbor domain class.
3. The unified model construction method according to claim 1 or 2, characterized in that The loss update in the training stage of the domain adaptation unified model includes two parts: One is to weighted sum the mean square error losses of the source data and the target data in the autoencoder, and the other is to weighted sum the mean square error loss and the triplet loss of the source data; The triplet loss is used to learn the compact representation of the source training samples in the feature space, train the model to distinguish samples of different classes, randomly select a sample in the source domain training sample dataset as the anchor sample for comparison with other samples, define the sample belonging to the same domain as the anchor sample as the positive sample, and the sample from a different domain from the anchor sample as the negative sample; the triplet loss is expressed as: TripletLoss = max(d(a,p) - d(a,n) + α, 0) where TripletLoss is the triplet loss value, d(a,p) represents the distance metric between the anchor sample and the positive sample, d(a,n) represents the distance metric between the anchor sample and the negative sample, and α is a margin or threshold that determines the minimum separation distance required between the positive sample and the negative sample; During the training process, the model learns to minimize the distance between the positive sample and the anchor sample, while maximizing the distance between the negative sample and the anchor sample; The goal of the autoencoder is to minimize the reconstruction error and make the reconstructed output x' equal to the input x, and this goal is achieved by minimizing the mean square error loss: where x′ represents the reconstructed output, and x and y represent the input image data and the features output by the encoder respectively, represents the predicted value of the i-th sample, represents the true value of the i-th sample, and n is the total number of samples.
4. The unified model construction method according to claim 3, wherein The tuple of the meta-features is defined as: [dist_ap, dist_an, avg_dist_ap, avg_dist_an, cv_ap, cv_an, mycv], and the parameters of the tuple respectively represent [the maximum within-class distance, the minimum between-class distance, the average maximum within-class distance of the whole sample, the minimum between-class distance of the whole sample, the coefficient of variation of the maximum within-class distance, the coefficient of variation of the minimum between-class distance, the deformation measurement parameter based on the coefficient of variation]; Among them, the calculation method of the coefficient of variation is as follows: Coefficient of variation of the distance of the anchored positive sample: where σ represents the standard deviation, μ is the mean, and AP represents the feature distance between the input image to be recognized and the positive sample of the candidate category; Coefficient of variation of the distance of the anchored negative samples: where σ represents the standard deviation, μ is the mean, and AN represents the feature distance between the input image to be recognized and the negative samples of the candidate class.
5. The unified model construction method according to claim 4, wherein The calculation method of the deformation measurement parameter mycv based on the coefficient of variation is as follows: mycv = (CV_AP - CV_AN) / (CV_AP + CV_AN), where mycv is the deformation measurement parameter, which normalizes the difference to the range of -1 to 1 through the deformation measurement parameter, normalizes the coefficient of variation between the target domain sample data and the source domain sample data with known labels, and measures the symmetric features in the source domain dataset and the target domain dataset.
Citation Information
Cited By
Federal learning distribution external generalization detection method based on local attention enhancement and singular vector global modeling
CN121413803A
A federated learning out-of-distribution generalization detection method based on local attention enhancement and singular vector global modeling
CN121413803B