Multi-domain data adaptive learning method, learning device, medium and learning system

Through the combination of cross-domain classification network and incremental learning network, knowledge distillation and feature discriminator are used to solve the problem of domain feature alignment in adaptive learning of multi-domain data, and the stable and efficient recognition of the model on multiple target domains is achieved.

CN120338038APending Publication Date: 2025-07-18中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510335712.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art cannot align the domain feature space due to too many target domains during the data feeding process of source domain and target domain, resulting in catastrophic forgetting, affecting the recognition accuracy and generalization ability of the model.

Method used

The cross-domain classification network is used to compare and learn the data of the source domain and the target domain, and combine the knowledge distillation and feature discriminator of the incremental learning network, and optimize the model parameters to achieve domain-invariant feature representation through fine-grained alignment and gradient inversion algorithms.

Benefits of technology

It improves the generalization ability of the model in the target domain, prevents the model from forgetting old knowledge when learning new knowledge, ensures the recognition accuracy and stability of the model on multiple target domains, and reduces the consumption of computing and storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338038A_ABST
    Figure CN120338038A_ABST
Patent Text Reader

Abstract

The invention provides a multi-domain data adaptive learning method, a learning device, a medium and a learning system. According to the learning method, data of a source domain and data of a target domain are subjected to comparative learning through a cross-domain classification network, so that a model can learn domain-invariant features, the generalization ability in the target domain is improved, model knowledge learned previously is distilled to a current model through a knowledge distillation mode of an incremental learning network, and the learning efficiency is improved. The method prevents the model from forgetting old knowledge when learning new knowledge, guarantees the continuity and stability of the model by storing and utilizing the model parameters trained in the last round, thereby avoiding disastrous forgetting, realizes fine-grained alignment of source domain features on a high-level semantic layer through a feature discriminator, and improves the accuracy of learning the new knowledge. According to the method, the model is ensured to be capable of learning feature representation with high discrimination and unchanged domains, so that the problem of disastrous forgetting caused by incapability of aligning domain feature spaces due to excessive target domains in a data feeding process of a source domain and a target domain in an existing scheme is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing. Specifically, it relates to a multi-domain data adaptive learning method, a learning device, a medium, and a learning system. Background Art

[0002] Currently, the banking industry generally uses computer vision technology to implement intelligent recognition services. Such technologies generally require the angles and lighting of the collected training samples and the samples to be recognized and tested to be the same. Once there are differences between the training set (source domain) and the test set (target domain), the accuracy of the intelligent recognition model will drop significantly. Traditional domain adaptation technologies generally transfer from a single source domain to a single target domain. However, in actual branch business scenarios, due to external factors such as weather and camera position changes, there are often spatial feature differences between the source domain and target domain data sets. Therefore, there are often certain application bottlenecks in only considering the cross-domain recognition task from a single source domain to a single target domain. At the same time, with the development of current AI face-swapping technology, many recognition technologies are difficult to learn underlying deep features, resulting in the emergence of situations where fakes can pass for real.

[0003] However, in the existing solutions during the data feeding process of the source domain and the target domain, due to too many target domains, it is impossible to align the domain feature spaces, resulting in catastrophic forgetting. Summary of the Invention

[0004] The main objective of the present application is to provide a multi-domain data adaptive learning method, a learning device, a medium, and a learning system, so as to at least solve the problem that in the existing solutions during the data feeding process of the source domain and the target domain, due to too many target domains, it is impossible to align the domain feature spaces, resulting in catastrophic forgetting.

[0005] To achieve the above object, according to one aspect of the present application, a multi-domain data adaptive learning method is provided. The learning method includes: processing the data of the source domain and the target domain by using a cross-domain classification network to obtain a first loss, where the first loss includes a source domain loss and a target domain loss. The source domain loss is a contrast feature loss between the data of the source domain and the source domain augmented data, and the target domain loss is a contrast feature loss between the data of the target domain and the target domain augmented data. The source domain augmented data is the data obtained by performing augmentation processing on the data of the source domain by using an augmentation model, and the target domain augmented data is the data obtained by performing augmentation processing on the data of the target domain by using an augmentation model; processing the data of the source domain by using an incremental learning network to obtain a second loss, where the second loss is the loss of the difference between the convolutional layer mapping result of the cross-domain classification network in the current iteration and the convolutional layer mapping result of the model parameters saved by the incremental learning network in the previous cross-domain classification network training; processing the data of the source domain by using a classifier to obtain a source domain self-supervised cross-entropy loss; processing the first source domain feature and the second source domain feature by using a feature discriminator to obtain a binary classification cross-entropy loss, and performing confusion processing on the feature discriminator by using a gradient reversal algorithm so that the feature discriminator cannot distinguish the first source domain feature and the second source domain feature. The first source domain feature is the source domain feature obtained after the data of the source domain is fed into the convolutional layer of the currently participating cross-domain classification network, and the second source domain feature is the source domain feature obtained after the data of the source domain is fed into the convolutional layer of the incremental learning network; determining a network total loss according to the first loss, the second loss, the source domain self-supervised cross-entropy loss, and the binary classification cross-entropy loss, and optimizing the parameters of the cross-domain classification network and the incremental learning network with the goal of minimizing the network total loss.

[0006] Optionally, determining the network total loss according to the first loss, the second loss, the source domain self-supervised cross-entropy loss, and the binary classification cross-entropy loss includes:

[0007] Determining a single-domain single-target domain loss according to the source domain self-supervised cross-entropy loss and the first loss;

[0008] According to L total = L STDA + L D + L bce (D(S prev ),0)+ L bce (D(S cur ),1), determining the network total loss;

[0009] Wherein, L total is the network total loss, L STDAFor the single-domain single-target domain loss, L D For the second loss, L bce (D(S prev ), 0) is the binary classification cross-entropy loss, L, when the network prediction value is D(S prev ), and the label value is 0 bce (D(S cur ), 1) is the binary classification cross-entropy loss, S, when the network prediction value is D(S cur ), and the label value is 1 cur For the first source domain feature, S prev For the second source domain feature.

[0010] Optionally, determining the single-domain single-target domain loss according to the source domain self-supervised cross-entropy loss and the first loss includes:

[0011] According to Determine the single-domain single-target domain loss;

[0012] Wherein, L STDA For the single-domain single-target domain loss, L ce For the source domain self-supervised cross-entropy loss, For the within-domain contrast loss of the source domain, For the within-domain contrast loss of the target domain, λ1 and λ2 are respectively the weights for balancing the contrast loss and the source domain self-supervised cross-entropy loss.

[0013] Optionally, processing the data of the source domain by using an incremental learning network to obtain a second loss, including:

[0014] According to L D = ||F s (x) - F'(x')|| 2 , determine the second loss;

[0015] Wherein, L D For the second loss, F s (x) is the convolutional layer mapping of the cross-domain classification network in the current iteration, and F'(x') is the convolutional layer mapping of the model parameters saved by the incremental learning network in the previous cross-domain classification network training.

[0016] Optionally, processing the first source domain feature and the second source domain feature by using a feature discriminator to obtain a binary classification cross-entropy loss, including:

[0017] According to L bce (pt, label) = -label * ln(pt) + (1 - label) * ln(1 - pt),

[0018] Determine the binary classification cross-entropy loss;

[0019] where L bce (pt, label) is the binary classification cross-entropy loss, label is the label value, and pt is the network prediction value.

[0020] Optionally, the method further includes: in the process of processing the data of the source domain and the target domain by using the cross-domain classification network each time to obtain the first loss, adding a target domain to perform a comparison process with the source domain.

[0021] Optionally, processing the data of the source domain by using a classifier to obtain the source domain self-supervised cross-entropy loss, including: inputting the data of the source domain into a feature extractor to obtain a feature vector; converting the feature vector into a representation vector through a multi-layer perceptron; using the classifier to process the representation vector to identify the category of the data of the source domain to obtain the prediction value of the data of the source domain; calculating the self-supervised cross-entropy loss between the prediction value and the true label of the data of the source domain to obtain the source domain self-supervised cross-entropy loss.

[0022] According to another aspect of the present application, a multi-domain data adaptive learning device is provided. The learning device includes: a first processing unit configured to process data of a source domain and a target domain by using a cross-domain classification network to obtain a first loss, where the first loss includes a source domain loss and a target domain loss. The source domain loss is a contrast feature loss between the data of the source domain and the source domain augmented data, and the target domain loss is a contrast feature loss between the data of the target domain and the target domain augmented data. The source domain augmented data is data obtained by performing augmentation processing on the data of the source domain by using an augmentation model, and the target domain augmented data is data obtained by performing augmentation processing on the data of the target domain by using the augmentation model; a second processing unit configured to process the data of the source domain by using an incremental learning network to obtain a second loss, where the second loss is a loss of the difference between the convolutional layer mapping result of the cross-domain classification network in the current iteration and the convolutional layer mapping result of the model parameters of the cross-domain classification network trained in the previous time saved by the incremental learning network; a third processing unit configured to process the data of the source domain by using a classifier to obtain a source domain self-supervised cross-entropy loss; a fourth processing unit configured to process a first source domain feature and a second source domain feature by using a feature discriminator to obtain a binary classification cross-entropy loss, and perform confusion processing on the feature discriminator by using a gradient reversal algorithm so that the feature discriminator cannot distinguish the first source domain feature and the second source domain feature. The first source domain feature is the source domain feature obtained after the data of the source domain is sent into the convolutional layer of the cross-domain classification network currently participating in training, and the second source domain feature is the source domain feature obtained after the data of the source domain is sent into the convolutional layer of the incremental learning network; a determination unit configured to determine a total network loss according to the first loss, the second loss, the source domain self-supervised cross-entropy loss, and the binary classification cross-entropy loss, and optimize the parameters of the cross-domain classification network and the incremental learning network with the goal of minimizing the total network loss.

[0023] According to another aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored program, where, when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the learning methods.

[0024] According to another aspect of the present application, a multi-domain data adaptive learning system is provided, including: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the learning methods.

[0025] Applying the technical solution of the present application, the contrastive learning of the data in the source domain and the target domain through the cross-domain classification network helps the model learn domain-invariant features, thereby improving the generalization ability in the target domain. Through the knowledge distillation method of the incremental learning network, the model knowledge learned previously is distilled to the current model to maintain the recognition ability of the model on the previous source domain and target domain, preventing the model from forgetting old knowledge when learning new knowledge. By saving and using the model parameters of the previous round of training, the continuity and stability of the model are ensured, thereby avoiding catastrophic forgetting. By accurately realizing classification on the source domain, reliable initial conditions are provided for the application of the model in the target domain. By implementing fine-grained alignment of the source domain features at the high-level semantic layer through the feature discriminator, it is ensured that the model can learn discriminative and domain-invariant feature representations. Finally, according to the first loss, the second loss, the source domain self-supervised cross-entropy loss, and the binary classification cross-entropy loss, the total network loss is determined, thereby solving the problem in the existing solution that during the data feeding process of the source domain and the target domain, due to too many target domains, the domain feature space cannot be aligned, resulting in catastrophic forgetting. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0027] Figure 1 FIG. shows a schematic flowchart of a multi-domain data adaptive learning method provided according to an embodiment of this application;

[0028] Figure 2 FIG. shows a schematic diagram of a multi-domain data adaptive learning framework provided in the embodiment of this application;

[0029] Figure 3 FIG. shows a schematic diagram of incremental training provided according to an embodiment of this application;

[0030] Figure 4 FIG. shows a structural block diagram of a multi-domain data adaptive learning device provided according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The following will refer to the accompanying drawings and combine the embodiments to detail this application.

[0032] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of this application here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0034] For the convenience of description, some nouns or terms related to the embodiments of this application are described below:

[0035] Cross-domain classification network: This network is based on contrastive learning, and its goal is to align the feature representations between the source domain and the target domain, enabling the model to effectively apply the features learned from the source domain to the target domain, even if the samples in the target domain have no labels. The cross-domain classification network learns domain-invariant feature representations by minimizing the feature differences between the source domain and the target domain, that is, these features remain consistent in different domains, thereby improving the generalization ability of the model and enabling it to perform well on unseen domains. Specifically, the network uses a feature extractor (such as ResNet50) to extract features from the input images, and then uses an MLP module to convert these features into representation vectors that can be used for contrastive learning. Through contrastive learning, the network optimizes its parameters to minimize the distance between positive sample pairs and maximize the distance between negative sample pairs, so as to learn to distinguish samples of different classes, not limited to the classes in the source domain.

[0036] Incremental Learning Network: The incremental learning network is designed to address the learning problem in multiple target domains. When the model is transferred from one source domain to multiple target domains, it may encounter the problem of catastrophic forgetting, that is, the model forgets the features of the previously learned old domain when learning the features of the new domain. To avoid this problem, the incremental learning network is designed to preserve the parameter weights of the network model at each learning stage (i.e., each time a new target domain is added for training) to ensure that the model does not forget the knowledge learned previously. At each incremental learning stage, the model first uses the samples from the source domain and the current target domain for training to obtain updated network weights. Then, these weights are copied to the incremental learning network and frozen to retain the feature alignment and classification ability obtained in the current learning stage. In this way, as the number of target domains increases, the model can gradually adapt to each new domain without forgetting the features of the previously learned domains, and finally achieve intelligent recognition from a single source domain to multiple target domains. In the present invention, the incremental learning network is also combined with the fine-grained alignment module to achieve finer-grained feature alignment through the feature discriminator, further ensuring the generalization ability and recognition accuracy of the model in multiple target domains.

[0037] Source Domain: A labeled dataset involved in model training.

[0038] Target Domain: An unlabeled dataset for testing the model performance (generally with different distribution characteristics from the source domain but the same classification).

[0039] Domain Adaptation: Domain Adaptation (DA) aims to solve the problem of the decline in model generalization performance due to data distribution differences. By using auxiliary means to learn the domain-invariant features between the source domain and the target domain, the performance of the model in the target domain is improved, and the accuracy loss caused by domain differences is reduced.

[0040] As introduced in the background art, traditional domain adaptation techniques generally transfer from a single source domain to a single target domain. However, in actual network business scenarios, due to external factors such as weather and camera position changes, there are often spatial feature differences between the source domain and target domain datasets. Therefore, there are often certain application bottlenecks in only considering the cross-domain recognition task from a single source domain to a single target domain. At the same time, with the development of current AI face-swapping technology, many recognition technologies are difficult to learn underlying deep features, resulting in the emergence of situations where fakes can pass for the real thing. To solve the problem that in the data feeding process of the existing solutions between the source domain and the target domain, due to too many target domains, it is impossible to align the domain feature space, resulting in catastrophic forgetting, the embodiments of the present application provide a multi-domain data adaptive learning method, learning device, medium, and learning system.

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0042] In this embodiment, a multi-domain data adaptive learning method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0043] Figure 1 It is a schematic flowchart of a multi-domain data adaptive learning method provided according to an embodiment of the present application.

[0044] As Figure 1 shown, the method includes the following steps:

[0045] Step S101, using a cross-domain classification network to process the data of the source domain and the target domain to obtain a first loss. The above first loss includes a source domain loss and a target domain loss. The above source domain loss is the contrast feature loss between the data of the source domain and the source domain enhanced data. The above target domain loss is the contrast feature loss between the data of the target domain and the target domain enhanced data. The above source domain enhanced data is the data obtained by enhancing the data of the source domain using an enhancement model. The above target domain enhanced data is the data obtained by enhancing the data of the target domain using an enhancement model;

[0046] Among them, a specific application scenario of step S101 is: in the intelligent operation scenario of a bank, the face recognition system is an important tool to improve customer experience and ensure security. However, different bank branches may be significantly affected by environmental conditions (such as lighting, angle changes, weather conditions, etc.), resulting in inconsistent distributions of the collected image data. For example, one branch may be located in a sunny area, while another branch may be in a basement with poor lighting conditions. In addition, the camera positions and models at each branch are also different, further increasing the difference in data distribution. In this scenario, using an intelligent recognition system based on a cross-domain classification network can effectively handle the inconsistency between the source domain and target domain data. The source domain can be a branch with relatively controllable conditions, which has collected a large amount of labeled face image data and is used as the training set of the model. The target domain can be other branches with changing conditions, which have collected unlabeled face image data. The cross-domain classification network can learn general face features from the source domain data through contrastive learning and adversarial learning techniques and try to apply them to the target domain data, even if the environmental conditions of the target domain are very different from those of the source domain.

[0047] Benefits of this specific application scenario: The cross-domain classification network can learn face features that are common across multiple outlets and various lighting conditions, thus maintaining a high recognition accuracy even in unseen outlet environments. For banks, this means that the face recognition system can provide consistent service quality regardless of which outlet the customer is at, enhancing the customer experience. The target domain data does not need to be manually labeled, and the model can automatically learn features from unlabeled data, significantly reducing the labor cost and time cost for banks in the data preparation phase. For banks, especially when deploying the face recognition system on a large scale, this can significantly reduce the operation cost. Through contrastive learning and adversarial learning, the cross-domain classification network can automatically adapt to and handle the impacts brought by environmental changes, such as changes in lighting, angle, or weather conditions. Even in harsh environments, such as strong light, rainy weather, or poor image quality, the system can still achieve high-precision face recognition, ensuring the operational safety of banks. By learning deep and invariant face features, the cross-domain classification network can reduce the impact of AI face swapping technology, improving security and trust. For banks, especially in financial transaction verification, this can effectively prevent fraud and protect the funds safety of customers.

[0048] Step S102, process the data in the source domain using an incremental learning network to obtain a second loss, where the second loss is the loss of the difference between the convolutional layer mapping result of the cross-domain classification network in the current iteration and the convolutional layer mapping result of the model parameters of the cross-domain classification network saved by the incremental learning network in the previous training.

[0049] Among them, processing the data in the source domain using an incremental learning network to obtain a second loss includes:

[0050] According to L D =||F s (x)-F'(x')|| 2 , determine the second loss;

[0051] Among them, L D is the second loss, F s (x) is the convolutional layer mapping of the cross-domain classification network in the current iteration, and F'(x') is the convolutional layer mapping of the model parameters of the cross-domain classification network saved by the incremental learning network in the previous training.

[0052] Specifically, in the process of step-by-step learning in multiple target domains, the incremental learning network can utilize the information by saving the model parameters after the previous round of cross-domain classification network training. This means that the model can not only acquire new knowledge (i.e., the features of the target domain), but also retain and utilize the old knowledge (i.e., the features of the previously learned source domain and target domain). The accumulation of this knowledge helps the model not to forget or destroy the existing domain feature representations when facing new domains, thus maintaining its recognition ability in all learned domains. In traditional machine learning and deep learning, when the model learns new tasks, it may "forget" the knowledge of old tasks, and this phenomenon is called catastrophic forgetting. By calculating the difference loss between the convolution layer mapping results of the cross-domain classification network in the current iteration and the convolution layer mapping results of the model parameters saved in the previous training of the incremental learning network, this strategy is equivalent to conducting a "review" process every time when learning a new domain. The model is forced to maintain consistency with the old knowledge while learning new knowledge, thus avoiding catastrophic forgetting and ensuring the stable performance of the model in all domains. The calculation of the second loss is based on the convolution layer mapping results of different iterations, which actually supervises the model to learn more general feature representations. By making the output of the model consistent in multiple training stages, the generalization ability of the model is cultivated. This means that even when facing an environment or data distribution that has never been encountered before, the model can utilize its accumulated and general feature representations to make accurate predictions. When dealing with a large number of target domains, directly training all source domain and target domain data together may lead to a huge consumption of computing resources and storage resources. By adopting the strategy of incremental learning network and knowledge distillation, it is allowed to gradually add target domains for training while only retaining the key model parameters, so that the performance of the model can be effectively improved and the computing and storage costs can be reduced in an environment with limited resources.

[0053] Step S103: Use a classifier to process the data of the source domain above to obtain the self-supervised cross-entropy loss of the source domain.

[0054] In an embodiment of the present application, using a classifier to process the data of the source domain above to obtain the self-supervised cross-entropy loss of the source domain includes: inputting the data of the source domain above into a feature extractor to obtain a feature vector; converting the feature vector into a characterization vector through a multi-layer perceptron; using the classifier above to process the characterization vector to identify the category of the data of the source domain to obtain the predicted value of the data of the source domain; calculating the self-supervised cross-entropy loss between the predicted value and the true label of the data of the source domain above to obtain the self-supervised cross-entropy loss of the source domain above.

[0055] Specifically, a feature extractor (such as ResNet50) is used to preprocess the source domain image data and extract the feature vectors that are important for recognition. This step can reduce the redundant information in the original image data, enhance the model's ability to capture key features in the image, make the model's decision more dependent on the image content rather than background noise or illumination changes, and thus improve the recognition robustness of the model in the new environment. Optimization of the representation vector: The feature vector is converted into a representation vector through a multi-layer perceptron (MLP). This process further optimizes the feature representation to make it more suitable for classification. The non-linear transformation ability of the MLP can help the model learn more advanced and abstract features, which are crucial for the model's adaptability in different domains and contribute to improving the generalization ability of the model. A classifier is used to process the representation vector to obtain the predicted value of the source domain data, and the self-supervised cross-entropy loss between the predicted value and the true label is calculated. The advantage of self-supervised learning is that it can use a large amount of unlabeled data for model training. Here, the self-supervised cross-entropy loss utilizes the guidance of the labeled data in the source domain, which can not only supervise the classification ability of the model but also help the model perform more effective unsupervised learning on the target domain data. By minimizing the classification error in the source domain, the model is prompted to learn more discriminative feature representations. Improving the generalization performance: The calculation of the self-supervised cross-entropy loss is based on the labeled data in the source domain. This supervised learning loss function helps the model avoid overfitting on the source domain data, enabling the model to better understand and adapt to the characteristics of the target domain data while learning the source domain features. By optimizing the classification task in the source domain, it can ensure that the model maintains high accuracy and generalization ability when processing source domain and target domain data.

[0056] Step S104, using a feature discriminator to process the first source domain feature and the second source domain feature to obtain a binary classification cross-entropy loss, and using the gradient reversal algorithm to confuse the above-mentioned feature discriminator so that the above-mentioned feature discriminator cannot distinguish the above-mentioned first source domain feature and the above-mentioned second source domain feature. The above-mentioned first source domain feature is the source domain feature obtained after the data in the source domain is sent into the convolutional layer of the current cross-domain classification network participating in training, and the above-mentioned second source domain feature is the source domain feature obtained after the data in the source domain is sent into the convolutional layer of the incremental learning network;

[0057] Among them, using a feature discriminator to process the first source domain feature and the second source domain feature to obtain a binary classification cross-entropy loss includes:

[0058] According to L bce (pt, label) = -label * ln(pt) + (1 - label) * ln(1 - pt),

[0059] determine the binary classification cross-entropy loss;

[0060] Among them, L bce (pt, label) is the binary classification cross-entropy loss, label is the label value, and pt is the network prediction value.

[0061] Specifically, through the discrimination function of the feature discriminator, the model can identify whether the features come from the cross-domain classification network or the incremental learning network. Then, through the gradient reversal algorithm, the discriminability of these two source features is confused, making it impossible for the feature discriminator to accurately distinguish. This process is actually feature alignment at the high-level semantic layer, ensuring that the source domain features learned by the model are not only aligned macroscopically (the entire feature space), but also consistent in detail, that is, fine-grained alignment. In multi-domain adaptive learning, the model is prone to forgetting the original label information when learning new domain features, that is, "catastrophic forgetting". By introducing the feature discriminator and the gradient reversal algorithm, while the model learns new domain features, it will maintain and optimize the representation of the original source domain features, thus building a bridge between old and new knowledge, reducing the forgetting of label information, and ensuring that when the model is applied in the new domain, it can maintain or even improve the recognition accuracy in the source domain. Fine-grained alignment can not only reduce the recognition error caused by domain differences, but also prompt the model to learn more general and invariant feature representations. These feature representations can provide more stable recognition performance when facing unknown domains or environmental changes, thereby improving the generalization ability of the model. For the banking industry, this means that the face recognition system can still maintain a high recognition rate even under different bank branches, changing weather conditions or different camera settings. Using the principle of the generative adversarial network (GAN), the combination of the feature discriminator and the gradient reversal algorithm creates an adversarial learning environment. This environment forces the model to continuously optimize during the process of generating and discriminating features, thereby enhancing the model's resistance to noise, transformation, and attacks. For banking operations, this helps prevent interference from technologies such as AI face swapping, improving the security and credibility of the recognition system.

[0062] It can be seen that the above steps form a domain adaptation solution. By adding an incremental learning network on top of the cross-domain adaptive image classification model based on contrast learning, the old knowledge obtained from the previous iterative learning is saved, and the new knowledge is incorporated into the old domain knowledge when continuing to train a new round of cross-domain classification modules. At the same time, a source domain feature fine-grained alignment module is added to the model. Using the generative adversarial idea, a feature discriminator D is introduced, and this discriminator preserves the alignment of the source domain high-level semantic layer to ensure that the label information will not be catastrophically forgotten.

[0063] Step S105, determine the total network loss according to the above first loss, the above second loss, the above source domain self-supervised cross-entropy loss, and the above binary classification cross-entropy loss, and optimize the parameters of the above cross-domain classification network and the above incremental learning network with the goal of minimizing the above total network loss.

[0064] In the above steps, the contrastive learning of the data in the source domain and the target domain through the cross-domain classification network helps the model learn domain-invariant features, thereby improving the generalization ability in the target domain. Through the knowledge distillation of the incremental learning network, the model knowledge learned previously is distilled to the current model to maintain the recognition ability of the model on the previous source domain and target domain, prevent the model from forgetting old knowledge when learning new knowledge, and ensure the continuity and stability of the model by saving and using the model parameters of the previous round of training, thereby avoiding catastrophic forgetting. By accurately implementing classification on the source domain, reliable initial conditions are provided for the application of the model in the target domain. The fine-grained alignment of the source domain features is realized by the feature discriminator at the high-level semantic layer to ensure that the model can learn discriminative and domain-invariant feature representations. Finally, according to the above first loss, the above second loss, the above source domain self-supervised cross-entropy loss, and the above binary classification cross-entropy loss, the total network loss is determined, thereby solving the problem in the existing solution that in the data feeding process of the source domain and the target domain, due to too many target domains, the domain feature space cannot be aligned, resulting in catastrophic forgetting.

[0065] In an embodiment of the present application, determining the total network loss according to the above first loss, the above second loss, the above source domain self-supervised cross-entropy loss, and the above binary classification cross-entropy loss includes:

[0066] According to determine the single-domain single-target domain loss;

[0067] wherein, L STDA is the single-domain single-target domain loss, L ce is the source domain self-supervised cross-entropy loss, is the intra-domain contrast loss of the source domain, is the intra-domain contrast loss of the target domain, and λ1 and λ2 are respectively the weights for balancing the contrast loss and the source domain self-supervised cross-entropy loss.

[0068] The source domain self-supervised cross-entropy loss ensures the accurate performance of the model on the source domain, that is, the model can perform effective supervised learning based on the labeled source domain data. The source domain loss and the target domain loss in the first loss respectively measure the gap between the predictions of the model on the source domain and the target domain and the true labels, reflecting the generalization ability of the model. Combining these two losses can ensure that while learning the source domain features, the model can effectively adapt to the target domain, balancing the accuracy of learning and the generalization of the model. Supervised learning relying only on source domain data may lead to overfitting of the model on the source domain data, that is, the model focuses too much on the details of the source domain data and ignores the features of the target domain. By adding the loss of the target domain, the model is forced to learn on the unlabeled target domain data, which helps the model extract more general features, avoid overfitting, and improve the recognition performance on unseen data. In the scenario of single domain and single target domain (STDA), the model needs to be able to adapt to the data distribution of a single target domain under the guidance of the source domain data. By simultaneously minimizing the source domain self-supervised cross-entropy loss and the first loss, the model can not only perform well on the source domain but also effectively align features on the target domain, improving its adaptability and recognition rate on the target domain.

[0069] According to L total = L STDA + L D + L bce (D(S prev ), 0) + L bce (D(S cur ), 1), determine the total network loss;

[0070] Wherein, L total is the total network loss, L STDA is the single domain and single target domain loss, L D is the second loss, L bce (D(S prev ), 0) is the binary classification cross-entropy loss when the network prediction value is D(S prev ) and the label value is 0, L bce (D(S cur ), 1) is the binary classification cross-entropy loss when the network prediction value is D(S cur ) and the label value is 1, S cur is the first source domain feature, S prev is the second source domain feature.

[0071] Combining different loss terms into the total network loss encourages the model to not only focus on the classification tasks of the source domain and the target domain during training but also consider the alignment between features and the preservation of knowledge. This comprehensive training approach helps the model learn more general and invariant feature representations, improving its generalization performance in unseen domains. For the face recognition system in the banking industry, this means maintaining a high recognition rate even under complex environmental changes, enhancing the customer experience and security.

[0072] In one embodiment of the present application, the above method further includes: during each process of using the cross-domain classification network to process the data of the source domain and the target domain to obtain the first loss, adding a target domain to perform a comparison process with the above source domain.

[0073] Specifically, by adding a target domain in each iteration, the model gradually exposes itself to more diverse data distributions and features, which forces the model to maintain and optimize its performance on previous domains (including the source domain and other learned target domains) while continuously adapting to new domains.

[0074] As Figure 2 shown, a multi-domain data adaptive learning framework is presented. In the cross-domain module based on contrast learning (the upper half), an intra-domain contrast learning framework is used. In the incremental learning framework (the lower half dashed box), the training parameters of the previous round are saved. The fine-grained alignment module (the middle part) performs an adversarial operation on the features obtained from the cross-domain module and the incremental features.

[0075] Explanation of the model framework module: Setting of the contrast learning sample pairs: Two random data augmentations are performed respectively to obtain a pair of positive samples, denoted as x i and x j , which serve as the positive sample pairs of the model. All other instances in the mini-batch are regarded as negative sample pairs. The data augmentations used include: random cropping, random horizontal flipping, image color jitter (brightness, contrast, saturation), and random grayscaling.

[0076] Feature extractor: Using ResNet50 as the backbone network framework of the feature extractor, feature vectors are extracted from the samples (x i and x j ) that have undergone data augmentation, obtaining f i and f j .

[0077] MLP module: According to SimClr, an MLP projection layer containing two fully connected layers is adopted and applied to f i and f j , respectively obtaining the representation vectors z i and z j .

[0078] Feature discriminator D and GRL module: The feature discriminator is a binary domain classifier used to determine whether the features come from the cross-domain module or the incremental module, aiming to clarify the source of the features. The GRL (Gradient Reversal Layer) module is used to confuse the two features so that the discriminator cannot determine the specific source of the features, thereby achieving a fine-grained feature alignment effect.

[0079] Detailed introduction of the three modules in this method framework:

[0080] Cross-domain classification module based on contrastive learning: Based on the SimClr theory, the contrastive learning method is applied to the source domain and the target domain respectively, and the source domain and the target domain are aligned by optimizing the contrastive loss within the two domains. The specific contrastive loss function is shown in Equation 1 below:

[0081]

[0082] where l [k≠i] is the indicator function, which takes the value of 0 when k = i;

[0083] represents the cosine similarity between the hidden representations z i and z j . N is the batch size, and t is the temperature parameter. z i and z j are the representation vectors of a pair of positive sample pairs, and the data flow is as follows: Samples x from the source domain or the target domain are randomly data-augmented to obtain a pair of positive samples, denoted as x i and x j , which are regarded as a positive sample pair. They are passed through a feature extractor G(ResNet50) and a multi-layer perceptron (MLP) with a non-linear fully connected layer to obtain the representation vectors z i and z j . The data augmentations used include: random cropping, random horizontal flipping, image color jitter (brightness, contrast, saturation), and random grayscale conversion.

[0084] Applying the contrastive loss to the source domain and the target domain respectively, the intra-domain contrastive loss function L con is as follows:

[0085]

[0086] z s1 and z s2 represent the representation vectors of the positive sample pairs in the source domain obtained through the feature extractor and the MLP. The same applies to the target domain. Combining the self-supervised cross-entropy loss in the source domain, the total loss function for a single-source and single-target domain (STDA) is shown in Equation 5 below:

[0087]

[0088] Among them, L ce represents the source domain self-supervised cross-entropy loss, and λ1 and λ2 are hyperparameters used to balance the weights of the contrast loss and the self-supervised loss.

[0089] One-to-many module based on incremental learning: The incremental learning network is added to align multiple target domains with the source domain, thereby improving the generalization ability of the classifier. On the basis of the cross-domain classification module based on contrast learning, in order to avoid the problem of forgetting the source domain label information and the knowledge of multiple target domains due to the gradual addition of target domains, this method introduces an incremental learning network to save the classification network parameter weights after each addition of a target domain and the completion of training, that is, the weight copy of the previous cross-domain classification network. It is also composed of a feature extractor and an MLP. Knowledge distillation is one of the important methods of incremental learning. In order to ensure the consistency of the semantic information output at the high level of the classification network, a knowledge distillation loss is added to the convolutional network at the high-level semantic layer to ensure that the source domain label information will not be lost due to the superposition of target domains. The distillation loss uses the classical L2 regularization loss function.

[0090] LD = ||F s (x) - F'(x')|| 2 (Formula 6);

[0091] F s is the convolutional layer mapping of the current contrast learning cross-domain classification module, and F' is the convolutional layer mapping of the model parameters of the previous cross-domain module training saved by the incremental learning network. Since the incremental learning network saves a weight copy of the previous training network model parameters. Therefore, before adding a new target domain for one-to-many training, the weight parameters of F s and F' are the same, that is, the parameter initialization consistency. In order to implement the classification task on multiple target domains and ensure the feature alignment between multiple target domains, and not lose the information of the previous target domains due to iterative training after adding a new target domain. Therefore, when adding a new target domain for training, it is necessary to fix the weights of the incremental learning network to remain unchanged to ensure the alignment of multiple domains when stacking domains and thus learn domain-invariant features. The optimized network weights and loss function are as follows:

[0092] L MTDA = L STDA + L D (Formula 7);

[0093] Feature alignment is achieved on the high-level semantic feature layer through the distillation loss, enabling the classification model to effectively fuse the knowledge of different domains during the learning process and using the strategy of progressive training to gradually adapt the network to multiple domains.

[0094] Source domain feature fine-grained alignment module: Since the distributions of the source domain and the target domains are different, and the distributions among multiple target domains are also different. Therefore, to achieve the multi-target domain classification task, it is necessary to align both the source domain and the target domains, and also to align among multiple target domains. However, it is difficult to achieve fine-grained alignment between the source domain and multiple target domains only using the distillation loss. Therefore, based on the incremental learning network, this method adds fine-grained alignment of the source domain features. This module feeds the source domain samples into the source domain features S obtained by learning through the current cross-domain convolutional layer participating in the training cur , and at the same time feeds this sample into the source domain features S obtained by the convolutional layer of the incremental learning network prev for fine-grained alignment. Since both features are generated from the same labeled source domain samples, it is more reasonable to perform fine-grained alignment on these two features. The specific method used is similar to DANN. The role of the feature discriminator in this module is to distinguish whether the current feature comes from the incremental learning network module or the cross-domain training module, and perform alignment during the adversarial learning process between the feature extractor and the domain classifier. The feature discriminator D is a binary classifier, and its loss is the binary classification cross-entropy loss. The specific formula is as follows:

[0095] L bce (pt, label) = -label * ln(pt) + (1 - label) * ln(1 - pt) (Formula 8);

[0096] where pt is the network prediction value and label is the label value. The final optimized objective function is as follows:

[0097] L total = L MTDA + L bce (D(S prev ), 0) + L bce (D(S cur ), 1) (Formula 9);

[0098] The working process of this framework is as follows: The source domain data and the k-th group of target domain data are input into the cross-domain classification module of the model. At the same time, the source domain data is fed into the training classifier, and the source domain samples are fed into the incremental module. The source domain sample features generated by the current cross-domain module and the source domain sample features generated by the incremental module that saves the training parameters of the k - 1-th round are fed into the feature discriminator together. The high-level semantic features of the cross-domain module and the high-level semantic features of the incremental module are subjected to knowledge distillation, and the parameters of the cross-domain module are updated through training iterations until the end of this round of iterative training. The parameters of the cross-domain module are saved to the incremental module. In the next round of iterative training, the parameters of this incremental module are fixed, and the (k + 1)-th group of target domain data is selected for iterative training until all target domain data is trained, and the final multi-domain adaptive model can be obtained.

[0099] The specific overall implementation process of incremental training is as follows: In the k-th round of this training iteration update, the cross-domain module parameter W is obtained. k The W of the cross-domain module in the k-th round k is saved to the incremental module. During the (k + 1)-th round of iterative training, the incremental module keeps the parameter W fixed k and does not update it. Only the cross-domain module is trained to update W. k+1 After the (k + 1)-th round of training, the W of the cross-domain module k+1 is still saved to the incremental module. And so on, until all N target domains have participated in the training. The finally obtained W n is the feature extraction part of the classification model.

[0100] In the initial stage (k = 0), this method uses the labeled source domain samples Ds and the unlabeled first target domain samples Dt1 to train the cross-domain module. By contrastive learning, the contrastive loss is minimized to obtain domain-invariant features. Finally, the initial convolutional layer weights are obtained and saved to the incremental learning network M0. Next, the incremental learning steps k = 1, 2, 3,... are executed. In each round, the source domain is used for contrastive learning cross-domain training with a new target domain. As Figure 3 shown, assume that the model is currently being trained to the k-th round. Then the model weights W k-1 in the (k - 1)-th round are stored in the incremental learning network. The samples from the source domain and the k-th group of target domains are input into the cross-domain network adaptively trained based on contrastive learning to obtain good network weights W k for training. Next, the weights W k are copied to the incremental learning network and frozen. This preserves the information learned in the source domain and target domain before the k-th round. Similarly, in the (k + 1)-th round, the source domain and the (k + 1)-th group of target domain samples are continuously used for adaptive training to obtain good training weights W k+1 . Finally, the knowledge of multiple target regions can be effectively learned through alternating training.

[0101] The model of the method of the present invention mainly studies how to better transfer the domain-invariant features learned from a single source domain to more target domains, so as to train an identification model that can achieve more accurate classification on multiple target domains. Therefore, the effect of stronger generalization and wider adaptation range is achieved. For example, each network point hopes to accurately identify the customer's face, risk profile, etc. even in rainy days, foggy days, and when the camera pixels and positions are offset, so as to improve the customer experience. At the same time, it is also hoped that customers can reduce the number of times of entering portraits and improve the efficiency of handling business. The model of the present invention introduces the mainstream framework of contrastive learning, and on this basis, an incremental learning network is incorporated to achieve incremental learning for multiple target domains, so as to avoid the problems of inability to align the domain feature space and catastrophic forgetting due to too many target domains. At the same time, the idea of adversarial learning is incorporated, and fine-grained alignment based on a feature discriminator is set up to better preserve the source domain label information, so as to achieve multi-domain adaptation.

[0102] Ablation experiments on the distillation loss, the binary classification loss of the feature discriminator, and the contrastive learning loss of the method of the present invention were carried out in the A→D and W tasks of the public dataset Office 31. The baseline of the ablation experiment is the recognition accuracy using the contrastive learning framework. The experimental results show that the recognition accuracy is improved by 5.6% only by adding the distillation loss, and the recognition accuracy is improved by 6.3% only by adding the classification loss of the feature discriminator. When the distillation loss and the loss of the feature discriminator are added at the same time, the accuracy of the present method is improved by 10.9%. The above data fully prove that the method of the present invention can indeed improve the generalization of the model in complex environments.

[0103] This application introduces a feature adversarial discriminator to preserve the non-loss of the source domain high-level semantic layer label information, so as to better align each feature space and learn domain-invariant features to complete the transfer task. At the same time, only a small amount of labeled source domain data is needed to train a more general identification model, reducing the labeling costs of manpower and material resources; the model is trained by using the distillation loss and the incremental learning strategy to achieve incremental learning for multiple target domains, so as to avoid the problems of inability to align the domain feature space and catastrophic forgetting due to too many target domains, thereby obtaining a one-to-many intelligent identification model.

[0104] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0105] The embodiments of the present application also provide a multi-domain data adaptive learning device. It should be noted that the multi-domain data adaptive learning device in the embodiments of the present application can be used to execute the multi-domain data adaptive learning method provided by the embodiments of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0106] The following introduces the multi-domain data adaptive learning device provided by the embodiments of the present application.

[0107] Figure 4 It is a structural block diagram of a multi-domain data adaptive learning device provided according to an embodiment of the present application. As Figure 4 shown, the learning device includes:

[0108] A first processing unit 41, configured to process the data of the source domain and the target domain by using a cross-domain classification network to obtain a first loss. The first loss includes a source domain loss and a target domain loss. The source domain loss is a contrast feature loss between the data of the source domain and the source domain enhanced data. The target domain loss is a contrast feature loss between the data of the target domain and the target domain enhanced data. The source domain enhanced data is the data obtained after enhancing the data of the source domain by using an enhancement model. The target domain enhanced data is the data obtained after enhancing the data of the target domain by using an enhancement model;

[0109] A second processing unit 42, configured to process the data of the source domain by using an incremental learning network to obtain a second loss. The second loss is a loss of the difference between the convolution layer mapping result of the cross-domain classification network in the current iteration and the convolution layer mapping result of the model parameters of the cross-domain classification network trained last time saved by the incremental learning network;

[0110] A third processing unit 43, configured to process the data of the source domain by using a classifier to obtain a source domain self-supervised cross-entropy loss;

[0111] A fourth processing unit 44, configured to process the first source domain feature and the second source domain feature by using a feature discriminator to obtain a binary classification cross-entropy loss, and use the gradient reversal algorithm to confuse the feature discriminator so that the feature discriminator cannot distinguish the first source domain feature and the second source domain feature. The first source domain feature is the source domain feature obtained after the data of the source domain is sent into the convolution layer of the cross-domain classification network currently participating in training. The second source domain feature is the source domain feature obtained after the data of the source domain is sent into the convolution layer of the incremental learning network;

[0112] A determination unit 45, configured to determine a total network loss according to the above-mentioned first loss, the above-mentioned second loss, the above-mentioned source domain self-supervised cross-entropy loss, and the above-mentioned binary classification cross-entropy loss, and optimize the parameters of the above-mentioned cross-domain classification network and the above-mentioned incremental learning network with the goal of minimizing the above-mentioned total network loss.

[0113] In an embodiment of the present application, the determination unit includes a first determination module and a second determination module;

[0114] The first determination module is configured to determine a single-domain single-target domain loss according to the above-mentioned source domain self-supervised cross-entropy loss and the above-mentioned first loss;

[0115] The second determination module is configured to determine the total network loss according to Formula 10:

[0116] L total = L STDA + L D + L bce (D(S prev ), 0) + L bce (D(S cur ), 1), where the total network loss is determined;

[0117] Wherein, L total is the total network loss, L STDA is the single-domain single-target domain loss, L D is the second loss, L bce (D(S prev ), 0) is the binary classification cross-entropy loss when the network prediction value is D(S prev ), and the label value is 0, L bce (D(S cur ), 1) is the binary classification cross-entropy loss when the network prediction value is D(S cur ), and the label value is 1, S cur is the first source domain feature, S prev is the second source domain feature.

[0118] In an embodiment of the present application, the first determination module includes a determination sub-module,

[0119] The determination sub-module is configured to determine the single-domain single-target domain loss according to where the single-domain single-target domain loss is determined;

[0120] Wherein, L STDA is the single-domain single-target domain loss, L ce is the source domain self-supervised cross-entropy loss, L s con is the intra-domain contrast loss of the source domain, L tcon is the in-domain contrast loss for the target domain, and λ1 and λ2 are weights for balancing the contrast loss and the source-domain self-supervised cross-entropy loss respectively.

[0121] In an embodiment of the present application, the second processing unit includes a third determination module.

[0122] The third determination module is used to determine the second loss according to L D = ||F s (x) - F'(x')|| 2 .

[0123] Wherein, L D is the second loss, F s (x) is the convolutional layer mapping of the cross-domain classification network in the current iteration, and F'(x') is the convolutional layer mapping of the model parameters saved by the incremental learning network in the previous cross-domain classification network training.

[0124] In an embodiment of the present application, the fourth processing unit includes a fourth determination module.

[0125] The fourth determination module is used to determine the binary classification cross-entropy loss according to Equation 8:

[0126] L bce (pt, label) = -label * ln(pt) + (1 - label) * ln(1 - pt),

[0127] to determine the binary classification cross-entropy loss.

[0128] Wherein, L bce (pt, label) is the binary classification cross-entropy loss, label is the label value, and pt is the network prediction value.

[0129] In an embodiment of the present application, the above learning device further includes a first processing module, which is used to add a target domain during the process of processing the data of the source domain and the target domain by using the cross-domain classification network each time to obtain the first loss, so as to perform a contrast process with the above source domain.

[0130] In an embodiment of the present application, the third processing unit includes a second processing module, a third processing module, a fourth processing module, and a fifth processing module. The second processing module is configured to input the data of the source domain into a feature extractor to obtain a feature vector; the third processing module is configured to convert the feature vector into a representation vector through a multi-layer perceptron; the fourth processing module is configured to process the representation vector by using the classifier to identify the category of the data of the source domain to obtain a predicted value of the data of the source domain; the fifth processing module is configured to calculate the self-supervised cross-entropy loss between the predicted value and the true label of the data of the source domain to obtain the self-supervised cross-entropy loss of the source domain.

[0131] The multi-domain data adaptive learning device includes a processor and a memory. The first processing unit, the second processing unit, the third processing unit, the fourth processing unit, the determination unit, etc. are all stored in the memory as program units, and the processor executes the program units stored in the memory to implement corresponding functions. All the above modules are located in the same processor; alternatively, the above modules are respectively located in different processors in any combination form.

[0132] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem of catastrophic forgetting caused by the inability to align the domain feature space due to too many target domains during the data feeding process in the source domain and the target domain in the existing solution can be solved.

[0133] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0134] An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the multi-domain data adaptive learning method.

[0135] An embodiment of the present invention provides a processor. The processor is used to run a program, wherein when the program runs, it executes the multi-domain data adaptive learning method.

[0136] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements at least the following steps: processing the data of the source domain and the target domain by using a cross-domain classification network to obtain a first loss, where the first loss includes a source domain loss and a target domain loss. The source domain loss is a contrast feature loss between the data of the source domain and the source domain enhanced data, and the target domain loss is a contrast feature loss between the data of the target domain and the target domain enhanced data. The source domain enhanced data is the data obtained by enhancing the data of the source domain by using an enhancement model, and the target domain enhanced data is the data obtained by enhancing the data of the target domain by using an enhancement model; processing the data of the source domain by using an incremental learning network to obtain a second loss, where the second loss is the loss of the difference between the convolution layer mapping result of the cross-domain classification network in the current iteration and the convolution layer mapping result of the model parameters saved by the incremental learning network in the previous cross-domain classification network training; processing the data of the source domain by using a classifier to obtain a source domain self-supervised cross-entropy loss; processing the first source domain feature and the second source domain feature by using a feature discriminator to obtain a binary classification cross-entropy loss, and using a gradient reversal algorithm to confuse the feature discriminator so that the feature discriminator cannot distinguish the first source domain feature and the second source domain feature. The first source domain feature is the source domain feature obtained after the data of the source domain is sent into the convolution layer of the currently participating cross-domain classification network, and the second source domain feature is the source domain feature obtained after the data of the source domain is sent into the convolution layer of the incremental learning network; determining a network total loss according to the first loss, the second loss, the source domain self-supervised cross-entropy loss, and the binary classification cross-entropy loss, and optimizing the parameters of the cross-domain classification network and the incremental learning network with the goal of minimizing the network total loss. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.

[0137] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps: processing data in a source domain and a target domain using a cross-domain classification network to obtain a first loss, where the first loss includes a source domain loss and a target domain loss, the source domain loss is a contrast feature loss between the data in the source domain and the source domain enhanced data, the target domain loss is a contrast feature loss between the data in the target domain and the target domain enhanced data, the source domain enhanced data is data obtained by enhancing the data in the source domain using an enhancement model, and the target domain enhanced data is data obtained by enhancing the data in the target domain using an enhancement model; processing the data in the source domain using an incremental learning network to obtain a second loss, where the second loss is a loss of the difference between the convolution layer mapping result of the cross-domain classification network in the current iteration and the convolution layer mapping result of the model parameters saved by the incremental learning network in the previous cross-domain classification network training; processing the data in the source domain using a classifier to obtain a source domain self-supervised cross-entropy loss; processing a first source domain feature and a second source domain feature using a feature discriminator to obtain a binary classification cross-entropy loss, and using a gradient reversal algorithm to confuse the feature discriminator so that the feature discriminator cannot distinguish the first source domain feature and the second source domain feature, where the first source domain feature is the source domain feature obtained after the data in the source domain is fed into the convolution layer of the currently participating cross-domain classification network for training, and the second source domain feature is the source domain feature obtained after the data in the source domain is fed into the convolution layer of the incremental learning network; determining a total network loss based on the first loss, the second loss, the source domain self-supervised cross-entropy loss, and the binary classification cross-entropy loss, and optimizing the parameters of the cross-domain classification network and the incremental learning network with the goal of minimizing the total network loss.

[0138] The present application also provides a multi-domain data adaptive learning system, including: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the above learning methods.

[0139] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program code executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0140] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0141] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the functions in the flowFigure 1 one or more processes and / or blocks Figure 1 steps of functions specified in one or more blocks

[0144] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0145] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0146] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0147] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-domain data adaptive learning method, characterized in that Including: Processing the data of the source domain and the target domain by using a cross-domain classification network to obtain a first loss, where the first loss includes a source domain loss and a target domain loss. The source domain loss is the contrast feature loss between the data of the source domain and the source domain augmented data, and the target domain loss is the contrast feature loss between the data of the target domain and the target domain augmented data. The source domain augmented data is the data obtained by performing augmentation processing on the data of the source domain by using an augmentation model, and the target domain augmented data is the data obtained by performing augmentation processing on the data of the target domain by using an augmentation model; Processing the data of the source domain by using an incremental learning network to obtain a second loss, where the second loss is the loss of the difference between the convolutional layer mapping result of the cross-domain classification network in the current iteration and the convolutional layer mapping result of the model parameters saved by the incremental learning network in the previous cross-domain classification network training; Processing the data of the source domain by using a classifier to obtain a source domain self-supervised cross-entropy loss; Processing the first source domain feature and the second source domain feature by using a feature discriminator to obtain a binary classification cross-entropy loss, and performing confusion processing on the feature discriminator by using a gradient reversal algorithm so that the feature discriminator cannot distinguish the first source domain feature and the second source domain feature. The first source domain feature is the source domain feature obtained after the data of the source domain is input into the convolutional layer of the currently participating cross-domain classification network, and the second source domain feature is the source domain feature obtained after the data of the source domain is input into the convolutional layer of the incremental learning network; Determining a network total loss according to the first loss, the second loss, the source domain self-supervised cross-entropy loss, and the binary classification cross-entropy loss, and optimizing the parameters of the cross-domain classification network and the incremental learning network with the goal of minimizing the network total loss.

2. The method according to claim 1, wherein Determining a network total loss according to the first loss, the second loss, the source domain self-supervised cross-entropy loss, and the binary classification cross-entropy loss, including: Determining a single-domain single-target domain loss according to the source domain self-supervised cross-entropy loss and the first loss; According to L total = L STDA + L D + L bce (D(S prev ), 0) + L bce (D(S cur ), 1), determine the total network loss; Among them, L total is the total network loss, L STDA is the single-domain single-target domain loss, L D is the second loss, L bce (D(S prev ), 0) is the binary classification cross-entropy loss when the network prediction value is D(S prev ), and the label value is 0, L bce (D(S cur ), 1) is the binary classification cross-entropy loss when the network prediction value is D(S cur ), and the label value is 1, S cur is the first source domain feature, S prev is the second source domain feature.

3. The method according to claim 2, wherein Determining a single-domain single-target domain loss according to the source domain self-supervised cross-entropy loss and the first loss, including: According to determine the single-domain single-target-domain loss; Among them, L STDA is the single-domain single-target domain loss, L ce is the source domain self-supervised cross-entropy loss, is the within-domain contrastive loss of the source domain, is the within-domain contrastive loss of the target domain, and λ1 and λ2 are the weights used to balance the contrastive loss and the source domain self-supervised cross-entropy loss respectively.

4. The method according to claim 1, wherein Processing the data of the source domain by using an incremental learning network to obtain a second loss, including: According to L D = ||F s (x) - F'(x')|| 2 , determine the second loss; where, L D is the second loss, F s (x) is the convolutional layer mapping of the cross-domain classification network in the current iteration, and F'(x') is the convolutional layer mapping of the model parameters saved by the incremental learning network for the previous training of the cross-domain classification network.

5. The method according to claim 1, wherein Processing the first source domain feature and the second source domain feature by using a feature discriminator to obtain a binary classification cross-entropy loss, including: According to L bce (pt, label) = -label * ln(pt) + (1 - label) * ln(1 - pt), Determining the binary classification cross-entropy loss; Among them, L bce (pt, label) is the binary classification cross-entropy loss, label is the label value, and pt is the network prediction value.

6. The method according to claim 1, wherein The method further includes: During the process of processing the data of the source domain and the target domain by using the cross-domain classification network each time to obtain a first loss, adding a target domain to perform contrast processing with the source domain.

7. The method according to any one of claims 1 to 6, characterized in that, Processing the data of the source domain by using a classifier to obtain a source domain self-supervised cross-entropy loss, including: Inputting the data of the source domain into a feature extractor to obtain a feature vector; Converting the feature vector into a representation vector through a multi-layer perceptron; Processing the representation vector by using the classifier to identify the category of the data of the source domain to obtain a predicted value of the data of the source domain. Calculate the self-supervised cross-entropy loss between the predicted value and the true label of the data in the source domain to obtain the source domain self-supervised cross-entropy loss.

8. A multi-domain data adaptive learning device, characterized in that, Comprising: A first processing unit configured to process the data in the source domain and the target domain using a cross-domain classification network to obtain a first loss, the first loss including a source domain loss and a target domain loss, the source domain loss being a contrast feature loss between the data in the source domain and the source domain augmented data, the target domain loss being a contrast feature loss between the data in the target domain and the target domain augmented data, the source domain augmented data being data obtained by performing augmentation processing on the data in the source domain using an augmentation model, and the target domain augmented data being data obtained by performing augmentation processing on the data in the target domain using an augmentation model; A second processing unit configured to process the data in the source domain using an incremental learning network to obtain a second loss, the second loss being the loss of the difference between the convolutional layer mapping result of the cross-domain classification network in the current iteration and the convolutional layer mapping result of the model parameters saved by the incremental learning network in the previous cross-domain classification network training; A third processing unit configured to process the data in the source domain using a classifier to obtain the source domain self-supervised cross-entropy loss; A fourth processing unit configured to process the first source domain feature and the second source domain feature using a feature discriminator to obtain a binary classification cross-entropy loss, and perform confusion processing on the feature discriminator using a gradient reversal algorithm so that the feature discriminator cannot distinguish between the first source domain feature and the second source domain feature, the first source domain feature being the source domain feature obtained after the data in the source domain is fed into the convolutional layer of the currently participating cross-domain classification network for training, and the second source domain feature being the source domain feature obtained after the data in the source domain is fed into the convolutional layer of the incremental learning network; A determination unit configured to determine a network total loss based on the first loss, the second loss, the source domain self-supervised cross-entropy loss, and the binary classification cross-entropy loss, and optimize the parameters of the cross-domain classification network and the incremental learning network with the goal of minimizing the network total loss.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the learning method according to any one of claims 1 to 7.

10. A multi-domain data adaptive learning system, characterized in that, Comprising: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing the learning method according to any one of claims 1 to 7.