Passive domain adaptive image classification method, system and device and storage medium
By introducing ViL model and low-rank adaptive LoRA method in passive domain adaptive image classification technology, a cross-model and cross-modal collaborative learning mechanism is adopted to form a teacher-student collaborative evolution training framework, solving the problems of model performance dependence, iterative training error accumulation and modal fragmentation in the existing technology, and achieving higher image classification accuracy and adaptability.
Patent Information
- Application Number
- CN202510204830.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
AI Technical Summary
The existing passive domain adaptive image classification technology faces the problems of excessive model performance dependence, error accumulation resulting from iterative training frameworks, and visual-language modal fragmentation, resulting in poor adaptability in the target domain.
By introducing ViL model as an auxiliary, a collaborative learning mechanism is used to cooperate across models and cross-modal collaboration, the ViL model is fine-tuned using the low-rank adaptive LoRA method, and the pre-trained target model is fine-tuned through knowledge transfer at the feature end and the prediction end to form a collaborative evolution training framework for teachers and students.
It effectively solves the over-dependence on model performance in traditional methods, improves the error accumulation phenomenon and modal fragmentation problems in the introduction of ViL model auxiliary methods, and improves the accuracy, adaptability and robustness of image classification.
Smart Images

Figure CN120147703A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and relates to a passive domain adaptation image classification method, system, device and storage medium. Background Art
[0002] Deep learning methods, which have achieved great success in the field of image processing and analysis, often rely on an important assumption that all data follows independent and identically distributed. However, in practical applications, due to factors such as image style, scene, acquisition device, modality, etc., there are inevitably distribution differences between data. This is usually referred to as the domain shift phenomenon, which greatly reduces the general applicability of deep learning models. The unsupervised domain adaptation task aims to solve the domain transfer problem and promote the generalization of the model to the target domain. Although excellent research results have been achieved, the problem with this task is that a large amount of source domain data is required in the training phase. In fact, this brings heavy storage and transmission burdens as well as data privacy issues. Therefore, the passive domain adaptation (SFDA) task has become the focus of recent research, aiming to adapt a source domain pre-trained model to the target domain without available source domain data.
[0003] Recently, some methods have introduced visual language multimodal large models (abbreviated as ViL models) into the SFDA task. This requires additional fine-tuning of the ViL model during training, allowing the customized ViL model to provide more reliable supervision signals and act as a "teacher" to guide the training of the "student" target model. The DIFO technology (Distilling Multimodal Foundation Model, multimodal foundation model distillation technology) uses mutual information to adjust the prediction distributions of the target model and the ViL model, so that both models can be adapted to the target domain. However, for technologies that do not introduce additional visual language multimodal large models: due to the lack of available labels in the target domain, existing traditional SFDA technologies mostly focus on using the predictions of "proximal" samples in the feature space as pseudo-labels for supervision. However, this technology is always limited by the performance of the model itself and is easily prone to training failure due to the unreliability of the pseudo-labels.
[0004] Meanwhile, for the technology of introducing an additional visual-linguistic multimodal large model (taking the state-of-the-art DIFO technology as an example), the following problems exist: 1. The mandatory iterative training framework leads to the accumulation of errors: The two models are iteratively updated in a unidirectional cycle, and the optimization objective is only provided by the mandatory pseudo-supervision information provided by the other model. When any one of the models performs poorly, its bias will be continuously amplified. This hinders the further exploration of the adaptation performance of the two models. 2. The disconnection between the visual and linguistic modalities: Only customizing the ViL model by optimizing the prefix prompts of the text encoder ignores the collaborative nature between the language and visual modalities, resulting in poor adaptability of the image encoder to the target domain. This is a very crucial idea in both the pre-training and fine-tuning stages of the ViL model. Consequently, the adaptability of the source-domain pre-trained model to the target domain is poor. Summary of the Invention
[0005] An object of the present invention is to overcome the above-mentioned disadvantages of the prior art and provide a source-domain-free adaptation image classification method, system, device, and storage medium.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In the first aspect of the present invention, a source-domain-free adaptation image classification method is provided, including: obtaining an image to be classified in the target domain; inputting the image to be classified in the target domain into a domain adaptation target model to obtain a classification result of the image to be classified in the target domain; wherein, the domain adaptation target model is obtained through the following training method: obtaining the prediction of the ViL model for the target domain training image as the first preliminary prediction, and obtaining the prediction of the pre-trained target model for the target domain training image as the second preliminary prediction; refining the second preliminary prediction according to the class prototype memory bank, and mixing the first preliminary prediction with the refined second preliminary prediction through an entropy uncertainty prediction mixing mechanism to obtain a mixed prediction; using the mixed prediction as a supervision signal, using the low-rank adaptation LoRA method to fine-tune the parameters of the ViL model, and using the ViL model after parameter fine-tuning as the alignment target for the feature level and prediction output level of the pre-trained target model to fine-tune the parameters of the pre-trained target model; iterating the above steps until a preset condition is reached, and taking the final pre-trained target model as the domain adaptation target model.
[0008] Optionally, the pre-trained target model is obtained through the following method: scaling the source domain training images to a unified shape and normalizing them using the mean and standard deviation of the ImageNet dataset to obtain preprocessed source domain training images; taking the minimization of the empirical risk on the source domain as the optimization objective, and pre-training the target model with the preprocessed source domain training images to obtain the pre-trained target model.
[0009] Optionally, refining the second preliminary prediction according to the class prototype memory bank includes: designing the class prototype memory bank as a parameterized matrix with uniformly distributed initialization where m j is the prototype of the j-th class, C is the number of classes; calculate the class attention of the second preliminary prediction through the following formula:
[0010] CA j = cos(p tar , m j )
[0011] where CA j is the class attention of the second preliminary prediction to the j-th class, cos(.) represents cosine similarity, and p tar is the second preliminary prediction.
[0012] Refine the second preliminary prediction through the following formula:
[0013]
[0014] where is the refined second preliminary prediction, T represents the matrix transpose operation.
[0015] Update the class prototype memory bank based on the weighted aggregation of p tar and the current class prototype memory bank according to the following formula:
[0016]
[0017] where α is a trade-off parameter, represents the indicator function, is the result class of the second preliminary prediction, is the result class of the first preliminary prediction.
[0018] Optionally, mixing the first preliminary prediction with the refined second preliminary prediction through the entropy uncertainty prediction mixing mechanism includes: obtaining the first preliminary prediction p clip based on the entropy-based prediction uncertainty u clip , and the entropy-based prediction uncertainty u of the refined second preliminary prediction tar :
[0019] u clip = -p clip · log(p clip )
[0020]
[0021] Obtain the mixed prediction p through the following formulamix :
[0022]
[0023] Among them, c tar = 1 - u tar is the prediction confidence of, c clip = 1 - u clip is the prediction confidence of p clip .
[0024] Optionally, the parameter fine-tuning of the ViL model using the low-rank adaptation LoRA method includes: deploying a trainable LoRA low-rank decomposition matrix as LoRA parameters for the ViL model, and introducing an anchor hint t a representing pure semantic information and an entanglement hint t e :
[0025] t a = a{class}
[0026] t e = a{domain}of a{class}
[0027] Both the anchor hint and the entanglement hint are an English sentence; where {class} is the class name with the highest prediction probability in the mixed prediction, and {domain} is the name of the domain.
[0028] Update the LoRA parameters of the ViL model by minimizing the LoRA parameter loss :
[0029]
[0030]
[0031] Among them, is the InfoNCE loss of the image-text pair, λ 1 and λ 2 are trade-off parameters, is the text similarity loss, represents taking the mathematical expectation, θ lora is the LoRA parameter, D t is the set of target domain training images, x i is the i-th target domain training image, is the sampling batch size, I i is the image feature output by the image encoder of x i based on the ViL model, and cos(.) represents the cosine similarity, Anchor hints for training images in the i-th target domain Text features output by the text encoder based on the ViL model, Train image anchor hints for the jth target domain Text features output by the text encoder based on the ViL model, is the absolute similarity loss, is the relative similarity loss, N s is the number of source domains, and x i The jth entanglement hint and the kth entanglement hint Text features output by the text encoder based on the ViL model.
[0032] Optionally, deploying a trainable LoRA low-rank decomposition matrix for the ViL model includes: deploying the LoRA low-rank decomposition matrix in the q matrix, k matrix and v matrix in all layers of the ViL model, and setting the low rank r of the LoRA low-rank decomposition matrix to 2.
[0033] Optionally, the fine-tuning of parameters of the pre-trained target model by using the fine-tuned ViL model as the alignment target of the feature level and the prediction output level of the pre-trained target model includes: obtaining the prediction of the target domain training image by the fine-tuned ViL model as the third preliminary prediction, and mixing the third preliminary prediction with the refined second preliminary prediction through an entropy uncertainty prediction mixing mechanism to obtain an accurate mixed prediction; using the accurate mixed prediction as a supervision signal, and fine-tuning the parameters of the pre-trained target model by minimizing the loss in the knowledge transfer stage according to the following formula:
[0034]
[0035]
[0036] in, is the loss in the knowledge transfer stage, is the mutual information loss, is the inner product loss, is the similarity-normalized learning loss, λ 3 and λ 4 is the trade-off parameter, θ t is the parameter of the pre-trained target model, E represents the mathematical expectation, and p tari Training image x for the i-th target domain i The second preliminary prediction is For x i The exact mixed prediction of t A training image set for the target domain, For xi The third preliminary prediction, where I() represents mutual information, is the anchor hint for the i-th target domain training image Based on the text features z output by the text encoder of the ViL model after parameter fine-tuning, i is x i Based on the intermediate image features of the pre-trained target model, cos(.) represents cosine similarity.
[0037] In the second aspect of the present invention, a passive domain adaptation image classification system is provided, including: an image acquisition module for acquiring the target domain image to be classified; an image classification module for inputting the target domain image to be classified into the domain adaptation target model to obtain the classification result of the target domain image to be classified; wherein, the domain adaptation target model is obtained through the following training method: obtaining the prediction of the ViL model for the target domain training image as the first preliminary prediction, and obtaining the prediction of the pre-trained target model for the target domain training image as the second preliminary prediction; refining the second preliminary prediction according to the category prototype memory bank, and mixing the first preliminary prediction with the refined second preliminary prediction through the entropy uncertainty prediction mixing mechanism to obtain a mixed prediction; using the low-rank adaptation LoRA method to fine-tune the parameters of the ViL model with the mixed prediction as the supervision signal, and using the ViL model with fine-tuned parameters as the alignment target for the feature level and prediction output level of the pre-trained target model to fine-tune the parameters of the pre-trained target model; iterating the above steps until the preset conditions, and using the final pre-trained target model as the domain adaptation target model.
[0038] In the third aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above passive domain adaptation image classification method are implemented.
[0039] In the fourth aspect of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above passive domain adaptation image classification method are implemented.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] The passive domain adaptation image classification method of the present invention classifies the images to be classified in the target domain based on the domain adaptation target model. Among them, the domain adaptation target model introduces the ViL model as an auxiliary, especially the core collaborative learning mechanism. This collaboration can be carried out across models to combine the advantages of the predictions of the two models, or across modalities. The low-rank adaptation LoRA method is used to readjust the visual and text features to achieve better task adaptation, effectively solving the over-reliance on the performance of the model itself in the traditional passive domain adaptation image classification self-training method, and improving the error accumulation phenomenon caused by the mandatory cyclic alignment iteration in the newer passive domain adaptation image classification method with the ViL model assistance and the disconnection phenomenon between the vision-language modalities in downstream task adaptation. Under the teacher-student co-evolution training framework of the present invention, cross-model collaboration provides reliable supervision for the evolution of the two models. Then, by encouraging cross-modal collaboration and finding domain-invariant features, the ViL model is fine-tuned. Then, the pre-trained target model is fine-tuned through knowledge transfer at the feature end and the prediction end, enabling it to learn more comprehensively and improve the classification performance in the target domain. The state-of-the-art results achieved on four major domain shift benchmark datasets prove that the present invention has higher accuracy, adaptability, and robustness in the passive domain adaptation image classification task. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flowchart of the passive domain adaptation image classification method according to an embodiment of the present invention.
[0043] Figure 2 It is a schematic diagram of the training principle of the domain adaptation target model according to an embodiment of the present invention.
[0044] Figure 3 It is a schematic diagram of the training process of the domain adaptation target model according to an embodiment of the present invention.
[0045] Figure 4 It is a graph showing the change of the classification accuracy of three different predictions generated during the training process of the teacher-student co-evolution training framework technology according to an embodiment of the present invention.
[0046] Figure 5 It is a graph showing the change of the classification accuracy during the training process under different ViL model fine-tuning strategies according to an embodiment of the present invention (taking the Others→Art task on the OfficeHome dataset as an example).
[0047] Figure 6This is a figure showing the t-distributed Stochastic Neighbor Embedding (t-SNE) distribution of the intermediate features of the domain adaptation target model obtained after training using the teacher-student co-evolution training framework technology of the embodiments of the present invention (taking the Others→Art task on the OfficeHome dataset as an example). Among them, Figure (a) shows the t-distributed Stochastic Neighbor Embedding distribution of the intermediate features of the domain adaptation target model pre-trained using the source domain empirical risk minimization technology; Figure (b) shows the t-distributed Stochastic Neighbor Embedding distribution of the intermediate features of the domain adaptation target model trained using the Source Hypothesis Transfer (SHOT) technology; Figure (c) shows the t-distributed Stochastic Neighbor Embedding distribution of the intermediate features of the domain adaptation target model trained using the multi-modal base model distillation (DIFO) technology; Figure (d) shows the t-distributed Stochastic Neighbor Embedding distribution of the intermediate features of the domain adaptation target model trained using the teacher-student co-evolution training framework (TSCEvo) technology of the embodiments of the present invention; Figure (e) shows the t-distributed Stochastic Neighbor Embedding distribution of the intermediate features of the domain adaptation target model trained using the oracle model (visible true labels in the target domain, for supervised training, as an upper limit).
[0048] Figure 7 This is a figure showing the t-distributed Stochastic Neighbor Embedding (t-SNE) distribution of the image features of the ViL model obtained after training using the teacher-student co-evolution training framework technology of the embodiments of the present invention (taking the Others→Art task on the OfficeHome dataset as an example). Among them, Figure (a) shows the t-distributed Stochastic Neighbor Embedding distribution of the zero-shot inference image features of the ViL model without using any fine-tuning technology; Figure (b) shows the t-distributed Stochastic Neighbor Embedding distribution of the image features of the ViL model obtained using the self-training fine-tuning technology; Figure (c) shows the t-distributed Stochastic Neighbor Embedding distribution of the image features of the ViL model trained using the teacher-student co-evolution training framework technology of the embodiments of the present invention; Figure (d) shows the t-distributed Stochastic Neighbor Embedding distribution of the image features of the ViL model trained using the oracle model.
[0049] Figure 8It is a t-distributed Stochastic Neighbor Embedding (t-SNE) distribution diagram of the text features of the ViL model obtained after training with the teacher-student co-evolution training framework technology in the embodiments of the present invention (taking the Others→Clipart task on the OfficeHome dataset as an example). Among them, Figure (a) is the t-distributed Stochastic Neighbor Embedding distribution of the zero-shot inference text features of the ViL model without using any fine-tuning technology; Figure (b) is the t-distributed Stochastic Neighbor Embedding distribution of the text features of the ViL model obtained using self-training fine-tuning technology; Figure (c) is the t-distributed Stochastic Neighbor Embedding distribution of the text features of the ViL model trained with the teacher-student co-evolution training framework technology in the embodiments of the present invention; Figure (d) is the t-distributed Stochastic Neighbor Embedding distribution of the text features of the ViL model trained using the prophecy model.
[0050] Figure 9 It is a probability visualization heat map of the category prototype memory bank in cross-model cooperation obtained after training with the teacher-student co-evolution training framework technology in the embodiments of the present invention (taking the Others→Clipart task on the OfficeHome dataset as an example).
[0051] Figure 10 It is a structural block diagram of the passive domain adaptation image classification system in the embodiments of the present invention. Detailed implementation manners
[0052] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0054] The present invention will be further described in detail below in conjunction with the accompanying drawings:
[0055] See Figure 1 , in one embodiment of the present invention, a passive domain adaptation image classification method is provided, specifically a teacher-student co-evolution passive domain adaptation image classification method in a data distribution shift scenario assisted by a vision-language multi-modal large model, which improves the accuracy, adaptability, and robustness of passive domain adaptation image classification.
[0056] Specifically, the passive domain adaptation image classification method of the present invention includes the following steps:
[0057] S1: Obtain the target domain images to be classified.
[0058] S2: Input the target domain images to be classified into the domain adaptation target model to obtain the classification results of the target domain images to be classified.
[0059] Among them, the domain adaptation target model is obtained through the following training method: Obtain the prediction of the ViL model for the target domain training images as the first preliminary prediction, and obtain the prediction of the pre-trained target model for the target domain training images as the second preliminary prediction; Refine the second preliminary prediction according to the category prototype memory bank, and mix the first preliminary prediction with the refined second preliminary prediction through the entropy uncertainty prediction mixing mechanism to obtain a mixed prediction; Use the mixed prediction as the supervision signal, use the low-rank adaptation LoRA method to fine-tune the parameters of the ViL model, and use the ViL model after parameter fine-tuning as the alignment target for the feature level and prediction output level of the pre-trained target model to fine-tune the parameters of the pre-trained target model; Iterate the above steps until the preset conditions are met, and use the final pre-trained target model as the domain adaptation target model.
[0060] The passive domain adaptation image classification method of the present invention classifies the images to be classified in the target domain based on the domain adaptation target model. Among them, the domain adaptation target model is assisted by introducing the ViL model, especially the core collaborative learning mechanism. This collaboration can be carried out across models to combine the advantages of the predictions of the two models, or across modalities, using the low-rank adaptation LoRA method to readjust the visual and text features to achieve better task adaptation. It effectively solves the over-reliance on the performance of the model itself in the traditional passive domain adaptation image classification self-training method, and improves the error accumulation phenomenon caused by the mandatory cyclic alignment iteration in the newer passive domain adaptation image classification method assisted by the ViL model and the disconnection phenomenon between the vision-language modalities in downstream task adaptation. Under the teacher-student co-evolution training framework of the present invention, cross-model collaboration provides reliable supervision for the evolution of the two models. Then, by encouraging cross-modal collaboration and finding domain-invariant features, the ViL model is fine-tuned. Then, through knowledge transfer at the feature end and prediction end, the pre-trained target model is fine-tuned to enable it to learn more comprehensively and improve the classification performance in the target domain. The state-of-the-art results achieved on four major domain shift benchmark datasets prove that the present invention has higher accuracy, adaptability, and robustness in the passive domain adaptation image classification task.
[0061] First, to explain the specific implementation of the present invention, the following notations need to be introduced and unifiedly described and defined here: The present invention involves two models: the target model f(x; θ) → y and the ViL model where x represents the image input, t represents the text input, is the classification label, C is the number of categories, and D represents the model parameters. Using CLIP as an instance of the ViL model, this model can be further decomposed into a parallel image encoder and a text encoder These encoders take x and t as inputs respectively and generate the image feature I and the text feature T. The teacher-student co-evolution training framework technology of the present invention will be based on the unlabeled target domain training images D t ={(x t )} to iteratively update the models f and All image data are preprocessed by shape scaling and ImageNet mean standard deviation normalization. For the target model, the initialization parameters θ s from multiple source domain ERM pre-training are gradually updated to θ t such that For the ViL model, additional effective parameters θ lora are added using LoRA such that After training, only the target model is used for inference.
[0062] In a possible implementation, see Figure 2 and3 , the training process of the domain adaptation target model of the present invention can be summarized into the following two training stages:
[0063] Stage 1: Evolution of the collaborative ViL model. At the beginning of each training epoch, first perform forward propagation through two models to obtain their respective predictions. Then input these predictions into the cross-model collaboration module to generate highly reliable hybrid predictions. Anchor cues and entanglement cues are constructed for each sample according to two types of templates and judgments from the hybrid predictions. All of the above are used to update the LoRA parameters in the two encoders of the ViL model through a carefully designed loss function.
[0064] Stage 2: Evolution of the pre-trained target model. The present invention freezes the ViL model and re-iterates all samples. Reuse cross-model collaboration to generate new guidance through the optimized ViL model after fine-tuning. Then update the pre-trained target model parameters by optimizing three types of knowledge transfer losses.
[0065] In a possible implementation manner, the pre-trained target model is obtained by the following method: scaling the source domain training images to a unified shape and normalizing them using the mean and standard deviation of the ImageNet dataset to obtain pre-processed source domain training images; taking the minimization of the empirical risk on the source domain as the optimization objective, and pre-training the target model with the pre-processed source domain training images to obtain the pre-trained target model.
[0066] In a possible implementation manner, the method for obtaining the prediction of the ViL model for the target domain training images as the first preliminary prediction and the prediction of the pre-trained target model for the target domain training images as the second preliminary prediction is as follows: scaling the target domain training images to a unified shape and normalizing them using the mean and standard deviation of the ImageNet dataset to obtain pre-processed target domain training images. For the target model, load the pre-trained model weight parameters and input the target domain training images to obtain the second preliminary prediction; for the ViL model, use the official pre-trained model weight parameters, construct a zero-shot linear classifier according to its official inference text prompt, and input the target domain training images into the image encoder and classifier in sequence to obtain the first preliminary prediction.
[0067] In a possible implementation manner, the refinement of the second preliminary prediction according to the category prototype memory bank includes: designing the category prototype memory bank as a parameterized matrix with uniformly distributed initialization where m j is the prototype of the j-th category, and C is the number of categories.
[0068] Calculate the category attention of the second preliminary prediction through the following formula:
[0069] CA j = cos(p tar , m j )
[0070] where CA j is the class attention of the second preliminary prediction for the j-th class, cos(.) represents cosine similarity, and p tar is the second preliminary prediction.
[0071] The second preliminary prediction is refined by the following formula:
[0072]
[0073] where is the refined second preliminary prediction, T represents the matrix transpose operation.
[0074] Based on the following formula, the class prototype memory bank is updated by the weighted aggregation of p tar and the current class prototype memory bank:
[0075]
[0076] where α is a trade-off parameter, represents the indicator function, is the class of the second preliminary prediction result, is the class of the first preliminary prediction result.
[0077] In a possible implementation, the mixing of the first preliminary prediction and the refined second preliminary prediction by the entropy uncertainty prediction mixing mechanism includes:
[0078] The first preliminary prediction p clip is obtained based on the prediction uncertainty u clip of entropy, and the prediction uncertainty u of entropy of the refined second preliminary prediction tar :
[0079] u clip = -p clip · log(p clip )
[0080]
[0081] The mixed prediction p mix is obtained by the following formula:
[0082]
[0083] where c tar = 1 - u tarFor the prediction confidence of, c clip = 1 - u clip is p clip the prediction confidence of.
[0084] Explanatory, the initial prediction of the pre-trained target model is refined through the class prototype memory bank for cross-model collaboration, and then the initial predictions from the two models are mixed through the cross-model collaboration entropy uncertainty prediction mixing mechanism to achieve a cooperative complementary effect. Specifically, for the initial prediction generated by the pre-trained target model, the class attention is calculated with the class prototype memory bank, and the class attention is used as the mixing weight to guide the generation of a robust refined prediction of the target model. Calculate the entropy uncertainty of the initial prediction of the pre-trained target model and the initial prediction of the ViL model, and use the entropy uncertainty as the mixing weight to generate a mixed prediction.
[0085] Specifically, in order to achieve better adaptation quality, more reliable supervision signals are needed. Therefore, the present invention deeply explores the complementary properties of the pre-trained target model and the ViL model.
[0086] Specifically, considering two predictions from the same target domain training image x t :
[0087] p tar = σ(f(x t ; θ t ))
[0088] p clip = σ(q(I, t c ))
[0089] where σ(·) represents the softmax function, q(·) is the zero-shot linear classifier of CLIP, synthesized by the learned text encoder, and t c represents C text prompts generated according to the official template of the ViL model: a photo of a {class}. Here, {class} is the name of each possible prediction class, and C is the number of classes.
[0090] Since the performance of p tar in the initial stage of training still needs to be optimized, which may lead to unstable results. Therefore, first use the class prototype memory bank to improve and adjust p tar . The class prototype memory bank is designed as a parameterized matrix M ∈ R C×C . This matrix represents the C-dimensional prediction prototype of each class, that is m j ∈ R C , j represents the class index. Given an input prediction p tar, the category prototype memory bank first generates category attention using stored past robust predictions:
[0091] CA j = cos(p tar ,m j )
[0092] where cos(·) represents cosine similarity. The category attention represents the judgment tendency for each category. Therefore, combining the category attention with the category prediction prototypes in the category prototype memory bank can obtain a more robust prediction, i.e., the refined second preliminary prediction: tar where T represents the matrix transpose operation.
[0093]
[0094] After that, to flexibly consider each sample, an entropy-based uncertainty prediction mixture is introduced to determine which model's prediction should bear a greater weight in the mixture process. Specifically, it is calculated through u = -p·log(p)
[0095] and the entropy-based prediction uncertainty u of p clip and u tar and u clip .
[0096] Then, the mixed prediction p mix is generated through the following formula:
[0097]
[0098] where the confidence c is defined as: c = 1 - u, c ∈ R C .
[0099] Finally, the category prototype memory bank is updated in a weighted ensemble manner, and the category prototype memory bank is updated through the weighted aggregation of p tar and the current category prototype:
[0100]
[0101] where α is a trade-off parameter, represents the indicator function, which is used to ensure that only reliable p tar can participate in the memory bank update process.
[0102] In this way, the collaboration mechanism can effectively utilize the combined advantages of the pre-trained target model and the ViL model even in challenging scenarios. This is crucial for the smooth and efficient training of the two models in the subsequent steps.
[0103] In a possible implementation, the parameter fine-tuning of the ViL model using the Low-Rank Adaptation (LoRA) method includes:
[0104] Deploying a trainable LoRA low-rank decomposition matrix as LoRA parameters for the ViL model, and introducing an anchor prompt t representing pure semantic information a and an entanglement prompt t representing style-semantic entanglement information e :
[0105] t a = a{class}
[0106] t e = a{domain} of a{class}
[0107] Both the anchor prompt and the entanglement prompt are an English sentence. Among them, {class} is the name of the category with the highest prediction probability in the mixed prediction, and {domain} is the name of the domain.
[0108] Updating the LoRA parameters of the ViL model by minimizing the LoRA parameter loss :
[0109]
[0110] Where is the InfoNCE loss of the image-text pair, λ 1 and λ 2 are trade-off parameters, is the text similarity loss, represents taking the mathematical expectation, θ lora is the LoRA parameter, D t is the set of target domain training images, x i is the i-th target domain training image, is the sampling batch size, I i is for x i the image feature output by the image encoder of the ViL model, cos(.) represents the cosine similarity, is the anchor prompt of the i-th target domain training image the text feature output by the text encoder of the ViL model, is the anchor prompt of the j-th target domain training image the text feature output by the text encoder of the ViL model, is the absolute similarity loss, is the relative similarity loss, N s is the number of source domains, and are for x ithe j-th entanglement prompt and the k-th entanglement prompt Text features output by the text encoder based on the ViL model.
[0111] Optionally, the deployable trainable LoRA low-rank decomposition matrix for the ViL model includes: deploying the LoRA low-rank decomposition matrix in the q matrix, k matrix, and v matrix of all layers of the ViL model, and setting the low rank r of the LoRA low-rank decomposition matrix to 2.
[0112] Explanatory, using the mixed prediction as the supervision signal, and using the low-rank adaptation LoRA method to fine-tune the ViL model in a cross-modal collaboration and domain-invariance encouraging manner; the present invention applies LoRA to the downstream task adaptation fine-tuning of the original ViL model, which means that all the original parameters of the ViL model remain unchanged, and only the low-rank decomposition matrix introduced by LoRA is trainable.
[0113] By deploying a trainable LoRA low-rank decomposition matrix for the ViL model. Under the guidance of cross-model collaborative mixed prediction, two types of text prompts are constructed: anchor prompts and entanglement prompts. By encouraging the absolute similarity and relative similarity between the anchor prompts and the entanglement prompts, the ViL model text encoder is prompted to acquire the ability to find domain-invariant features. The cross-modal collaboration is achieved by calculating the contrastive loss based on the anchor prompt features and the target domain image features. The above two jointly guide the LoRA parameter fine-tuning of the ViL model.
[0114] In the teacher-student co-evolution framework of the present invention, the mixed prediction obtained through cross-model collaboration is used as reliable supervision. Specifically, there are two optimization objectives to guide the fine-tuning process:
[0115] One is text similarity encouragement. By introducing the name of the source domain to slightly relax the SFDA task constraints, the text encoder of the ViL model can extract domain-invariant features by encouraging feature consistency.
[0116] The present invention designs two types of prompt templates: anchor prompt t a , representing pure semantic information, and entanglement prompt t e , representing style-semantic entanglement information:
[0117] t a = a{class}
[0118] t e = a{domain} of a{class}
[0119] where {class} represents p mixThe class name with the highest predicted probability in the prediction, and {domain} represents the name of the domain. In the leave-one-domain strategy, the present invention assigns the name of each source domain to {domain}. Therefore, for a target domain training image x i ∈D t , a series of entangled prompts can be obtained N s represents the number of source domains.
[0120] First, optimize the absolute similarity loss:
[0121]
[0122] where T a or represents the text feature output by the text encoder of the ViL model after inputting the corresponding prompt t a or t e . In this way, it can be encouraged that the entangled prompts have similar feature outputs to the corresponding anchor prompts, thereby pulling the entangled text features towards the semantic subspace.
[0123] Next, optimize the relative similarity loss, which will help encourage the entangled prompts belonging to different domains to have similar feature outputs and explicitly promote domain invariance:
[0124]
[0125] The overall text similarity loss is defined as:
[0126]
[0127] where λ 1 is a trade-off parameter.
[0128] Second, it is cross-modal contrast alignment. It is far from enough to only adapt the language modality of the ViL model to downstream tasks. More importantly, this method may even damage the visual-language alignment property of the ViL model. Therefore, the present invention deploys LoRA for both the image and text encoders of CLIP to explore cross-modal collaboration in the adaptation of the ViL model to downstream tasks. Specifically, the InfoNCE loss of the image-text pair is reused as the optimization objective:
[0129]
[0130] where, represents the image feature of the i-th sample, is the sampling batch size. Since the anchor prompts are constructed according to the mixed prediction p mix , this will provide strong cross-model collaboration supervision for the LoRA parameters.
[0131] Finally, in the fine-tuning stage, the following loss function was optimized to update all LoRA parameters:
[0132]
[0133] where λ 2 is a trade-off parameter.
[0134] In one possible implementation, the parameter fine-tuning of the pre-training target model by using the ViL model with fine-tuned parameters as the alignment target at the feature level and the prediction output level includes:
[0135] Obtaining the prediction of the ViL model with fine-tuned parameters for the training images in the target domain as the third preliminary prediction, and mixing the third preliminary prediction with the refined second preliminary prediction through an entropy uncertainty prediction mixing mechanism to obtain an accurate mixed prediction; using the accurate mixed prediction as a supervision signal, and fine-tuning the parameters of the pre-training target model by minimizing the loss in the knowledge transfer stage according to the following formula:
[0136]
[0137] where is the loss in the knowledge transfer stage, is the mutual information loss, is the inner product loss, is the similarity regularization learning loss, λ 3 and λ 4 are trade-off parameters, θ t are the parameters of the pre-training target model, represents taking the mathematical expectation, is the second preliminary prediction of the i-th training image x i in the target domain, is the accurate mixed prediction of x i , ⊙ represents the inner product operation, D t is the set of training images in the target domain, is the third preliminary prediction of x i , I() represents the mutual information, is the anchor hint of the i-th training image in the target domain based on the text features output by the text encoder of the ViL model with fine-tuned parameters, z i is the intermediate image feature of x i based on the pre-training target model, and cos(.) represents the cosine similarity.
[0138] Explanatorily, using the mixed prediction as a supervision signal and combining the fine-tuned ViL model as the alignment target at the feature level and the prediction output level to achieve knowledge transfer from the ViL model to the pre-training target model.
[0139] Using cross-model collaborative hybrid prediction as the supervised pseudo-label, without losing prior information, encourages the pre-trained target model to predict approaching the hybrid prediction; uses mutual information as a metric to encourage the pre-trained target model's prediction to align with the ViL model's prediction; encourages the cross-modal similarity regularization learning process between the intermediate features of the pre-trained target model and the anchor prompt features. Finally, the above three jointly guide the training of the target model.
[0140] Through carefully designed co-evolution, the ViL model now has excellent capabilities for specific target domains. The next goal is to transfer knowledge from the ViL model to the pre-trained target model. Therefore, keep the ViL model parameters unchanged and use a new optimizer to update the parameters of the pre-trained target model.
[0141] Specifically, the present invention re-iterates the target domain training image x t ∈D t , and uses the currently fine-tuned ViL model to generate more accurate CLIP predictions p ′ clip . Subsequently, reuse cross-model collaboration to obtain p ′ mix .
[0142] To make full use of the prior knowledge represented by the prediction probability of each class in p ′ mix , in this embodiment, it is chosen not to use the argmax operation to convert p ′ mix into a standard one-hot label, because this assigns equal importance to all negative classes. Instead, directly minimize the negative value of the inner product between the two prediction vectors as the optimization objective:
[0143]
[0144] where ⊙ represents the inner product operation. This enables the pre-trained target model to make full use of the supervision provided by p ′ mix , which contains valuable knowledge obtained from the fine-tuned vision-language model.
[0145] In addition, mutual information is a powerful means to align the prediction distributions. Therefore, the following optimization objective is introduced:
[0146]
[0147] where I(.) represents mutual information. By maximizing the mutual information, the pre-trained target model can be prompted to mimic the prediction probability distribution of the vision-language model.
[0148] Meanwhile, a feature distance loss is introduced to estimate cross-modal feature alignment. The anchor text feature Ta Rich in general semantic knowledge of this category. Therefore, matching the intermediate image feature z learned by the pre-trained target model encoder with T a can standardize the learning process of the target model. Specifically, there are the following optimization objectives:
[0149]
[0150] Among them, is generated using the text encoder of the currently fine-tuned ViL model.
[0151] In summary, the overall optimization objective in the knowledge transfer stage is defined as follows:
[0152]
[0153] Among them, λ 3 and λ 4 are trade-off parameters.
[0154] To verify the image classification performance in the SFDA task scenario of the present invention, the following will compare with the currently most advanced SFDA method in the industry and conduct ablation studies to analyze various components of the present invention. In addition, further model analysis is carried out to verify some effects mentioned in the present invention.
[0155] First, the experimental settings are described as follows:
[0156] Passive Domain Adaptation Setup: The present invention explores the more common SFDA domain split setup in practical applications, namely the leave-one-domain cross-validation evaluation protocol, aiming to adapt models pre-trained from multiple source domains to a single target domain. In terms of the class label space problem, the most common closed-set setup is adopted, where the training and testing phases share the same class space. Datasets: The present invention evaluates four standard domain shift task benchmarks: PACS, Office-Home, VLCS, and TerraIncognita. Among them, TerraIncognita is a highly challenging real-world dataset that contains images captured by camera traps deployed for monitoring animal populations, and its domain is defined by different camera positions. Baseline Methods: The technology described in the present invention is compared with 13 existing excellent SFDA technologies. These technical methods can be divided into three major categories: (1) The first group consists of baseline models that do not contain any domain adaptation technology: ERM (pre-trained on the source domain), CLIP (zero-shot). (2) The second group includes the current state-of-the-art methods that do not introduce the ViL model as an auxiliary: SHOT, NRC, GKD, AdaContrast, CoWA, SCLM, UniDG, TPDS, and PLUE. (3) The third group includes the state-of-the-art methods that introduce the ViL model, including PromptStyler and DIFO. To ensure a fair comparison, the results of the above methods were reproduced using the provided code in the SFDA setup described in the present invention. However, PromptStyler is an exception because it is not an open-source work. Therefore, only the results reported in its original paper are presented.
[0157] Implementation Details: Unless otherwise specified, all target models in the experimental results reported in the present invention use ResNet-50, and all ViL models use CLIP with a ViT-B / 16 image encoder. The present invention includes five trade-off hyperparameters, among which α is set to 0.5, λ 1 and λ 2 are set to 1.0, and λ 3 and λ 4 are searched within the range [0.1, 10.0]. Unless otherwise specified, in all experiments of the present invention, LoRA is applied to the q, k, v attention matrices in all layers of CLIP, and the low rank r is set to 2. In all experiments of the present invention, the SGD optimizer is used for and the AdamW optimizer is used for
[0158] Both optimizers are configured with a cosine annealing learning rate scheduler. All experiments are conducted using the PyTorch deep learning framework on a single NVIDIA RTX 4090 GPU.
[0159] 1. Comparison results of baseline methods. Refer to Table 1, which shows the comparison results of the method of the present invention (TSCEvo) and all baseline methods on three synthetic datasets (PACS, OfficeHome, and VLCS). Compared with the state-of-the-art method DIFO, the average accuracy of the present invention has increased by 1.3% (on PACS), 8.0% (on OfficeHome), 1.9% (on VLCS), and 3.7% (overall average). Compared with the most powerful method without the assistance of the ViL model, the method of the present invention has increased the average accuracy by 4.9% (compared with TPDS on PACS), 11.1% (compared with TPDS on OfficeHome), and 3.7% (compared with UniDG on VLCS), respectively. Even the weaker version (*) using ViT-B / 32 as the CLIP image encoder also shows impressive performance.
[0160] Table 1
[0161]
[0162] Refer to Table 2, which shows the comparison results of the method of the present invention and baseline methods on the challenging real-world task dataset TerraIncognita. Since the images in this dataset are much more complex, the zero-shot CLIP that has never seen this dataset performs even worse than the ERM target model (31.5% vs. 48.0%). This contradicts the original intention of introducing the ViL model as a reliable teacher. Therefore, first fine-tune CLIP on the source domain, denoted as CLIP(ft), which will improve its performance on the target domain. Then, use CLIP(ft) to initialize the ViL model and continue the training process (*).
[0163] Compared with the state-of-the-art method, the method of the present invention has significantly increased the average accuracy on the TerraIncognita dataset by 12.6%. In addition, it is found that DIFO* performs worse than CLIP(ft). It is speculated that this is because it over-relies on a poor target model (average accuracy of only 48.0%) when customizing the ViL model, and the insufficient representation ability of prefix prompt learning is difficult to adapt to more complex tasks.
[0164] Table 2
[0165]
[0166]
[0167] 2. Ablation experiment results. Refer to Table 3. Through ablation experiments, the effectiveness of each component in the technology of the present invention is demonstrated. Comparing #1 with #2, #3, and #4, it can be seen that the introduction of three types of knowledge transfer losses all makes positive contributions. It is worth noting that the widely used mutual information loss performs poorly in dealing with the complexity of the Terra dataset, while the prior knowledge-based in the technology of the present invention effectively solves this limitation. Comparing #4 with #5 and #6 shows that fine-tuning the ViL model is beneficial to the target model because it can obtain more reliable guidance. The results of #7 and #8 prove the effectiveness of the two components in the cross-model collaboration module proposed in the technology of the present invention. Here, #8 represents the result without calculating uncertainty, but assigning the same mixing weight to the two predictions.
[0168] Table 3
[0169]
[0170] 3. Results of the reliability verification experiment for hybrid prediction. Refer to Figure 4 , which shows the accuracy changes of three different predictions during the training process. It can be observed that the p mix obtained through cross-model collaboration always reaches the highest accuracy, proving its reliability as a supervision signal. In addition, as the training progresses, the accuracy difference between the predictions gradually narrows, indicating that the evolution mechanism in the teacher-student co-evolution training framework technology of the present invention is effective for both the ViL model and the target model.
[0171] 4. Results of the ViL model performance experiment under different fine-tuning strategies. After the teacher-student co-evolution training framework technology of the present invention achieved such an impressive performance improvement, an attempt was made to explore the root cause. A possible factor is that the fine-tuning strategy under the technology framework of the present invention provides a more powerful ViL model as the teacher. Therefore, refer to Table 4 to examine the performance of the ViL model under different fine-tuning strategies in the SFDA task.
[0172] The results show that, except for the oracle method (visible target domain true labels, supervised training, used as the upper limit), the technology of the present invention indeed achieves the best ViL model fine-tuning performance. Compared with the DIFO method, this not only proves that the collaboration-based fine-tuning strategy of the present invention is more effective than prefix prompt learning in the SFDA task, but also confirms that a better teacher will lead to a better student. Refer to Figure 5 , which also visually shows the accuracy changes of the ViL model under different fine-tuning methods during training.
[0173] Table 4
[0174]
[0175] 5. Experimental results of tuning related parameters of Low-Rank Adaptation (LoRA). The application position of LoRA in the ViL model and the low rank r are two hyperparameters in the method of the present invention. To explore their effects and determine the best LoRA application specific to the SFDA task, the effects of the position and low rank were studied under the constraint of a fixed additional model parameter budget (about 1MB in FP16 data type). Referring to Table 5, the default setting of the present invention (applying LoRA to the q, k, and v matrices in the attention layer and setting the low rank r = 2) achieved the best average performance. Using a higher low rank at a single position generally performed worse than applying a lower low rank at more positions, which is consistent with the results reported in previous related studies for natural language processing tasks.
[0176] Table 5
[0177]
[0178]
[0179] 6. Visualization experiment of feature distribution. The t-distributed Stochastic Neighbor Embedding (t-SNE) technique was used to visualize the distribution of the intermediate features z of the target model under different methods. As Figure 6 shown, although DIFO is the best-performing method at present, its feature distribution is still quite chaotic. In contrast, the TSCEvo method proposed by the present invention exhibits a more convincing feature distribution, even very similar to the oracle model.
[0180] Referring to Figure 7 , a visual comparison of the t-SNE feature distributions of the CLIP image encoder under different fine-tuning strategies is shown. Obviously, if no downstream task adaptation is performed on the image encoder of the ViL model, its zero-shot inference will not produce an ideal image feature distribution. This will inevitably hinder the performance of the ViL model in the SFDA task. This proves the necessity of introducing a cross-modal collaborative fine-tuning strategy in the technology of the present invention. In addition, a weak fine-tuning strategy, such as self-training, also cannot produce a good image feature distribution of the ViL model. In contrast, the TSCEvo method of the present invention deeply explores the potential ability of the ViL model in the SFDA task and achieves a feature distribution close to that of the oracle model. This proves the excellent ability of the technology of the present invention in achieving cross-modal collaboration. Referring to Figure 8, it shows the visual comparison of the t-SNE feature distributions of the CLIP text encoder under different fine-tuning strategies. To obtain these text features, each domain and class name is first assigned to the template "a {domain} of a {class}" to construct prompts, and then they are input into the CLIP text encoder under different fine-tuning strategies. By carefully observing the results under the zero-shot and self-training strategies, it can be found that the text features of the target domain (Clipart) are difficult to adapt to the decision boundary constructed by the text features of the source domains (Art, RealWorld, and Product). In contrast, the results under the TSCEvo method of the present invention show clear classification boundaries, and the features of samples from the same class but different domains show obvious overlap. This intuitively indicates that the ViL model fine-tuned under the training framework of the present invention has found domain-invariant features, overcoming the adverse effects brought by the domain shift problem. Therefore, the present invention has indeed learned better feature representations than any other competitors, thus supporting its achievement of the new state-of-the-art SFDA task performance.
[0181] 7. Visualization experiment of the class prototype memory bank. See Figure 9 , which presents the class prototype memory bank in the cross-model collaboration part of the TSCEvo framework of the present invention in the form of a probability heatmap. Each row in the figure represents the predicted prototype of the corresponding class, and each column represents the predicted probability of this class in the prototype. Obviously, the higher probability values are concentrated on the diagonal, which indicates that each predicted prototype accurately captures the general information belonging to this class.
[0182] By observing some prototypes with relatively low predicted probabilities on the diagonal, such as Figure 9 the prototype of the "Maker (marker)" class marked with a yellow dashed line in
[0183] The following is the device embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For the details not disclosed in the device embodiment, please refer to the method embodiment of the present invention.
[0184] See Figure 10, in another embodiment of the present invention, a passive domain adaptation image classification system is provided, which can be used to implement the above-mentioned passive domain adaptation image classification method. Specifically, the passive domain adaptation image classification system includes an image acquisition module and an image classification module.
[0185] Among them, the image acquisition module is used to acquire the target domain images to be classified; the image classification module is used to input the target domain images to be classified into the domain adaptation target model to obtain the classification results of the target domain images to be classified. Among them, the domain adaptation target model is obtained through the following training method: obtaining the prediction of the ViL model for the target domain training images as the first preliminary prediction, and obtaining the prediction of the pre-trained target model for the target domain training images as the second preliminary prediction; refining the second preliminary prediction according to the category prototype memory bank, and mixing the first preliminary prediction with the refined second preliminary prediction through the entropy uncertainty prediction mixing mechanism to obtain a mixed prediction; using the low-rank adaptation LoRA method to fine-tune the parameters of the ViL model with the mixed prediction as the supervision signal, and using the ViL model after parameter fine-tuning as the alignment target of the feature level and prediction output level of the pre-trained target model to fine-tune the parameters of the pre-trained target model; iterating the above steps until the preset conditions are met, and using the final pre-trained target model as the domain adaptation target model.
[0186] All relevant contents of each step involved in the embodiment of the foregoing passive domain adaptation image classification method can be cited in the function description of the corresponding functional modules of the passive domain adaptation image classification system in the embodiment of the present invention, and will not be repeated here.
[0187] The division of modules in the embodiments of the present invention is illustrative, only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present invention, each functional module can be integrated in one processor, or can exist independently physically, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0188] In another embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used for the operation of the passive domain adaptation image classification method.
[0189] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space that stores the operating system of the terminal. And in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the passive domain adaptation image classification method in the above embodiment.
[0190] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0191] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0192] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A passive domain adaptation image classification method, characterized in that: include: Obtain the target domain image to be classified; Adapt the input domain of the target domain image to be classified to the target model, and obtain the classification result of the target domain image to be classified; Among them, the domain adaptation target model is obtained through the following training method: Obtaining the prediction of the ViL model for the target domain training image as a first preliminary prediction, and obtaining the prediction of the pre-trained target model for the target domain training image as a second preliminary prediction; Refining the second preliminary prediction according to the category prototype memory bank, and mixing the first preliminary prediction with the refined second preliminary prediction through an entropy uncertainty prediction mixing mechanism to obtain a mixed prediction; Using the mixed prediction as the supervisory signal, the low-rank adaptive LoRA method is used to fine-tune the parameters of the ViL model, and the fine-tuned ViL model is used as the alignment target of the feature level and prediction output level of the pre-trained target model to fine-tune the parameters of the pre-trained target model; Iterate the above steps to the preset conditions and use the final pre-trained target model as the domain adaptation target model.
2. The passive domain adaptation image classification method according to claim 1, characterized in that: The pre-trained target model is obtained in the following way: The source domain training images are scaled to a uniform shape and normalized using the mean standard deviation of the ImageNet dataset to obtain preprocessed source domain training images. Taking the empirical risk minimization on the source domain as the optimization goal, the target model is pre-trained by preprocessing the source domain training images to obtain a pre-trained target model.
3. The passive domain adaptation image classification method according to claim 1, characterized in that: The refining of the second preliminary prediction according to the category prototype memory library comprises: Designing the Category Prototype Memory as a Parameterized Matrix with Uniform Initialization Among them, m j is the prototype of the jth category, and C is the number of categories; The category attention of the second preliminary prediction is calculated as follows: THAT j =cos(p tar ,m j ) Among them, CA j is the second preliminary prediction of the class attention for the jth class, cos(.) represents the cosine similarity, and p tar For the second preliminary prediction; The second preliminary prediction is refined by the following formula: in, is the second preliminary forecast after refinement, T represents the matrix transpose operation; According to the following formula based on p tar Update the category prototype memory by weighted aggregation of the current category prototype memory: Among them, α is the trade-off parameter, represents the indicator function, is the second preliminary prediction result category, It is the first preliminary prediction result category.
4. The passive domain adaptation image classification method according to claim 1, characterized in that: The mixing of the first preliminary prediction with the refined second preliminary prediction by the entropy uncertainty prediction mixing mechanism comprises: The first preliminary prediction p is obtained by clip Entropy-based prediction uncertainty u clip , and the refined second preliminary prediction The entropy-based prediction uncertainty u tar : u clip =-p clip ·log(p clip ) The hybrid prediction p is obtained by the following formula mix : Among them, c tar =1-u tar for The prediction confidence, c clip =1-u clip For p clip prediction confidence.
5. The passive domain adaptation image classification method according to claim 1, characterized in that: The use of the low-rank adaptive LoRA method to fine-tune the parameters of the ViL model includes: The ViL model deploys a trainable LoRA low-rank decomposition matrix as LoRA parameters and introduces anchor prompts t representing pure semantic information. a and the entanglement hint t representing the style semantic entanglement information e : t a =a{class} t e =a{domain}of a{class} Both the anchor prompt and the entanglement prompt are an English sentence; where {class} is the name of the class with the highest prediction probability in the mixed prediction, and {domain} is the name of the domain; By minimizing the LoRA parameter loss To update the LoRA parameters of the ViL model: in, is the InfoNCE loss of the image-text pair, λ1 and λ2 are the trade-off parameters, is the text similarity loss, represents the mathematical expectation, θ lora is the LoRA parameter, D t The training image set for the target domain, x i is the i-th target domain training image, |B| is the sampling batch size, I i For x i Image features output by the image encoder based on the ViL model, cos(.) represents cosine similarity, Anchor hints for training images in the i-th target domain Text features output by the text encoder based on the ViL model, Train image anchor hints for the jth target domain Text features output by the text encoder based on the ViL model, is the absolute similarity loss, is the relative similarity loss, N s is the number of source domains, and x i The jth entanglement hint and the kth entanglement hint Text features output by the text encoder based on the ViL model.
6. The passive domain adaptation image classification method according to claim 5, characterized in that: The described deployment of a trainable LoRA low-rank decomposition matrix for the ViL model includes: The LoRA low-rank decomposition matrix is deployed in the q matrix, k matrix and v matrix in all layers of the ViL model, and the low rank r of the LoRA low-rank decomposition matrix is set to 2.
7. The passive domain adaptation image classification method according to claim 1, characterized in that: The fine-tuning of the parameters of the pre-trained target model by using the parameter-fine-tuned ViL model as the alignment target of the feature level and the prediction output level of the pre-trained target model comprises: The prediction of the target domain training image by the parameter-fine-tuned ViL model is obtained as the third preliminary prediction, and the third preliminary prediction is mixed with the refined second preliminary prediction through the entropy uncertainty prediction mixing mechanism to obtain an accurate mixed prediction; The accurate hybrid prediction is used as the supervision signal, and the parameters of the pre-trained target model are fine-tuned by minimizing the loss in the knowledge transfer phase according to the following formula: in, is the loss in the knowledge transfer stage, is the mutual information loss, is the inner product loss, is the similarity norm learning loss, λ3 and λ4 are trade-off parameters, and θ t are the parameters of the pre-trained target model, represents the mathematical expectation, Training image x for the i-th target domain i The second preliminary prediction is For x i The exact mixed prediction of t A training image set for the target domain, For x i The third preliminary prediction of , I() represents the mutual information, Anchor hints for training images in the i-th target domain Text features output by the text encoder of the ViL model after fine-tuning parameters, z i For x i Based on the intermediate image features of the pre-trained target model, cos(.) represents the cosine similarity.
8. A passive domain adaptation image classification system, characterized in that: include: An image acquisition module is used to acquire the image to be classified in the target domain; An image classification module is used to adapt the input domain of the target domain image to be classified to the target model to obtain the classification result of the target domain image to be classified; Among them, the domain adaptation target model is obtained through the following training method: Obtaining the prediction of the ViL model for the target domain training image as a first preliminary prediction, and obtaining the prediction of the pre-trained target model for the target domain training image as a second preliminary prediction; Refining the second preliminary prediction according to the category prototype memory bank, and mixing the first preliminary prediction with the refined second preliminary prediction through an entropy uncertainty prediction mixing mechanism to obtain a mixed prediction; Using the mixed prediction as the supervisory signal, the low-rank adaptive LoRA method is used to fine-tune the parameters of the ViL model, and the fine-tuned ViL model is used as the alignment target of the feature level and prediction output level of the pre-trained target model to fine-tune the parameters of the pre-trained target model; Iterate the above steps to the preset conditions and use the final pre-trained target model as the domain adaptation target model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the passive domain adaptation image classification method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the passive domain adaptation image classification method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Industrial zero sample anomaly detection method and system based on cross-modal prompt learning
CN120726400A
Characterization fine tuning method and system in training during graph data node classification test
CN121599043A