Knowledge enhanced multistage transfer calibration algorithm for passive domain visual scene domain transfer
By training multiple models of different structures in the source domain model library and calibrating with transferability information across multiple scales, the problem of inability to effectively estimate instance-level transferability between multiple different architecture models in the prior art is solved, and more accurate cross-model instance-level transferability comparison and model selection are achieved.
Patent Information
- Application Number
- CN202510147788.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-13
AI Technical Summary
Existing transferability estimation methods are not applicable in models of multiple different architectures, and cannot effectively estimate instance-level transferability between different models in the source domain model library.
A knowledge-enhanced multi-level transfer calibration algorithm for passive domain visual scene domain migration is proposed. By training multiple models of different structures in the source domain model library, the transferability information across multiple scales is used for calibration, and multi-level instance transferability calibration is achieved.
The cross-model instance-level transferability comparison in multi-model passive domain video domain migration scenario is realized, providing more accurate instance-level passive video transferability measurement, and improving the accuracy of model selection.
Smart Images

Figure CN119992262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of passive domain visual scene domain migration calibration algorithm, and in particular to a knowledge-enhanced multi-stage transfer calibration algorithm for passive domain visual scene domain migration. Background Art
[0002] When multiple pre-trained source domain models are available, it is particularly important to evaluate the transferability of each source domain model to the target domain. The traditional approach is to predict the performance of the source model after fine-tuning it on the target domain in a supervised manner. However, the requirement of target labels in these methods hinders their application in a wider context. Recently, some methods have been proposed to estimate the transferability of passive unsupervised tasks. These methods can be roughly divided into two categories: distribution-guided methods and uncertainty-based methods. Distribution-guided methods measure transferability by developing and evaluating some distribution-related assumptions. For example, SUTE proposes assumptions about the distribution of the dataset, including individual certainty, semantic consistency, and global discreteness. MDE proposes to remove the energy assumption and transform the information of all samples into a statistical probability distribution. Uncertainty-based methods use uncertainty methods to estimate transferability, including entropy, temporal and spatial interference, based on the assumption that highly transferable models have low uncertainty. This technique has broad application prospects and can estimate the transferability of models, and can be used for model selection and application in multi-model scenarios.
[0003] Existing transferability estimation methods can be roughly divided into two categories: distribution-inducing methods, which develop and evaluate some distribution-related assumptions, and uncertainty-based methods, which are based on the assumption that highly transferable models have low uncertainty.
[0004] While these methods demonstrate effectiveness in measuring image model transferability, our experimental results in Table 1 show that both methods are not applicable to situations where there are multiple models with different architectures in the source domain model library.
[0005] Limitations of distribution-guided methods: These methods exploit dataset-level information for cross-model transferability estimation, but ignore the fact that source domain models may show different preferences for individual instances. Unlike images, videos consist of a series of consecutive frames, containing both spatial and temporal information. The temporal dimension introduces additional complexity and variability. In addition, videos usually have greater content complexity than images, resulting in greater differences between video instances than between image instances. Ignoring instance-level transferability estimation may hinder further improvements in adaptation performance.
[0006] Limitations of uncertainty-based methods: These methods can estimate instance-level transferability within a single model. However, when the source domain model library contains models with different architectures, the differences between these models can introduce bias at the dataset level. Without addressing these cross-model differences, directly estimating instance-level transferability between different source domain models may not provide accurate results.
[0007] To this end, we propose a knowledge-enhanced multi-level transfer calibration algorithm for passive domain visual scene domain migration to solve the above problems. Summary of the invention
[0008] The purpose of the present invention is to solve the shortcomings existing in the prior art and propose a knowledge-enhanced multi-stage transfer calibration algorithm for passive domain visual scene domain migration.
[0009] In order to achieve the above object, the present invention adopts the following technical solution: a knowledge-enhanced multi-stage transfer calibration algorithm for passive domain visual scene domain migration, comprising the following steps:
[0010] S1. Based on the mmaction2 framework, the pre-trained model provided by the mmaction2 framework for the video action recognition task is selected in the source domain training and source domain model library, and the training is performed on the source domain dataset. The training parameters in the source domain model library can be customized according to the needs. After completing the training on the source domain dataset, a source domain model library containing multiple different structural models is obtained.
[0011] S2, by obtaining the feature f(V i ); The model output h(V i ); predict semantics c h c (V i ); The pseudo-label y of each target data is calculated by applying a clustering-based pseudo-labeling technique based on the extracted features and predicted probabilities, and calculating the passive domain transferable estimation index (SUTE) based on the distribution-guided method, according to the formula:
[0012] S3. An intermediate grouping, called “group level”, is introduced to enhance transferability calibration; the group level transferability calculation method is based on the formula:
[0013]
[0014] S4. Calculate the individual certainty of each source model instance using the following formula:
[0015] T I =H(h(V i ))
[0016] S5. Calibrate the instance-level transferability using information from multiple levels through a calibration function, where the formula of the calibration function is:
[0017]
[0018] Final Multi-Level Instance Transferability Calibration Algorithm (MITC):
[0019] MITC=Φ(T I ,Φ(T G ,T D )).
[0020] Compared with the prior art, the main purpose of this application is to solve the problem that the existing methods use dataset-level information to estimate cross-model transferability, but ignore the problem that the source domain model may show different preferences for a single instance; the present invention uses transferability information across multiple scales to calibrate instance-level transferability; achieves more fine-grained instance-level transferability estimation across models; and provides better guidance for model selection.
[0021] Preferably, h in S2 c (V i ) represents the predicted probability on class c; E represents mathematical expectation, H represents entropy, D represents the source domain dataset, and D T Represents the target domain dataset; SUTE provides information about the dataset scale, using T D express.
[0022] Preferably, in S3, d and represents calculating the Euclidean distance of the two video predictions, argsort k It means taking the first k distances, and after experiments, k=5 is the final value.
[0023] Preferably, the group level in S3 is a more fine-grained level compared to the dataset level, and is also a more coarse-grained level compared to the instance level.
[0024] Preferably, in said S2, c (V i ) represents the predicted probability of class c.
[0025] The beneficial effects of the present invention are:
[0026] 1. This application designs a multi-scale calibration function for the multi-model passive domain video domain migration scenario; through the multi-scale calibration function, it focuses on the significant differences between video instances and the differences between different architecture models; based on the characteristics of information at multiple scales, a multi-scale calibration function suitable for the multi-model passive domain video domain migration scenario is designed to compare the transferability across model instances;
[0027] 2. This application proposes to perform multi-level calibration to achieve cross-model comparison of instance-level transferability across models; use transferability information across multiple scales to calibrate instance-level transferability; this method can more accurately measure instance-level passive video transferability. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 A flow chart of source domain model library training for the knowledge-enhanced multi-stage transfer calibration algorithm for passive domain visual scene domain migration proposed by the present invention;
[0029] Figure 2 Schematic diagram of the flow of the multi-level instance transferability calibration algorithm in the passive domain to video domain migration scenario in the present invention. DETAILED DESCRIPTION
[0030] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0031] Reference Figure 1-2 , the knowledge-enhanced multi-level transfer calibration algorithm for passive domain visual scene domain migration includes the following steps:
[0032] S1. Training the source domain model library in the source domain. This embodiment is based on the mmaction2 framework, selects the pre-trained model provided by it for the video action recognition task, and performs training on the source domain dataset. The training parameters in the source domain model library can be customized according to the needs; after completing the training on the source domain dataset, a source domain model library containing multiple different structural models is obtained;
[0033] S2, by obtaining the feature f(V i ); The model output h(V i ); predict semantics c h c (V i ), where h c (V i ) represents the predicted probability on class c; the pseudo-label y of each target data is calculated by applying a clustering-based pseudo-labeling technique based on the extracted features and the predicted probability, and the passive domain transferable estimation index (SUTE) based on the distribution-guided method is calculated according to the formula:
[0034]
[0035] Where E represents mathematical expectation, H represents entropy, D represents the source domain dataset, and D T Represents the target domain dataset; SUTE provides information about the dataset scale, using TD express;
[0036] S3. We further introduce an intermediate grouping, called “group level”, to enhance transferability calibration; this level is a more fine-grained level compared to the dataset level, and it is also a more coarse-grained level compared to the instance level; the group-level transferability calculation method is based on the formula:
[0037]
[0038] Where d represents the Euclidean distance between two video predictions, argsort k It means taking the first k distances, where k=5 is the final value after experiments;
[0039] S4. Calculate the individual certainty of each source model instance using the following formula:
[0040] T I =H(h(V i ))
[0041] S5. Calibrate the instance-level transferability using information from multiple levels through a calibration function, where the formula of the calibration function is:
[0042]
[0043] Final Multi-Level Instance Transferability Calibration Algorithm (MITC):
[0044] MITC=Φ(T I ,Φ(T G ,T D )).
[0045] Reference Figure 1 , 2 ;in Figure 2 a Current uncertainty-based methods focus on model transferability but lack cross-model generalization; b Distribution-induced methods estimate data-level transferability between models but ignore instance-level transferability within models; c To overcome these limitations;
[0046] Through the multi-layer instance transferability calibration algorithm, cross-model instance transferability estimation can be achieved under the condition of multiple model architectures in the source domain, and accurate instance-level passive video transferability measurement can be performed; the effectiveness of the proposed transferability calibration algorithm is experimentally verified on the large-scale public dataset Daily-DA for video action recognition; DailyDA is another large-scale cross-domain action recognition benchmark; it includes 4 datasets, namely ARID (A), HMDB51 (H), Moments-in-Time (M) and Kinetics (K); videos from 8 shared classes are used for cross-domain evaluation; because the models in the source domain model library are pre-trained in Kinetics, in order to ensure the accuracy of performance, the Kinetics dataset is chosen to be removed, and after removal, there are three datasets left, resulting in a total of 6 tasks; for instance-level transferability estimation, it is compared with several methods based on distribution guidance methods; these methods include negative mutual information (NMI), pseudo-label-based methods LEEP / LogME (called LEEP* / LogME* at the instance level), energy-based methods meta-distribution energy (MDE) and passive unsupervised transferability estimation metric (SUTE);
[0047] Our approach is compared with direct uncertainty-based methods, including entropy, temporal consistency, and spatial consistency; instance-level transferability is assessed by computing the Spearman rank correlation coefficient between the cross entropy of each source model on target domain instances and the estimated transferability of each instance.
[0048] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A knowledge-enhanced multi-level transfer calibration algorithm for passive domain visual scene domain migration, characterized by: The following steps are involved: S1. Based on the mmaction2 framework, the pre-trained model provided by the mmaction2 framework for the video action recognition task is selected in the source domain training and source domain model library, and the training is performed on the source domain dataset. The training parameters in the source domain model library can be customized according to the needs. After completing the training on the source domain dataset, a source domain model library containing multiple different structural models is obtained. S2, by obtaining the feature f(V i ); The model output h(V i ); predict semantics The pseudo-label ˜y of each target data is calculated by applying a clustering-based pseudo-labeling technique based on the extracted features and predicted probabilities, and calculating the passive domain transferable estimation index (SUTE) based on the distribution-guided method according to the formula: S3. An intermediate grouping, called "group level", is introduced to enhance transferability calibration; the group level transferability calculation method is based on the formula: S4. Calculate the individual certainty of each source model instance using the following formula: T I =H(h(V i )) S5. Calibrate the instance-level transferability using information from multiple levels through a calibration function, where the formula of the calibration function is: Final Multi-Level Instance Transferability Calibration Algorithm (MITC): MITC=Φ(T I ,Φ(T G ,T D ))。 2. The knowledge-enhanced multi-stage transfer calibration algorithm for passive domain visual scene domain migration according to claim 1 is characterized by: The h in S2 c (V i ) represents the predicted probability on class c; E represents mathematical expectation, H represents entropy, D represents the source domain dataset, and D T Represents the target domain dataset; SUTE provides information about the dataset scale, using T D express.
3. The knowledge-enhanced multi-stage transfer calibration algorithm for passive domain visual scene domain migration according to claim 1 is characterized in that: The S3 d represents the calculation of the Euclidean distance between the two video predictions, argsort k It means taking the first k distances, and after experiments, k=5 is the final value.
4. The knowledge-enhanced multi-stage transfer calibration algorithm for passive domain visual scene domain migration according to claim 1, characterized in that: The group level in S3 is a finer-grained level compared to the dataset level, and is a coarser-grained level compared to the instance level.
5. The knowledge-enhanced multi-stage transfer calibration algorithm for passive domain visual scene domain migration according to claim 1, characterized in that: The S2h c (V i ) represents the predicted probability of class c.