Image classification method and system based on cross-domain collaborative model technology
Through the image classification method of cross-domain collaborative model technology, pre-trained models are used to calculate clustering features and characterize vector distances, cross-domain collaborative learning is realized, model forgetting and storage needs are reduced, and learning efficiency and generalization ability are improved.
Patent Information
- Application Number
- CN202510166732.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-18
AI Technical Summary
Cross-domain collaborative models tend to forget the learned knowledge when learning new data, resulting in performance degradation, and traditional methods increase computing and storage pressure, posing a risk of data privacy.
The image classification method using cross-domain collaborative model technology is used to calculate clustering features and characterize vector distances using the trained pre-trained model, and image classification is performed by weighted mixing probability, reducing adjustments to the original model parameters, and only prompt parameters and clustering features of specific fields are stored.
It effectively alleviates catastrophic forgetting, reduces computing and storage costs, and improves the learning efficiency and generalization capabilities of the model.
Smart Images

Figure CN120339667A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence, domain incremental learning, and prompt learning, and particularly relates to an image classification method and system for a cross-domain collaborative model technology. Background Art
[0002] In order for image classification models to be applied in critical systems in the long term, they must operate robustly in various different environments. Taking the object detection task of autonomous driving as an example, the model needs to process the collected images in real time to identify and classify obstacles on the road. However, these models are usually trained using image sets collected under clear weather conditions. In actual driving, weather changes frequently, especially in bad weather, which can lead to a decline in image quality and may significantly reduce the performance of the model, thus posing serious safety hazards. Image data sets with different feature distributions are divided into different domains, and there is data isolation between different domains. Obviously, the performance of a model trained separately for each domain is limited and has poor generalization ability, which cannot meet the requirements. There is a need to be able to link and utilize data from other domains to improve the service. An ideal solution is to train a cross-domain collaborative model that can adapt to all environments and different data distributions.
[0003] From the perspective of the scalability of cross-domain collaborative models, as more data domains are added to the collaborative modeling architecture, the intelligent model needs to continuously evolve. This is domain incremental learning, which means that the data sets of these domains are sequentially input into the same model for training. The main problem faced is the catastrophic forgetting problem for the domains that have been learned. The emergence of this knowledge forgetting phenomenon is because the training of traditional machine learning models is based on an assumption that the distribution of input data is fixed or stable. In reality, when the model sequentially learns a series of isolated domains, the data distributions of these domains are heterogeneous and appear at different time intervals, resulting in a non-stationary overall data distribution. This causes the model to potentially forget and overwrite the learned knowledge when learning new data, leading to a rapid decline in performance on old data. For example, an image classification model first learns in an indoor environment and then transfers to an outdoor environment. Due to changes in environmental contexts such as lighting conditions, the knowledge learned by the model at different stages may conflict. At the same time, under cross-domain collaborative modeling, additional conditions such as security restrictions on the original data not leaving the domain, limited storage resources, and limited transmission volume are added, further increasing the difficulty of solving the catastrophic forgetting problem.
[0004] In domain incremental learning, there have been many methods to address the problem of catastrophic forgetting. Replay-based methods are a simple and effective strategy that filters and retains key samples from previous tasks for review during new task training, such as LRCIL, ER, and iCaRL. Additionally, training a generative model to simulate the data distribution of previous tasks is also an effective means of generating review samples. To reduce the over-reliance on old samples, some technical strategies aim to reduce the mutual influence between new and old tasks to mitigate the risk of overfitting. However, these methods may increase computational and storage pressure, as well as data privacy risks. Regularization-based methods limit the changes in model parameters by introducing regularization terms into the training loss, reducing the impact of new tasks on the knowledge of old tasks, such as EWC. More flexible control can be achieved by evaluating the importance of each parameter in the network and adjusting the parameter update magnitude accordingly, such as LUCIR. Knowledge distillation techniques have also been used to guide the network to maintain the stability of the output for old task categories while leaving room for learning new categories. However, such methods may weaken the learning ability of personalized knowledge and limit performance improvement. Parameter isolation techniques generally do not limit the model scale. By retaining the key parameters of previous tasks and introducing new parameters, they enhance the adaptability to new tasks and reduce parameter interference between domains. Freezing and separating parameters are the main means to achieve this goal, such as GDumb and BiC. By learning a mask or path for each parameter or each layer, it is possible to dynamically determine which parameters are activated during new task learning, achieving more flexible parameter isolation. However, as the number of tasks increases, the complexity and cost of the network architecture increase, and it may ignore the cross-task correlations, resulting in a decline in the adaptability to unknown tasks. Compared with the traditional three methods, the domain incremental learning method based on prompt learning shows significant advantages in terms of model performance and efficiency. Prompt learning effectively reduces time and cost consumption by adjusting the insertable prompts instead of significantly changing the backbone model weights. For example, L2P designs a shared prompt pool, and S-Prompts learns personalized prompts for each domain. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems in the related art to some extent.
[0006] The present invention proposes an image classification method for cross-domain collaborative model technology to solve the problem of catastrophic forgetting faced by the cross-domain collaborative model in the above background art.
[0007] Another object of the present invention is to propose an image classification system for cross-domain collaborative model technology.
[0008] To achieve the above object, on the one hand, the present invention proposes an image classification method for cross-domain collaborative model technology, including:
[0009] Using the image encoder of the trained pre-trained model, calculate the clustering features of each domain based on the prompt combinations of the corresponding domains;
[0010] Calculate the representation vector corresponding to the sample through the trained pre-trained model, and calculate the distance between the representation vector and the clustering features of the i-th domain, so as to take the minimum value of all distances as the distance between the sample and the i-th domain;
[0011] Calculate the relative distances between the sample and domain i respectively for measurement to obtain a factor set;
[0012] Normalize the factor set and calculate the normalized weights of each factor, and perform weighted mixing with the prediction probability set generated by each domain prompt model to obtain the final classification mixing probability for image classification prediction.
[0013] The image classification method of the cross-domain collaboration model technology in the embodiments of the present invention may also have the following additional technical features:
[0014] In an embodiment of the present invention, the pre-trained model is CLIP; if the currently trained domain is the S-th domain, the training data set used is Initialize the visual-side prompt parameters of this domain and the language-side prompt parameters where L PV represents the length of the visual-side prompt, L PL is the length of the language-side prompt, and D is the alignment length of the model output vector.
[0015] In an embodiment of the present invention, training CLIP includes:
[0016] On the visual side, a 2D picture in the training data set is sliced into a 1D sequence, and each slice is mapped into a one-dimensional vector through a linear mapping, and a position vector is added at the same time;
[0017] The vector is input into the 12-layer multi-head self-attention mechanism of the visual encoder embedded with the prompt parameters to output an image vector to complete the feature aggregation with prompts for all image slices.
[0018] In an embodiment of the present invention, it further includes:
[0019] On the language side, copy the prompt parameters U times, and splice them on the prefix of the text encoding vectors of all categories to form a new set of input vectors Where U is the total number of categories, and C is the length of the sentence containing placeholders formed by each category;
[0020] Input e into the language encoder to obtain a text vector group for all categories
[0021] In one embodiment of the present invention, based on the image vector and the text vector group Construct the cosine similarity The following formula:
[0022]
[0023] Obtain the normalized probability set of the sample x for each category The formula is as follows:
[0024]
[0025] Where k is the Boltzmann constant, T is the temperature, and H s (x) is the cosine similarity group output by the model, and H s (x)[y] is the cosine similarity corresponding to the y-th classification, and U is the number of classification categories.
[0026] In one embodiment of the present invention, optimize the prompt parameters in the current field Read data in batches from the data within the field. Assume the read data set Calculate the total loss Consists of two parts, and λ is used to adjust the proportion of these two parts of the loss:
[0027]
[0028] Where, is the categorical cross-entropy loss, aiming to minimize the difference between the probability distribution predicted by the model and the probability distribution of the true label. The formula is as follows:
[0029]
[0030] In one embodiment of the present invention, is the distribution range of the partition function value for normalizing the output vector of the prompt model. Extract the denominator of P s (y|x) as the partition function, and unify the mean value of its numerical distribution to E. The formula is as follows:
[0031]
[0032] After completing the learning of the S-th field, cumulatively obtain the optimal combination of prompt parameters
[0033] In one embodiment of the present invention, the method further includes:
[0034] For the i-th domain, based on all the training data within the domain Generate a set of image feature vectors And use the K-means algorithm to extract K clustering centers Each clustering center The calculation formula is as follows:
[0035]
[0036] Wherein, Represents the process of the model extracting features, Is the image encoder of the pre-trained model,
[0037] Is the data used for the j-th clustering in the i-th domain, Represents its quantity;
[0038] For the sample x, calculate its representation vector z v (x) on the pre-trained model; use the L1-norm to calculate z v (x) and the representation M of the clustering center in the i-th domain i The minimum value of the distance is taken as the distance D i (x) between the sample and the i-th domain:
[0039]
[0040] Wherein, K represents the number of clustering points, Represents the j-th clustering center in the i-th domain, z v (x) is the representation vector of the sample;
[0041] Calculate the relative distance between the sample x and the domain i respectively for measurement, and obtain a factor set The calculation formula is as follows:
[0042]
[0043] In the incremental learning dataset And obtain the hint combination {P 1 , P 2 , …, P s} of the corresponding domain. After that, for the predicted picture x without the domain label, input it into the CLIP image Transformer without hints to obtain an image vector, and calculate its distance from the clustering points of each domain to obtain a factor set Take DA i(x) Softmax normalization weights and the set of prediction probabilities generated by the domain-specific prompt models Perform weighted mixing to obtain the final classification mixing probability P m (x) to perform the image classification prediction task, and the formula is as follows:
[0044]
[0045] To achieve the above object, on the other hand, the present invention proposes an image classification system based on cross-domain collaborative model technology, including:
[0046] A clustering feature calculation module, configured to calculate the clustering features of each domain based on the corresponding domain's prompt combination by using the image encoder of the trained pre-trained model;
[0047] A distance calculation module, configured to calculate the representation vector corresponding to the sample through the trained pre-trained model, and calculate the distance between the representation vector and the clustering feature of the i-th domain, and take the minimum value of all distances as the distance between the sample and the i-th domain;
[0048] A factor set calculation module, configured to calculate the relative distance between the sample and domain i respectively for measurement to obtain a factor set;
[0049] A classification mixing probability calculation module, configured to perform normalization processing on the factor set, calculate the normalized weight of each factor, and perform weighted mixing with the set of prediction probabilities generated by the domain-specific prompt models to obtain the final classification mixing probability for image classification prediction.
[0050] The image classification method and system based on cross-domain collaborative model technology according to the embodiments of the present invention, through the prompt learning technology, the cross-domain collaborative intelligent model continuously learns the models under various data distributions without adjusting the parameters of the original pre-trained model, greatly reducing the training cost. The two prompts help the model to extract knowledge concisely, and the reasonable weighted sum of the model prediction probabilities alleviates the catastrophic forgetting phenomenon. The present invention only needs to store the prompt parameters and clustering features of specific domains, so it shows extremely high efficiency in terms of computational cost and memory consumption. This method avoids storing the entire model parameters or any data samples in different domains, thus significantly reducing the learning cost.
[0051] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The above and / or additional aspects and advantages of the present invention will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0053] Figure 1Flowchart of an image classification method for cross - domain collaborative model technology according to an embodiment of the present invention;
[0054] Figure 2 Case diagram of parallel cross - domain collaborative modeling based on prompt learning according to an embodiment of the present invention;
[0055] Figure 3 Training architecture diagram of an image classification method according to an embodiment of the present invention;
[0056] Figure 4 Inference architecture diagram of an image classification method according to an embodiment of the present invention;
[0057] Figure 5 Structure diagram of an image classification system for cross - domain collaborative model technology according to an embodiment of the present invention. Detailed implementation manners
[0058] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0059] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0060] The image classification method and system for cross - domain collaborative model technology proposed according to an embodiment of the present invention will be described below with reference to the drawings.
[0061] Figure 1 Flowchart of an image classification method for cross - domain collaborative model technology according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0062] S1. Using the image encoder of the trained pre - trained model to calculate the clustering features of each domain based on the prompt combination of the corresponding domain;
[0063] S2. Calculating the representation vector corresponding to the sample through the trained pre - trained model, and calculating the distance between the representation vector and the clustering features of the i - th domain, and taking the minimum value of all distances as the distance between the sample and the i - th domain;
[0064] S3. Calculating the relative distance between the sample and the domain i respectively to perform measurement to obtain a factor set;
[0065] S4. Normalize the factor set, calculate the normalized weight of each factor, and perform weighted mixing with the set of prediction probabilities generated by the domain-specific prompt models to obtain the final classification mixing probability for image classification prediction.
[0066] As Figure 2 shown, the present invention supports each domain to train personalized prompts based on local data, and then summarize and save them. Since the number of prompt parameters is small, the transmission or storage occupancy generated is small. Then, combined with the spatial distance, weights are assigned to the domain-specific prompt models to reuse knowledge and reduce forgetting. Secondly, the present invention adopts a visual-language dual-modal pre-trained model, and the prompts take effect on the encoders of each modality, while the specific embedding methods adopted are different. Through the prompt learning technology, the cross-domain collaborative intelligent model continuously learns various data distributions without adjusting the parameters of the original pre-trained model, greatly reducing the training cost. The two types of prompts help the model to extract knowledge concisely, and the reasonable weighted sum of the model prediction probabilities alleviates the catastrophic forgetting phenomenon. The present invention only needs to store the prompt parameters and clustering features of specific domains, so it shows extremely high efficiency in terms of computational cost and memory consumption. This method avoids storing the entire model parameters or any data samples of different domains, thus significantly reducing the learning cost.
[0067] It can be understood that the cross-domain collaborative learning scenario studied by the present invention covers multiple domains, and each domain contains a dataset with highly heterogeneous data feature distributions to jointly solve the image classification task with the same classification objective. Assume there are S domains, and the image classification model starts training from , can only access one domain at a time, and sequentially learns . When switching to a new domain, the model cannot retain the samples of the old domain, but is allowed to retain a small number of prompt parameters of the old domain. Even when learning , the model can still solve the task problem in , and will not forget the knowledge of the learned old domain due to learning the knowledge of the new domain. At the same time, the model should be able to provide inferences for the unknown domain . After training, it should be able to give the best decision result without knowing the data source.
[0068] Specifically, the dataset of the i-th domain is composed of image-text pairs, denoted as where represents the j-th image sample in the i-th domain, the image size is W×H, the number of channels is C, and N is the total number of samples. t i,j ∈{a, b, … , z}|t| It is the text of the class name associated with the current sample, consisting of lowercase English letters, and |t| represents the length of the class name. After quantifying the classification label, it can be expressed as y i,j ∈{1,…,U}, where U is the total number of classes. Therefore, the dataset can be expressed as
[0069] The present invention adopts prompt learning technology to optimize learning performance. Since the training is carried out in isolation between different domains, the ultimate goal of the present invention is to find a set of optimal prompt parameters P = {P 1 , P 2 ,…, P s} to minimize the overall objective loss:
[0070]
[0071] Here, is the prediction projection matrix of all images within the domain under the action of the pre-trained image classification model and the fine-tuning prompt P i in the current i-th domain. is the within-domain loss in the i-th domain, which is calculated based on the difference between the prediction projection matrix and the true classification matrix .
[0072] As Figure 3 shown, it is the training architecture diagram of the image classification method of the cross-domain collaboration model technology based on vision-language bimodal prompt learning of the present invention. This architecture adopts prompt technology for learning and selects CLIP as the pre-trained model. CLIP is an image classification model trained on a large-scale image-text pair, and it performs excellently in various downstream multi-modal tasks. The input modalities include vision and language. Each domain locally initializes personalized prompt parameters and embeds them into the CLIP model. These parameters are only optimized independently on the within-domain data, while the parameters of the pre-trained model remain unchanged. In the loss calculation, only these few prompt parameters participate in the backpropagation optimization.
[0073] Assume that the currently trained domain is the S-th domain, and the dataset used is The visual-side prompt parameters of this domain have been initialized and the language-side prompt parameters where L PV represents the length of the visual-side prompt, L PL is the length of the language-side prompt, and D is the alignment length of the model output vector. Next, the training process will be introduced in detail.
[0074] In the visual side, a 2D image in the training set is sliced into a 1D sequence, and each slice is mapped to a one-dimensional vector through a linear mapping (convolutional layer), while adding a position vector. Then, the vector is input into the 12-layer multi-head self-attention mechanism of the visual encoder Transformer embedded with the prompt parameter The output image vector
[0075] completes the feature aggregation with prompts for all image slices, where the prompt embedding method follows the prefix-one-tuning method. In the language side, the prompt parameter is copied U times and concatenated on the prefix of the text encoding vectors of all categories to form a new set of input vectors
[0076] where U is the total number of categories and C is the length of the sentence containing placeholders formed by each category. Then, e is input into the language encoder Transformer to obtain the text vector group of all categories Based on the image vector and the text vector group The cosine similarity is constructed as follows:
[0077]
[0078] Thus, the present invention can further obtain the normalized probability set of the sample x on each category The formula is as follows:
[0079]
[0080] where k is the Boltzmann constant, T is the temperature, H s (x) is the cosine similarity group output by the model, and H s (x)[y] is the cosine similarity corresponding to the y-th classification, and U is the number of classification categories.
[0081] To optimize the prompt parameter in the current field Data is read in batches from the in-field data. Assume the read dataset Calculate the total loss which consists of two parts, and λ is used to adjust the proportion of the losses of these two parts:
[0082]
[0083] where It is the categorical cross-entropy loss, aiming to minimize the difference between the probability distribution predicted by the model and the probability distribution of the true labels. The formula is as follows:
[0084]
[0085] where k is the Boltzmann constant, T is the temperature, and H s (x) is the cosine similarity group output by the model, and H s (x)[y] is the cosine similarity corresponding to the y-th classification, and U is the number of classification categories.
[0086] And is the distribution range of the partition function value that normalizes the output vector of the prompt model. Extract the denominator of P s (y|x) as the partition function, and unify the mean value of its numerical distribution to E. The formula is as follows:
[0087]
[0088] where k is the Boltzmann constant, T is the temperature, and H s (x) is the cosine similarity group output by the model, and H s (x)[y] is the cosine similarity corresponding to the y-th classification, and U is the number of classification categories.
[0089] After completing the learning in the S domain, the optimal prompt parameter combination will be accumulated This way reduces the dependence on other domains, enables the training architecture to support domain parallelism. Each domain can train the local personalized prompt model in parallel separately. The order independence makes the cross-domain collaborative intelligent modeling architecture of this method more scalable and flexible.
[0090] As Figure 4 shown, it is the inference architecture diagram of the image classification method of the cross-domain collaborative model technology based on visual language bimodal prompt learning in the present invention. In the inference stage, the present invention constructs weights based on the distance in the common feature space to achieve cross-domain joint inference and knowledge reuse, and reduce forgetting. Feature extraction can be performed on the datasets of each domain in the pre-trained model CLIP without prompts to construct a feature space, in which the data of each domain shows a certain degree of discriminative clustering. Based on the assumption that "the sample is closer to the similar domain in the feature space", the distance between the sample and a certain domain is used as a scalar to reflect the closeness of the sample to that domain. In the unified feature space, the distance factors between the sample and the clustering points of different domains are within the same data range. Therefore, the relative similarity between the sample and each domain can be evaluated by comparing the magnitudes of these distances. Therefore, the distance can be used as a powerful prior knowledge to quantify the cross-domain similarity and assign weights to the prompt models of each domain.
[0091] Specifically, the image encoder of the pre-trained model is utilized Without introducing any prompts, calculate the clustering features in each domain. For the i-th domain, based on all the training data within the domain Generate a set of image feature vectors And apply the K-means algorithm to extract K clustering centers Each clustering center The calculation formula is as follows:
[0092]
[0093] Here, Represents the process of the model extracting features, Is the image encoder of the pre-trained model,
[0094] Is the data used for the j-th clustering in the i-th domain, Represents its quantity.
[0095] For the test sample x, calculate its representation vector z v (x) on the pre-trained model. Use the L1-norm to calculate the distance between z v (x) and the representation M of the clustering center in the i-th domain i Take the minimum value of these distances as the distance D of the sample from the i-th domain i (x):
[0096]
[0097] Here, K represents the number of clustering points, Represents the j-th clustering center in the i-th domain, and z v (x) is the representation vector of the test sample.
[0098] By comparing the distances of the sample from each domain, it can be used to evaluate the similarity of the sample to each domain. Ideally, for the sample x belonging to the i-th domain i , the inter-domain distance relationship can be expressed as:
[0099] D i (x i ) < D j (x i ), where j ≠ i.
[0100] Furthermore, calculate the relative distances between the sample x and the domains i (where i ranges from 1 to s) respectively for measurement, and obtain a set of factors The calculation formula is as follows:
[0101]
[0102] Then, after incrementally learning the dataset and obtaining the hint combination {P 1 , P 2 , …, P s} in the corresponding field, for the predicted image x without the field label, its image vector is obtained by inputting it into the CLIP image Transformer without hints, and the distances between it and the clustering points in each field are calculated to obtain the factor set Then, the Softmax normalization weights of DA i (x) and the prediction probability set generated by the hint models in each field are weighted and mixed to obtain the final classification mixed probability P m (x) to perform the image classification prediction task, and the formula is as follows:
[0103]
[0104] According to the image classification method of the cross-domain collaborative model technology in the embodiment of the present invention, by reforming and upgrading the cross-domain collaborative model technology method, a hybrid architecture based on visual-linguistic bimodal hint learning is designed and proposed. During training, it supports parallel training in each field, reduces the dependence between fields, and uses different embedding methods for each field to fine-tune the visual and linguistic side hints. The obtained hints can provide services independently or perform joint reasoning based on the weights calculated by distance. At the same time, it adapts to the personalized and generalized image classification requirements. In the training design of the bimodal hint, this solution proposes a loss formula based on similarity vector normalization and classification performance optimization. Without the need for data sample replay in other fields, the loss value is calculated only through the output cosine similarity vector of a specific hint model, the hint parameters are optimized to maximize the similarity of the sample target, and the distribution range of the vector partition function value is normalized to reduce the deviation in cross-domain joint calculation. And in the weight design of the hybrid weighting, this solution constructs the weight based on the cross-domain distance in the common feature space, uses the pre-trained model without hints to extract the differences in the data distribution between fields, realizes cross-domain joint reasoning and knowledge reuse, and through experiments and research, it is shown to be effective in reducing catastrophic forgetting, proving that this weight calculation method is an effective method for measuring the similarity between test samples and each field.
[0105] To implement the above embodiment, as Figure 5 shown, this embodiment also provides an image classification system 10 of the cross-domain collaborative model technology, including:
[0106] A clustering feature calculation module 100, configured to calculate the clustering features of each field based on the hint combination in the corresponding field by using the image encoder of the trained pre-trained model;
[0107] A distance calculation module 200 is configured to calculate a representation vector corresponding to a sample through a trained pre-trained model, and calculate the distance between the representation vector and the clustering features of the i-th domain, and use the minimum value of all distances as the distance between the sample and the i-th domain;
[0108] A factor set calculation module 300 is configured to calculate the relative distance between the sample and the i-th domain respectively for measurement to obtain a factor set;
[0109] A classification mixing probability calculation module 400 is configured to perform normalization processing on the factor set, calculate the normalized weight of each factor, and perform weighted mixing with the prediction probability set generated by each domain prompt model to obtain the final classification mixing probability for image classification prediction.
[0110] The image classification system of the cross-domain collaborative model technology according to the embodiment of the present invention designs and proposes a hybrid architecture based on vision-language bimodal prompt learning by reforming and upgrading the cross-domain collaborative model technology method. During training, it supports parallel training in each domain, reduces the dependence between domains, and uses different embedding methods for each domain to fine-tune the vision-side and language-side prompts. The obtained prompts can provide services independently or perform joint reasoning based on the weights calculated by distance. At the same time, it adapts to the personalized and generalized image classification requirements. In the training design of the bimodal prompt, this solution proposes a loss formula based on similarity vector normalization and classification performance optimization. Without the need for data sample replay in other domains, the loss value is calculated only through the output cosine similarity vector of a specific prompt model, the prompt parameters are optimized to maximize the sample target similarity, and the distribution range of the vector partition function value is standardized to reduce the deviation in cross-domain joint calculation. And in the weight design of the hybrid weighting, this solution constructs the weight based on the cross-domain distance of the common feature space, uses the pre-trained model without prompts to extract the differences in the data distributions between domains, realizes cross-domain joint reasoning and knowledge reuse, and experiments and research show that it is effective in reducing catastrophic forgetting, proving that this weight calculation method is an effective method for measuring the similarity between test samples and each domain.
[0111] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0112] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
Claims
1. An image classification method for cross-domain collaborative model technology, characterized in that, Including: Using the image encoder of the trained pre-trained model to calculate the clustering features of each domain based on the hint combination of the corresponding domain; Calculating the representation vector corresponding to the sample through the trained pre-trained model, and calculating the distance between the representation vector and the clustering features of the i-th domain, taking the minimum value of all distances as the distance between the sample and the i-th domain; Calculating the relative distance between the sample and domain i respectively for measurement to obtain a factor set; Performing normalization processing on the factor set, calculating the normalized weight of each factor, and performing weighted mixing with the prediction probability set generated by the hint models of each domain to obtain the final classification mixing probability for picture classification prediction.
2. The method according to claim 1, wherein The pre-trained model is CLIP; if the current training domain is the S-th domain, the training dataset used is Initialize the visual-side prompt parameters for this domain and the language-side prompt parameters where L PV represents the length of the visual-side prompt, L PL is the length of the language-side prompt, and D is the alignment length of the model output vector.
3. The method according to claim 2, wherein Training CLIP, including: In the visual side, a 2D image in the training dataset is sliced into a 1D sequence, and each slice is mapped to a one-dimensional vector through a linear mapping while adding a position vector; Input the vector into the 12-layer multi-head self-attention mechanism of the visual encoder embedded with the prompt parameter to output an image vector so as to complete the feature aggregation with prompts for all image slices.
4. The method according to claim 3, characterized in that, Also including: On the language side, the prompt parameter is copied U times and concatenated onto the prefixes of the text encoding vectors of all categories to form a new set of input vectors where U is the total number of categories and C is the length of the sentence containing placeholders formed by each category; Input e into the language encoder to obtain a text vector group for all categories 5. The method according to claim 4, wherein Based on the image vector and the text vector group Construct the cosine similarity The following formula: Obtain the normalized probability set of the sample x for each category The formula is as follows: where k is the Boltzmann constant, T is the temperature, and H s (x) is the cosine similarity group output by the model, and H s (x)[y] is the cosine similarity corresponding to the y-th classification, and U is the number of classification categories.
6. The method according to claim 5, wherein Optimize the hint parameters in the current field Read data in batches from the in-field data, assuming the read data set Calculate the total loss It consists of two parts, and λ is used to adjust the proportion of these two parts of losses: Among them, is the categorical cross-entropy loss, which aims to minimize the difference between the probability distribution predicted by the model and the probability distribution of the true labels. The formula is as follows:
7. The method according to claim 6, wherein is the distribution range of the partition function values that normalize the output vectors of the prompt model. Extract the denominator of P s (y|x) as the partition function and uniformly fix the mean value of its numerical distribution to E. The formula is as follows: After completing the learning in the S domain, the optimal hint parameter combination is cumulatively obtained 8. The method according to claim 1, wherein The method also includes: For the i-th domain, based on all the training data within the domain Generate a set of image feature vectors And use the K-means algorithm to extract K cluster centers Each cluster center The calculation formula is as follows: Among them, represents the process of the representative model extracting features, is the image encoder of the pre-trained model, is the data used for the j-th clustering in the i-th domain, represents its quantity; For the sample x, calculate its representation vector z on the pre-trained model v (x); calculate the distance between z v (x) and the cluster center representation M of the i-th domain i , and take the minimum value of the distance as the distance D between the sample and the i-th domain i (x): where K represents the number of clustering points, represents the j-th clustering center in the i-th domain, and z v (x) is the feature vector of the sample; The relative distance between the sample x and the domain i is calculated separately for measurement, and a factor set is obtained The calculation formula is as follows: In the incremental learning dataset and obtain the hint combination {P 1 , P 2 ,..., P s} in the corresponding field. After that, for the predicted picture x without the field label, it is input into the CLIP image Transformer without hints to obtain the image vector, and the distances between it and the clustering points in each field are calculated to obtain the factor set The Softmax-normalized weights of DA i (x) and the set of prediction probabilities generated by the hint models in each field are weighted and mixed to obtain the final classification mixed probability P m (x) to perform the picture classification prediction task. The formula is as follows:
9. An image classification system using cross - domain collaborative model technology, characterized in that, Including: A clustering feature calculation module, configured to use the image encoder of the trained pre-trained model to calculate the clustering features of each domain based on the hint combination of the corresponding domain; A distance calculation module, configured to calculate the representation vector corresponding to the sample through the trained pre-trained model, and calculate the distance between the representation vector and the clustering features of the i-th domain, taking the minimum value of all distances as the distance between the sample and the i-th domain; A factor set calculation module, configured to calculate the relative distance between the sample and domain i respectively for measurement to obtain a factor set; A classification mixing probability calculation module, configured to perform normalization processing on the factor set, calculate the normalized weight of each factor, and perform weighted mixing with the prediction probability set generated by the hint models of each domain to obtain the final classification mixing probability for picture classification prediction.