A method for homework search and reuse for image classification tasks
Patent Information
- Application Number
- CN202410684501.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-05-30
AI Technical Summary
然而,现有模型复用方法均假设有一个或多个可复用的模型已经给定,在实际应用中,如何从海量候选模型中选择最适合目标任务的模型是一个非常困难的问题
[0021]一种计算机设备,该计算机设备包括存储器、处理器及存储在存储器上并可在处理器上运行的计算机程序,处理器执行上述计算机程序时实现如上所述的针对图像分类任务的学件查搜与复用的方法。
Smart Images

Figure CN118691874B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for learning object retrieval and reuse for image classification tasks, belonging to the fields of machine learning technology and image classification technology. Background Technology
[0002] In the field of image classification, the mainstream approach to quickly build machine learning models for new image classification tasks, avoiding the need to collect large amounts of data from scratch for model training, involves pre-training large pre-trained models with a large number of parameters on massive image classification datasets, and then fine-tuning these models using specific downstream task data. However, these large pre-trained models have a huge number of parameters, and even fine-tuning them on new tasks requires significant computational resources. Therefore, this method faces problems of poor applicability and inefficiency in resource-constrained tasks. Furthermore, while large pre-trained models have strong generality, their performance on specific downstream tasks may not be optimal. Therefore, using large pre-trained models for fine-tuning is not the best choice for specific downstream image classification tasks. Meanwhile, the development of machine learning has generated a large number of image classification datasets and domain-specific models for solving various image classification tasks. When faced with new image classification tasks, effectively reusing existing domain-specific models can greatly reduce resource requirements and improve performance on new tasks. However, existing model reuse methods all assume that one or more reusable models are already given. In practical applications, selecting the most suitable model for the target task from a massive number of candidate models is a very difficult problem. Some existing works have considered constructing surrogate metrics to estimate the model's performance on the target task, thereby assisting in model selection. However, these works require predicting each candidate model on the target task data one by one and obtaining the output, which is difficult to apply when the number of candidate models is large. Learning-related research proposes to describe the model's function by model assignment specification and to achieve efficient model search by matching the new task with the existing model specification. However, existing model specifications only consider the distribution of the model's training data. The reusability of a model is affected not only by the data distribution but also by the model's performance itself. How to consider both data distribution and model performance simultaneously is a shortcoming of existing learning-related search techniques. In addition, existing learning-related search techniques are mainly for data with isomorphic features. How to transform image classification models from different data distributions and heterogeneous feature spaces into a unified feature space for representation remains an unsolved problem. Summary of the Invention
[0003] Purpose of the invention: To address the problems and shortcomings of existing technologies, this invention provides a method for searching and reusing learning materials for image classification tasks. The aim is to construct specifications that accurately characterize the functions of image classification models, collect a large number of image classification models, generate specifications for them, and build a learning material library. This allows users to search and reuse models within the learning material library, thereby solving new image classification tasks better and faster.
[0004] The core technologies of this invention include: 1) a text representation method for image classification model performance, which obtains the predicted categories and corresponding confidence scores of the model on training data and converts them into text form; 2) an image classification model reduction construction method that simultaneously considers the distribution of training data and the performance of the model. This method is based on a text-image pre-trained model, transforming heterogeneous training image data from different tasks into a new feature space using an image encoder, and transforming the model's prediction results and confidence scores on their respective training data into the new feature space using a text encoder; then, aligning the image space and text space, and constructing a reduction using Reduced Kernel Mean Embedding (RKME) in the aligned text-image representation space, so as to simultaneously characterize the distribution of training data and the model's performance on the training data in the new feature space; 3) In addition, this invention also provides a detailed method for constructing a learning resource library and searching for learning resources. This invention, targeting image classification tasks, by simultaneously considering the distribution of training data and the performance of the model, can efficiently and accurately enable users to accurately search for and reuse models for new downstream tasks.
[0005] Technical solution: A method for learning object retrieval and reuse for image classification tasks, comprising:
[0006] 1) Obtain models trained on various image classification tasks and corresponding training datasets, wherein the models are obtained after sufficient training on the training datasets;
[0007] 2) By using the corresponding trained model to make predictions on the training data, the model prediction results and confidence scores are obtained and converted into text form;
[0008] 3) Use the image encoder of the text-image pre-trained model to perform feature transformation on the training data to obtain image feature representation in a unified representation space. Use the text encoder to perform feature transformation on the text obtained from the model prediction results and confidence transformation to obtain text feature representation in a unified representation space.
[0009] 4) The RKME (Reduced Kernel Mean Embedding) technique is used to construct the image and text feature representations in the unified representation space obtained in step 3) as model reductions to describe the model's function. These models cover different image classification tasks, enabling users to find the most suitable model from the learning library when faced with new image classification tasks.
[0010] 5) By constructing specifications for a large number of image classification models, a learning library is formed, which contains multiple image classification models, each of which has a specification that can describe the function of the model;
[0011] 6) For the new image classification task, the user uses the image encoder of the text-image pre-trained model to perform feature transformation on the training data images of the new task to obtain image feature representation, and transforms the corresponding ground truth labels of the images to obtain text feature representation, and uses step 4) to construct a reduction for the new task;
[0012] 7) Select the candidate model with the smallest distance to the new task specification by calculating the maximum mean difference distance between the candidate model reduction in the learning material library and the new task specification;
[0013] 8) Adapt and reuse the selected model in new tasks to achieve rapid response and efficient solution to new tasks.
[0014] In step 2), natural language is used to describe the model's prediction results and confidence levels, thereby characterizing the model's performance.
[0015] The step of adapting and reusing the selected model on a new task refers to using the matched model as the initial model, and then using its own data to further train or adjust it. This eliminates the need to collect a large amount of data and consume a lot of resources to train data from scratch, thus achieving rapid response and efficient solution to new tasks.
[0016] The acquisition of models and corresponding training data trained in various image classification tasks includes collecting representative image classification models and corresponding training datasets from multiple fields and application scenarios. These models should be fully trained on their training datasets.
[0017] The image encoder and text encoder of the text-image pre-trained model can convert image and text data into feature vectors. These feature vectors reside in a unified feature space, which serves as the basis for subsequent model reduction. The text-image pre-trained model is a multimodal pre-trained model trained using a large amount of image-text pair data through contrastive learning. It includes an image encoder and a text encoder, which can map images and text into representations in a unified feature space, facilitating more accurate reduction generation and model search.
[0018] The abbreviated kernel mean embedding technique can accurately capture the relationship between model functionality and data distribution, thereby generating more accurate and comprehensive reductions for each model. The step of integrating the model's training data distribution and performance into a unified model reduction using abbreviated kernel mean embedding involves effectively mapping the model's training data distribution and performance using this technique, allowing this information to be represented in a unified and information-rich reduction, thus enhancing the accuracy and efficiency of model retrieval and reuse.
[0019] Users can construct corresponding specifications using the methods described above, based on the specific requirements of downstream tasks. Then, they search the learning library for the most suitable machine learning model for that task. This step allows users to quickly locate the model best suited to their specific needs, thereby improving the efficiency and accuracy of model reuse.
[0020] After selecting a model, users will adapt and reuse it for downstream tasks. The key to this step is to effectively apply the capabilities of the selected model to new tasks, achieving the goals of rapid response and efficient problem-solving.
[0021] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for learning object retrieval and reuse for an image classification task as described above.
[0022] A computer-readable storage medium storing a computer program that performs the method of learning object retrieval and reuse for an image classification task as described above.
[0023] This invention provides a method for learning object retrieval and reuse for image classification tasks. By reducing the image classification model to generate a learning object library, and by searching in the learning object library, the reuse performance of the model in different downstream tasks is effectively improved, and the model retrieval and deployment process is significantly optimized. Attached Figure Description
[0024] Figure 1 This is a flowchart of a method according to an embodiment of the present invention;
[0025] Figure 2 This is a flowchart of the specification construction proposed in the embodiments of the present invention. Detailed Implementation
[0026] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0027] like Figure 1 As shown, the methods for learning object retrieval and reuse for image classification tasks mainly include:
[0028] 101. Collect the trained image classification models;
[0029] Obtaining image classification models and their corresponding training data involves collecting representative image classification models and their corresponding training datasets from multiple fields and application scenarios. These models should be fully trained on their training datasets. In this embodiment of the invention, the model can be any machine learning model used for image classification tasks, without being limited by the model's structure or the training method used.
[0030] 102. Use the collected models to obtain predicted categories and confidence scores on their respective training data, and convert them into text format;
[0031] Suppose a model is f, and its corresponding training dataset is Where, x i For image content, y i Image annotation. First, obtain the model's annotation for each training sample x. i The predicted probability, i.e., q i =f(y|x i Then, the argmax function is applied to the probability distribution to generate a one-hot encoded prediction label. Where max(q) i This can be viewed as the model predicting x. i For tags The confidence level. After this, a new transformed dataset can be obtained. This reflects the distribution of the training data and the functionality of the model. Based on the text-image pre-trained model, the model can be applied to each sample x. i Prediction results and prediction confidence max(q) i Convert ) to text t(x) i For example, if the model predicts that an image belongs to the category "dog" and the confidence level is 0.9, the text "a photo of a dog, confidence: 0.9" can be obtained.
[0032] 103. Use a text-image pre-trained model to project the images in the training data and the model's output on the corresponding data into a unified feature space;
[0033] Given each training sample x i and its corresponding predicted text t(x) i By projecting the data into a unified space using the image encoder EI(·) and text encoder ET(·) in the text-image pre-trained model, unified features are obtained.
[0034]
[0035] 104. Use Reduced Kernel Mean Embedding (RKME) to construct the model's reduction from the unified feature representation obtained in the above steps;
[0036] Kernel Mean Embedding (KME) will be defined in probability distribution on Mapped to an element of the regenerated Hilbert space (RKHS), it is represented as:
[0037]
[0038] in It is a kernel function, assumed to be continuous, bounded, and positive definite, and related to RHKS. By using kernel functions, such as the Gaussian kernel function, the distribution... In mapping No information will be lost afterward. In machine learning tasks, we can only access information from the distribution. The training dataset obtained by sampling Therefore, the usual practice is to use the empirical approximation of KME:
[0039]
[0040] This can be The rate converges to the true KME.
[0041] KME's advantageous properties make it a potential specification for describing model functionality; however, its dependence on the original training data can raise data privacy issues. To address this, abbreviated sets are used... Obtain an approximate original training dataset The RKME of the empirical KME, where m is the size of the abbreviated set, u is the number of samples in the abbreviated set, and β is the weight of each sample in the abbreviated set. This is combined with the unified feature representation obtained in step 103, which contains information about the training data distribution and model functionality. RKME reduction for model f is generated by minimizing the following objective.
[0042]
[0043] The above minimization problem can be solved using an alternating optimization algorithm. The convergence rate of the empirical KME on the original training dataset using RKME is...
[0044] 105. Package each model and its corresponding reduction to build a learning library;
[0045] In this invention, it is assumed that the learning materials library has M available models. Where f m In the dataset The feature space of the training is The label space is {1,…,K} m}. N m ,d m and K m These represent the number of training samples, the dimension of the feature space, and the dimension of the label space, respectively. When a model developer submits a model to the learning repository, a specification S is assigned to each model according to the steps described in section 104. m .
[0046] 106. Users construct specifications for the target task and then search for the most suitable model in the learning library;
[0047] When users want to utilize the learning materials library When solving their own tasks, the target dataset for the new task is... Its characteristic space is The label space is {1,…,K} τ First, a text-image pre-trained model maps the new task image data and corresponding data annotations to a unified representation space. Then, a thumbnail set is generated using thumbnail kernel mean embedding technology, which serves as the user's specification S. τ After receiving the user's specification, it compares the user's specification with the model specification in the learning library and returns a model that matches S. τ The model with the highest reduction similarity.
[0048] 107. Use the searched models as machine learning models for downstream tasks for further fine-tuning and reuse.
[0049] After receiving the model retrieved from the learning library, users can deploy the model on their own tasks using fine-tuning or other model reuse methods.
[0050] Obviously, those skilled in the art should understand that the steps of the image classification task learning object search and reuse method described in the above embodiments of the present invention can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. Furthermore, in some cases, the steps shown or described can be performed in a different order than presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the embodiments of the present invention are not limited to any particular hardware and software combination.
Claims
1. A method for learning object retrieval and reuse for image classification tasks, characterized in that, include: 1) Obtain the models trained on various image classification tasks and their corresponding training datasets; 2) By using the corresponding trained model to make predictions on the training data, the model prediction results and confidence scores are obtained and converted into text form; 3) Use the image encoder of the text-image pre-trained model to perform feature transformation on the training data to obtain image feature representation in a unified representation space. Use the text encoder to perform feature transformation on the text obtained from the model prediction results and confidence transformation to obtain text feature representation in a unified representation space. 4) Using the abbreviated kernel mean embedding technique, the image and text feature representations obtained in step 3) are constructed into a model reduction describing the model's function. 5) By constructing specifications for image classification models, a learning library is formed, which contains multiple image classification models, each of which has specifications that can describe the model's functionality; 6) For the new image classification task, the user uses the image encoder of the text-image pre-trained model to perform feature transformation on the training data of the new task to obtain the image feature representation, and transforms the corresponding real annotations of the data to obtain the text feature representation, and uses step 4) to construct a reduction for the new task; 7) Select the candidate model with the smallest distance to the new task specification by calculating the maximum mean difference distance between the candidate model reduction in the learning material library and the new task specification; 8) Adapt and reuse the selected model in new tasks; In step 2), natural language is used to describe the model's prediction results and confidence levels, thereby characterizing the model's performance. Steps 2)-3) use the image encoder and text encoder of the text-image pre-trained model to convert the training data and model prediction results into image representations and text representations, and map them to a unified feature space. The text-image pre-trained model is a multimodal pre-trained model trained using image-text pair data through a contrastive learning method. It includes an image encoder and a text encoder to map images and text into representations in a unified feature space, so as to facilitate more accurate reduction generation and model search.
2. The method for learning object retrieval and reuse for image classification tasks according to claim 1, characterized in that, Users construct a specification description for a new image classification task using a specification construction method, and then search for and reuse the most suitable machine learning model. This includes constructing a specification using local data and matching it with model specifications in the learning library to determine which models are best suited for the current task. The matching process refers to calculating the similarity between the user's specification and the specifications of each model in the learning library, and selecting one or more models with the highest similarity as the matching result.
3. The method for learning object retrieval and reuse for image classification tasks according to claim 1, characterized in that, The step of adapting and reusing the selected model on a new task refers to using the matched model as the initial model and then further training or adjusting it using its own data.
4. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for learning and reusing learning materials for an image classification task as described in any one of claims 1-3.
5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that performs the method of learning and reusing software for an image classification task as described in any one of claims 1-3.
Citation Information
Patent Citations
Universal learning document searching and multiplexing method and device
CN116542136A
Searching and multiplexing method for heterogeneous feature spatial learning elements
CN116629374A