Method for finely adjusting image classification model and computing equipment
By introducing a prompt generator and a pretrained model, the method of generating and embedding representations of sample images is solved, and the pretrained model has high fine-tuning computing resources and small amount of private domain data is achieved, and the fine-tuning effect of efficient image classification model is achieved.
Patent Information
- Application Number
- CN202510401549.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing pretrained image classification model consumes high computing resources during fine-tuning and is difficult to achieve insufficient distinction from the pretrained model during fine-tuning.
The prompt generator is combined with the pre-trained classification model. By generating the first prompt distribution and embedding representations with the sample image, feature extraction and classification are used for the encoding layer and the output layer, the pre-trained model parameters are frozen, and the prompt generator parameters are only adjusted to reduce calculation consumption and improve fitting ability.
While reducing the computing power consumption in the fine-tuning process, the classification accuracy and adaptability of the model are improved, achieving good fine-tuning effects.
Smart Images

Figure CN120259772A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the technical field of data processing, and in particular, relate to a method for fine-tuning an image classification model and a computing device. Background Art
[0002] Nowadays, large pre-trained models have been widely applied in the field of image classification. Generally, a pre-trained image classification model is trained using a public data set containing a large amount of image data and can complete basic image classification tasks. Thus, users can further fine-tune the publicly available image recognition model using private domain data, so that a personalized model that adapts to their own scenario requirements and has a higher classification accuracy for specific types of images can be obtained with only a small amount of training resources consumed.
[0003] However, on the one hand, when the parameter scale of the pre-trained model is large, even if only a small-scale sample data is used to fine-tune the pre-trained model, parameter adjustment still requires a large amount of computing resources; on the other hand, when the amount of private domain data used by the user for fine-tuning is small, it is difficult for the fine-tuned personalized model to have a large difference from the pre-trained model. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for fine-tuning an image classification model and a computing device, including:
[0005] A method for fine-tuning an image classification model is provided in the first aspect of this specification. The method is executed using a prompt generator and a pre-trained classification model. The pre-trained classification model includes an encoding layer and an output layer. The method includes:
[0006] Obtain training samples, where the training samples include sample images and classification labels corresponding to the sample images;
[0007] Use the prompt generator to generate a first prompt distribution according to the sample images;
[0008] Randomly sample in the first prompt distribution to determine a first prompt;
[0009] Concatenate the embedding representation of the sample image with the first prompt to obtain a first comprehensive embedding representation;
[0010] Use the encoding layer to determine a first classification feature corresponding to the first comprehensive embedding representation;
[0011] Use the output layer to determine a first classification result distribution corresponding to the first classification feature;
[0012] Train at least the prompt generator according to the first classification result distribution and the classification labels.
[0013] The second aspect of this specification provides a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in the first aspect is implemented.
[0014] An embodiment of this specification provides a method and a computing device for fine-tuning an image classification model. By means of an external prompt generator, the pre-trained classification model is fine-tuned. On the one hand, the computing power consumption in the fine-tuning process can be reduced. On the other hand, the first prompt distribution generated by the prompt generator can better fit the classification information in the image. Even under the condition of freezing the parameters of the pre-trained model, good fine-tuning effects can still be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 is a schematic structural diagram of a fine-tuning system for an image classification model in an embodiment of this specification;
[0017] Figure 2 is a schematic flowchart of a method for fine-tuning an image classification model in an embodiment of this specification;
[0018] Figure 3 is a schematic diagram of a method for masking the first image feature in an embodiment of this specification;
[0019] Figure 4(A) is a schematic flowchart of a method for image classification using a fine-tuned image classification model in an embodiment of this specification;
[0020] Figure 4(B) is a schematic flowchart of a method for image classification using a fine-tuned image classification model in an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0022] Generally, during the process of fine-tuning a pre-trained model, to save the computational overhead of fine-tuning, users will choose to freeze some model parameters and only adjust some model layers of the pre-trained model. However, when the parameter scale of the pre-trained model is large, adjusting some model layers of the pre-trained model is still a large expense. Visual Prompt Tuning (VPT) can significantly reduce the storage cost during the fine-tuning process by introducing a small number of trainable prompt parameters (usually less than 1% of the model parameters) in the input space of the pre-trained model while keeping the backbone network of the pre-trained model frozen. However, how to make the introduced prompt parameters have stronger generalization ability is still an urgent problem to be solved.
[0023] Figure 1 FIG. is a schematic structural diagram of a fine-tuning system for an image classification model provided in this specification, and the fine-tuning system can be deployed in a computing device. As shown in the figure, the fine-tuning system can include a feature extraction module, a prompt generator, an image processing module, a pre-trained classification model, and a training module. Among them, the feature extraction module can be used to extract the first image feature of the sample image; the prompt generator can be used to determine the corresponding first prompt distribution according to the first image feature output by the feature extraction module; the image processing module can be used to convert the input sample image into an embedding representation, sample a first prompt from the first prompt distribution output by the prompt generator, and splice the first prompt with the embedding representation of the sample image to obtain a first comprehensive embedding representation; the pre-trained classification model can output the first classification result distribution corresponding to the image according to the input first comprehensive embedding representation; the training module can use a preset objective function to adjust at least the parameters of the prompt generator according to the first classification result distribution and the classification label corresponding to the sample image.
[0024] In some implementation manners, the feature extraction module can be a deep learning model, such as a convolutional neural network; it can also be a computing module with a built-in feature extraction algorithm, and the feature extraction algorithm can be, for example, the Histogram of Oriented Gradient (HOG) feature extraction method or the Scale-Invariant Feature Transform (SIFT) algorithm, which is not limited in this specification. The prompt generator can be a deep learning model, such as a convolutional neural network. It should be noted that the output space of the feature extraction module is the same as the input space of the prompt generator. The pre-trained classification model can be a pre-trained Vision Transformer (ViT) model, and the pre-trained ViT model can include an encoding layer and an output layer. The encoding layer can extract the corresponding classification features according to the embedding representation of the input image, and the output layer can further extract the corresponding classification results according to the classification features.
[0025] Figure 2 The figure shows the process of a method for fine-tuning an image classification model in this specification, including:
[0026] S201: Obtain training samples, where the training samples include sample images and classification labels corresponding to the sample images.
[0027] First, a computing device deployed with a fine-tuning system can obtain training samples from a pre-existing device memory or other communicable databases. The device memory or database may store a training sample set composed of several training samples. A training sample may include a sample image and the classification label corresponding to the sample image.
[0028] In some implementation manners, in the application scenario of personalized federated learning, each user can use the method as shown in Figure 2 to obtain training samples from their respective private training sample sets, and based on the same pre-trained classification model, fine-tune to obtain a personalized image classification model adapted to the characteristics of their own training sample sets.
[0029] S203: Use the prompt generator to generate a first prompt distribution according to the sample image.
[0030] After obtaining the training image, the probability distribution of the classification information corresponding to the sample image - the first prompt distribution - can be generated according to the sample image.
[0031] Specifically, the fine-tuning system can first use the feature extraction module to encode the sample image to determine the first image feature corresponding to the sample; subsequently, input the first image feature into the prompt generator, and use the prompt generator to generate the first prompt distribution corresponding to the sample image. The first prompt distribution is the probability distribution of the prompt information corresponding to the sample image.
[0032] It should be noted that, since in the subsequent step S207, the first prompt needs to be concatenated with the embedding representation of the sample image, and in step S209, the first comprehensive embedding representation obtained by concatenating the first prompt and the embedding representation of the sample image is input into the encoding layer of the pre-trained classification model. Therefore, the dimension of the first prompt needs to be adapted to the input space of the encoding layer. For example, when the input space of the encoding layer is a one-dimensional vector with a hidden layer dimension of d, the first prompt can be a one-dimensional vector with a hidden layer dimension of d.
[0033] On the other hand, the user can pre-set the target profile of the probability distribution generated by the prompt generator. Thus, in step S203, the output result of the prompt generator is the parameters of the target profile. For example, when the target profile is a Gaussian distribution, the output result of the prompt generator can be the mean and variance of the Gaussian distribution; when the target profile is a uniform distribution, the output result of the prompt generator can be the endpoints of the interval of the uniform distribution. Among them, this specification does not limit the specific type of the target profile.
[0034] S205: Randomly sample from the first prompt distribution to determine a first prompt.
[0035] After determining the first prompt distribution, the image processing module in the fine-tuning system can perform random sampling in the first prompt distribution and determine the first prompt according to the sampling result. Specifically, random sampling can be performed once in the first prompt distribution, and the sampling result of the random sampling can be directly used as the first prompt; random sampling can also be performed several times in the first prompt distribution, and the first prompt can be further determined according to the sampling results of the several random samplings, which is not limited in this specification.
[0036] Therefore, compared with directly generating the first prompt according to the sample image, in the process of determining the first prompt using steps S203-S205, the idea of Bayesian optimization is used to regard the probability distribution that accurately represents the classification information of the sample image as the target distribution, and the process of obtaining the first prompt by sampling the first prompt distribution generated by the prompt generator and then sampling the first prompt according to the first prompt distribution is regarded as an approximate fit to the process of sampling in the target distribution. When the appropriate target probability is determined in advance, compared with directly generating the first prompt, the classification information of the sample image has stronger fitting ability and expression ability. On the other hand, the process of determining the first prompt in step S205 has a certain degree of randomness, which can reduce the risk of overfitting in the training process of subsequent steps.
[0037] It should be noted that, although the first prompt includes classification information in the sample image, the classification information in the first prompt may be a hidden feature and does not need to correspond to text or images with clear semantics.
[0038] S207: Concatenate the embedded representation of the sample image with the first prompt to obtain a first comprehensive embedded representation.
[0039] Before executing step S207, the fine-tuning system can determine the embedded representation of the sample image based on the input sample image in advance, so that after determining the first prompt, the embedded representation of the sample image can be spliced with the first prompt to obtain a first comprehensive embedded representation.
[0040] Specifically, when the pre-trained classification model is a ViT model, the sample image can be divided into image patches of a preset size, and then the pixel points corresponding to each image patch are flattened and mapped to a one-dimensional vector with a hidden layer dimension of d. Finally, the mapped one-dimensional vectors are concatenated to obtain the embedding representation corresponding to the sample image. In some implementation manners, the mapped one-dimensional vectors may further include position encodings indicating the positions of the image patches in the original sample image.
[0041] It should be noted that when the pre-trained classification model is a ViT model, the embedding representation corresponding to the sample image can be a one-dimensional vector with a hidden layer dimension of d and a length of a. Correspondingly, the first prompt can be a one-dimensional vector with a hidden layer dimension of d and a length of b. Concatenating the embedding representation of the sample image and the first prompt can obtain a first comprehensive embedding representation with a hidden layer dimension of d and a length of a + b.
[0042] S209: Using the encoding layer, determine the first classification feature corresponding to the first comprehensive embedding representation.
[0043] After determining the first comprehensive embedding representation, the fine-tuning system can use the encoding layer of the pre-trained classification model to process the first comprehensive embedding representation, and the output result of the encoding layer is the first classification feature corresponding to the first comprehensive embedding representation.
[0044] Since the first comprehensive embedding representation is obtained by concatenating the first prompt and the embedding representation of the sample image, when the prompt generator has the ability to extract classification information in the sample image, the weight of the classification information in the first comprehensive embedding representation is greater than the weight of the classification information in the embedding representation of the sample image. Further, when the encoding layers are the same, when extracting features from the first comprehensive embedding representation, compared with extracting features from the embedding representation of the sample image, the classification information in the input result of the encoding layer should also have a higher weight.
[0045] In some implementation manners, when the pre-trained classification model is a ViT model, the encoding layer may include several transformer layers.
[0046] S211: Using the output layer, determine the first classification result distribution corresponding to the first classification feature.
[0047] After determining the first classification feature, the fine-tuning system can continue to input the first classification feature into the output layer of the pre-trained classification model. The output layer can reduce the dimension of the first classification feature and map the reduced-dimensional first classification feature to a first classification result distribution. Among them, the first classification result distribution can be the probability distribution of each classification category corresponding to the sample image.
[0048] In some implementations, when the pre-trained classification model is a ViT model, the output layer may include several fully connected layers, and an activation layer may be included between any two adjacent fully connected layers.
[0049] S213: Train at least the prompt generator according to the first classification result distribution and the classification label.
[0050] After determining the first classification result distribution, the fine-tuning system can input the classification label corresponding to the sample image and the first classification result distribution into the training module, so that the training module determines the posterior distribution of the parameters to be adjusted according to the preset objective function and the current first classification result part and the classification label, and then adjusts at least the parameters of the prompt generator.
[0051] Specifically, the idea of Bayesian optimization can be used:
[0052]
[0053] Where X represents the observed data, that is, the sample image and the label corresponding to the sample image; Z represents the latent variable, including the adjustable parameters and non-adjustable parameters of the model, and the adjustable parameters at least include the parameters of the prompt generator; P(Z) represents the prior distribution of the latent variable, which can be determined according to the parameters corresponding to the latent variable in the model before executing step S213; P(X|Z) represents the probability distribution of the observed data under the current latent variable condition. Similarly, it can be directly determined according to the parameters corresponding to the latent variable in the model; P(X) represents the marginal probability distribution of the observed data. Generally, P(X) = ∫P(Z,X)dψ =
[0054] ∫P(X|Z)·P(Z)dψ is determined. The optimization goal is to determine the adjustable parameters in the posterior distribution P(Z|X) of the latent variable.
[0055] Specifically, the posterior distribution P(Z|X) of the latent variable can be directly determined after determining P(Z), P(X|Z), and P(X), and then at least the parameters of the prompt generator can be adjusted according to P(Z|X); or the parameter - expected distribution Q(Z) of the prompt generator can be set, and further according to the foregoing determination Taking the minimum difference between Q(Z) and P(Z|X) as the goal, at least adjust the prompt generator. This specification does not limit the method for training the prompt generator.
[0056] Since the structure of the prompt generator is independent of the pre-trained classification model, generally, the parameter scale of the prompt generator is much smaller than that of each model layer in the pre-trained classification model. Adjusting the parameters of the prompt generator consumes much less computing power than adjusting the original model layers in the pre-trained classification model. On the other hand, in the form of splicing the first prompt with the embedded representation of the sample image, even if the parameter scale of the prompt generator is small, it can have a relatively significant impact on the first classification result distribution.
[0057] As Figure 2 shown in a method for fine-tuning an image classification model, the pre-trained classification model is fine-tuned by means of an external prompt generator. On the one hand, the computing power consumption during the fine-tuning process can be reduced. On the other hand, the first prompt distribution generated by the prompt generator can better fit the classification information in the image. Even under the condition of freezing the parameters of the pre-trained model, good fine-tuning results can still be obtained.
[0058] In some implementation manners, in step S203 as Figure 2 shown, the sample image is encoded to determine the first image feature corresponding to the sample image, the first image feature is randomly masked to determine the first masked feature corresponding to the first image feature, and the first masked feature is input into the prompt generator to generate the first prompt distribution corresponding to the sample image.
[0059] Specifically, the process of masking the first image feature can be as Figure 3 shown. First, a first mask adapted to the dimension of the first image feature is randomly generated, and the first mask is multiplied bit by bit with the first image feature to determine the first masked feature corresponding to the first image feature. Among them, each element in the first mask may include a 0 element or a 1 element. As Figure 3 shown, the black squares in the figure represent taking 0 elements, and the white squares represent 1 elements. After the first image feature is multiplied bit by bit with the first mask, the elements corresponding to the 0 elements are masked, and the elements corresponding to the 1 elements are retained.
[0060] In some implementation manners, sampling can be performed on a pre-determined Bernoulli distribution of parameters, and then a first mask adapted to the dimension of the first image feature is generated.
[0061] Thus, for the same sample image, the first masked features obtained during training with the sample image in different rounds are also different. On the other hand, to prevent the prompt generator from falling into an overfitting state due to over-relying on the image features corresponding to some positions, the robustness of the trained prompt generator can be improved.
[0062] In some implementations, the first hint distribution is a Gaussian distribution. In step S203 as shown in Figure 3 , the mean and variance of the first hint distribution are generated.
[0063] When the first hint distribution is a Gaussian distribution, a unique first hint distribution can be uniquely represented by the determined mean and variance.
[0064] In some implementations, in step S203 as shown in Figure 2 , a number of initial hint distributions are generated according to the sample image, and a number of first hint distributions are determined according to the number of initial hint distributions; in step S205 as shown in Figure 2 , random sampling is performed in each first hint distribution to determine the initial hint corresponding to each first hint distribution, and the first hint is determined according to each initial hint.
[0065] Among them, the number of the above initial hint distributions has nothing to do with the number of first hint distributions. Each first hint distribution is determined according to each initial hint distribution, and each first hint distribution can be determined by weighted summing the parameters of each initial hint distribution with different weights.
[0066] When each initial hint distribution is a Gaussian distribution, the first hint distribution is also a Gaussian distribution.
[0067] Furthermore, each initial hint can be determined by random sampling in the first hint distribution, and the first hint can be expressed as a weighted sum of each initial hint.
[0068] As described above, the target distribution is the true distribution of the classification information of the sample image, and it is also the optimization target of the process of using the hint generator to determine the first hint. Generally, the actual distribution function of the target distribution is irregular and relatively complex. By the method of generating the first hint distribution by the hint generator and then sampling in the first hint distribution to obtain the first hint, the distribution of the first hint can be extended beyond the exponential distribution family, making this process have a stronger fitting ability for the target distribution. Furthermore, the trained hint generator can generate hint information containing classification information more accurately.
[0069] In some implementations, in step S207 as shown in Figure 2 , the embedded representation of the sample image, the first hint, and the classifier are concatenated.
[0070] Among them, when the pre-trained classification model is a ViT model, the output space of the encoding layer is usually the same as the input space dimension of the encoding layer, that is, when a vector of length n is input, its output is also a vector of length n.
[0071] Thus, when splicing to obtain the first comprehensive embedding representation, the embedding representation of the sample image, the first prompt, and the classifier can be spliced. Among the first comprehensive embedding representations corresponding to different sample images, each classifier is located at the same position. Thus, in the output result of the encoding layer, the elements corresponding to each classifier are also located at the same position. Thus, the elements corresponding to the classifier in the output result can be extracted as the first classification feature, so that it is not necessary to input the complete output result of the encoding layer into the subsequent output layer.
[0072] In some implementation manners, in step S213 as shown in Figure 2 According to the first classification result distribution and the classification label, the prompt generator and the output layer are trained.
[0073] According to the foregoing introduction to step S213, when the output layer needs to be trained, the parameters of the output layer are regarded as part of the latent variable Z, and the method corresponding to step S213 is used to determine the posterior distribution of the latent variable, so as to adjust the parameters of the output layer.
[0074] Generally, the parameter scale of the output layer is much smaller than that of the encoding layer. In step S213, on the basis of adjusting the parameters of the prompt generator, the parameters of the output layer are also adjusted, which can further improve the matching degree between the model and the training samples, improve the classification effect of the fine-tuned pre-trained classification model, and at the same time, does not consume too much computing power.
[0075] In some implementation manners, in step S213 as shown in Figure 2 According to the first classification result distribution and the classification label, with the goal of maximizing a preset objective function, at least the prompt generator is trained, where the objective function represents the evidence lower bound of the observed data, and the observed data is the sample image and the classification label.
[0076] According to the foregoing formula (1), generally, in the process of training the model, it is difficult to directly determine P(Z|X) - a large amount of computing power is required for the step of determining P(X) = ∫P(Z,X)dψ = ∫P(X|Z)·P(Z)dψ (when the latent variable, that is, the model parameters are too many, integrating the value of each model parameter requires extremely high computing power).
[0077] Thus, a desired distribution can be preset Using the evidence lower bound of the observed data to measure The difference from the true posterior distribution Pθ(X,Z), when the evidence lower bound reaches the maximum value, it means that The difference from Pθ(X,Z) is the smallest. According to this The parameters of the corresponding latent variables are used to adjust the parameters of the prompt generator, and the prompt generator can accurately generate the classification information corresponding to any image as the prompt of the image. It should be noted that in Bayesian inference, we usually need to maximize the log-likelihood function logPθ(X) of the observed data X to adjust the parameters of the model. Maximizing the log-likelihood function also means determining the model parameter θ that makes the current observed data X appear with the highest probability. Usually, the log-likelihood function uses the following formula:
[0078]
[0079] The difficulty of integrating Pθ(X,Z) has been explained above. Fitting Pθ(Z|X), substituting into formula (2) yields:
[0080]
[0081] again but Therefore, It is called the evidence lower bound. When the evidence lower bound is the largest, the log-likelihood function logPθ(X) of the observed data X can achieve its maximum value.
[0082] Specifically, X represents the sample image and Y represents the classification label. Indicates prompt information. is the mean value of the prompt information, is the variance of the prompt information, The probability distribution of When S initial prompt distributions ψ1,…,ψS are generated according to the sample image, and j first prompt distributions ψ1,…,ψJ are determined, then j initial prompts are obtained by sampling each first prompt distribution When It can be expressed as:
[0083]
[0084] Among them, qφ(ψ) represents the mixed distribution, which can correspond to the target parameter of the prompt generator; q(V|ψ,X)q(ψ|X)=
[0085] q(V,ψ|X) represents the joint probability distribution of the parameters of the first prompt distribution based on the sample image and the initial prompt, which can be determined according to the parameters preset by the prompt generator and the aforementioned process of determining the initial prompt; represents the joint distribution of the classification label based on the sample graph and the initial prompt, where Represents the probability distribution of the classification result based on the initial prompt and the sample image. Since the sample image and the first prompt (determined according to the initial prompt) are the input data of the pre-trained model, It can be determined by the current parameters of the pre-trained model, Represents the distribution of the initial prompt based on the sample image. Since the initial prompt is randomly sampled from the first prompt distribution, and the first prompt distribution is generated by the prompt generator according to the sample image, It can be determined according to the current parameters of the prompt generator; Is the probability distribution of the initial prompt based on the first prompt distribution. After determining the first prompt distribution, the corresponding distribution of the initial prompt can be determined; Is the joint distribution of the initial prompt distribution and the initial prompt. This joint distribution can be determined according to the target parameters of the prompt generator and the steps of determining the first prompt distribution according to the initial prompt distribution and sampling the initial prompt from the first prompt distribution.
[0086] Thus, the random gradient ascent method can be used to determine the parameters of the prompt generator that can make Ascend as the parameters of the updated prompt generator.
[0087] Figure 4(A) shows a schematic flowchart of an image classification method provided in this specification. This method is executed using a prompt generator and a pre-trained classification model. The pre-trained classification model includes an encoding layer and an output layer. The method includes:
[0088] Step S401A: Obtain the image to be classified.
[0089] Step S403A: Use the prompt generator to generate the mean of the second prompt distribution according to the image to be classified as the second prompt corresponding to the image to be classified.
[0090] Specifically, the description corresponding to step S203 can be referred to for generating the second prompt distribution. However, after generating the mean of the second prompt distribution, there is no need to continue sampling, and the mean of the second prompt distribution is directly used as the second prompt corresponding to the image to be classified.
[0091] It should be noted that only one second prompt needs to be generated for the image to be classified in step S403A.
[0092] Step S405A: Determine the classification result of the image to be classified according to the embedded representation of the image to be classified and the second prompt.
[0093] Specifically, refer to the descriptions corresponding to steps S207 - S211. According to the embedding representation of the image to be classified and the second hint, determine the second classification result distribution of the image to be classified, and use the classification category with the highest corresponding probability in the second classification result distribution as the classification result corresponding to the image to be classified.
[0094] Figure 4(B) also shows a schematic flowchart of image classification using a fine-tuned image classification model provided in this specification. This method is executed using a hint generator and a pre-trained classification model. The pre-trained classification model includes an encoding layer and an output layer. This method includes:
[0095] Step S401B: Obtain the image to be classified.
[0096] Step S401B is the same as step S401A and will not be elaborated here.
[0097] Step S403B: Encode the image to be classified to determine the second image feature corresponding to the image to be classified.
[0098] Specifically, refer to the description corresponding to S203 and will not be elaborated here.
[0099] Step S405B: Perform n random masks on the second image feature to determine n second masked features corresponding to the image feature.
[0100] The method of performing random masking on the second image feature can refer to the method of masking the first image feature described above. Of course, the Bernoulli distribution used for random masking of the second image feature and the parameter of the Bernoulli distribution used for masking the first image feature - that is, the probabilities of 0 elements and 1 elements appearing in the mask - may not be the same.
[0101] Thus, several different second masked features can be obtained from one second image feature.
[0102] Step S407B: Input the n second masked features into the hint generator respectively, and generate the mean of the n second hint distributions corresponding to the image to be classified as the n second hints corresponding to the image to be classified.
[0103] Specifically, for any second masked feature, refer to step S207 to determine the second hint corresponding to the second masked feature, and thus determine n second hints.
[0104] Step S409B: Concatenate the n second hints with the embedding representation of the image to be classified respectively to obtain n second comprehensive embedding representations.
[0105] Specifically, for any second prompt, refer to step S209 to determine the second comprehensive embedding representation corresponding to the second prompt, thereby determining n second comprehensive embedding representations.
[0106] Step S411B: Determine the n second classification result distributions respectively corresponding to the n second comprehensive embedding representations.
[0107] Specifically, for any second comprehensive embedding representation, refer to step S211 to determine the second classification result distribution corresponding to the second comprehensive embedding representation.
[0108] Step S413B: Determine the classification result of the image to be classified according to the n second classification result distributions.
[0109] After determining each second classification result distribution, the classification result of the image to be classified can be determined by taking the classification category with the highest total probability based on the probabilities corresponding to each classification category in each second classification result distribution.
[0110] Since each second prompt is generated according to different mask features, further, step S411B determines the second classification result distribution corresponding to each second prompt, which can reduce the influence of specific positions in the image to be classified. On the other hand, it can also achieve the effect of repeating experiments multiple times and improve the accuracy of the classification result of the image to be classified.
[0111] In the 1990s, it was obvious to distinguish whether an improvement to a technology was a hardware improvement (e.g., improvement to circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement to method flows). However, with the development of technology, many improvements to method flows today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program by themselves to "integrate" a digital system on a piece of PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called Hardware Description Language (HDL). And there is not only one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0112] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0113] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers for implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.
[0114] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among many execution orders of steps and does not represent the only execution order. When the actual device or terminal product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, product or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device including the said elements. For example, if terms such as first and second are used to denote names, they do not denote any specific order.
[0115] For convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing one or more of this specification, the functions of each module may be implemented in the same or multiple software and / or hardware, or the modules implementing the same function may be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some interfaces. The indirect coupling or communication connection of the device or unit may be in electrical, mechanical or other forms.
[0116] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for realizing the function specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0117] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the acts Figure 1 and / or boxes Figure 1 specified in one or more of the acts and / or boxes
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing steps for implementing the functions specified in one or more of the acts Figure 1 and / or boxes Figure 1 specified in one or more of the acts and / or boxes
[0119] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory
[0120] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium
[0121] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves
[0122] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0123] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0124] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For related parts, reference can be made to the description of the method embodiments. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0125] The above description is only for the embodiments of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included within the scope of the claims.
Claims
1. A method for fine-tuning an image classification model, the method being performed using a prompt generator and a pre-trained classification model, the pre-trained classification model including an encoding layer and an output layer, the method comprising: Obtaining training samples, the training samples including sample images and classification labels corresponding to the sample images; Using the prompt generator to generate a first prompt distribution according to the sample images; Randomly sampling in the first prompt distribution to determine a first prompt; Concatenating the embedding representation of the sample image with the first prompt to obtain a first comprehensive embedding representation; Using the encoding layer to determine a first classification feature corresponding to the first comprehensive embedding representation; Using the output layer to determine a first classification result distribution corresponding to the first classification feature; Training at least the prompt generator according to the first classification result distribution and the classification labels.
2. The method according to claim 1, using the prompt generator to generate a first prompt distribution according to the sample images, specifically comprising: Encoding the sample images to determine first image features corresponding to the sample images; Randomly masking the first image features to determine first masked features corresponding to the first image features; Inputting the first masked features into the prompt generator to generate a first prompt distribution corresponding to the sample images.
3. The method according to claim 1, wherein the first prompt distribution is a Gaussian distribution; Generating a first prompt distribution, specifically comprising: Generating the mean and variance of the first prompt distribution.
4. The method according to claim 1, generating a first prompt distribution according to the sample images, specifically comprising: Generating a plurality of initial prompt distributions according to the sample images; Determining a plurality of first prompt distributions according to the plurality of initial prompt distributions; Randomly sampling in the first prompt distribution to determine a first prompt, specifically comprising: Randomly sampling in each first prompt distribution to determine an initial prompt corresponding to each first prompt distribution; Determining a first prompt according to the respective initial prompts.
5. The method according to claim 1, concatenating the embedding representation of the sample image with the first prompt, specifically comprising: Concatenating the embedding representation of the sample image, the first prompt, and a classifier.
6. The method according to claim 1, training at least the prompt generator according to the first classification result distribution and the classification labels, specifically comprising: Training the prompt generator and the output layer according to the first classification result distribution and the classification labels.
7. The method according to claim 3, training at least the prompt generator according to the first classification result distribution and the classification labels, specifically comprising: Training at least the prompt generator with the goal of maximizing a preset objective function according to the first classification result distribution and the classification labels, wherein the objective function represents the evidence lower bound of the observed data, and the observed data is the sample image and the classification label.
8. The method according to claim 3, further comprising: Obtaining an image to be classified; Using the hint generator, generate the mean of the second hint distribution according to the image to be classified as the second hint corresponding to the image to be classified. Determine the classification result of the image to be classified according to the embedding representation of the image to be classified and the second hint.
9. The method according to claim 8, using the hint generator to generate the mean of the second hint distribution according to the image to be classified as the second hint corresponding to the image to be classified, specifically including: Encode the image to be classified to determine the second image feature corresponding to the image to be classified. Perform n random masks on the second image feature to determine n second masked features corresponding to the image feature. Input the n second masked features into the hint generator respectively to generate the means of n second hint distributions corresponding to the image to be classified as n second hints corresponding to the image to be classified. Determine the classification result of the image to be classified according to the embedding representation of the image to be classified and the second hint, specifically including: Concatenate the n second hints with the embedding representation of the image to be classified respectively to obtain n second comprehensive embedding representations. Determine n second classification result distributions corresponding to the n second comprehensive embedding representations respectively. Determine the classification result of the image to be classified according to the n second classification result distributions.
10. A computing device, including a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-9 is implemented.
Citation Information
Patent Citations
Long tail distribution-oriented vision-language model prompt learning framework
CN118917276A
Text embedding adapter
US20250005807A1