Long-tail target classification model training method and device based on Prompt-Tuning algorithm
By using the Prompt-Tuning algorithm in the long-tail target classification model, learningable Prompt shared tokens and Prompt class tokens are introduced, and a multi-expert model network is built, which solves the problem of poor generalization ability of fine-tuning pre-trained models in the existing technology, achieving better generalization ability and reducing the risk of overfitting.
Patent Information
- Application Number
- CN202510139680.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the generalization ability of fine-tuning pre-trained models is poor and overfitting is prone to occur, resulting in poor performance in long-tail target classification tasks.
The long-tail target classification model training method based on the Prompt-Tuning algorithm is adopted. By constructing pre-trained models and long-tail distribution data sets, the pre-trained models are frozen and learnable Prompt shared tokens and Prompt class tokens are introduced to build a multi-expert model network to optimize the model's domain adaptability and generalization capabilities on the long-tail distribution data set.
The generalization ability of the long-tail target classification model is improved, the probability of the model being overfitted in the target domain is reduced, and the classification performance of the model on the long-tail distribution dataset is significantly improved.
Smart Images

Figure CN120070976A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing models, and in particular to a method and device for training a long-tail target classification model based on the Prompt-Tuning algorithm. Background Art
[0002] In view of the excellent performance of large models in the fields of natural language processing, computer vision, etc., the long-tail target classification method technology field often uses supervised or self-supervised training examples based on large models and large-scale data sets to obtain a pre-trained model with strong discriminative ability; then, based on the pre-trained model and the long-tail distribution data set, methods such as resampling or reweighting are used to improve the adaptive ability of the model to the long-tail distribution data set. However, in actual application scenarios, fine-tuning the pre-trained model is often limited by computing resources, resulting in an increase in training costs; secondly, fine-tuning the pre-trained model often leads to a decline in the generalization ability of the model, causing the model to overfit in the target domain.
[0003] Therefore, providing a method and device for training a long-tail target classification model based on the Prompt-Tuning algorithm to improve the generalization ability of the long-tail target classification method and reduce the probability of the model overfitting in the target domain has become an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0004] To this end, the embodiments of the present invention provide a method and device for training a long-tail target classification model based on the Prompt-Tuning algorithm to solve the problem that the generalization ability of fine-tuning the pre-trained model in the prior art is poor and overfitting is likely to occur, thereby improving the generalization ability of the long-tail target classification method and reducing the probability of the model overfitting in the target domain.
[0005] To achieve the above object, the embodiments of the present invention provide the following technical solutions:
[0006] The present invention provides a method for training a long-tail target classification model based on the Prompt-Tuning algorithm, the method comprising:
[0007] Construct a pre-trained model and a long-tail distribution data set, and divide the long-tail distribution data set into a training set, a validation set, and a test set;
[0008] In an image processing task, construct an optimization strategy for the pre-trained model, and build a multi-expert model network based on the optimization strategy;
[0009] Use the training set of the long-tail distribution data set to train the multi-expert model network to obtain an initial classification model;
[0010] Use the validation set and test set of the long-tailed distribution dataset to validate and test the initial classification model respectively to obtain the final long-tailed target classification model.
[0011] In some embodiments, constructing the pre-trained model specifically includes:
[0012] Based on a large model and an open dataset, obtain the pre-trained model through supervised training or self-supervised training.
[0013] In some embodiments, the optimization strategy for constructing the pre-trained model specifically includes:
[0014] Freeze the pre-trained model, and optimize the pre-constructed supervised training network by introducing learnable prompt sharing patches using the long-tailed distribution dataset.
[0015] In some embodiments, freezing the pre-trained model and optimizing the pre-constructed supervised training network by introducing learnable prompt sharing patches using the long-tailed distribution dataset specifically includes:
[0016] The input image passes through the patch vectorization layer to obtain the feature patch Z 0 and the class patch C 0 of Z 0 ;
[0017] The pre-trained model contains L Transformer Blocks, and the i-th Block is denoted as Block i , and its input is C i-1 , Z i-1 and the prompt sharing patch U i-1 , input q i-1 = [C i-1 , Z i-1 , k i-1 = [C i-1 , Z i-1 , U i-1 , v i-1 = [C i-1 , Z i-1 , U i-1 , where U i-1 is the learnable prompt sharing patch corresponding to the i-th Block, and the updated expression of (C i , Z i ) is:
[0018] (C i , Z i ) = Block i (q i-1 , k i-1 , v i-1 )
[0019] The discriminative feature of the image is C L , and the logits s of the image can be expressed as:
[0020] s = f θ (C L )
[0021] where f θ is a learnable cosine classifier;
[0022] The loss function uses the cross-entropy loss function and can be expressed as:
[0023] L loss = L cross-entropy (s, y)
[0024] where y is the label information of the image.
[0025] In some embodiments, a multi-expert model network is built based on the optimization strategy, which specifically includes:
[0026] Find the most matching head-expert category hint patch, middle-expert category hint patch, and tail-expert category hint patch in the head-expert category hint, middle-expert category hint, and tail-expert category hint respectively;
[0027] The expert model introduces respectively in the last L - K Blocks of the supervised training network where, represents the most matching head-expert category hint patch, represents the most matching middle-expert category hint patch, represents the most matching tail-expert category hint patch.
[0028] In some embodiments, the index for finding the most matching head-expert category hint patch in the head-expert category hint is expressed as:
[0029] i h = top-1(<C L , [r h0 ,..., r hN-1 >)
[0030] where r hi represents the head-expert category hint patch with the category index i + 1, N is the number of categories, and < > represents the cosine similarity operation.
[0031] In some embodiments, the loss function of the expert model is:
[0032]
[0033]
[0034] Where ||D|| is the number of samples in a batch, is the logical value of the i-th sample in the batch for the k-th category, y i is the label information of the i-th sample, λ is a hyperparameter designed to differentiate the capabilities of the expert models. Specifically, when λ > 1, the expert model is the tail expert; when 1 > λ > 0, the expert model is the middle expert; when λ < 0, the expert model is the head expert, x i represents the input image of the i-th sample, θ represents the training parameters of the model, p(x i , θ) represents the probability distribution of the model predicting the i-th sample, and n represents the number of samples. represents the logical value of the i-th sample in the batch for the c-th category.
[0035] The present invention also provides a long-tail object classification model training device based on the Prompt-Tuning algorithm, and the device includes:
[0036] A dataset construction unit for constructing a pre-trained model and a long-tail distribution dataset, and dividing the long-tail distribution dataset into a training set, a validation set, and a test set;
[0037] A network construction unit for constructing an optimization strategy for the pre-trained model in an image processing task, and building a multi-expert model network based on the optimization strategy;
[0038] A model training unit for training the multi-expert model network using the training set of the long-tail distribution dataset to obtain an initial classification model;
[0039] A model optimization unit for validating and testing the initial classification model using the validation set and the test set of the long-tail distribution dataset respectively to obtain a final long-tail object classification model.
[0040] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned method are implemented.
[0041] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method are implemented.
[0042] The long-tail target classification model training method and device provided by the present invention construct a pre-trained model and a long-tail distribution data set, and divide the long-tail distribution data set into a training set, a validation set, and a test set; in an image processing task, construct an optimization strategy for the pre-trained model, and build a multi-expert model network based on the optimization strategy; use the training set of the long-tail distribution data set to train the multi-expert model network to obtain an initial classification model; use the validation set and the test set of the long-tail distribution data set to verify and test the initial classification model respectively to obtain a final long-tail target classification model.
[0043] In this way, the method first obtains a pre-trained model with strong representation ability based on a large model and a large-scale data set, using supervised or self-supervised training examples; secondly, freeze the pre-trained model, and improve the domain adaptation ability and class discrimination ability of the model in the long-tail distribution data set by introducing learnable Prompt shared tokens; finally, introduce Prompt class tokens to establish the network structure and loss of the multi-expert model, reduce the bias and variance of model prediction, and thus improve the generalization ability of the model for the long-tail distribution data set, thereby solving the problem of poor generalization ability of fine-tuning the pre-trained model in the prior art and prone to overfitting, improving the generalization ability of the long-tail target classification method, and reducing the probability of overfitting of the model in the target domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained according to the provided drawings.
[0045] The structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have technical essence. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that the technical content disclosed by the present invention can cover.
[0046] Figure 1 It is a flowchart of the long-tail target classification model training method provided by the present invention based on the Prompt-Tuning algorithm;
[0047] Figure 2One of the schematic diagrams of the model network provided by the present invention;
[0048] Figure 3 Another schematic diagram of the model network provided by the present invention;
[0049] Figure 4 Block diagram of the structure of the long-tail target classification model training device based on the Prompt-Tuning algorithm provided by the present invention;
[0050] Figure 5 Block diagram of the structure of a computer device provided by the present invention. Detailed implementation manners
[0051] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] The following combines Figures 1 - 4 to introduce the long-tail target classification model training method and device based on the Prompt-Tuning algorithm provided by the present invention.
[0053] Please refer to Figure 1 , Figure 1 Flowchart of the long-tail target classification model training method based on the Prompt-Tuning algorithm provided by the present invention.
[0054] In a specific implementation manner, the present invention provides a long-tail target classification model training method based on the Prompt-Tuning algorithm, and the method includes the following steps:
[0055] S110: Construct a pre-trained model and a long-tail distribution data set, and divide the long-tail distribution data set into a training set, a validation set, and a test set; that is, based on a large model and an open data set, obtain the pre-trained model through supervised training or self-supervised training; it should be understood that this long-tail distribution data set is a data set in a natural scene, such as Image Net-LT, iNaturalist 2018; the division ratio of the training set, the validation set, and the test set can be 7:2:1 or 8:1:1.
[0056] S120: In the image processing task, an optimization strategy for constructing a pre-trained model is built, and a multi-expert model network is built based on the optimization strategy; based on the good feature extraction ability of the pre-trained large model in the downstream long-tail classification task, a two-stage training method is adopted to gradually improve the classification performance of the model for long-tail distribution data. In the first stage, the domain adaptation ability of the model for the long-tail distribution data set is improved by introducing Promptshared token. In the second stage, based on the characteristics of discriminative features being compact within classes and dispersed between classes in the first stage, a learnable Prompt class token is introduced in combination with the network structure of the multi-expert model, thereby improving the generalization ability of the model for the long-tail distribution data set; in the inference stage, the model adopts a two-stage network structure, synthesizes the prediction results of each expert, reduces the bias and variance of the model prediction, and improves the long-tail classification performance of the model.
[0057] S130: Use the training set of the long-tail distribution data set to train the multi-expert model network to obtain an initial classification model; that is, use the training set to train the constructed deep learning network to obtain a preliminary classification model.
[0058] S140: Use the validation set and test set of the long-tail distribution data set to validate and test the initial classification model respectively to obtain the final long-tail target classification model.
[0059] In some embodiments, the optimization strategy for constructing a pre-trained model specifically includes:
[0060] Freeze the pre-trained model, and by introducing a learnable Prompt shared token (prompt shared patch), optimize the pre-constructed supervised training network using the long-tail distribution data set.
[0061] Aim to improve the feature discrimination ability of the model for the long-tail distribution data set in the first stage by freezing the pre-trained model, introducing a learnable Prompt shared token and a cosine classification head based on the good feature discrimination ability of the pre-trained model, thereby significantly improving the domain adaptation ability of the model for the long-tail distribution data set; accordingly, optimizing the long-tail distribution data set using the pre-constructed supervised training network specifically includes:
[0062] The input image passes through the Patch Embeding (patch vectorization) layer to obtain the Patch token (feature patch) Z 0 and Z 0 's Class token (category patch) C 0 ;
[0063] The pre-trained model contains L Transformer Blocks. The i-th Block is denoted as Block i , and its inputs are C i-1 , Z i-1 , and the Prompt shared token U i-1 . The input q i-1 = [C i-1 , Z i-1 , k i-1 = [C i-1 , Z i-1 , U i-1 , v i-1 = [C i-1 , Z i-1 , U i-1 , where U i-1 is the learnable Prompt shared token corresponding to the i-th Block. The updated expression of (C i , Z i ) is:
[0064] (C i , Z i ) = Block i (q i-1 , k i-1 , v i-1 )
[0065] The discriminative feature of the image is C L . The logical value s of the image can be expressed as:
[0066] s = f θ (C L )
[0067] where f θ is a learnable cosine classifier;
[0068] The loss function uses the cross-entropy loss function and can be expressed as:
[0069] L loss = L cross-entropy (s, y)
[0070] where y is the label information of the image.
[0071] The one-stage model significantly improves the domain adaptation ability of the model through supervised training. That is, the discriminative features of the one-stage model are characterized by intra-class compactness and inter-class dispersion. However, the characteristics of the long-tail dataset lead to an unbalanced spatial distribution of the discriminative features, that is, the head-class samples compress the spatial distribution of the tail-class samples, resulting in the phenomenon that the model misclassifies the tail-class samples as head-class samples, which hinders the improvement of the model performance. Therefore, learnable Prompt classtokens are introduced, specifically the head expert class prompt patch, the middle expert class prompt patch, and the tail expert class prompt patch, aiming to provide the model with priors about the class discriminative features, thereby improving the generalization ability of the model for the long-tail distribution dataset.
[0072] Adopt the network structure of the multi-expert model. The Head expert class prompt (head expert class prompt), Middle expert class prompt (middle expert class prompt), and Tail expert class prompt (tail expert class prompt) respectively provide fine-grained discriminative features for the Head expert (head expert), Middle expert (middle expert), and Tail expert (tail expert). The prediction results integrate the knowledge of each expert, reducing the bias and variance of the model prediction, thereby improving the generalization ability of the model for the long-tail distribution dataset. The network structure of the two-stage model is as Figure 3 shown. Based on this, build a multi-expert model network based on the optimization strategy, specifically including:
[0073] Complete the supervised training of the one-stage model, which already has the characteristics of intra-class compactness and inter-class dispersion. Based on the good class feature discrimination ability, according to the cosine similarity information, find the most matching Head expert prompt class token (head expert class prompt patch), Middle expert prompt class token (middle expert class prompt patch), and Tail expert prompt class token (tail expert class prompt patch) in the Head expert class prompt, Middle expert class prompt, and Tail expert class prompt respectively; for example, in the Head expert class prompt
[0074] The index for finding the most matching Head expert prompt class token can be expressed as:
[0075] i h = top-1(<CL , [r h0 ,..., r hN-1 )
[0076] Among them, r hi represents the Head expert prompt class token with a category index of i + 1, N is the number of categories, and < > represents the cosine similarity operation.
[0077] In the last L - K blocks of the supervised training network, the expert model respectively introduces Combined with the network structure of the one - stage model, the network structure of the two - stage model is constructed. As Figure 3 shown, specifically, on the basis of maintaining the network structure of the one - stage model, in its last L - K blocks, the expert model respectively introduces to provide fine - grained knowledge within the domain for the expert model. In addition, the fine - grained knowledge within the domain only comes from and the optimization of the weight parameters of the cosine classification head. Taking the Head expert as an example below, the network structure of the expert model is elaborated. For K <= i <= L - 1, (C i , Z i ) can be updated as:
[0078] (C i , Z i ) = Block i (q i-1 , k i-1 , v i-1 )
[0079] In the above formula, q i-1 = [C i-1 , Z i-1 ,
[0080] The setting of the loss aims to differentiate the capabilities of the expert model, thereby reducing the bias and variance of model prediction. The loss can be expressed as:
[0081]
[0082] Among them, ||D|| is the number of samples in the batch, is the logical value of the i - th sample in the batch for the k - th category, y i is the label information of the i - th sample, λ is a hyperparameter aimed at differentiating the capabilities of the expert model. Specifically, when λ > 1, the expert model is the tail expert; when 1 > λ > 0, the expert model is the middle expert; when λ < 0, the expert model is the head expert, x i represents the input image of the i - th sample, and θ represents the training parameters of the model, p(xi , θ) represents the probability distribution of the model predicting the i-th sample, n represents the number of samples, It represents the logical value of the i-th sample in the batch for the c-th category.
[0083] In the above specific implementation manner, the method for training a long-tail target classification model based on the Prompt-Tuning algorithm provided by the present invention constructs a pre-trained model and a long-tail distribution data set, and divides the long-tail distribution data set into a training set, a validation set, and a test set; in an image processing task, an optimization strategy for the pre-trained model is constructed, and a multi-expert model network is built based on the optimization strategy; the training set of the long-tail distribution data set is used to train the multi-expert model network to obtain an initial classification model; the validation set and the test set of the long-tail distribution data set are used to optimize the initial classification model respectively to obtain a final long-tail target classification model.
[0084] In this way, the method first obtains a pre-trained model with strong representation ability based on a large model and a large-scale data set, using supervised or self-supervised training examples; secondly, the pre-trained model is frozen, and by introducing learnable Prompt shared tokens, the domain adaptation ability and class discrimination ability of the model in the long-tail distribution data set are improved; finally, Prompt class tokens are introduced to establish the network structure and loss of the multi-expert model, reducing the bias and variance of model prediction, thereby improving the generalization ability of the model for the long-tail distribution data set, thus solving the problem that the generalization ability of fine-tuning the pre-trained model in the prior art is poor and overfitting is likely to occur, enhancing the generalization ability of the long-tail target classification method, and reducing the probability of the model overfitting in the target domain.
[0085] In addition to the above method, the present invention also provides a device for training a long-tail target classification model based on the Prompt-Tuning algorithm, as Figure 4 shown, the device includes:
[0086] A data set construction unit 410, configured to construct a pre-trained model and a long-tail distribution data set, and divide the long-tail distribution data set into a training set, a validation set, and a test set;
[0087] A network construction unit 420, configured to construct an optimization strategy for the pre-trained model in an image processing task, and build a multi-expert model network based on the optimization strategy;
[0088] A model training unit 430, configured to use the training set of the long-tail distribution data set to train the multi-expert model network to obtain an initial classification model;
[0089] The model optimization unit 440 is configured to optimize the initial classification model by using the validation set and the test set of the long-tail distribution dataset respectively, so as to obtain the final long-tail target classification model.
[0090] In some embodiments, constructing the pre-trained model specifically includes:
[0091] Based on a large model and an open dataset, the pre-trained model is obtained through supervised training or self-supervised training.
[0092] In some embodiments, the optimization strategy for constructing the pre-trained model specifically includes:
[0093] Freeze the pre-trained model, and by introducing a learnable Prompt shared token, optimize the pre-constructed supervised training network by using the long-tail distribution dataset.
[0094] In some embodiments, freeze the pre-trained model, and by introducing a learnable Prompt shared token, optimize the long-tail distribution dataset by using the pre-constructed supervised training network, which specifically includes:
[0095] The input image passes through the Patch Embeding layer to obtain the Patch token Z 0 and the Class token C of Z 0 ; 0 ;
[0096] The pre-trained model contains L TransformerBlocks, and the i-th Block is denoted as Block i , and its input is C i-1 , Z i-1 and the Prompt shared token U i-1 , the input q i-1 = [C i-1 , Z i-1 , k i-1 = [C i-1 , Z i-1 , U i-1 , v i-1 = [C i-1 , Z i-1 , U i-1 , where U i-1 is the learnable Prompt shared token corresponding to the i-th Block, and the updated expression of (C i , Z i ) is:
[0097] (Ci , Z i ) = Block i (q i-1 , k i-1 , v i-1 )
[0098] The discriminative feature of the image is C L , and the logits s of the image can be expressed as:
[0099] s = f θ (C L )
[0100] where f θ is a learnable cosine classifier;
[0101] The loss function uses the cross-entropy loss function and can be expressed as:
[0102] L loss = L cross-entropy (s, y)
[0103] where y is the label information of the image.
[0104] In some embodiments, a multi-expert model network is built based on the optimization strategy, specifically including:
[0105] Find the most matching Head expert prompt class token, Middle expert prompt class token, and Tail expert prompt class token in Head expert class prompt, Middle expert class prompt, and Tail expert class prompt respectively;
[0106] The expert model introduces respectively in the last L - K Blocks of the supervised training network
[0107] In some embodiments, the index representation of finding the most matching Head expert prompt class token in Head expert class prompt is:
[0108] i h = top-1(<C L , [r h0 ,..., r hN-1 >)
[0109] where rh iDenote the Head expert prompt class token with class index \(i + 1\), \(N\) is the number of classes, and \(\langle\rangle\) represents the cosine similarity operation.
[0110] In some embodiments, the loss function of the expert model is:
[0111]
[0112] where \(\vert\vert D\vert\vert\) is the number of samples in the batch, is the logical value of the \(i\)-th sample in the batch for the \(k\)th class, \(y\) i is the label information of the \(i\)-th sample, \(\lambda\) is a hyperparameter designed to differentiate the capabilities of the expert model. Specifically, when \(\lambda>1\), the expert model is a tail expert; when \(1>\lambda>0\), the expert model is an intermediate expert; when \(\lambda<0\), the expert model is a head expert, \(x\) i represents the \(i\)-th sample input image, \(\theta\) represents the training parameters of the model, \(p(x\) i , \(\theta)\) represents the probability distribution of the model predicting the \(i\)-th sample, and \(n\) represents the number of samples, represents the logical value of the \(i\)-th sample in the batch for the \(c\)th class.
[0113] In the above specific implementation, the long-tail target classification model training device provided by the present invention constructs a pre-trained model and a long-tail distribution dataset, and divides the long-tail distribution dataset into a training set, a validation set, and a test set; in the image processing task, constructs an optimization strategy for the pre-trained model, and builds a multi-expert model network based on the optimization strategy; uses the training set of the long-tail distribution dataset to train the multi-expert model network to obtain an initial classification model; uses the validation set and test set of the long-tail distribution dataset to verify and test the initial classification model to obtain the final long-tail target classification model.
[0114] In this way, the device first obtains a pre-trained model with strong representation ability based on a large model and a large-scale dataset using supervised or self-supervised training examples; secondly, freezes the pre-trained model and improves the domain adaptation ability and class discrimination ability of the model in the long-tail distribution dataset by introducing learnable Prompt shared tokens; finally, introduces Prompt class tokens, establishes the network structure and loss of the multi-expert model, reduces the bias and variance of model prediction, and thereby improves the generalization ability of the model for the long-tail distribution dataset, thus solving the problem that the generalization ability of fine-tuning the pre-trained model in the prior art is poor and overfitting is likely to occur, enhancing the generalization ability of the long-tail target classification method, and reducing the probability of overfitting of the model in the target domain.
[0115] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structural diagram may be as shown in Figure 5 . The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store static information and dynamic information data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiment.
[0116] Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0117] Corresponding to the above embodiment, an embodiment of the present invention further provides a computer storage medium, which contains one or more program instructions. Among them, the one or more program instructions are used to be executed in the method as described above.
[0118] The present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the above method.
[0119] In an embodiment of the present invention, the processor may be an integrated circuit chip with signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0120] The various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. The processor reads the information in the storage medium and combines its hardware to complete the steps of the above method.
[0121] The storage medium can be a memory, for example, it can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.
[0122] Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.
[0123] The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).
[0124] The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memories.
[0125] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by a combination of hardware and software. When applying software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transmission of a computer program from one place to another. The storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0126] The above specific implementation manners have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only the specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included in the protection scope of the present invention.
Claims
1. A long-tail target classification model training method based on the Prompt-Tuning algorithm, characterized in that: The method comprises: Constructing a pre-training model and a long-tail distribution data set, and dividing the long-tail distribution data set into a training set, a validation set, and a test set; In image processing tasks, construct an optimization strategy for pre-trained models and build a multi-expert model network based on the optimization strategy; Using the training set of the long-tail distribution data set to train the multi-expert model network to obtain an initial classification model; The initial classification model is verified and tested respectively using the verification set and the test set of the long-tail distribution data set to obtain the final long-tail target classification model.
2. The long-tail target classification model training method based on the Prompt-Tuning algorithm according to claim 1 is characterized in that: Building a pre-trained model specifically includes: Based on a large model and an open data set, the pre-trained model is obtained through supervised training or self-supervised training.
3. The long-tail target classification model training method based on the Prompt-Tuning algorithm according to claim 1 is characterized in that: Optimization strategies for building pre-trained models include: The pre-trained model is frozen and the pre-built supervised training network is optimized using a long-tail distribution dataset by introducing a learnable cue-sharing patch.
4. The long-tail target classification model training method based on the Prompt-Tuning algorithm according to claim 3 is characterized in that: Freeze the pre-trained model, introduce learnable hint sharing patches, and optimize the pre-built supervised training network using a long-tail distribution dataset, including: The input image passes through the patch vectorization layer to obtain the feature patch Z0 and the category patch C0 of Z0; The pre-trained model contains L attention modules Block, and the i-th Block is represented as Block i , whose input is C i-1 , Z i-1 And prompt sharing patch U i-1 , enter q i-1 =[C i-1 ,Z i-1 ],k i-1 =[C i-1 ,Z i-1 ,U i-1 ],v i-1 =[C i-1 ,Z i-1 ,U i-1 ], where U i-1 is the learnable hint shared patch corresponding to the i-th Block, (C i ,Z i ) is updated as follows: (C i ,Z i )=Block i (q i-1 ,k i-1 ,v i-1 ) The discriminant feature of the image is C L , the original prediction value s of the image is expressed as: s=f θ (C L ) Among them, f θ is a learnable cosine classifier; Loss function L loss Using the cross entropy loss function L cross-entropy (s, y), expressed as: L loss =L cross-entropy (s,y) Among them, y is the label information of the image.
5. The long-tail target classification model training method based on the Prompt-Tuning algorithm according to claim 1 is characterized in that: Building a multi-expert model network based on the optimization strategy specifically includes: Find the best matching head expert category prompt patch, middle expert category prompt patch and tail expert category prompt patch in the head expert category prompt, middle expert category prompt and tail expert category prompt respectively; The expert model is introduced in the last LK blocks of the supervised training network. in, indicates the best matching head expert category hint patch, indicates the best matching intermediate expert category hint patch, Indicates the best matching tail expert category hint patch.
6. The long-tail target classification model training method based on the Prompt-Tuning algorithm according to claim 5 is characterized in that: The index of the head expert category hint patch that finds the best match in the head expert category hint is represented as: i h =top-1(<C L ,[r h0 ,...,r hN-1 ]>) Among them, r hi represents the head expert category hint patch with category index i+1, N is the number of categories, and <> represents the cosine similarity operation.
7. The long-tail target classification model training method based on the Prompt-Tuning algorithm according to claim 5 is characterized in that: The loss function L of the expert model ep for: Among them, ||D|| is the number of samples in the batch, is the logical value of the i-th sample in the batch in the k-th category, y i is the label information of the i-th sample, λ is a hyperparameter, when λ>1, the expert model is a tail expert, when 1>λ>0, the expert model is an intermediate expert, when λ<0, the expert model is a head expert, x i represents the i-th sample input image, θ represents the training parameters of the model, p(x i , θ) represents the probability distribution of the model predicting the i-th sample, n represents the number of samples, Represented as the logical value of the i-th sample in the batch in category c.
8. A long-tail target classification model training device based on the Prompt-Tuning algorithm, characterized in that: The device comprises: A data set construction unit, used to construct a pre-training model and a long-tail distribution data set, and divide the long-tail distribution data set into a training set, a validation set, and a test set; A network construction unit, used to construct an optimization strategy for a pre-trained model in an image processing task, and to build a multi-expert model network based on the optimization strategy; A model training unit, used for training the multi-expert model network using a training set of the long-tail distribution data set to obtain an initial classification model; The model optimization unit is used to verify and test the initial classification model using the verification set and the test set of the long-tail distribution data set, so as to obtain the final long-tail target classification model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.