Reinforcement learning-based classification model training method, system, device and medium
Patent Information
- Application Number
- CN202311385573.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-10-23
AI Technical Summary
但是,无论是手动裁剪还是基于规则的裁剪,在裁剪的过程中都忽略了模型的性能问题,从而导致裁剪的效率和准确度降低
[0015]本申请提出的基于强化学习的分类模型训练方法、系统、设备及介质,通过将第一样本数据输入到预先训练好的样本筛选模型中,以对第一样本数据的冗余数据进行确定并筛除,从而减少训练过程中的噪声和干扰,提高对第一样本数据进行分类的效率。之后,根据筛选后的第一样本数据对分类模型进行训练,使得分类模型在训练的过程中越加专注于重要和有用的数据,进一步提升分类模型的准确性和泛化能力,使得分类环境得到改善。
Smart Images

Figure CN117541788B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, system, device and medium for training a classification model based on reinforcement learning. Background Technology
[0002] Cropping technology refers to the process of cutting or cropping images or text to remove redundant parts and retain the areas of interest.
[0003] In related technologies, image or text cropping is generally performed in two ways: manual cropping and rule-based cropping. Specifically, manual cropping involves manually specifying the cropping ratio for each stage and then cropping the image or text according to that ratio. Rule-based cropping, on the other hand, uses a lightweight predictor to predict scores and then crops at a fixed ratio. However, both manual and rule-based cropping neglect model performance during the cropping process, leading to reduced efficiency and accuracy. Summary of the Invention
[0004] The main objective of this application is to propose a method, system, device, and medium for training a classification model based on reinforcement learning, which can improve the efficiency and accuracy of data pruning by improving the performance of the model.
[0005] To achieve the above objectives, a first aspect of this application proposes a method for training a classification model based on reinforcement learning. The method includes: acquiring a first sample dataset, wherein the first sample dataset includes multiple first sample data; inputting the multiple first sample data into a pre-trained sample selection model to determine and filter out redundant data in the multiple first sample data, obtaining multiple filtered first sample data; training a classification model based on the multiple filtered first sample data; wherein the sample selection model is trained based on a first loss, the first loss is determined based on a first predicted classification result and a mask label, the mask label comes from a second sample dataset, the second sample dataset includes multiple second sample data and the mask label corresponding to each second sample data, and the first predicted classification result is obtained by inputting the multiple second sample data filtered by the sample selection model into the classification model.
[0006] According to some embodiments of this application, the sample selection model is trained through the following steps: obtaining the second sample dataset, wherein the second sample dataset includes multiple second sample data and mask labels corresponding to each second sample data; inputting the multiple second sample data into the sample selection model respectively, identifying and filtering out redundant data in the multiple second sample data, and obtaining multiple filtered second sample data; inputting the multiple filtered second sample data into the classification model to obtain a first predicted classification result for each second sample data; using the first predicted classification result as feedback information in the reinforcement learning process of the sample selection model, determining a first loss based on the first predicted classification result and the mask labels, and training the sample selection model based on the first loss.
[0007] According to some embodiments of this application, determining and filtering redundant data in a plurality of second sample data to obtain a plurality of filtered second sample data includes: determining a corresponding first token vector and a plurality of second token vectors for each second sample data; wherein the first token vector is used to predict the category label; generating a comprehensive decision vector based on the first token vector and the plurality of second token vectors; determining redundant data in each second sample data according to the comprehensive decision vector, and pruning the redundant data to obtain the filtered second sample data.
[0008] According to some embodiments of this application, generating a comprehensive decision vector based on the first token vector and a plurality of second token vectors includes: multiplying the first token vector and the plurality of second token vectors accordingly to generate a first evaluation vector; performing a linear transformation on the second token vector and concatenating it with the first evaluation vector to obtain a comprehensive evaluation vector corresponding to the second sample data; and performing normalization processing on the comprehensive evaluation vector to obtain a comprehensive decision vector; wherein the comprehensive decision vector represents the probability value of the existence of the redundant data.
[0009] According to some embodiments of this application, the method further includes: acquiring an unfiltered first sample dataset; inputting the first sample dataset into the classification model to obtain a second predicted classification result for each of the first sample data, and calculating a second loss based on the second predicted classification result and the mask label; inputting the first sample data into the sample filtering model, and pruning any one of the second token vectors through the sample filtering model to obtain filtered third sample data; inputting the filtered third sample data into the classification model to obtain a third predicted classification result for the third sample data, and calculating a third loss based on the third predicted classification result and the mask label; comparing the second loss and the third loss to determine a reward / penalty mechanism for the sample filtering model; and training the sample filtering model based on the reward / penalty mechanism.
[0010] According to some embodiments of this application, training a classification model based on multiple filtered first sample data includes: merging the first loss and the second loss to obtain a total loss value; and training the classification model based on the total loss value.
[0011] According to some embodiments of this application, the method further includes: generating a redundant vector based on the redundant data, and generating a training vector with the same size and dimension as the redundant vector; appending the training vector to the redundant vector and performing single-layer calculation to generate a transfer vector; and transferring the transfer vector to the filtered first sample data.
[0012] To achieve the above objectives, a second aspect of this application proposes a classification model training system based on reinforcement learning. The system includes: a first sample dataset acquisition module for acquiring a first sample dataset, wherein the first sample dataset includes multiple first sample data; a first sample data acquisition module for inputting the multiple first sample data into a pre-trained sample selection model to determine and filter out redundant data from the multiple first sample data, obtaining multiple filtered first sample data; and a classification model training module for training a classification model based on the multiple filtered first sample data. The sample selection model is trained based on a first loss, which is determined based on a first predicted classification result and a mask label. The mask label comes from a second sample dataset, which includes multiple second sample data and the mask label corresponding to each second sample data. The first predicted classification result is obtained by inputting the multiple filtered second sample data from the sample selection model into the classification model.
[0013] To achieve the above objectives, a third aspect of the present application provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the reinforcement learning-based classification model training method described in any one of the embodiments of the first aspect of the present application.
[0014] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the reinforcement learning-based classification model training method described in any one of the embodiments of the first aspect of the present application.
[0015] The reinforcement learning-based classification model training method, system, device, and medium proposed in this application identify and remove redundant data from the first sample data by inputting it into a pre-trained sample selection model. This reduces noise and interference during training, improving the efficiency of classifying the first sample data. Subsequently, the classification model is trained using the selected first sample data, allowing it to focus more on important and useful data during training, further improving the accuracy and generalization ability of the classification model and ultimately improving the classification environment. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the structure of a classification model training system based on reinforcement learning provided in an embodiment of this application;
[0017] Figure 2 This is a flowchart of a classification model training method based on reinforcement learning provided in an embodiment of this application;
[0018] Figure 3 This is a flowchart illustrating how the sample screening model provided in this application masks redundant data.
[0019] Figure 4 This is a flowchart illustrating the training process of the sample selection model provided in this application embodiment;
[0020] Figure 5 This is a flowchart of obtaining second sample data provided in an embodiment of this application;
[0021] Figure 6 This is a flowchart of the process for generating a comprehensive decision vector provided in an embodiment of this application;
[0022] Figure 7 This is a flowchart of an embodiment of the present application showing how an agent performs occlusion and cropping operations;
[0023] Figure 8 This is a flowchart of step S302;
[0024] Figure 9 This is another flowchart of the reinforcement learning-based classification model training method provided in the embodiments of this application;
[0025] Figure 10 This is a flowchart illustrating the training process using a three-channel mechanism provided in an embodiment of this application;
[0026] Figure 11 yes Figure 2 Flowchart of step S103;
[0027] Figure 12 This is another flowchart of the reinforcement learning-based classification model training method provided in the embodiments of this application;
[0028] Figure 13 This is a schematic diagram of the functional modules of the reinforcement learning-based classification model training system provided in the embodiments of this application;
[0029] Figure 14 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0031] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0033] In recent years, pre-trained models have gradually become mainstream in many application areas, such as using pre-trained language models (Bidirectional Encoder Representations from Transformers, BERT) for Natural Language Processing (NLP) tasks and using vision processing models (Vision Transformers, ViT) for Computer Vision (CV) tasks. However, with the increasing length of input text for pre-trained language models and the demands of computer vision tasks for processing high-resolution images, traditional pre-trained models face challenges in terms of computational accuracy and memory requirements when processing longer sequences.
[0034] In related technologies, token pruning is generally performed using the following methods: First, by introducing class attention to estimate token scores and employing a slow-fast token evolution method to improve the throughput of the visual processing model. This method requires manually specifying the pruning ratio for each stage, resulting in low pruning efficiency. Second, by determining the importance of tokens through additional metrics. For example, after pre-determining the important positions of text or images, a fixed pruning ratio is set for tokens. However, in long texts or images, key information may be distributed in different locations. The fixed-ratio pruning method cannot capture the accurate location of key information, leading to the pruning result potentially leaving redundant information while losing important information. Therefore, related technologies that require operator feedback to optimize the model also place too much emphasis on optimizing the performance of the model used for pruning.
[0035] However, in reality, the environment plays a crucial role in reinforcement learning during the image or text cropping process. Therefore, the classification model can be viewed as an environment that needs optimization, and the overall system quality can be improved by continuously refining the model.
[0036] Based on this, embodiments of this application provide a method, system, device, and medium for training a classification model based on reinforcement learning, which can improve the efficiency and accuracy of image cropping by optimizing the environment of the classification model.
[0037] The reinforcement learning-based classification model training method, system, device, and medium provided in this application are specifically described through the following embodiments. First, the reinforcement learning-based classification model training system in this application is described.
[0038] Please refer to Figure 1 In some embodiments, the reinforcement learning-based classification model training system includes a server 101, a controller 102, a sample selection model 103, a classification model 104, and a terminal 105.
[0039] In some embodiments, server 101 is responsible for managing user requests issued by terminal 105, receiving and processing data input, and returning the results to terminal 105. Server 101 can use existing web server architectures, such as web servers or cloud servers.
[0040] For example, controller 102 can be the nerve center and command center of the system. Controller 102 can generate operation control signals based on instruction opcodes and timing signals to control the fetching and execution of instructions. For instance, controller 102 can execute training instructions for sample screening model 103 and classification model 104, as well as classification instructions for the data to be classified, and continuously optimize classification model 104 to provide a good environment for data classification, thereby improving the efficiency and accuracy of image or text classification.
[0041] In some embodiments, the system can train the sample screening model 103 to improve the ability of the sample screening model 103 to recognize redundant data of images or texts, and screen the identified redundant data. The screened image or text data is then transmitted to the classification model 104 for classification, so as to continuously optimize the classification model 104 through reinforcement learning.
[0042] It should be noted that the classification model 104 can be used to classify the input data. The classification model 104 can be built based on a reinforcement learning algorithm, and the accuracy of token pruning and the classification effect of the data can be continuously optimized through repeated iterative training. In some embodiments, the terminal 105 is responsible for transmitting data to the server 101 for processing. The terminal 105 can be a computer, mobile phone, sensor or other device, and can be used to send data classification requests, view data classification results, etc.
[0043] It is understandable that the server 101, controller 102, sample screening model 103, and classification model 104 work together to classify and process images or text, such as image data and text data, and interact with the terminal 105.
[0044] The reinforcement learning-based classification model training method in this application can be illustrated through the following embodiments.
[0045] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user will be obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent will the necessary user-related data for the normal operation of the embodiments of this application be obtained.
[0046] Figure 2 This is an optional flowchart of a reinforcement learning-based classification model training method provided in the embodiments of this application. Figure 2 The method may include, but is not limited to, steps S101 to S103.
[0047] Step S101: Obtain the first sample dataset, wherein the first sample dataset includes multiple first sample data.
[0048] In some embodiments, the first sample data is uncropped image data or text data; that is, the first sample data includes both redundant and non-redundant data. It is understood that different first sample datasets need to be selected based on different classification tasks. For example, for natural language processing tasks, the first sample dataset consists of text data; for computer vision processing tasks, the first sample dataset consists of image data.
[0049] In some embodiments, after obtaining the first sample dataset, the first sample data needs to be preprocessed before being input into the pre-trained sample selection model. Specifically, the preprocessing process for image or text data is as follows: For natural language processing tasks, text data needs to be loaded from the first sample dataset, then the text data is divided into multiple words or sub-words, and then the mask labels are converted into numerical forms. For computer vision processing tasks, the image size of the loaded image needs to be adjusted, and the pixel values are mapped to the range between 0 and 1 to adjust the image data to the same scale. Then, the image data format is converted according to different models, such as converting it to a three-color image or a grayscale image, etc.
[0050] Understandably, the first sample dataset can be randomly divided, for example, 20% of the first sample data can be used as the training set, and 80% of the first sample data can be used as the test set, etc. The division ratio can be flexibly adjusted as needed.
[0051] For example, the first sample data set may come from public data sets for natural language processing, which can be obtained through public channels, such as the GLUE data set, and computer vision data sets, such as the Cifar-100 data set, the FC100 data set and the Mini-imagenet data set, etc., and may also come from other data sets, such as a data set collected by technicians themselves, etc., which is not specifically limited in the embodiments of the present application.
[0052] Step S102: inputting a plurality of first sample data into a pre-trained sample screening model respectively, determining and screening out redundant data in the plurality of first sample data, to obtain a plurality of first sample data after screening.
[0053] In some embodiments, the sample screening model is configured to identify and screen out redundant data in image data or text data, and transmit the data after screening to a classification model, so as to train the classification model. It can be understood that after the redundant data is identified by the sample screening model, the redundant data can be cropped, that is, parts that are unimportant and not beneficial to data classification are cropped, so as to improve cropping efficiency and cropping accuracy.
[0054] It should be noted that the sample screening model may be an agent, a model based on visual features, an N-gram model, a TF-IDF model, a model based on deep learning, and the like. Specifically, an agent refers to an entity with perception, decision-making and action capabilities, which may be a neural network model or an automatic program, etc.
[0055] Please refer to Figure 3 , Figure 3 , which shows that when the sample screening model processes image data, it identifies and screens out redundant data. In Figure 3 , oblique lines are used to represent masked parts or parts to be cropped (i.e., redundant data), and the masked parts from Figure 3 's subgraph a to subgraph d increase sequentially. It can be understood that the more and more accurate the masked parts are, the better the performance of the sample screening model is, which is more beneficial to the optimization of the subsequent classification model. In some embodiments, the sample screening model can mask redundant parts (such as the oblique line part in Figure 3 ), so that the subsequent classification model can directly ignore the masked parts and improve classification efficiency. In some embodiments, for text data, the text data can be directly cropped, for example, removing "的" from "今天的天气真好", or masking redundant parts, etc., which is not specifically limited in the embodiments of the present application.
[0056] Step S103: Train a classification model based on multiple first sample data after screening; wherein, the sample screening model is trained based on the first loss, the first loss is determined based on the first predicted classification result and the mask label, the mask label comes from the second sample dataset, the second sample dataset includes multiple second sample data and the mask label corresponding to each second sample data, and the first predicted classification result is obtained by inputting multiple second sample data after screening by the sample screening model into the classification model.
[0057] It should be noted that during the optimization of the classification model, the sample selection model can be trained to continuously improve its ability to identify redundant data. After the sample selection model trims the redundant data, the trimmed data is then input into the pre-trained classification model, enabling the classification model to continuously adapt to the corresponding recognition environment and improve the overall quality of the system.
[0058] Specifically, the second sample dataset is used to train the sample selection model. The second sample data in the second sample dataset is unselected; that is, it includes redundant data and data beneficial for classification. During the training of the sample selection model, the model filters the second sample data, that is, it trims the redundant data. The trimmed second sample data is then input into the classification model, which classifies the second sample data to obtain a first predicted classification result. Further, based on a pre-set loss function and using the mask label as a reference, a first loss is calculated on the first predicted classification result. The effectiveness of the sample selection model in trimming redundant data can then be determined based on the first loss. In some embodiments, the parameters of the sample selection model can be continuously adjusted based on the newly calculated first loss until a preset number of training iterations or the convergence condition of the sample selection model is reached, at which point training of the sample selection model can be stopped.
[0059] In some embodiments, the first sample data after filtering is data that has been accurately identified and pruned from redundant data. By inputting multiple first sample data after filtering into the classification model, the classification model can focus more on key data during the continuous training process, that is, more on data that is conducive to rapid classification.
[0060] In some embodiments, since natural language processing models typically have 12 layers, corresponding to the third, sixth, and ninth layers respectively from the perspectives of phrase, linguistics, and semantics, if the classification model is a natural language processing model, then the sample selection model can prune or mask redundant data at the third, sixth, and ninth layers of the classification model. It is understandable that for different classification models, the categories of the classification model can be analyzed, and based on the layer distribution of the classification model and the function of each layer, a corresponding sample selection model can be set to filter redundant data, thereby improving the efficiency and accuracy of the classification model in classifying data.
[0061] The reinforcement learning-based classification model training method, system, device, and medium proposed in this application identify and remove redundant data from the first sample data by inputting it into a pre-trained sample selection model. This reduces noise and interference during training, improving the efficiency of classifying the first sample data. Subsequently, the classification model is trained using the selected first sample data, allowing it to focus more on important and useful data during training, further improving the accuracy and generalization ability of the classification model and ultimately improving the classification environment.
[0062] Please refer to Figure 4 In some embodiments, the sample selection model can be trained through the following steps S201 to S204:
[0063] Step S201: Obtain the second sample dataset, wherein the second sample dataset includes multiple second sample data and the mask label corresponding to each second sample data.
[0064] It is understandable that the second sample data is used to train the sample selection model. The unselected second sample data includes redundant data. By inputting the second sample data into the sample selection model, the redundant data can be identified and pruned to obtain the second sample data with non-critical data removed.
[0065] In some embodiments, the mask label refers to the classification label or grouping label corresponding to the second sample data. The mask label can be created manually or automatically generated according to an algorithm. For example, the corresponding label can be automatically generated by machine learning algorithms, such as clustering algorithms, decision tree algorithms, etc. For example, the mask label can be an emotion classification label, an item classification label, etc., and this application embodiment does not impose specific limitations on this.
[0066] Step S202: Input multiple second sample data into the sample screening model respectively, identify and screen out redundant data in the multiple second sample data, and obtain multiple screened second sample data.
[0067] Specifically, redundant data refers to information that is excessive or repetitive in text or images. For example, for text data, redundant data could be words that are repeatedly used in sentences and paragraphs, such as "He felt very sad today," in which case "felt" could be considered redundant data; or, words that have no effect on classification could be considered redundant data, such as "of," etc. This application does not impose specific limitations on this. For images, redundant data could include the same objects or scenes that appear repeatedly in the image, unnecessary backgrounds, and noisy data, and so on. It is understandable that identifying and removing redundant data can improve the efficiency of classifying sample data and enable sample selection and classification models to better identify and select key data.
[0068] In some embodiments, the sample selection model can select data that meets the requirements from the given data according to specific conditions and rules. It is understood that any model that can identify and filter redundant data in the second sample data can be used as a sample selection model, and the embodiments of this application do not impose specific limitations on this.
[0069] Step S203: Input the filtered second sample data into the classification model to obtain the first predicted classification result of each second sample data.
[0070] In some embodiments, the second sample data after filtering out redundant data can be input into the classification model for classification to obtain the first predicted classification result of the second sample data. It is understood that if the sample filtering model can correctly identify and trim redundant data, then the classification model can correctly classify the second sample data based on the remaining untrimmed second sample data, and the first predicted classification result can correspond to the mask label. For example, if the first predicted classification result predicts the object as a cat, then the mask label of the second sample data is also a cat. In this case, the sample filtering model can be rewarded to continuously optimize it. If the sample filtering model cannot correctly identify and trim redundant data, for example, by mistakenly trimming important information, then the classification model may not be able to correctly identify the second sample data. For example, if the first predicted classification result predicts the object as a cat, but the mask label of the second sample data is a dog, this indicates that the sample filtering model has made a trimming error. In this case, the parameters of the sample filtering model can be adjusted. In some embodiments, if the sample filtering model makes a trimming error, it can be penalized to reduce the occurrence of such errors in subsequent training, thereby improving the efficiency of training the sample filtering model.
[0071] Step S204: Use the first predicted classification result as feedback information in the reinforcement learning process of the sample selection model, determine the first loss based on the first predicted classification result and the mask label, and train the sample selection model based on the first loss.
[0072] In some embodiments, the first predicted classification result can be used as feedback information in the reinforcement learning process of the sample selection model. A first loss is calculated based on the first predicted classification result and the mask label according to a pre-set loss function, and this first loss is passed as a feedback signal to the sample selection model. This allows the sample selection model to adjust its parameters based on the first loss, thereby continuously improving the accuracy of the sample selection model in identifying redundant data. It is understood that the loss function used to calculate the first loss can be a cross-entropy loss function, a logarithmic loss function, a squared loss function, an absolute value loss function, etc.
[0073] For example, the first loss is calculated using the cross-entropy loss function. The expression for the cross-entropy loss function is as follows:
[0074]
[0075] Among them, L ce Let y represent the first loss, p represent the first predicted classification result, and i represent different data. For example, when i=1, y1 represents the mask label with sequence number 1, when i=2, y2 represents the mask label with sequence number 2, and so on. N is the number of corresponding task types. For example, if the corresponding emotion types are positive, neutral, and negative, then the value of N is 3.
[0076] In some embodiments, if the second sample data is image data, the Dice loss function can also be used to calculate the first loss.
[0077] Please refer to Figure 5 In some embodiments, redundant data in a plurality of second sample data are identified and filtered out to obtain a plurality of filtered second sample data, including steps S301 to S303:
[0078] Step S301: Determine the corresponding first token vector and multiple second token vectors based on each second sample data; wherein, the first token vector is used to predict the category label.
[0079] In some embodiments, a first token vector and multiple second token vectors can be determined from second sample data. Specifically, the first token vector can be a classification token vector (CLS vector), used to represent the task that the sample selection model needs to perform, such as determining whether the sentiment is positive given a passage, or determining whether a picture contains a cat, etc. Therefore, the first token vector can be used to predict the category label.
[0080] Understandably, the second token vector refers to the information included in the second sample data. For example, for the text data "He is happy today", the second token vector can be "He is happy today", or it can be "Today", "He", "Happy". The specific generation of the second token vector can be based on preset rules. The rules can be formulated by technical personnel or automatically identified by the sample screening model after training.
[0081] Step S302: Generate a comprehensive decision vector based on the first token vector and multiple second token vectors.
[0082] In some embodiments, the first token vector and each second token vector can be processed through different fully connected layers, and then element-wise multiplied to generate a vector for evaluating the importance of the second token vector to the first token vector, thereby generating a corresponding first evaluation vector according to the classification task. Further, each second token vector can be processed through a linear layer and then concatenated with the first evaluation vector to form a comprehensive evaluation vector. In some embodiments, the first token vector can also be processed through a linear layer and then concatenated with the first evaluation vector to form a comprehensive evaluation vector.
[0083] In some embodiments, the comprehensive evaluation vector can be passed through a fully connected layer and a softmax function to generate a comprehensive decision vector, which indicates whether the corresponding token needs to be retained. It is understood that there is at least one comprehensive decision vector, and each comprehensive decision vector corresponds to the probability that a second token vector is redundant data. Alternatively, there may be only one comprehensive decision vector, summing the probabilities of all token vectors being redundant data; this embodiment does not impose specific limitations on this.
[0084] Step S303: Determine the redundant data in each second sample data according to the comprehensive decision vector, and trim the redundant data to obtain the filtered second sample data.
[0085] In some embodiments, the comprehensive decision vector can represent the probability that each piece of data in image data or text data is redundant or not redundant. For example, if the task for an image is to identify whether there is a cat in the image, then the clouds, grass, etc. around the cat are all redundant data. In this case, the probability that the clouds and grass are redundant data can be marked as 90%, or the probability that the clouds and grass are valuable data can be marked as 10%, etc. The embodiments of this application do not impose specific limitations on this.
[0086] In some embodiments, the sample selection model can decide how to handle redundant data based on instructions or pre-set parameters. For example, it can directly crop the redundant data to reduce memory usage, or mask the redundant data to avoid interfering with the classification model's recognition of the target image or target text.
[0087] Please refer to Figure 6 and Figure 7 The process of generating decision vectors using a sample selection model is described in general. In some embodiments, the first token vector and the second token vector are each processed through a fully connected layer and then multiplied element-wise to generate a vector used to evaluate the importance of the first token vector to the second token vector, i.e., the first evaluation vector. Then, the first evaluation vector is processed through a linear layer and concatenated with the second token vector to generate a comprehensive evaluation vector. Further, the comprehensive evaluation vector is input into a fully connected layer and a softmax layer to obtain the comprehensive decision vector.
[0088] Understandably, Softmax is an activation function that normalizes a numerical vector into a probability distribution vector, where the sum of all probabilities is 1. The composite vector indicates which token vectors are redundant and which are important, thus instructing the sample selection model to make decisions about whether to mask or prune the corresponding redundant data.
[0089] Please refer to Figure 8 In some embodiments, step S302 may include steps S401 to S403:
[0090] Step S401: Multiply the first token vector and multiple second token vectors accordingly to generate the first evaluation vector.
[0091] Understandably, by multiplying the first token vector by multiple second token vectors, the resulting first evaluation vector can be used to measure the correlation or weight between the information represented by the first token vector and the second token vectors. Understandably, by multiplying the first token vector by each second token vector, each second token vector carries task information, thus ensuring that no matter how the second token vectors are subsequently pruned, it will not affect the classification model's classification of the second sample data.
[0092] For example, suppose the first token vector represents the task of identifying sadness as the emotion, and the second token vector represents the emotional statement. Then, by multiplying the first token vector and the second token vector respectively, we can obtain the tendency of the second token vector to be sad.
[0093] Step S402: After linearly transforming the second token vector, concatenate it with the first evaluation vector to obtain the comprehensive evaluation vector corresponding to the second sample data.
[0094] Understandably, each second token vector can be linearly transformed and then concatenated with the first evaluation vector to obtain a comprehensive evaluation vector. This comprehensive evaluation vector can also take into account the relationship between the first and second token vectors, as well as the features of the second sample data.
[0095] Step S403: Normalize the comprehensive evaluation vector to obtain the comprehensive decision vector; wherein, the comprehensive decision vector represents the probability value of the existence of redundant data.
[0096] Understandably, normalizing the comprehensive evaluation vector can transform it into a series of probability values, representing the probability of redundant data existing. This can be used to determine whether redundant data exists and the likelihood of its existence.
[0097] Please refer to Figure 9 In some embodiments, the method may further include steps S501 to S506:
[0098] Step S501: Obtain the first sample dataset that has not been filtered.
[0099] In some embodiments, any sample dataset that has not undergone sample screening is acceptable, and is not limited to the first sample dataset. For example, unscreened image datasets and text datasets, etc., can be selected.
[0100] Step S502: Input the first sample dataset into the classification model to obtain the second predicted classification result of each first sample data, and calculate the second loss based on the second predicted classification result and the mask label.
[0101] In some embodiments, the first sample data is directly input into the classification model for classification without undergoing the sample screening model to identify and trim redundant data, thereby obtaining the second predicted classification result for each first sample data, and calculating the second loss based on the second predicted classification result and the mask label.
[0102] Step S503: Input the first sample data into the sample screening model, and use the sample screening model to prune any second token vector to obtain the screened third sample data.
[0103] In some embodiments, any second token vector can be pruned using a sample filtering model, i.e., random single-time pruning, to obtain filtered third sample data.
[0104] Step S504: Input the filtered third sample data into the classification model to obtain the third predicted classification result of the third sample data, and calculate the third loss based on the third predicted classification result and the mask label.
[0105] In some embodiments, the classification model can accurately classify the first sample data in the first sample dataset, obtain a third predicted classification result, and calculate a third loss based on the third predicted classification result and the mask label, thereby determining the recognition accuracy of the sample screening model based on the third loss.
[0106] Step S505: Based on the comparison between the second loss and the third loss, determine the reward and punishment mechanism for the sample selection model.
[0107] Step S506: Train the sample selection model based on the reward and punishment mechanism.
[0108] Understandably, a three-channel mechanism can be used to train the sample selection model and the classification model. This can generate rewards or penalties based on the pruning accuracy of the sample selection model, thereby further guiding the decision-making of the sample selection model. At the same time, it can train the reasoning ability of the classification model when the pruned data is incomplete, so as to optimize the performance of the classification model.
[0109] Specifically, the three-channel mechanism introduces a contrastive stream and a single-pruning stream to evaluate the effectiveness of the sample selection model's pruning action on the sample data. The contrastive stream represents the standard pattern, while the single-pruning stream restricts pruning to only one token vector at a time. By comparing the second loss obtained from the contrastive stream and the third loss obtained from the single-pruning stream, the difference in losses is obtained, and the effectiveness of the sample selection model's pruning behavior is determined by this difference. The sample selection model is then trained using the cross-entropy loss function.
[0110] It is understandable that the loss function of the three-channel mechanism can be composed of the cross-entropy loss function, or other loss functions. The cross-entropy loss function is used to calculate during the training of the sample selection model and the classification model, and is used for backpropagation based on the value of the cross-entropy loss function to update and optimize the parameters of the sample selection model and the classification model.
[0111] In some embodiments, the expression for the cross-entropy loss function is as follows:
[0112]
[0113] Among them, L ceLet y represent the first loss, p represent the first predicted classification result, and i represent different data. For example, when i=1, y1 represents the mask label with sequence number 1, when i=2, y2 represents the mask label with sequence number 2, and so on. N is the number of corresponding task types. For example, if the corresponding emotion types are positive, neutral, and negative, then the value of N is 3.
[0114] Please refer to Figure 10 For example, an embodiment is provided to illustrate the process of training a sample selection model using a three-channel mechanism.
[0115] In some embodiments, the process of optimizing the classification model and calculating the second loss by contrastive stream is as follows: First, the original input feature T_in is input into the classification model, and T_in is encoded into T_0 through a Transformer layer without pruning T_0. T_0 is then processed through a series of Transformer layers to obtain the classification result. The second loss La is calculated based on the classification result. La reflects the performance of the classification model without pruning.
[0116] It is understandable that a Transformer layer is a neural network layer with a self-attention mechanism. Multiple Transformer layers make up a Transformer model. The Transformer model is an encoder-decoder model based on an attention mechanism, used for sequence-to-sequence tasks, such as machine translation and text summarization.
[0117] In some embodiments, the process of optimizing the classification model and calculating the second loss through a single pruning flow is as follows: The original input feature T_in is input into the sample selection model and encoded into T_0 through a Transformer layer. Then, a token vector is randomly pruned through the sample selection model to obtain S_1. S_1 is then input into the classification model corresponding to the comparison flow for processing to obtain the third predicted classification result. The third loss Lb is calculated through the third classification prediction result. By comparing La and Lb, it is determined whether the sample selection model has accurately pruned redundant data.
[0118] Specifically, if Lb≤La, meaning the third loss is less than or equal to the second loss, it indicates that the redundant data pruned in a single pruning process is superfluous and uninformative. This will serve as a reward signal for the sample selection model, continuously optimizing its performance. If Lb>La, meaning the third loss is greater than the second loss, it indicates that the redundant data pruned in a single pruning process is critical and informative. This will serve as a penalty signal for the sample selection model, continuously optimizing its performance. In the process, this also improves the generalization ability of the classification model and its ability to identify incomplete information.
[0119] Please refer to Figure 11 In some embodiments, step S103 may include steps S601 to S602:
[0120] Step S601: Combine the first loss and the second loss to obtain the total loss value.
[0121] Step S602: Train the classification model based on the total loss value.
[0122] In some embodiments, the three-channel mechanism introduces a contrastive stream and a model pruning stream to evaluate the effectiveness of the classification model's actions in pruning sample data. Specifically, the second loss calculated through the contrastive stream and the first loss calculated through the model pruning stream can be added to obtain a total loss value, forming a supervisory signal for the classification model, thereby continuously improving the model's performance. An exemplary embodiment illustrates the process of optimizing a classification model using the three-channel mechanism.
[0123] Please refer to Figure 10 For example, an embodiment is provided to illustrate the process of training a sample selection model using a three-channel mechanism.
[0124] In some embodiments, the process of optimizing the classification model and calculating the second loss by contrastive stream is as follows: First, the original input feature T_in is input into the classification model, and T_in is encoded into T_0 through a Transformer layer without pruning T_0. T_0 is then processed through a series of Transformer layers to obtain the classification result. The second loss La is calculated based on the classification result. La reflects the performance of the classification model without pruning.
[0125] In some embodiments, the process of optimizing the classification model and calculating the first loss (representing the loss data obtained by the sample selection model at a specific stage) through model pruning flow is as follows: First, the original input features T_in are input into the classification model, and T_in is encoded as T_0 through a Transformer layer. After the sample selection model independently determines multiple redundant data and pruning positions, the identified redundant data are pruned and then input into the classification model. After classification by the classification model, the first loss Lc is output. It should be noted that, in order to improve the robustness and generalization performance of the model, La and Lc are merged during model training to obtain the total loss. The total loss is used as the supervision signal for gradient update to train the classification model, so as to continuously improve the robustness and generalization performance of the classification model.
[0126] Understandably, after both the sample selection model and the classification model have been trained, a test dataset can be used to evaluate the classification performance of the classification model. Specifically, this can be evaluated by calculating an accuracy metric. In some embodiments, a baseline model can be introduced, and the performance of the extreme model and the classification model after introducing model pruning can be compared, along with visual analysis of the results, to assess whether the sample selection model helps improve the efficiency of the classification model and thus enhance its performance.
[0127] In some embodiments, the test data of the reinforcement learning classification model in natural language, computer vision, etc., and the optimization data of the classification model itself in terms of inference computation, memory usage, and acceleration are represented by the following test results. The classification accuracy of the classification model is represented by the accuracy of the classification results. The computational complexity is verified by analyzing the changes in floating-point operations to confirm the improvement in classification model speed and the reduction in computational resource consumption. Simultaneously, all image data in the test set were cropped using model inference and sample selection models, and the effectiveness of the sample selection model in detecting and removing noise was visualized and verified.
[0128] The following are the test results in natural language processing (the datasets used in this application are all commonly used datasets that are publicly available, and F1 is a comprehensive metric calculated based on precision and recall):
[0129]
[0130] As can be seen from the above test results in natural language processing, the accuracy of this application is improved to a certain extent after reinforcement learning, compared with the accuracy obtained by processing the dataset using the original method. In other words, the classification model of this application performs better after reinforcement learning.
[0131] The following are the test results of different classification models in computer vision before and after reinforcement learning, with accuracy as the metric:
[0132]
[0133] As can be seen from the above data, the classification model after reinforcement learning in this application has a higher accuracy in computer vision than the classification model without reinforcement learning.
[0134] The following are the optimizations in terms of inference computation, memory usage, and speedup after applying the reinforcement learning method of this application to the classification model:
[0135]
[0136] Experiments and statistics show that the classification model obtained by reinforcement learning using the method in this application has significantly improved the accuracy of classification of datasets, and has remarkable effects in terms of inference acceleration and reduction of inference computation, and has strong generalization ability and ease of use.
[0137] Please refer to Figure 12 In some embodiments, the method further includes steps S701 to S703:
[0138] Step S701: Generate a redundant vector based on the redundant data, and generate a training vector with the same size and dimension as the redundant vector.
[0139] In some embodiments, for the token vector to be pruned, a trainable vector of the same size and dimension as the pruned token vector can be appended. Specifically, the trainable vector is a pre-trained training vector that can adaptively learn important features contained in redundant information. The training vector can be used to retain some information related to non-redundant data to overcome the information loss problem caused by token pruning.
[0140] Step S702: After attaching the training vector to the redundant vector, perform single-layer calculation to generate the transfer vector.
[0141] Understandably, training vectors of the same dimension are appended to redundant vectors and used as input data for single-layer computation. Specifically, the input data is processed through a single neural network layer. This neural network layer can be a linear layer, convolutional layer, recurrent layer, etc., chosen according to the specific task requirements. Single-layer computation allows for further feature extraction and information compression of the input data, reducing the number of parameters and computational cost of the sample selection model, thereby improving its information extraction capabilities.
[0142] Step S703: Pass the transfer vector to the filtered first sample data.
[0143] In some embodiments, after single-layer computation, valuable information from the pruned redundant data can be transferred to the first sample data, or the valuable information can be transferred to the next processing layer. For example, if the processing was originally performed at the third layer of the classification model, the valuable information can now be transferred to the next processing layer, such as the ninth layer, so that the valuable information is not pruned along with the redundant data, and the performance and effectiveness of the sample selection model and the classification model are improved.
[0144] Please see Figure 13 This application also provides a reinforcement learning-based classification model training system, which can implement the above-mentioned reinforcement learning-based classification model training method. The reinforcement learning-based classification model training system includes:
[0145] The first sample dataset acquisition module 1301 is used to acquire the first sample dataset, wherein the first sample dataset includes multiple first sample data;
[0146] The first sample data acquisition module 1302 is used to input multiple first sample data into a pre-trained sample screening model, identify and screen out redundant data in the multiple first sample data, and obtain multiple first sample data after screening.
[0147] The classification model training module 1303 is used to train a classification model based on multiple filtered first sample data. The sample filtering model is trained based on a first loss, which is determined based on a first predicted classification result and a mask label. The mask label comes from a second sample dataset, which includes multiple second sample data and the mask label corresponding to each second sample data. The first predicted classification result is obtained by inputting multiple filtered second sample data into the classification model.
[0148] The specific implementation of this reinforcement learning-based classification model training system is basically the same as the specific embodiments of the reinforcement learning-based classification model training method described above, and will not be repeated here. Subject to meeting the requirements of the embodiments of this application, the reinforcement learning-based classification model training system may also be equipped with other functional modules to implement the reinforcement learning-based classification model training method in the above embodiments.
[0149] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described reinforcement learning-based classification model training method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0150] Please see Figure 14 , Figure 14 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0151] The processor 1401 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0152] The memory 1402 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1402 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1402 and is called and executed by the processor 1401 to execute the reinforcement learning-based classification model training method of the embodiments of this application.
[0153] The input / output interface 1403 is used to implement information input and output;
[0154] The communication interface 1404 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0155] Bus 1405 transmits information between various components of the device (e.g., processor 1401, memory 1402, input / output interface 1403, and communication interface 1404);
[0156] The processor 1401, memory 1402, input / output interface 1403 and communication interface 1404 are connected to each other within the device via bus 1405.
[0157] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described reinforcement learning-based classification model training method.
[0158] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0159] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0160] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0161] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0162] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0163] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0164] It should be understood that in this application, "at least one" and "several" refer to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0165] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0166] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0167] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0168] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0169] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for training a classification model based on reinforcement learning, characterized in that, include: Obtain a first sample dataset, wherein the first sample dataset includes multiple first sample data, and the first sample data is uncropped image data; Multiple sets of first sample data are input into a pre-trained sample selection model to identify and filter out redundant data, resulting in multiple sets of filtered first sample data. The sample selection model is trained through the following steps: obtaining a second sample dataset, which includes multiple sets of second sample data and corresponding mask labels; inputting the multiple sets of second sample data into the sample selection model to identify and filter out redundant data, resulting in multiple sets of filtered second sample data; inputting the multiple sets of filtered second sample data into an image classification model to obtain a first predicted classification result for each set of second sample data; using the first predicted classification result as feedback information in the reinforcement learning process of the sample selection model, determining a first loss based on the first predicted classification result and the mask labels, and training the sample selection model based on the first loss. The step of determining and filtering redundant data from multiple second sample data to obtain multiple filtered second sample data includes: determining a corresponding first token vector and multiple second token vectors for each second sample data; wherein the first token vector is used to predict the category label; generating a comprehensive decision vector based on the first token vector and multiple second token vectors; determining redundant data in each second sample data according to the comprehensive decision vector, and pruning the redundant data to obtain the filtered second sample data; The image classification model is trained based on multiple filtered first sample data. The sample selection model is trained based on a first loss, which is determined based on a first prediction classification result and a mask label. The mask label comes from a second sample dataset, which includes multiple second sample data and the mask label corresponding to each second sample data. The first prediction classification result is obtained by inputting the multiple second sample data filtered by the sample selection model into the image classification model.
2. The method for training a classification model based on reinforcement learning according to claim 1, characterized in that, The step of generating a comprehensive decision vector based on the first token vector and multiple second token vectors includes: A first evaluation vector is generated by multiplying the first token vector and multiple second token vectors accordingly. After performing a linear transformation on the second token vector, it is concatenated with the first evaluation vector to obtain the comprehensive evaluation vector corresponding to the second sample data. The comprehensive evaluation vector is normalized to obtain the comprehensive decision vector; wherein the comprehensive decision vector represents the probability value of the existence of the redundant data.
3. The method for training a classification model based on reinforcement learning according to claim 1, characterized in that, The method further includes: Obtain the first sample dataset that has not been filtered; The first sample dataset is input into the image classification model to obtain the second predicted classification result for each of the first sample data, and the second loss is calculated based on the second predicted classification result and the mask label. The first sample data is input into the sample filtering model, and the sample filtering model is used to prune any second token vector to obtain the filtered third sample data. The filtered third sample data is input into the image classification model to obtain the third predicted classification result of the third sample data, and the third loss is calculated based on the third predicted classification result and the mask label. The reward and punishment mechanism for the sample screening model is determined by comparing the second loss and the third loss. The sample selection model is trained based on the aforementioned reward and punishment mechanism.
4. The method for training a classification model based on reinforcement learning according to claim 3, characterized in that, The training of the image classification model based on multiple filtered first sample data includes: The first loss and the second loss are combined to obtain the total loss value; The image classification model is trained based on the total loss value.
5. The method for training a classification model based on reinforcement learning according to claim 1, characterized in that, The method further includes: Based on the redundant data, a redundant vector is generated, and a training vector with the same size and dimension as the redundant vector is generated. After appending the training vector to the redundant vector, a single-layer calculation is performed to generate the transfer vector; The transfer vector is then passed to the filtered first sample data.
6. A classification model training system based on reinforcement learning, characterized in that, The system includes: The first sample dataset acquisition module is used to acquire the first sample dataset, wherein the first sample dataset includes multiple first sample data, and the first sample data is uncropped image data; The first sample data acquisition module is used to input multiple sets of first sample data into a pre-trained sample filtering model to identify and filter out redundant data from the multiple sets of first sample data, thereby obtaining multiple sets of filtered first sample data. The sample filtering model is trained through the following steps: acquiring a second sample dataset, wherein the second sample dataset includes multiple sets of second sample data and mask labels corresponding to each set of second sample data; inputting the multiple sets of second sample data into the sample filtering model to identify and filter out redundant data from the multiple sets of second sample data, thereby obtaining multiple sets of filtered second sample data; and inputting the multiple sets of filtered second sample data into an image classification model to obtain a first predicted classification result for each set of second sample data. The first predicted classification result is used as feedback information in the reinforcement learning process of the sample selection model. A first loss is determined based on the first predicted classification result and the mask label, and the sample selection model is trained based on the first loss. The step of determining and filtering redundant data from multiple second sample data to obtain multiple filtered second sample data includes: determining a corresponding first token vector and multiple second token vectors for each second sample data; wherein the first token vector is used to predict the category label; generating a comprehensive decision vector based on the first token vector and multiple second token vectors; determining redundant data in each second sample data according to the comprehensive decision vector, and pruning the redundant data to obtain the filtered second sample data. The classification model training module is used to train the image classification model based on multiple filtered first sample data; wherein, the sample filtering model is trained based on a first loss, the first loss is determined based on a first predicted classification result and a mask label, the mask label comes from a second sample dataset, the second sample dataset includes multiple second sample data and the mask label corresponding to each second sample data, and the first predicted classification result is obtained by inputting multiple filtered second sample data into the image classification model.
7. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the reinforcement learning-based classification model training method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the reinforcement learning-based classification model training method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Image classification model training method, image classification method, equipment and medium
CN115496955A
Systems and method for automatically configuring machine learning models
US20190272479A1