Training sample screening method and device and computer readable storage medium
By dividing data clusters and calculating contributions on multiple sets of prediction results of training samples, the problem of insufficient evaluation and screening of training samples in the prior art is solved, and the training effect of artificial intelligence models is improved.
Patent Information
- Application Number
- CN202311657649.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology lacks effective quality evaluation and screening methods in the preprocessing of training samples, resulting in poor training of artificial intelligence models.
By obtaining the multiple sets of prediction results of the training samples under multiple sets of parameters of the artificial intelligence model, dividing them into multiple data clusters, and calculating the contribution of each data cluster to model training, the target training sample is selected.
The performance of the trained artificial intelligence model is improved, and efficient screening of training samples is achieved through more refined data cluster division and contribution calculation.
Smart Images

Figure CN120105083A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and in particular to a method and device for screening training samples, and a computer-readable storage medium. Background Art
[0002] With the comprehensive deepening of the digitalization of the real economy, data-driven intelligent decision-making systems are widely used in the production and processing of data. Data-driven technologies are, for example, classification technologies for text, images or table data. Classification technology based on artificial intelligence models is one of the most commonly used classification technologies.
[0003] In the related art, the preprocessing process for classification tasks is heuristic, that is, preprocessing rules are constructed based on intuition or experience. For example, for text classification, preprocessing techniques mainly include: removing text data that is too short, truncating text data that is too long according to the maximum length; performing grammatical correction on texts from social media or informal sources; removing words that appear infrequently in the corpus; and normalizing abbreviations, abbreviations, and morphologically variable words. For image data, preprocessing techniques mainly include: denoising the noise (such as water mist, glass reflection, etc.) introduced into the image by factors such as acquisition equipment and environment; and performing feature transformation on the image to enhance its attributes (such as lighting, color, etc.). For tabular data, preprocessing techniques mainly include: completing missing data items in the table. Summary of the invention
[0004] According to the first aspect of the present disclosure, a method for screening training samples is provided, including: obtaining multiple groups of prediction results for each training sample in a set of training samples under multiple groups of parameters of an artificial intelligence model; dividing multiple training samples in the set of training samples into multiple data clusters according to the multiple groups of prediction results of each training sample in the set of training samples; calculating the contribution of each data cluster of the multiple data clusters to the training of the artificial intelligence model; and selecting a target training sample from the multiple training samples according to the contribution of the multiple data clusters.
[0005] In some embodiments, calculating the contribution of each data cluster of the multiple data clusters to the training of the artificial intelligence model includes: sorting the multiple data clusters; training the artificial intelligence model in turn using each data cluster of the multiple data clusters according to the sorting of the multiple data clusters, and obtaining evaluation information of the artificial intelligence model after each training; calculating the contribution of the first data cluster according to the corresponding evaluation information of each first data cluster in the multiple data clusters and the corresponding evaluation information of all second data clusters sorted before the first data cluster.
[0006] In some embodiments, the contribution of each of the multiple data clusters to the training of the artificial intelligence model is calculated, including: looping the steps of sorting the multiple data clusters, obtaining evaluation information, and calculating the contribution of each first data cluster; calculating the moving average of the contribution of the first data cluster in this loop according to the contribution of the first data cluster in this loop and the contribution of the first data cluster in the previous loop; when the change of the moving average of the contribution of the first data cluster in this loop relative to the moving average of the contribution obtained in the previous loop is less than or equal to a first threshold, stopping the loop and taking the contribution of the first data cluster obtained in this loop as the contribution of the first data cluster.
[0007] In some embodiments, a moving average of the contribution of the first data cluster in this cycle is calculated based on the contribution of the first data cluster in this cycle and the contribution of the first data cluster in the previous cycle, including: calculating a moving average of the contribution of the first data cluster in this cycle based on the contribution of the first data cluster in this cycle and the moving average of the first data cluster in the previous cycle.
[0008] In some embodiments, the initial value of the moving average value of the contribution of the first data cluster is determined according to the contribution of the first data cluster obtained in the first cycle.
[0009] In some embodiments, selecting a target training sample from the multiple training samples according to the contribution of the multiple data clusters includes: selecting the target training sample from multiple training samples corresponding to data clusters whose contribution exceeds a second threshold among the multiple data clusters.
[0010] In some embodiments, based on the multiple groups of prediction results of each training sample in the set of training samples, multiple training samples in the set of training samples are divided into multiple data clusters, including: calculating the confidence of the multiple groups of prediction results of each training sample; based on the confidence of the multiple groups of prediction results of each training sample, the multiple training samples in the set of training samples are divided into multiple data clusters.
[0011] In some embodiments, the multiple training samples in the set of training samples are divided into multiple data clusters according to the confidence of the multiple groups of prediction results of each training sample, including: calculating the standard deviation of the multiple groups of prediction results of each training sample; and dividing the multiple training samples in the set of training samples into multiple data clusters according to the confidence and the standard deviation.
[0012] In some embodiments, multiple training samples in a set of training samples are divided into multiple data clusters based on confidence and standard deviation, including: using the confidence and standard deviation of the multiple groups of prediction results of each training sample as the coordinates of each training sample, dividing the multiple training samples into multiple grids based on the coordinates, and treating the training samples in a grid as a data cluster.
[0013] In some embodiments, dividing a plurality of training samples in a set of training samples into a plurality of data clusters according to confidence and standard deviation includes: clustering the plurality of training samples according to confidence and standard deviation to obtain a plurality of data clusters.
[0014] In some embodiments, based on multiple groups of prediction results, multiple training samples in a set of training samples are divided into multiple data clusters, including: selecting training samples with confidence and standard deviation within a specified range from the set of training samples; and dividing the selected multiple training samples into multiple data clusters.
[0015] In some embodiments, each set of parameters in the multiple sets of parameters of the artificial intelligence model is obtained through each iterative training of multiple iterative trainings during the training of the artificial intelligence model.
[0016] In some embodiments, multiple sets of prediction results are obtained for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model, including: obtaining a set of prediction results for each training sample in the set of training samples under each set of parameters of the artificial intelligence model, so as to obtain multiple sets of prediction results under multiple sets of parameters.
[0017] In some embodiments, multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model are obtained, including: based on all features of each training sample in the set of training samples, multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model are obtained.
[0018] In some embodiments, the contribution is a Shapley value.
[0019] In some embodiments, each of the training samples includes at least one of text, image, and table.
[0020] According to a second aspect of the present disclosure, a device for screening training samples is provided, comprising: an acquisition module, configured to acquire multiple groups of prediction results for each training sample in the set of training samples under multiple groups of parameters of an artificial intelligence model; a division module, configured to divide multiple training samples in the set of training samples into multiple data clusters according to the multiple groups of prediction results of each training sample in the set of training samples; a calculation module, configured to calculate the contribution of each data cluster of the multiple data clusters to the training of the artificial intelligence model; and a selection module, configured to select a target training sample from the multiple training samples according to the contribution of the multiple data clusters.
[0021] According to a third aspect of the present disclosure, a training sample screening device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the training sample screening method according to any embodiment of the present disclosure based on instructions stored in the memory.
[0022] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the instructions are executed by a processor, the method for screening training samples according to any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0024] The present disclosure may be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:
[0025] Figure 1 A flowchart showing a method for screening training samples according to some embodiments of the present disclosure;
[0026] Figure 2(a)-Figure 2(c) A schematic diagram showing the division of training samples according to some embodiments of the present disclosure;
[0027] Figure 3 A block diagram showing a device for screening training samples according to some embodiments of the present disclosure;
[0028] Figure 4 A block diagram showing a device for screening training samples according to other embodiments of the present disclosure;
[0029] Figure 5 A block diagram of a computer system for implementing some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0030] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.
[0031] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0032] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0033] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.
[0034] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0035] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0036] The performance of artificial intelligence models depends to a large extent on high-quality training data. In classification scenarios, a large number of "raw" data samples can be obtained. Such samples are large in number and noisy, making it difficult to fully cover the knowledge (i.e., features) on which correct decisions depend, or to reflect the concepts represented by the target classification task (i.e., the mapping relationship from input to category).
[0037] In related technologies, the heuristic sample preprocessing process ignores the evaluation of the quality of training data, making it difficult to effectively screen the training data, which will affect the effect of artificial learning model training.
[0038] The present disclosure provides a method and device for screening training samples, and a computer-readable storage medium, which can effectively screen training samples, thereby improving the performance of the trained artificial intelligence model.
[0039] Figure 1 A flowchart of a method for screening training samples according to some embodiments of the present disclosure is shown.
[0040] like Figure 1 As shown, the method for screening training samples includes steps S1 to S4. In some embodiments, the method for screening training samples is performed by a screening device for training samples.
[0041] In step S1, multiple groups of prediction results are obtained for each training sample in a set of training samples under multiple groups of parameters of an artificial intelligence model. The artificial intelligence model is, for example, a neural network model.
[0042] For example, the set of training samples is recorded as D = (x j ,y j ), j = 1, 2, ... M. x j Denotes the jth training sample d j The characteristics of y j Represents the training sample d j The label of , M and j are positive integers.
[0043] In some embodiments, each training sample includes at least one of text, image and table, wherein the table is, for example, a data table, and the elements in the table are composed of numbers.
[0044] In some embodiments, each set of parameters in the multiple sets of parameters of the artificial intelligence model is obtained through each iterative training in multiple iterative training during the training of the artificial intelligence model.
[0045] Learning dynamics are statistics of the AI classifier during training, such as the probability of predicting the correct label. In order to obtain multiple sets of parameters, first θ Perform N steps of mini-batch stochastic gradient descent, save the model parameters once every k steps of stochastic gradient descent (i.e., each iteration, also called each parameter update), and record the model parameters saved at the i×k-th iteration checkpoint as θ i , then i=1..N / K, where N, i and K are all positive integers.
[0046] In some embodiments, obtaining multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model includes: obtaining a set of prediction results for each training sample in the set of training samples under each set of parameters in the multiple sets of parameters of the artificial intelligence model, so as to obtain multiple sets of prediction results under multiple sets of parameters.
[0047] For the training set D = (x j ,y j ) and take any training sample in it and assign its feature x j Enter the artificial intelligence model corresponding to the i-th group of parameters After that, forward calculation is performed to obtain the prediction result of the training sample. j ,y j Respectively represent the j-th training sample d jThe features and labels of θ i Represents the artificial intelligence classifier M θ Perform N steps of mini-batch stochastic gradient descent, and save the model parameters once every K steps of stochastic gradient descent, the model parameters of the checkpoint of the i×K-th iteration (i.e., the i-th checkpoint). i = 1..N / K, N, i and K are all positive integers.
[0048] In some embodiments, multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model are obtained, including: based on all features of each training sample in the set of training samples, multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model are obtained.
[0049] Text data is inherently a discrete high-dimensional data. Its characteristic dimension is, for example, the size of the vocabulary in the text. The characteristic dimension of image data is the number of pixels it contains (for example, 128×128). The characteristic dimension of tabular data is often hundreds or even thousands of dimensions. The outlier removal in related technologies processes single variables in the features rather than processing high-dimensional features as a whole, ignoring the impact of the dependency between variables on the training sample screening results.
[0050] According to some embodiments of the present disclosure, in the process of characterizing and screening data samples, the characteristics of the training samples are regarded as a whole, processed and predicted results are obtained, thereby fully considering the feature dependencies of each dimension of high-dimensional data, overcoming the difficulties in characterizing high-order correlations of features when performing data diagnosis, and improving the accuracy of screening training samples.
[0051] In some embodiments, the prediction result is that the artificial intelligence model predicts that the category of the jth training sample is y at the i-th parameter checkpoint j The probability of y j It is also the jth training sample d j .
[0052] For example, for N-step mini-batch stochastic gradient descent, the model parameters are saved every K-step stochastic gradient descent, and each of the M training samples is scored with N / K sets of model parameters to obtain M×N / K probability values, where N, K, and N / K are all positive integers.
[0053] In step S2, multiple training samples in the set of training samples are divided into multiple data clusters according to multiple groups of prediction results of each training sample in the set of training samples.
[0054] For example, according to the learning dynamics, the training samples are clustered, and the samples in the same data cluster have similar effects on the model training process. Given the number of data clusters k, multiple training samples are divided into k data clusters, and the k data clusters are recorded as Among them, i ranges from 1 to k, After expansion, it is recorded as {D 1 , D 2 , ...D k}.
[0055] In some embodiments, based on multiple groups of prediction results of each training sample in the set of training samples, multiple training samples in the set of training samples are divided into multiple data clusters, including: calculating the confidence of the multiple groups of prediction results of each training sample; based on the confidence of the multiple groups of prediction results of each training sample, multiple training samples in the set of training samples are divided into multiple data clusters.
[0056] In some embodiments, the confidence of multiple groups of prediction results is the average value of the multiple groups of prediction results.
[0057] For example, for N-step mini-batch stochastic gradient descent, the model parameters are saved every K-step stochastic gradient descent. Then the training sample d j Multiple prediction results The confidence level conf(x j ) is calculated as follows.
[0058]
[0059] in, Indicates that the category of the training sample is predicted to be y using the model parameters saved for the i-th time j probability.
[0060] In some embodiments, multiple training samples in a set of training samples are divided into multiple data clusters based on the confidence of the multiple groups of prediction results of each training sample, including: calculating the standard deviation of the multiple groups of prediction results of each training sample; and dividing the multiple training samples in the set of training samples into multiple data clusters based on the confidence and the standard deviation.
[0061] For example, for N-step mini-batch stochastic gradient descent, the model parameters are saved every K-step stochastic gradient descent. The change in the confidence of the training sample is represented by the standard deviation of the prediction results under N / K models, where N, K, and N / K are all positive integers. The standard deviation is calculated as follows.
[0062]
[0063] Among them, x j Denotes the jth training sample d j The characteristics of yj Denotes the jth training sample d j Tags, Indicates that the category of the jth training sample is predicted to be y using the model parameters saved for the i-th time j probability.
[0064] After obtaining the confidence and standard deviation, the confidence and standard deviation are used as the coordinates of each training sample to construct a data map and realize the two-dimensional visualization of the training data set. For example, the standard deviation of the training sample is used as the horizontal coordinate and the confidence as the vertical coordinate, and all the training samples in the data set D are plotted in a two-dimensional plane map of (var(x,y),conf(x,y)) to facilitate data cluster division.
[0065] Since the learning dynamics are characterized by the confidence of data points and the change in confidence during model training, samples with similar confidence have similar effects on the training effect of the model, thus enabling a more reasonable division of training samples.
[0066] Figure 2(a)-Figure 2(c) A schematic diagram showing the division of training samples according to some embodiments of the present disclosure.
[0067] As shown in Figure 2(a), the training sample (x j ,y j ) reflects the certainty of the average decision on the training sample during the learning process of the machine learning model, that is, the higher the confidence level of the training sample, the easier it is to learn; conversely, the lower the confidence level of the training sample, the more difficult it is to learn.
[0068] The standard deviation of a training sample reflects whether the decision on the training sample fluctuates greatly at different stages during the learning process of the machine learning model. The greater the fluctuation, the higher the ambiguity or noise in the training sample. Using confidence and standard deviation, we can characterize the difficult-to-learn data points and noisy singular data points distributed in the lower and right sides of the data map. Moreover, through standard deviation and confidence, two types of statistics that are highly correlated with the model learning process, we can more accurately divide data clusters for samples with similar contributions.
[0069] In some embodiments, based on multiple groups of prediction results, multiple training samples in a set of training samples are divided into multiple data clusters, including: selecting training samples with confidence and standard deviation within a specified range from the set of training samples; and dividing the selected multiple training samples into multiple data clusters.
[0070] In some embodiments, the specified range is that the confidence is between the first confidence threshold and the second confidence threshold, and the standard deviation is between the first standard deviation threshold and the second standard deviation threshold. These four thresholds can be set at a relatively high level so that the screened samples have the characteristics of easy learning, high noise, and difficult learning.
[0071] For example, given the first confidence threshold conf low With the second confidence threshold conf high , the first standard deviation threshold var low With the second standard deviation threshold var high , for conf(x j ,y j )>conf high And var(x j ,y j ) low The samples are recorded as set D easy , which is an easy-to-learn training sample. Where high and low represent the upper and lower limits of the specified range, respectively. high >conf low , var high >var low .
[0072] For conf(x j ,y j ) <conf low And var(x j ,y j ) low The training samples are recorded as set D hard , which is considered to be a hard-to-learn sample. high With conf low but with a high standard deviation (var(x j ,y j )>var high ) are considered to be samples with large ambiguity or noise samples that are difficult for the model to fit, and are recorded as set D noisy . D easy , D noisy , D hard The samples outside are moderate samples, denoted as D medium , in D medium The contribution of each data cluster is calculated.
[0073] In some embodiments, dividing multiple training samples in the set of training samples into multiple data clusters according to the confidence and the standard deviation includes: clustering the multiple training samples according to the confidence and the standard deviation to obtain multiple data clusters.
[0074] For example, as shown in Figure 2(b), the k-means clustering algorithm is used to cluster the training data in the coordinate system (var(xj ,y j ),conf(x j ,y j )) is divided into k data clusters.
[0075] In some embodiments, multiple training samples in a set of training samples are divided into multiple data clusters based on confidence and standard deviation, including: using the confidence and standard deviation of multiple groups of prediction results of each training sample as the coordinates of each training sample, dividing the multiple training samples into multiple grids based on the coordinates, and treating the training samples in a grid as a data cluster.
[0076] For example, as shown in Figure 2(c), the training data is divided into k grids directly in the data map through straight lines horizontally and vertically to the coordinate axis, and the training samples in each grid belong to the same data cluster.
[0077] In step S3, the contribution of each data cluster of the multiple data clusters to the training of the artificial intelligence model is calculated.
[0078] Calculating the contribution of each training sample separately requires a large amount of calculation. The present invention calculates the contribution of each cluster based on the data cluster, which can improve the calculation efficiency of the contribution.
[0079] In some embodiments, the contribution of the data cluster is a Shapley value. For example, the data quality of the model is accurately characterized in fine granularity through the Shapley value of the data cluster. The Shapley value of the data cluster means the marginal contribution of the data cluster to the performance of the model trained using the original data set of the training sample.
[0080] In some embodiments, calculating the contribution of each data cluster of multiple data clusters to the training of an artificial intelligence model includes: sorting the multiple data clusters; training the artificial intelligence model in turn using each data cluster of the multiple data clusters according to the sorting of the multiple data clusters, and obtaining evaluation information of the artificial intelligence model after each training; calculating the contribution of the first data cluster according to the corresponding evaluation information of each first data cluster in the multiple data clusters and the corresponding evaluation information of all second data clusters sorted before the first data cluster.
[0081] For example, given the data cluster Among them, k represents the total number of data clusters, j ranges from 1 to k, After expansion, it is recorded as {D 1 , D 2 , ...D k}, j∈{1,2..,k}, i, j and k are positive integers. For simplicity, each data cluster is numbered and denoted as j. The data cluster arrangement is sampled to obtain the arrangement σ, and the data cluster D j The position in the arrangement σ is denoted by σ j, that is, σ[σ j ] = j. According to the order of arrangement of the data clusters, starting from the initialized model (i.e., the model with randomized parameters), k mini-batch stochastic gradient descents are performed to obtain the weights of k models. In other words, the data clusters are input into the machine learning model in sequence according to the order, and the parameters of the model are updated and saved each time.
[0082] Relative to the number σ[σ j ] (i.e., the first data cluster), and all data clusters before the first data cluster (i.e., the second data cluster) are recorded as σ[1:σ j -1]. Use σ j -The weight of the model trained with 1 data cluster is θ 0 , using σ j The weight of the model trained with a set of data clusters is θ 1 , the data cluster D is calculated as follows j contribution value.
[0083] Δ(V,σ,j,D′)=V(θ 1 ,D′)-V(θ 0 ,D′)
[0084] Among them, V represents the evaluation information, which is usually an indicator to measure the performance of the model, such as accuracy, precision, recall, etc. D′ represents the set of test samples, σ represents the arrangement of data clusters, V(θ 1 ,D′) means using the test sample D′, in the model parameter θ 1 Under this condition, we can obtain the evaluation information. V(θ 0 ,D′) means using the test sample D′, in the model parameter θ 0 Next, obtain the evaluation information.
[0085] In some embodiments, a moving average of the contribution of the first data cluster in this cycle is calculated based on the contribution of the first data cluster in this cycle and the contribution of the first data cluster in the previous cycle, including: calculating a moving average of the contribution of the first data cluster in this cycle based on the contribution of the first data cluster in this cycle and the moving average of the first data cluster in the previous cycle.
[0086] For example, for the tth cycle, the order obtained by sampling is recorded as σ t , and obtain the Shapley value of each data cluster in this cycle, where t is a positive integer. The Shapley value of this cycle and the Shapley value of the same data cluster in the previous t-1 cycles are combined by moving average. The calculation formula of the moving average is as follows.
[0087]
[0088] Among them, V represents the evaluation information, D j Represents the data cluster numbered j.
[0089] In some embodiments, the initial value of the moving average value of the contribution of the first data cluster is determined according to the contribution of the first data cluster obtained in the first cycle.
[0090] For example, the initial value of the moving average value of the contribution of the first data cluster is equal to the contribution of the first data cluster obtained in the first cycle, that is, the calculation formula of the moving average value of the first cycle of the data cluster is as follows.
[0091] S(D j ,σ 1 )=Δ(V,σ 1 ,j,D′)
[0092] Among them, D j represents the data cluster numbered j, σ 1 represents the sorting of the data clusters sampled in the first cycle, V represents the evaluation information, and D′ represents the set of test samples.
[0093] In some embodiments, the contribution of each data cluster of multiple data clusters to the training of the artificial intelligence model is calculated, including: looping through a step of sorting multiple data clusters, a step of obtaining evaluation information, and a step of calculating the contribution of each first data cluster; calculating a moving average of the contribution of the first data cluster in this loop based on the contribution of the first data cluster in this loop and the contribution of the first data cluster in the previous loop; when the change in the moving average of the contribution of the first data cluster in this loop relative to the moving average of the contribution obtained in the previous loop is less than or equal to a first threshold, stopping the loop and using the contribution of the first data cluster obtained in this loop as the contribution of the first data cluster.
[0094] For example, according to the data cluster D j The moving average of the contribution obtained in this cycle is consistent with the cluster D j The moving average of the contribution obtained in the last cycle and the difference between the two determine the change ΔS (D j ), and the calculation formula is as follows.
[0095] ΔS(D j )=S(D j ,σ t )-S(D j ,σ t-1 )
[0096] Among them, D jrepresents the data cluster numbered j, σ t Represents the sorting obtained by sampling in the tth cycle.
[0097] When the change degree of some or all data clusters ΔS(D j ) is less than the preset tolerance ΔS, the loop is stopped, that is, no new arrangement σ is sampled, and the contribution of each data cluster calculated in this loop is used as the final contribution.
[0098] In step S4, a target training sample is selected from a plurality of training samples according to the contribution of a plurality of data clusters.
[0099] For example, the contribution of each data cluster is sorted from high to low, and the training samples in the data cluster ranked first are selected as target training samples for training the classification model.
[0100] When screening data, it is difficult to learn samples D hard Or ambiguous sample D noisy , can be directly removed. For the easy-to-learn sample D easy , just keep a certain number. For a moderate sample D medium By trying to select samples ranked in the top n% of Shapley values as a set of training samples, the sample screening steps of the present disclosure are repeated until a suitable sample is found, where n>0.
[0101] By calculating the contribution value of each data sample and sorting them, samples can be selected from the top-ranked data clusters for training better classification models. Compared with the data clusters constructed by heuristic methods in related technologies, the method disclosed in this disclosure for clustering based on similar learning dynamics has a stronger perception of the model training process and the final performance (reflected by the V value), thereby screening out more suitable samples.
[0102] In some embodiments, selecting a target training sample from a plurality of training samples according to the contribution of a plurality of data clusters includes: selecting the target training sample from a plurality of training samples corresponding to a data cluster whose contribution exceeds a second threshold value in the plurality of data clusters. For example, selecting a training sample in a data cluster whose contribution exceeds the second threshold value as the target training sample.
[0103] The screening method of training samples disclosed in the present invention includes: obtaining multiple groups of prediction results for each training sample in the set of training samples under multiple groups of parameters of an artificial intelligence model; dividing multiple training samples in the set of training samples into multiple data clusters according to the multiple groups of prediction results of each training sample in the set of training samples; calculating the contribution of each data cluster of the multiple data clusters to the training of the artificial intelligence model; and selecting a target training sample from the multiple training samples according to the contribution of the multiple data clusters.
[0104] The above-mentioned technical solution disclosed in the present invention calculates the contribution of training samples to model training based on the prediction results of the artificial intelligence model for training samples, and uses the information during model training to diagnose the training samples in turn, so that data diagnosis and model training form a closed loop, and can efficiently screen training samples, improve the efficiency of training artificial intelligence models, and improve the performance of trained artificial intelligence models. Compared with data clusters constructed in a heuristic way, the training sample screening method disclosed in the present invention has a stronger perception of the model training process and final performance. In addition, the present invention proposes a method for calculating contribution by clustering. Using the prediction results, samples with similar contributions are divided into a cluster. Compared with the method of calculating the contribution for each training sample, the amount of calculation is reduced and the calculation efficiency is improved.
[0105] The following uses the classification of comments on items as an example to introduce the application of the training sample screening method disclosed in the present invention.
[0106] In the review scenario, by analyzing users' comments on items, we can extract users' evaluations of certain attributes of items (such as quality, size, user experience, etc.). For example, if a user comments that "the brightness can be adjusted, the maximum brightness is still not bright enough", we can determine that the comment is a description of the "lighting effect" by analyzing the comment.
[0107] According to the above description, the analysis task of item reviews is a classification problem, that is, given user reviews (training samples) and attributes of various dimensions that an item may have (all labels), the reviews are matched with specific labels.
[0108] Items of different categories sometimes have the same attributes, and some categories sometimes have attributes unique to that category. Even if two categories A and B have different attributes, the training data of category A can also play a role in transfer learning for the training of the classification model of category B (target category). In addition, the labeling of the attribute dimension of the item to which the user comment belongs is based on rules or weak labels obtained through the transfer of comments from other categories. Therefore, the set of training samples includes weakly supervised training samples of M attributes of category B and comment data of M attributes of category B.
[0109] Using a set of training samples, the machine learning model is trained for multiple epochs to obtain multiple sets of model parameters. Then, for each training sample, multiple sets of parameters are used to generate multiple sets of prediction results. Based on the prediction results, the confidence and variance of each training sample are calculated. Based on the confidence and variance, the noise training samples are removed, and the remaining training samples are clustered to obtain k data clusters. Sample a permutation of k data clusters, and update the model with the training data of each data cluster in sequence. Based on the model before and after each data cluster is updated, the change in the F1 (F1-score) score of the development set is calculated, and the contribution and moving average of each data cluster are calculated based on the F1 score. When the contribution of each data cluster converges, the contribution of each data cluster is output, and based on the contribution of each data cluster, the training samples that are more helpful for model training are screened.
[0110] Training sample diagnosis strategies include: (1) screening out training samples with high contribution from other categories; (2) excluding abnormal training samples in the weakly supervised corpus of the target category, for example, training samples with confidence and standard deviation outside the specified range and training samples with low contribution.
[0111] According to the training sample screening method disclosed in the present invention, samples that are useful for the review analysis of the target category can be selected from samples of other categories. In addition, when the labels of the training samples of the target category are weak labels obtained by migrating reviews of other categories and there is a lot of noise, the training samples with a lot of noise can be excluded from the training samples of the target category, thereby constructing a clean test set.
[0112] Figure 3 A block diagram showing a device for screening training samples according to some embodiments of the present disclosure.
[0113] like Figure 3 As shown, the training sample screening device 3 includes an acquisition module 31 , a division module 32 , a calculation module 33 and a selection module 34 .
[0114] The acquisition module 31 is configured to obtain multiple sets of prediction results of each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model, for example, by executing Figure 1 Step S1 is shown.
[0115] The partitioning module 32 is configured to partition multiple training samples in the set of training samples into multiple data clusters according to multiple groups of prediction results of each training sample in the set of training samples, for example, by performing the following steps: Figure 1 Step S2 is shown.
[0116] The calculation module 33 is configured to calculate the contribution of each data cluster of the plurality of data clusters to the training of the artificial intelligence model, for example, by performing the following steps: Figure 1Step S3 shown.
[0117] The selection module 34 is configured to select a target training sample from a plurality of training samples according to the contribution of a plurality of data clusters, for example, by performing the following steps: Figure 1 Step S4 is shown.
[0118] In some embodiments, the calculation module 33 is further configured to calculate the contribution of each data cluster of the multiple data clusters to the training of the artificial intelligence model, including: sorting the multiple data clusters; training the artificial intelligence model in turn using each data cluster of the multiple data clusters according to the sorting of the multiple data clusters, and obtaining evaluation information of the artificial intelligence model after each training; calculating the contribution of the first data cluster according to the corresponding evaluation information of each first data cluster in the multiple data clusters and the corresponding evaluation information of all second data clusters sorted before the first data cluster.
[0119] In some embodiments, the calculation module 33 is further configured to calculate the contribution of each of the multiple data clusters to the training of the artificial intelligence model, including: looping the sorting step of the multiple data clusters, the acquisition step of evaluation information, and the step of calculating the contribution of each first data cluster; calculating the moving average of the contribution of the first data cluster in this cycle according to the contribution of the first data cluster in this cycle and the contribution of the first data cluster in the previous cycle; when the change of the moving average of the contribution of the first data cluster in this cycle relative to the moving average of the contribution obtained in the previous cycle is less than or equal to a first threshold, stop the loop and use the contribution of the first data cluster obtained in this cycle as the contribution of the first data cluster.
[0120] In some embodiments, the calculation module 33 is further configured to calculate the moving average of the contribution of the first data cluster in this cycle based on the contribution of the first data cluster in this cycle and the contribution of the first data cluster in the previous cycle, including: calculating the moving average of the contribution of the first data cluster in this cycle based on the contribution of the first data cluster in this cycle and the moving average of the first data cluster in the previous cycle.
[0121] In some embodiments, the initial value of the moving average value of the contribution of the first data cluster is determined according to the contribution of the first data cluster obtained in the first cycle.
[0122] In some embodiments, the selection module 34 is further configured to select a target training sample from the multiple training samples based on the contribution of the multiple data clusters, including: selecting a target training sample from multiple training samples corresponding to data clusters whose contribution exceeds a second threshold in the multiple data clusters.
[0123] In some embodiments, the division module 32 is further configured to divide the multiple training samples in the set of training samples into multiple data clusters according to the multiple groups of prediction results of each training sample in the set of training samples, including: calculating the confidence of the multiple groups of prediction results of each training sample; dividing the multiple training samples in the set of training samples into multiple data clusters according to the confidence of the multiple groups of prediction results of each training sample.
[0124] In some embodiments, the division module 32 is further configured to divide the multiple training samples in the set of training samples into multiple data clusters according to the confidence of the multiple groups of prediction results of each training sample, including: calculating the standard deviation of the multiple groups of prediction results of each training sample; dividing the multiple training samples in the set of training samples into multiple data clusters according to the confidence and the standard deviation.
[0125] In some embodiments, the division module 32 is further configured to divide multiple training samples in the set of training samples into multiple data clusters according to confidence and standard deviation, including: using the confidence and standard deviation of the multiple groups of prediction results of each training sample as the coordinates of each training sample, dividing the multiple training samples into multiple grids according to the coordinates, and treating the training samples in a grid as a data cluster.
[0126] In some embodiments, the partitioning module 32 is further configured to partition multiple training samples in the set of training samples into multiple data clusters according to confidence and standard deviation, including: clustering the multiple training samples according to confidence and standard deviation to obtain multiple data clusters.
[0127] In some embodiments, the partitioning module 32 is further configured to partition multiple training samples in the set of training samples into multiple data clusters based on multiple groups of prediction results, including: selecting training samples whose confidence and standard deviation are within a specified range from the set of training samples; and partitioning the selected multiple training samples into multiple data clusters.
[0128] In some embodiments, each set of parameters in the multiple sets of parameters of the artificial intelligence model is obtained through each iterative training of multiple iterative trainings during the training of the artificial intelligence model.
[0129] In some embodiments, the acquisition module 31 is further configured to obtain multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model, including: obtaining a set of prediction results for each training sample in the set of training samples under each set of parameters in the multiple sets of parameters of the artificial intelligence model, so as to obtain multiple sets of prediction results under multiple sets of parameters.
[0130] In some embodiments, the acquisition module 31 is further configured to obtain multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model, including: obtaining multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model based on all features of each training sample in the set of training samples.
[0131] In some embodiments, the contribution is a Shapley value.
[0132] In some embodiments, each of the training samples includes at least one of text, image, and table.
[0133] Figure 4 A block diagram showing a device for screening training samples according to some other embodiments of the present disclosure.
[0134] like Figure 4 As shown, the training sample screening device 4 includes a memory 41; and a processor 42 coupled to the memory 41, and the memory 41 is used to store and execute the training sample screening method. The processor 42 is configured to execute the training sample screening method in any of the embodiments of the present disclosure based on the instructions stored in the memory 41.
[0135] Figure 5 A block diagram of a computer system for implementing some embodiments of the present disclosure is shown.
[0136] like Figure 5 As shown, the computer system 50 may be embodied in the form of a general purpose computing device. The computer system 50 includes a memory 510, a processor 520, and a bus 500 that connects the various system components.
[0137] The memory 510 may include, for example, a system memory, a non-volatile storage medium, etc. The system memory may store, for example, an operating system, an application program, a boot loader, and other programs. The system memory may include a volatile storage medium, such as a random access memory (RAM) and / or a cache memory. The non-volatile storage medium may store, for example, instructions for executing the method for screening training samples in any of the embodiments of the present disclosure. The non-volatile storage medium may include, but is not limited to, a disk memory, an optical memory, a flash memory, etc.
[0138] The processor 520 can be implemented by a general processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistors, etc. Discrete hardware components. Accordingly, each module such as the judgment module and the determination module can be implemented by a central processing unit (CPU) running instructions in a memory that execute corresponding steps, or can be implemented by a dedicated circuit that executes corresponding steps.
[0139] The bus 500 may use any of a variety of bus architectures, including, but not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, and a Peripheral Component Interconnect (PCI) bus.
[0140] The computer system 50 may also include an input / output interface 530, a network interface 540, a storage interface 550, etc. These interfaces 530, 540, 550, the memory 510, and the processor 520 may be connected via a bus 500. The input / output interface 530 may provide a connection interface for input / output devices such as a display, a mouse, and a keyboard. The network interface 540 may provide a connection interface for various networked devices. The storage interface 550 may provide a connection interface for external storage devices such as a floppy disk, a USB flash drive, and an SD card.
[0141] Here, various aspects of the present disclosure are described with reference to flowcharts and / or block diagrams of methods, devices, and computer program products according to embodiments of the present disclosure. It should be understood that each frame of the flowchart and / or block diagram and the combination of frames can be implemented by computer-readable program instructions.
[0142] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable device to produce a machine, so that the processor executes the instructions to produce means for implementing the functions specified in one or more blocks in the flowchart and / or block diagram.
[0143] These computer-readable program instructions can also be readable and stored in a computer-readable memory, which cause the computer to work in a specific manner to produce a product, including instructions for implementing the functions specified in one or more blocks in the flowchart and / or block diagram.
[0144] The present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.
[0145] Through the training sample screening method and device and computer-readable storage medium in the above-mentioned embodiments, training samples can be efficiently screened and the efficiency of training artificial intelligence models can be improved.
[0146] So far, the method and device for screening training samples and the computer-readable storage medium according to the present disclosure have been described in detail. In order to avoid obscuring the concept of the present disclosure, some details known in the art are not described. Based on the above description, those skilled in the art can fully understand how to implement the technical solution disclosed herein.
Claims
1. A method for screening training samples. include: Obtain multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model; According to the multiple groups of prediction results of each training sample in the set of training samples, the multiple training samples in the set of training samples are divided into multiple data clusters; Calculating the contribution of each of the multiple data clusters to the training of the artificial intelligence model; A target training sample is selected from the plurality of training samples according to the contribution degrees of the plurality of data clusters.
2. The method for screening training samples according to claim 1, in, Calculating the contribution of each of the multiple data clusters to the artificial intelligence model training includes: sorting the plurality of data clusters; According to the sorting of the multiple data clusters, using each of the multiple data clusters, sequentially train the artificial intelligence model, and obtain evaluation information of the artificial intelligence model after each training; The contribution of the first data cluster is calculated according to the evaluation information corresponding to each first data cluster in the plurality of data clusters and the evaluation information corresponding to all second data clusters ranked before the first data cluster.
3. The method for screening training samples according to claim 2, in, Calculating the contribution of each of the multiple data clusters to the artificial intelligence model training includes: cyclically executing the steps of sorting the plurality of data clusters, acquiring the evaluation information, and calculating the contribution of each first data cluster; Calculate a moving average of the contribution of the first data cluster in this cycle according to the contribution of the first data cluster in this cycle and the contribution of the first data cluster in the previous cycle; When the change of the moving average of the contribution of the first data cluster obtained in this cycle relative to the moving average of the contribution obtained in the previous cycle is less than or equal to the first threshold, the loop is stopped and the contribution of the first data cluster obtained in this cycle is used as the contribution of the first data cluster.
4. The method for screening training samples according to claim 3, in, Calculating a moving average of the contribution of the first data cluster in the current cycle according to the contribution of the first data cluster in the current cycle and the contribution of the first data cluster in the previous cycle, including: According to the contribution of the first data cluster in the current cycle and the moving average of the first data cluster in the previous cycle, a moving average of the contribution of the first data cluster in the current cycle is calculated.
5. The method for screening training samples according to claim 4, in, The initial value of the moving average value of the contribution of the first data cluster is determined according to the contribution of the first data cluster obtained in the first cycle.
6. The method for screening training samples according to claim 1, in, Selecting a target training sample from the plurality of training samples according to the contribution of the plurality of data clusters includes: A target training sample is selected from a plurality of training samples corresponding to a data cluster whose contribution exceeds a second threshold value among the plurality of data clusters.
7. The method for screening training samples according to claim 1, in, According to the multiple groups of prediction results of each training sample in the set of training samples, multiple training samples in the set of training samples are divided into multiple data clusters, including: Calculating the confidence of multiple groups of prediction results of each training sample; According to the confidence of the multiple groups of prediction results of each training sample, the multiple training samples in the set of training samples are divided into multiple data clusters.
8. The method for screening training samples according to claim 7, in, According to the confidence of the multiple groups of prediction results of each training sample, the multiple training samples in the set of training samples are divided into multiple data clusters, including: Calculating the standard deviation of the multiple groups of prediction results of each training sample; The plurality of training samples in the set of training samples are divided into a plurality of data clusters according to the confidence and the standard deviation.
9. The method for screening training samples according to claim 8, in, According to the confidence and standard deviation, multiple training samples in the set of training samples are divided into multiple data clusters, including: The confidence and standard deviation of the multiple groups of prediction results of each training sample are used as the coordinates of each training sample, and the multiple training samples are divided into multiple grids according to the coordinates, and the training samples in one grid are regarded as a data cluster.
10. The method for screening training samples according to claim 8, in, According to the confidence and standard deviation, multiple training samples in the set of training samples are divided into multiple data clusters, including: The multiple training samples are clustered according to the confidence and the standard deviation to obtain multiple data clusters.
11. The method for screening training samples according to claim 8, in, According to the confidence and the standard deviation, the plurality of training samples in the set of training samples are divided into a plurality of data clusters, including: From the set of training samples, select training samples whose confidence and standard deviation are within the specified range; The selected multiple training samples are divided into multiple data clusters.
12. The method for screening training samples according to any one of claims 1 to 11, in, Each set of parameters in the multiple sets of parameters of the artificial intelligence model is obtained through each iterative training in multiple iterative training during the artificial intelligence model training.
13. The method for screening training samples according to any one of claims 1 to 11, in, Obtain multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model, including: A set of prediction results is obtained for each training sample in the set of training samples under each set of parameters in the multiple sets of parameters of the artificial intelligence model to obtain multiple sets of prediction results under the multiple sets of parameters.
14. The method for screening training samples according to any one of claims 1 to 11, in, Obtain multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model, including: According to all the features of each training sample in the set of training samples, multiple groups of prediction results are obtained for each training sample in the set of training samples under multiple groups of parameters of the artificial intelligence model.
15. The method for screening training samples according to any one of claims 1 to 11, in, The contribution is the Shapley value.
16. The method for screening training samples according to any one of claims 1 to 11, in, Each of the training samples includes at least one of text, image and table.
17. A training sample screening device, include: An acquisition module is configured to obtain multiple sets of prediction results for each training sample in the set of training samples under multiple sets of parameters of the artificial intelligence model; A partitioning module is configured to partition a plurality of training samples in the set of training samples into a plurality of data clusters according to a plurality of groups of prediction results of each training sample in the set of training samples; A calculation module, configured to calculate the contribution of each of the multiple data clusters to the artificial intelligence model training; The selection module is configured to select a target training sample from the multiple training samples according to the contribution of the multiple data clusters.
18. A training sample screening device, include: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the method for screening training samples according to any one of claims 1 to 16 based on instructions stored in the memory.
19. A computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implements the method for screening training samples according to any one of claims 1 to 16.
Citation Information
Cited By
High-voltage cable classification method and device, terminal equipment and storage medium
CN121302115A
A high-voltage cable classification method, device, terminal equipment and storage medium
CN121302115B