The present application relates to the model training sample screening field, for solving the problem that the data cannot be effectively screened in the model training process, resulting in the training cost is improved and the
training effect is reduced, specifically for the model data sample intelligent
screening method and
system for knowledge
distillation; The present application carries out multistage screening to the basic data for training when training the model, so as to fully remove the data completely ineffective for training, false interference data and high
repeatability homogeneous data, simultaneously, the large amount of data is processed by dichotomization, the homogeneous data can be efficiently selected by
jumping, both the removal rate of homogeneous data is guaranteed and the
workload when removing homogeneous data is reduced, the data for model training can play an optimizing role for model training, meanwhile, the negative interference of false data is reduced, also can avoid a large amount of homogeneous data to make repeated training, improve the
data quality, reduce the training cost.