Training sample processing method, mail identification method and computing device
By dividing the training sample set and adding non-degraded samples, the problem of low efficiency and accuracy of degraded samples in the prior art is solved, and the stability and accuracy of model performance are improved.
Patent Information
- Application Number
- CN202510504305.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, it is determined that the deteriorated sample efficiency and accuracy of the training sample set are low, resulting in a degraded model performance.
By dividing the training sample sets, the preset models are trained separately, the deteriorated samples that cause the model performance to deteriorate, and the deteriorated samples are determined as deteriorated samples when the number of deteriorated samples is less than the threshold, or non-deteriorated samples are added when the number is greater than the threshold to meet the model input quantity and proportional requirements.
The determination efficiency and accuracy of degraded samples are improved, ensuring the stability and improvement of model performance.
Smart Images

Figure CN120342888A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of data processing, and in particular, to a method for processing training samples, a method for email recognition, and a computing device. Background Art
[0002] When training a model based on a training sample set, if it is found that the model performance of the trained model fails to be effectively improved, or even the model performance deteriorates, it indicates that there are likely to be degraded samples in the training sample set that cause the model performance to decline. Therefore, to improve the quality of the training sample set and the model performance of the trained model, it is necessary to identify the degraded samples in the training sample set and remove the degraded samples. However, the efficiency and accuracy of the methods for identifying degraded samples in the prior art are relatively low and still need further improvement. Summary of the Invention
[0003] The embodiments of the present application provide a method for processing training samples, a method for email recognition, and a computing device, which can identify the degraded samples that cause the model performance of a preset model to decline in a training sample set.
[0004] To achieve the above object, the embodiments of the present application adopt the following technical solutions:
[0005] In a first aspect, the embodiments of the present application provide a method for processing training samples, the method including: obtaining a first target sample set that causes the model performance of a preset model to be lower than a preset performance, and dividing the first target sample set into a plurality of first sample subsets; training the preset model respectively based on each of the plurality of first sample subsets, and after the training is completed, determining a second target sample set that causes the model performance of the preset model to be lower than the preset performance in the plurality of first sample subsets; in the case where the number of training samples in the second target sample set is less than or equal to a preset threshold, determining the training samples in the second target sample set as the degraded samples that cause the model performance of the preset model to be lower than the preset performance.
[0006] Based on this implementation, in the case where the training sample set causes the model performance of the preset model to decline, the range of the degraded samples can be narrowed by dividing the training sample set, so as to identify the degraded samples that cause the model performance of the preset model to decline, thereby improving the efficiency and accuracy of identifying the degraded samples.
[0007] In a possible implementation, the above method further includes: when the number of training samples in the above second target sample set is greater than the above preset threshold, dividing the above second target sample set into multiple second sample subsets; based on the number of training samples in each of the above second sample subsets in the above multiple second sample subsets and the minimum input amount of the training samples corresponding to the above preset model, training the above preset model, and after the training is completed, determining a third target sample set in the above multiple second sample subsets that causes the model performance of the above preset model to be lower than the above preset performance.
[0008] Based on this implementation, when the number of training samples in the target sample set is greater than the preset threshold, the target sample set can be further divided to obtain multiple sample subsets; and when training the preset model based on the multiple sample subsets, the number of training samples in each second sample subset can be ensured to meet the minimum input amount of the training samples corresponding to the preset model by adding training samples first, and then the preset model can be trained based on the multiple second sample subsets after the addition.
[0009] In another possible implementation, the training of the above preset model based on the number of training samples in each of the above second sample subsets in the above multiple second sample subsets and the minimum input amount of the training samples corresponding to the above preset model includes: if the number of training samples in each of the above second sample subsets is greater than or equal to the above minimum input amount, training the above preset model based on each of the above second sample subsets; if there is a target sample subset in the above multiple second sample subsets where the number of training samples is less than the above minimum input amount, determining the number of samples to be added corresponding to the above target second sample subset based on the above minimum input amount and the number of training samples in the above target sample subset; obtaining non-degraded samples with the same number of samples to be added as the above target sample subset, and adding the above non-degraded samples to the above target sample subset to obtain multiple second sample subsets after the addition, and training the above preset model based on the multiple second sample subsets after the addition; where the number of training samples in the above target second sample subset is greater than or equal to the above minimum input amount.
[0010] Based on this implementation, it can be ensured that the second sample subset meets the minimum input amount of the training samples corresponding to the preset model, and since the added samples are non-degraded samples, the determination process of the degraded samples will not be affected.
[0011] In yet another possible implementation, obtaining non-deteriorated samples with the same number as the samples to be added corresponding to the target sample subset, and adding the non-deteriorated samples to the target sample subset to obtain multiple second sample subsets after addition, includes: determining the number of positive samples and negative samples in the target sample subset; determining the number of positive samples to be added and the number of negative samples to be added in the samples to be added corresponding to the target sample subset according to the preset sample ratio corresponding to the preset model, the number of samples to be added corresponding to the target sample subset, the number of positive samples and the number of negative samples in the target sample subset; wherein the preset sample ratio is the ratio between the number of positive samples and the number of negative samples input into the preset model; obtaining non-deteriorated positive samples with the same number as the positive samples to be added corresponding to the target sample subset, and non-deteriorated negative samples with the same number as the negative samples to be added corresponding to the target sample subset, and adding the non-deteriorated positive samples and the non-deteriorated negative samples to the target sample subset to obtain the multiple second sample subsets after addition.
[0012] Based on this implementation, while enabling the number of training samples in the target sample subset to meet the minimum input amount of the training samples corresponding to the preset model, it can also meet the preset sample ratio corresponding to the preset model, thereby avoiding the degradation of the model performance of the preset model caused by the training samples not meeting the minimum input amount or the preset sample ratio of the training samples corresponding to the preset model, and further affecting the determination of deteriorated samples.
[0013] In yet another possible implementation, the above method further includes: obtaining a training sample set, and training the above preset model based on the above training sample set; wherein, the above training sample set includes a plurality of training samples; after the training of the above preset model is completed, if the model performance of the preset model corresponding to the above training sample set is lower than the above preset performance, then the above training sample set is determined as the above first target sample set; wherein, in the case that the test score of the preset model corresponding to the above training sample set is lower than the preset test score of the above preset model, the model performance of the preset model corresponding to the above training sample set is lower than the above preset performance; the above test score and the above preset test score are used to characterize the model performance of the above preset model; the above determining the second target sample set that causes the model performance of the above preset model to be lower than the above preset performance among the above plurality of first sample subsets after the training is completed includes: obtaining test samples, and respectively inputting the above test samples into the preset models corresponding to each of the above first sample subsets, so that the preset models corresponding to each of the above first sample subsets process the above test samples and obtain processing results; obtaining the processing results of the preset models corresponding to each of the above first sample subsets, and scoring the processing results of the preset models corresponding to each of the above first sample subsets to obtain the test scores of the preset models corresponding to each of the above first sample subsets; based on the test scores of the preset models corresponding to each of the above first sample subsets and the preset test score of the above preset model, determining the first sample subsets among the above plurality of first sample subsets whose corresponding preset model test scores are lower than the above preset test score as the above second target sample set.
[0014] Based on this implementation, the performance of the trained preset models corresponding to each sample subset (such as the first sample subset or the second sample subset) can be tested based on the test samples, so as to determine whether the model performance corresponding to the preset models trained based on each sample subset has decreased, thereby the target sample set (such as the first target sample set or the second target sample set) containing deteriorated samples can be determined among the plurality of sample subsets.
[0015] In yet another possible implementation, the above dividing the first target sample set into a plurality of first sample subsets based on a preset division value includes: determining a first combination mode corresponding to the plurality of training samples in the first target sample set based on the above preset division value, the preset sample quantity, or the minimum input quantity of the training samples corresponding to the above preset model, and the plurality of training samples in the above first target sample set; based on the above first combination mode, dividing the above first target sample set into the above plurality of first sample subsets.
[0016] Based on this implementation manner, when dividing the target sample set based on a preset division value, a preset sample quantity, or the minimum input quantity of training samples corresponding to a preset model, any combination method of multiple training samples in the target sample set can be determined, and thus the training samples can be divided into different sample subsets based on this combination method.
[0017] In one implementation manner, the above method further includes: if the second target sample set does not exist in the multiple first sample subsets, determining whether there is a second combination method corresponding to the first target sample set other than the first combination method based on the preset division value and the multiple training samples in the first target sample set; where the second combination method is an unused combination method; if the second combination method exists, re-dividing the first target sample set into multiple third sample subsets based on the second combination method, and continuing to train the preset model based on each of the multiple third sample subsets until the target sample set is determined or there is no unused combination method corresponding to the first target sample set; if the second combination method does not exist, determining all the training samples in the first target sample set as the deteriorated samples.
[0018] Based on this implementation manner, when the training sample set includes deteriorated samples, the target sample set can be divided based on different combination methods to determine a group of deteriorated samples that jointly affect the model performance of the preset model in the training sample set. Thus, it is possible to avoid the situation where the group of deteriorated samples cannot be determined due to dividing the target sample set based on one combination method, and further improve the accuracy of determining the deteriorated samples.
[0019] In another possible implementation manner, the above method further includes: deleting the deteriorated samples from the training sample set to obtain a target training sample set; training the preset model based on the target training sample set, and after the training of the preset model is completed, if the model performance of the preset model corresponding to the target training sample set is lower than the preset performance, continuing to determine the deteriorated samples in the target training sample set until the model performance of the preset model trained based on the training sample set after deleting the deteriorated samples is higher than or equal to the preset performance.
[0020] Based on this implementation manner, it is possible to ensure that all the deteriorated samples in the training sample set are determined, so that the preset model can be trained based on the training sample set after deleting all the deteriorated samples, so as to improve or maintain the stability of the performance of the preset model.
[0021] Second aspect, an embodiment of the present application further provides a method for email recognition, which includes: obtaining a training data set, and processing the training data set based on the processing method of the training samples in the first aspect above to determine the degraded email data in the training data set; wherein, the training data set includes multiple email data; deleting the degraded email data from the training data set to obtain a target training data set; training a to-be-trained email recognition model based on the training data set to obtain a trained email recognition model; obtaining the email data to be recognized, and recognizing the email data to be recognized based on the trained email recognition model to determine that the email data to be recognized is a malicious email or determine that the email data to be recognized is a non-malicious email.
[0022] Based on this implementation, the degraded email data in the training data set can be recognized, so as to obtain the target training data after deleting the degraded email data. Training the email recognition model based on this target training data can ensure the performance of the trained email recognition model, thereby improving the accuracy of recognizing malicious emails.
[0023] Third aspect, an embodiment of the present application further provides a processing device for training samples, which is characterized by including: a division module configured to obtain a first target sample set that causes the model performance of a preset model to be lower than a preset performance, and divide the first target sample set into multiple first sample subsets; a first training module configured to train the preset model based on each of the first sample subsets in the multiple first sample subsets, and determine a second target sample set that causes the model performance of the preset model to be lower than the preset performance in the multiple first sample subsets after the training is completed; a first determination module configured to, when the number of training samples in the second target sample set is less than or equal to a preset threshold, determine the training samples in the second target sample set as the degraded samples that cause the model performance of the preset model to be lower than the preset performance.
[0024] In a possible implementation, the division module is further configured to: when the number of training samples in the second target sample set is greater than the preset threshold, divide the second target sample set into multiple second sample subsets; the first training module is further configured to: train the preset model based on the sample numbers of the training samples in each of the second sample subsets in the multiple second sample subsets and the minimum input amount of the training samples corresponding to the preset model, and determine a third target sample set that causes the model performance of the preset model to be lower than the preset performance in the multiple second sample subsets after the training is completed.
[0025] In another possible implementation, the first training module is specifically configured as follows: If the number of training samples in each of the above second sample subsets is greater than or equal to the above minimum input amount, then train the above preset model based on each of the above second sample subsets; If there is a target sample subset in the above multiple second sample subsets where the number of training samples is less than the above minimum input amount, then determine the number of samples to be added corresponding to the target second sample subset based on the above minimum input amount and the number of training samples in the target sample subset; Obtain non-degraded samples with the same number of samples to be added as the target sample subset, and add the non-degraded samples to the target sample subset to obtain the multiple second sample subsets after addition, and train the above preset model based on the multiple second sample subsets after addition; where the number of training samples in the above target second sample subset is greater than or equal to the above minimum input amount.
[0026] In yet another possible implementation, the first training module is specifically configured as follows: Determine the number of positive samples and negative samples in the target sample subset; According to the preset sample ratio corresponding to the above preset model, the number of samples to be added corresponding to the target sample subset, the number of positive samples and negative samples in the target sample subset, determine the number of positive samples to be added and the number of negative samples to be added in the samples to be added corresponding to the target sample subset; where the above preset sample ratio is the ratio between the number of positive samples and the number of negative samples input into the above preset model; Obtain non-degraded positive samples with the same number of positive samples to be added as the target sample subset, and non-degraded negative samples with the same number of negative samples to be added as the target sample subset, and add the non-degraded positive samples and the non-degraded negative samples to the target sample subset to obtain the above multiple second sample subsets after addition.
[0027] In yet another possible implementation, the device further includes a second determination module; the first training module is further configured to: obtain a training sample set and train the preset model based on the training sample set; wherein the training sample set includes a plurality of training samples; the second determination module is configured to: after the training of the preset model is completed, if the model performance of the preset model corresponding to the training sample set is lower than the preset performance, determine the training sample set as the first target sample set; wherein, when the test score of the preset model corresponding to the training sample set is lower than the preset test score of the preset model, the model performance of the preset model corresponding to the training sample set is lower than the preset performance; the test score and the preset test score are used to characterize the model performance of the preset model; the first training module is specifically configured to: obtain test samples and input the test samples into the preset models corresponding to the respective first sample subsets, so that the preset models corresponding to the respective first sample subsets process the test samples and obtain processing results; obtain the processing results of the preset models corresponding to the respective first sample subsets, score the processing results of the preset models corresponding to the respective first sample subsets, and obtain the test scores of the preset models corresponding to the respective first sample subsets; based on the test scores of the preset models corresponding to the respective first sample subsets and the preset test score of the preset model, determine the first sample subsets whose test scores of the preset models corresponding thereto are lower than the preset test score among the plurality of first sample subsets as the second target sample set.
[0028] In yet another possible implementation, the partitioning module is specifically configured to: based on a preset partitioning value, a preset sample quantity, or the minimum input quantity of the training samples corresponding to the preset model, and the plurality of training samples in the first target sample set, determine a first combination mode corresponding to the plurality of training samples in the first target sample set; based on the first combination mode, partition the first target sample set into the plurality of first sample subsets.
[0029] In yet another possible implementation, the device further includes a third determination module; the third determination module is configured to: if the second target sample set does not exist in the multiple first sample subsets, determine whether there is a second combination mode corresponding to the first target sample set other than the first combination mode based on the preset division value and the multiple training samples in the first target sample set; wherein, the second combination mode is an unused combination mode; the division module is further configured to: if the second combination mode exists, re-divide the first target sample set into multiple third sample subsets based on the second combination mode, and continue to train the preset model based on each of the multiple third sample subsets until a target sample set is determined or there is no unused combination mode corresponding to the first target sample set; the first determination module is further configured to: if the second combination mode does not exist, determine all the training samples in the first target sample set as the deteriorated samples.
[0030] In yet another possible implementation, the device further includes a first deletion module; the first deletion module is configured to: delete the deteriorated samples from the training sample set to obtain a target training sample set; the first training module is further configured to: train the preset model based on the target training sample set, and after the training of the preset model is completed, if the model performance of the preset model corresponding to the target training sample set is lower than the preset performance, continue to determine the deteriorated samples in the target training sample set until the model performance of the preset model trained based on the training sample set after deleting the deteriorated samples is higher than or equal to the preset performance.
[0031] In a fourth aspect, an embodiment of the present application further provides a mail recognition device, the device includes: a processing module, configured to obtain a training data set and process the training data set based on a processing method of training samples to determine deteriorated mail data in the training data set; wherein, the training data set includes multiple mail data; a second deletion module, configured to delete the deteriorated mail data from the training data set to obtain a target training data set; a second training module, configured to train a mail recognition model to be trained based on the training data set to obtain a trained mail recognition model; an identification module, configured to obtain mail data to be identified and identify the mail data to be identified based on the trained mail recognition model to determine that the mail data to be identified is a malicious mail or determine that the mail data to be identified is a non-malicious mail.
[0032] Fifth aspect, an embodiment of the present application further provides a computing device, including: a processor and a memory; the processor and the memory are coupled; the memory is used to store program instructions; the processor is used to execute the program instructions to execute the method according to any one of the above first aspect, or the method according to any one of the above second aspect.
[0033] Sixth aspect, an embodiment of the present application provides a chip, and the chip is used to execute the method according to any one of the above first aspect, or the method according to any one of the above second aspect.
[0034] Seventh aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a computer, the method according to any one of the first aspect or the method according to any one of the above second aspect is implemented.
[0035] Eighth aspect, an embodiment of the present application provides a program product, including a computer program, and when the computer program is executed by a processor, the method according to any one of the first aspect or the method according to any one of the above second aspect is implemented. Description of the Drawings
[0036] Figure 1 It is a flowchart of a method for processing training samples provided by an embodiment of the present application;
[0037] Figure 2 It is a schematic diagram of a method for processing training samples provided by an embodiment of the present application;
[0038] Figure 3 It is a flowchart of a method for determining a first target sample set provided by an embodiment of the present application;
[0039] Figure 4 It is a flowchart of a method for training a preset model based on multiple second sample subsets provided by an embodiment of the present application;
[0040] Figure 5 It is a flowchart of a method for adding non-degraded samples provided by an embodiment of the present application;
[0041] Figure 6 It is a flowchart of a method for determining a first target sample set provided by an embodiment of the present application;
[0042] Figure 7 It is a flowchart of a method for determining a second target sample set provided by an embodiment of the present application;
[0043] Figure 8 It is a flowchart of another method for determining multiple first sample subsets provided by an embodiment of the present application;
[0044] Figure 9A flowchart for determining a deterioration sample provided by an embodiment of the present application;
[0045] Figure 10 A schematic diagram for determining a deterioration sample provided by an embodiment of the present application;
[0046] Figure 11 Another schematic diagram for determining a deterioration sample provided by an embodiment of the present application;
[0047] Figure 12 Another flowchart for determining a deterioration sample provided by an embodiment of the present application;
[0048] Figure 13 A flowchart of a method for processing email data for model training provided by an embodiment of the present application;
[0049] Figure 14 A scenario diagram of a method for processing email data for model training provided by an embodiment of the present application;
[0050] Figure 15 A flowchart for determining a first target data set provided by an embodiment of the present application;
[0051] Figure 16 A flowchart for determining a second target data set provided by an embodiment of the present application;
[0052] Figure 17 A flowchart of a method for email recognition provided by an embodiment of the present application;
[0053] Figure 18 A flowchart of another method for processing training samples provided by an embodiment of the present application;
[0054] Figure 19 A schematic diagram of a device for processing training samples provided by an embodiment of the present application;
[0055] Figure 20 A schematic diagram of a device for email recognition provided by an embodiment of the present application;
[0056] Figure 21 A schematic diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners
[0057] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. For the convenience of clearly describing the technical solutions in the embodiments of the present application, the first, second, etc. descriptions that appear in the embodiments of the present application are only for schematic and distinguishing the described objects, without an order, and do not represent a special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.
[0058] The processing method for training samples provided by the embodiments of the present application can be applied to a computing device, which can be a server or a terminal device. When training a preset model deployed in the computing device based on a training sample set, if the model performance of the trained preset model deteriorates, the training sample set can be recursively divided and the model can be trained to identify deteriorated samples (or called problem samples) that cause the model performance of the preset model to deteriorate in the training sample set. Thus, after removing the deteriorated samples from the training sample set, the preset model can be retrained based on the optimized training sample set (i.e., the training sample set without deteriorated samples) to improve the model performance of the preset model.
[0059] Figure 1 FIG. is a flowchart of a method for processing training samples provided by the embodiments of the present application. As Figure 1 shown, the method includes steps 110 to 140.
[0060] Step 110, obtain a first target sample set that causes the model performance of the preset model to be lower than the preset performance, and divide the first target sample set into multiple first sample subsets.
[0061] In some embodiments, the preset model can be a machine learning model (such as a support vector machine, a decision tree) or a large-scale pre-trained model based on deep learning (such as Transformer, BERT), etc. The specific type of the preset model is not limited in this embodiment. When training the preset model, if there are training samples (hereinafter referred to as deteriorated samples) in the training sample set that interfere with the model training, such as samples with incorrect labels, samples irrelevant to model training, and error samples, then after training the preset model based on this training sample set, the model performance of the preset model will decrease compared to the model performance before training based on this training sample set (hereinafter referred to as the preset performance). That is to say, after training the preset model based on the training sample set containing deteriorated samples, it will cause the model performance of the preset model to be lower than the preset performance (i.e., the preset model deteriorates).
[0062] In some embodiments, if the training sample set (such as the first target sample set) causes the model performance of the preset model to be lower than the preset performance, the deteriorated samples in the first target sample set can be determined by recursively dividing the first target sample set, and thus the deteriorated samples can be removed from the first target sample set. It should be noted that the first target sample set can be any sample set that causes the model performance of the preset model to be lower than the preset performance. For example, the first target sample set can be the training sample set, or any target sample set determined after one or more recursive divisions, etc., which is not limited here.
[0063] Exemplarily, the first target sample set can be divided into multiple first sample subsets based on a preset division value, or a preset number of samples, or the minimum input amount of the training samples corresponding to a preset model, etc. Among them, the value of the preset division value is a positive integer greater than or equal to 2; the preset number of samples refers to the number of training samples in the sample subsets obtained after division, and the preset number of samples is a positive integer greater than or equal to 1; the minimum input amount of the training samples corresponding to the preset model is a positive integer greater than or equal to 1. In some examples, to improve the speed of determining deteriorated samples, the first target sample set can be recursively divided by the dichotomy method, that is, at this time the preset division value is equal to 2, and after dividing the first target sample set, two first sample subsets are obtained.
[0064] Reference Figure 2 A schematic diagram of a method for processing training samples as shown in Figure 2 As shown, the training sample set can be divided into multiple sample subsets, such as sample subset A, sample subset B, sample subset C, sample subset D, etc., based on a preset division value, or a preset number of samples, or the minimum input amount of the training samples corresponding to a preset model. If the target sample set is divided based on the preset division value and the preset division value is equal to 2 (that is, the first target sample set is divided by the dichotomy method), then after dividing the target sample set based on the preset division value, the corresponding two sample subsets are obtained. For example, after dividing sample subset C by the dichotomy method, two sample subsets, sample subset A1 and sample subset A2, are obtained; after dividing sample subset A1 by the dichotomy method, two sample subsets, sample subset B1 and the first sample subset B2, are obtained.
[0065] Step 120: Train the preset model based on each of the multiple first sample subsets respectively, and after the training is completed, determine a second target sample set in the multiple first sample subsets that causes the model performance of the preset model to be lower than the preset performance.
[0066] In some embodiments, after obtaining the above multiple first sample subsets, each first sample subset is respectively input into the preset model to train the preset model based on each first sample subset, so as to obtain the preset model trained based on each first sample subset. Among them, when training the preset model based on each first sample subset respectively, the preset model is the preset model before being trained based on the training sample set.
[0067] Exemplarily, when training a preset model based on different training sample sets, the parameters of the preset model can be exported and saved after each training is completed. Thus, if the current training sample set causes the model performance of the preset model to decline, the model parameters saved after training based on the previous training sample set can be obtained, that is, the model parameters of the preset model before deterioration (hereinafter referred to as the preset model parameters). During the process of recursively partitioning the training sample set, the preset model parameters can be imported into the current preset model, so that the preset models trained based on each first sample subset are all those that have not deteriorated. In this way, the model performance of the preset models obtained after training based on each first sample subset can be compared with their preset performance to determine whether the performance of the preset model has declined.
[0068] In some embodiments, after determining the preset models obtained by training based on each first sample subset, the performance of the preset models corresponding to each first sample subset can be tested to determine the model performance of the preset models corresponding to each first sample subset. Exemplarily, the performance of the preset model can be tested by inputting test samples, and the embodiment does not limit the method of performance testing.
[0069] After determining the model performance of the preset models corresponding to each first sample subset, it can be determined whether there is a preset model whose model performance is lower than the preset performance. If so, the first sample subset corresponding to the preset model is determined as the second target sample set containing deteriorated samples. That is to say, the first sample subset that causes the model performance of the preset model to be lower than the preset performance is determined as the second target sample set.
[0070] As Figure 2 shown, after obtaining the first sample subset A1 and the first sample subset A2, the first sample subset A1 and the first sample subset A2 are respectively input into the preset model to train the preset model. If the model performance of the preset model obtained after training based on the first sample subset A1 is lower than the preset performance, while the model performance of the preset model obtained after training based on the first sample subset A2 is higher than (or equal to) the preset performance, the first sample subset A1 can be determined as the second target sample set containing deteriorated samples.
[0071] Step 130, when the number of training samples in the second target sample set is less than or equal to the preset threshold, the training samples in the second target sample set are determined as the deteriorated samples that cause the model performance of the preset model to be lower than the preset performance.
[0072] In some embodiments, to ensure the accuracy of determining the deteriorated samples, the value of the preset threshold can be set to 1. Exemplarily, when the second target sample set includes one training sample, that is, the second target sample set cannot be further divided, so that it can be determined that the training sample in the second target sample set is the deteriorated sample that causes the degradation of the model performance of the preset model.
[0073] As Figure 2 shown, if the finally determined target sample set containing deteriorated samples (such as sample subset X1) includes one training sample, the training sample in the sample subset X1 can be determined as the deteriorated sample.
[0074] Through the above solution, when the training sample set causes the degradation of the model performance of the preset model, the range of the deteriorated samples can be continuously narrowed by dividing the training sample set, so as to determine the deteriorated samples that cause the degradation of the model performance of the preset model, thereby improving the efficiency and accuracy of determining the deteriorated samples.
[0075] Figure 3 The following is a flowchart of a method for determining a first target sample set provided by an embodiment of the present application. As Figure 3 shown, the above method further includes steps 310 to 320.
[0076] Step 310, when the number of training samples in the second target sample set is greater than the preset threshold, divide the second target sample set into multiple second sample subsets.
[0077] In some embodiments, if the determined number in the second target sample set is greater than the preset threshold, for example, the second target sample set includes multiple training samples, then the second target sample set can be further divided to obtain multiple second sample subsets. As Figure 2 shown, if the second target sample set A1 includes multiple training samples, the second target sample set A1 can be further divided by the dichotomy method to obtain a second sample subset B1 and a second sample subset B2.
[0078] Step 320, based on the sample quantity of the training samples in each of the multiple second sample subsets and the minimum input quantity of the training samples corresponding to the preset model, train the preset model, and after the training is completed, determine in the multiple second sample subsets the third target sample set that causes the model performance of the preset model to be lower than the preset performance.
[0079] In some embodiments, then, the preset model is further trained based on each of the multiple second sample subsets, and after the training is completed, a third target sample set containing deteriorated samples is determined from the multiple second sample subsets based on the model performance of the preset model, and so on, until the number of training samples in the determined target sample set is less than or equal to a preset threshold.
[0080] As Figure 2 shown, if it is determined that the second sample subset B2 is the third target sample set containing deteriorated samples, the dichotomy method can be continued to divide the third target sample set B2, and so on, until the determined target sample set cannot be divided any further.
[0081] In some embodiments, when training the preset model based on each of the second sample subsets, it can be first determined whether the number of training samples in each of the second sample subsets meets the minimum input amount of the training samples corresponding to the preset model. It can be understood that the preset model usually has requirements for the minimum input amount of training samples. Therefore, the number of training samples input to the preset model each time needs to be greater than or equal to the minimum input amount of the training samples. The minimum input amount of the training samples is related to factors such as the model structure and the difficulty level of the training task. Thus, if the number of training samples in each of the second sample subsets in the multiple second sample subsets is greater than or equal to the target sample subset of the minimum input amount of the training samples corresponding to the preset model, the preset model is continued to be trained based on each of the second sample subsets; while if there is a target sample subset in the multiple second sample subsets where the number of training samples is less than the minimum input amount of the training samples corresponding to the preset model, the number of training samples in the target sample subset can be made greater than or equal to the minimum input amount of the training samples by adding training samples to the target sample subset. Then, the preset model is trained based on the multiple second sample subsets after the addition.
[0082] Through the above solution, when the number of training samples in the target sample set is greater than the preset threshold, the target sample set can be continuously divided to obtain multiple sample subsets; and when training the preset model based on the multiple sample subsets, it can be first ensured that the number of training samples in each of the second sample subsets meets the minimum input amount of the training samples corresponding to the preset model by adding training samples, and then the preset model is trained based on the multiple second sample subsets after the addition.
[0083] Figure 4 The figure is a flowchart of training a preset model based on multiple second sample subsets provided by an embodiment of the present application. As Figure 4 shown, the above step 320 includes steps 410 to 430.
[0084] Step 410: If the number of training samples in each second sample subset is greater than or equal to the minimum input amount, then train the preset model based on each second sample subset.
[0085] In some embodiments, after dividing the second target sample set into multiple second sample subsets, if the number of training samples in each second sample subset is greater than or equal to the minimum input amount of the training samples corresponding to the preset model, then the preset model can be continuously trained based on each second sample subset.
[0086] Step 420: If there is a target sample subset in the multiple second sample subsets where the number of training samples is less than the minimum input amount, then determine the number of samples to be added corresponding to the target sample subset based on the minimum input amount and the number of training samples in the target sample subset.
[0087] In some embodiments, after dividing the second target sample set into multiple second sample subsets, if there is a second sample subset (i.e., the target sample subset) in the multiple second sample subsets where the number of training samples is less than the minimum input amount, then to ensure that the number of training samples in each second sample subset among the multiple second sample subsets meets the minimum input amount of the training samples corresponding to the preset model, the number of training samples in the target sample subset can be first determined, and based on the minimum input amount of the training samples corresponding to the preset model and the number of training samples in the target sample subset, determine the number of training samples that need to be added to the target sample subset (i.e., the number of samples to be added corresponding to the target sample subset), so that the number of training samples in the target sample subset after adding the training samples is greater than or equal to the minimum input amount of the training samples corresponding to the preset model.
[0088] Step 430: Obtain non-degraded samples with the same number as the number of samples to be added corresponding to the target sample subset, and add the non-degraded samples to the target sample subset to obtain multiple added second sample subsets, and train the preset model based on the multiple added second sample subsets.
[0089] In some embodiments, after determining the number of samples to be added corresponding to the target sample subset, multiple high-quality training samples (hereinafter referred to as non-degraded samples) can be obtained, and the non-degraded samples are added to the target sample subset according to the number of samples to be added corresponding to the target sample subset. Among them, the non-degraded samples are training samples that will not cause the model performance of the preset model to be lower than the preset performance. It should be noted that the non-degraded samples added to different target sample subsets can be the same or different, that is, different non-degraded samples can be added to different target sample subsets, or the same degraded samples can be added.
[0090] Through the above solution, it is possible to ensure that the second sample subset meets the minimum input amount of the training samples corresponding to the preset model, and since the added samples are non-degraded samples, the determination process of the degraded samples will not be affected.
[0091] Figure 5 The figure is a flowchart of adding non-degraded samples provided by an embodiment of the present application. As Figure 5 shown, the above step 430 includes steps 510 to 530.
[0092] Step 510, determine the number of positive samples and the number of negative samples in the target sample subset.
[0093] In some embodiments, while ensuring that the sample subset can meet the minimum input amount of the training samples corresponding to the preset model by adding non-degraded samples, it can also be ensured that after adding the non-degraded samples, the ratio of positive and negative samples in the sample subset can meet the preset sample ratio corresponding to the preset model. Exemplarily, the number of positive samples and the number of negative samples in the above target sample subset can be determined first. Among them, the preset sample ratio corresponding to the preset model is the ratio between the number of positive samples input to the preset model and the number of negative samples during each training. Positive samples are samples that are consistent with the true sample labels, and negative samples are samples that are inconsistent with the true sample labels. For example, if the preset model is a model for identifying malicious emails (hereinafter referred to as the malicious email identification model), then the positive samples for training the malicious email identification model are malicious emails, and the negative samples are non-malicious emails. It can be understood that the positive and negative samples in the training sample set can be determined according to the label information corresponding to each training sample. It should be noted that the total number of training samples in each second sample subset can be equal or approximately the same.
[0094] Step 520, according to the preset sample ratio corresponding to the preset model, the number of samples to be added corresponding to the target sample subset, the number of positive samples and the number of negative samples in the target sample subset, determine the number of positive samples to be added and the number of negative samples to be added in the samples to be added corresponding to the target sample subset.
[0095] In some embodiments, after determining the number of positive samples and the number of negative samples existing in the target sample subset, and the number of samples to be added corresponding to the target sample subset, the number of positive samples to be added and the number of negative samples to be added corresponding to the target sample subset can be determined according to the preset sample ratio corresponding to the preset model.
[0096] Step 530, obtain non-degraded positive samples with the same number as the number of positive samples to be added corresponding to the target sample subset, and non-degraded negative samples with the same number as the number of negative samples to be added corresponding to the target sample subset, and add the non-degraded positive samples and non-degraded negative samples to the target sample subset to obtain multiple added second sample subsets.
[0097] In some embodiments, then, non-degraded positive samples and non-degraded negative samples can be added to the target sample subset according to the number of positive samples to be added and the number of negative samples to be added corresponding to the above target sample subset.
[0098] For example, taking the minimum input amount of the training samples corresponding to the preset model as 200 and the preset sample ratio corresponding to the preset model as positive sample:negative sample = 3:1 as an example, if the number of training samples in both the target sample subset A and the target sample subset B is 100, the number of positive samples in the target sample subset A is 80 and the number of negative samples is 20, and the number of positive samples in the target sample subset B is 70 and the number of negative samples is 30, then the number of non-degraded positive samples added to the target sample subset A can be 70, and the number of non-degraded negative samples can be 30; the number of non-degraded positive samples added to the target sample subset B can be 80, and the number of non-degraded negative samples can be 20.
[0099] Through the above solution, it is possible to make the number of training samples in the target sample subset meet the minimum input amount of the training samples corresponding to the preset model, and at the same time, it can also meet the preset sample ratio corresponding to the preset model, thereby avoiding the degradation of the model performance of the preset model caused by the training samples not meeting the minimum input amount or the preset sample ratio of the training samples corresponding to the preset model, and further affecting the determination of degraded samples.
[0100] Figure 6 The flowchart of a method for determining the first target sample set provided by the embodiments of the present application is as Figure 6 shown, and the above method further includes steps 610 to 620.
[0101] Step 610, obtain a training sample set and train a preset model based on the training sample set.
[0102] In some embodiments, during the process of training the preset model, the preset model can be trained based on multiple training sample sets in sequence to gradually optimize or stabilize the model performance of the preset model, so that the trained preset model can obtain accurate data processing results during the inference process. Among them, the training sample set includes multiple training samples.
[0103] Step 620, after the training of the preset model is completed, if the model performance of the preset model corresponding to the training sample set is lower than the preset performance, then determine the training sample set as the first target sample set.
[0104] In some embodiments, after the preset model is trained based on the current training sample set, the performance of the preset model can be tested to determine the model performance of the preset model corresponding to the training sample set. Exemplarily, the model performance of the preset model corresponding to the training sample set can be tested by test samples. For example, the test samples can be input into the preset model corresponding to the training sample set, and the processing result of the preset model can be obtained; then, the processing result can be scored to obtain a test score, and the test score can be compared with a preset test score. If the test score is lower than the preset test score, it indicates that the model performance of the preset model is lower than the preset performance, then the training sample set contains deteriorated samples. Among them, the model performance of the preset model can be the model performance of the preset model trained based on the previous training sample set; the preset test score can be the test score determined after testing the performance of the preset model trained based on the previous training sample set.
[0105] Exemplarily, in the case where the model performance of the preset model corresponding to the training sample set is lower than the preset performance, the training sample set can be determined as the first target sample set containing deteriorated samples. Then, the above steps 110 to 130 are executed on the first target sample set to determine the deteriorated samples in the current training sample set.
[0106] Through the above solution, after the preset model is trained based on the training sample set, the model performance of the preset model can be tested to determine whether the training sample set contains deteriorated samples.
[0107] Figure 7 The flowchart for determining the second target sample set provided by the embodiments of the present application is as Figure 7 shown, and the above step 120 includes steps 710 to 730.
[0108] Step 710, obtain test samples, and input the test samples into the preset models corresponding to each first sample subset respectively, so that the preset models corresponding to each first sample subset process the test samples and obtain processing results.
[0109] In some embodiments, after the corresponding trained preset model is obtained based on the above training sample set, first sample subset or second sample subset, the model performance of the preset model can be determined by inputting test samples. That is to say, test samples can be obtained and input into the above trained preset model, so that the trained preset model processes the test samples, thereby obtaining the processing result of the test samples. It can be understood that the test samples input into the preset model do not contain label information.
[0110] For example, if the preset model is a malicious email recognition model, the test samples can be multiple email data, which can include multiple malicious emails and multiple non-malicious emails. After inputting the multiple email data into the trained malicious email recognition model, the malicious email recognition model recognizes the malicious emails in the multiple email data, so as to obtain the recognition results of the malicious emails.
[0111] Exemplarily, taking the determination of the model performance of the preset model corresponding to each first sample subset as an example, the model performance of the preset model corresponding to each first sample subset can be tested by test samples respectively, so as to obtain the processing results of the preset model corresponding to each first sample subset for the test samples.
[0112] Step 720, obtain the processing results of the preset models corresponding to each first sample subset, and score the processing results of the preset models corresponding to each first sample subset to obtain the test scores of the preset models corresponding to each first sample subset.
[0113] In some embodiments, after the trained preset model finishes processing the test samples, the processing results of the trained preset model can be scored to obtain the test scores, so as to evaluate the model performance of the trained preset model based on the test scores. Exemplarily, taking the determination of the model performance of the preset model corresponding to each first sample subset as an example, the processing results of the preset models corresponding to each first sample subset for the test samples can be obtained, and the processing results of the preset models corresponding to each first sample subset can be scored, so as to obtain the test scores of the preset models corresponding to each first sample subset. Among them, the test scores are used to characterize the model performance of the preset model.
[0114] In some examples, the test scores of the trained preset model can be determined based on True Positive (TP), True Negative (TN), False Positive (FP), False Negative (FN), Accuracy, Precision, Recall, F1-Score, etc. corresponding to the processing results of the preset model.
[0115] For example, take the case where the preset model is a binary classification model, the number of test samples is 100, the number of positive samples is 30, and the number of negative samples is 70. If after the preset model corresponding to the first sample subset A1 processes this test sample, the obtained processing result is: the number of negative samples is 50, and the number of positive samples is 50, then it can be determined that the TP (the number of samples predicted as positive and actually positive) corresponding to this processing result = 30, TN (the number of samples predicted as negative and actually negative) = 50, FP (the number of samples predicted as positive but actually negative) = 20, FN (the number of samples predicted as negative but actually positive) = 0; furthermore, it can be determined that the accuracy corresponding to this processing result = the number of correctly predicted samples / the total number of samples = (TP + TN) / the total number of samples = (30 + 50) / 100 = 80%. Exemplarily, the test score of the preset model corresponding to the first sample subset A1 can be determined to be 80 points based on this accuracy rate.
[0116] If after the preset model corresponding to the first sample subset A2 processes this test sample, the obtained processing result is: the number of negative samples is 65, and the number of positive samples is 35, then it can be determined that the TP (the number of samples predicted as positive and actually positive) corresponding to this processing result = 30, TN (the number of samples predicted as negative and actually negative) = 65, FP (the number of samples predicted as positive but actually negative) = 5, FN (the number of samples predicted as negative but actually positive) = 0; furthermore, it can be determined that the accuracy corresponding to this processing result = the number of correctly predicted samples / the total number of samples = (TP + TN) / the total number of samples = (30 + 65) / 100 = 95%. Exemplarily, the test score of the preset model corresponding to the first sample subset A2 can be determined to be 95 points based on this accuracy rate.
[0117] It should be noted that the above specific method for determining the test score corresponding to the trained preset model is only an example, and this is not limited in this embodiment.
[0118] Step 730, based on the test scores of the preset models corresponding to each first sample subset and the preset test score of the preset model, determine the first sample subsets whose test scores of the preset models are lower than the preset test score among the multiple first sample subsets as the second target sample set.
[0119] In some embodiments, after obtaining the test scores corresponding to the preset model after the above training, the test scores corresponding to the preset model after training can be compared with the preset test scores of the preset model, and it can be determined whether the test scores corresponding to the preset model after training are lower than the preset test scores. Among them, the preset test score is used to characterize the preset performance of the preset model, that is, after training the preset model based on the previous training sample set, the preset test score obtained after testing the preset model corresponding to the previous training sample set based on the test samples; the method for determining the preset test score can refer to the method for predicting the score above, which will not be elaborated here.
[0120] Exemplarily, after obtaining the test scores of the preset models corresponding to the above-mentioned first sample subsets, based on the test scores of the preset models corresponding to the first sample subsets and the preset test scores of the preset model, the first sample subsets in which the test scores of the preset models corresponding to the multiple first sample subsets are lower than the preset test scores can be determined as the second target sample sets containing deteriorated samples.
[0121] Continuing with the above example, if the preset test score of the preset model is 90 points, the test score (80 points) of the preset model corresponding to the first sample subset A1 is lower than the preset test score of the preset model, and the test score (95 points) of the preset model corresponding to the first sample subset A2 is higher than the preset test score of the preset model, then the first sample subset A1 can be determined as the second target sample set containing deteriorated samples.
[0122] Through the above solution, the performance of the trained preset models corresponding to each sample subset (such as the first sample subset or the second sample subset) can be tested based on the test samples, so as to determine whether the model performance corresponding to the preset models trained based on each sample subset has decreased, so that the target sample sets (such as the first target sample set or the second target sample set) containing deteriorated samples can be determined among multiple sample subsets.
[0123] Figure 8 Another flowchart for determining multiple first sample subsets provided by the embodiments of the present application is as Figure 8 shown, and the above step 110 includes steps 810 to 820.
[0124] Step 810, based on the preset division value, the preset sample quantity, or the minimum input quantity of the training samples corresponding to the preset model, and the multiple training samples in the first target sample set, determine the first combination method corresponding to the multiple training samples in the first target sample set.
[0125] In some embodiments, the first target sample set can be divided into multiple first sample subsets according to any one of the division methods of the preset division value, the preset sample quantity, or the minimum input quantity of the training samples corresponding to the preset model.
[0126] Exemplarily, taking the division of the first target sample set based on a preset division value as an example, first, based on the preset division value and the first target sample set, the combination modes corresponding to multiple training samples in the first target sample set can be determined, so as to determine how to divide the multiple training samples in the first target sample set into multiple first sample subsets. Among them, there can be multiple combination modes corresponding to the multiple training samples in the first target sample set, and any one of the combination modes (such as the first combination mode) can be used to divide the multiple training samples in the first target sample set.
[0127] Exemplarily, taking the preset division value as 2, and the first target sample set A includes training sample a1, training sample a2, training sample a3, and training sample a4 as an example, then it is necessary to divide the above 4 training samples in the first target sample set A into two first sample subsets. Then, there can be multiple combination modes corresponding to the above 4 training samples. For example:
[0128] Combination mode 1: {training sample a1, training sample a2}, {training sample a3, training sample a4};
[0129] Combination mode 2: {training sample a1, training sample a3}, {training sample a2, training sample a4};
[0130] Combination mode 3: {training sample a1, training sample a4}, {training sample a2, training sample a3};
[0131] Combination mode 4: {training sample a1}, {training sample a2, training sample a3, training sample a4};
[0132] Combination mode 5: {training sample a2}, {training sample a1, training sample a3, training sample a4};
[0133] Combination mode 6: {training sample a3}, {training sample a1, training sample a2, training sample a4};
[0134] Combination mode 7: {training sample a4}, {training sample a1, training sample a2, training sample a3};
[0135] Therefore, any one of the combination modes 1 to 7 above can be selected as the first combination mode corresponding to the multiple training samples in the first target sample set A.
[0136] Step 820: Based on the first combination mode, divide the first target sample set into multiple first sample subsets.
[0137] In some embodiments, subsequently, based on the first combination method corresponding to multiple training samples in the above first target sample set, the first target sample set may be divided into multiple first sample subsets.
[0138] Continuing with the above example, taking the division of multiple training samples in the first target sample set A into the first sample subset A1 and the first sample subset A2 based on the above first combination method 1 as an example, the training samples in the obtained first sample subset A1 may include training sample a1 and training sample a2, and the training samples in the first sample subset A2 may include training sample a3 and training sample a4.
[0139] It should be noted that the above steps 810 to 820 are not only applicable to the first target sample set, but also applicable to all target sample sets determined in the recursive division process. For the method of dividing other target sample sets based on the combination method, reference may be made to the descriptions of the above steps 810 to 820, which will not be elaborated here.
[0140] Through the above solution, when dividing the target sample set based on the preset division value, the preset number of samples, or the minimum input amount of the training samples corresponding to the preset model, any combination method of multiple training samples in the target sample set can be determined, and thus the training samples can be divided into different sample subsets based on this combination method.
[0141] Figure 9 The following is a flowchart of a method for determining deteriorated samples provided by an embodiment of the present application. As Figure 9 shown, after the above step 120 "determine the second target sample set in multiple first sample subsets after training, which causes the model performance of the preset model to be lower than the preset performance", the above method further includes steps 910 to 930.
[0142] Step 910, if there is no second target sample set in multiple first sample subsets, determine whether there is a second combination method corresponding to the first target sample set other than the first combination method based on the preset division value and multiple training samples in the first target sample set.
[0143] Referring to Figure 10 a schematic diagram of a method for determining deteriorated samples shown in Figure 10 shown (in the figure, the black rectangles represent deteriorated samples, and the white rectangles represent non-deteriorated legal samples), in the case where there is one deteriorated sample in the training sample set, this deteriorated sample is the training sample that alone causes the model performance of the preset model to decline. Then, there must be a target sample set that causes the model performance of the preset model to decline in the sample subsets obtained after dividing each target sample set. Thus, by dividing the training sample set into sample subsets containing one training sample until the end, the deteriorated samples that cause the model performance of the preset model to decline can be determined.
[0144] Reference Figure 11 Another schematic diagram for determining a deteriorated sample is shown as Figure 11 shown (the shaded rectangles in the figure represent deteriorated samples, and the white rectangles represent non-deteriorated legal samples). However, in the case where the training sample set includes multiple deteriorated samples, it is possible that each of the deteriorated samples in the multiple deteriorated samples is a training sample that separately causes a decrease in the model performance of the preset model. In this case, it is only necessary to divide the training sample set until a sample subset containing one training sample is obtained; it is also possible that the multiple deteriorated samples are training samples that jointly cause a decrease in the model performance of the preset model (hereinafter referred to as a deteriorated sample group). In this case, during the recursive division of the training sample set, if the deteriorated samples in the deteriorated sample group are not all divided into the same sample subset, then it will occur that none of the divided sample subsets will cause a decrease in the model performance of the preset model, thus making it impossible to determine the target sample set containing deteriorated samples.
[0145] Therefore, if, after dividing the first target sample set into multiple first sample subsets based on the above first combination method and training the preset model, none of the multiple first sample subsets causes a decrease in the model performance of the preset model, that is, there is no second target sample set among the multiple first sample subsets, then in this case, the first target sample set can be re-divided based on other combination methods (such as the second combination method) until the target sample set is determined or all the combination methods for the first target sample set have been used. Among them, the second combination method is a combination method that has not been used.
[0146] Step 920, if there is a second combination method, then re-divide the first target sample set into multiple third sample subsets based on the second combination method, and continue to train the preset model based on each of the third sample subsets in the multiple third sample subsets until the target sample set is determined or there is no unused combination method corresponding to the first target sample set.
[0147] In some embodiments, in the case where there is an unused second combination method in the first target sample set, the first target sample set can be re-partitioned into multiple third sample subsets based on the second combination method, and the preset model can be trained respectively based on each of the third sample subsets in the multiple third sample subsets. If the deteriorated sample group is partitioned into the same third sample subset based on the second combination method, then a target sample set containing the deteriorated samples can be determined from the multiple third sample subsets; if the deteriorated sample group fails to be partitioned into the same third sample subset based on the second combination method, then a target sample set containing the deteriorated samples cannot be determined from the multiple third sample subsets. At this time, the first target sample set can be re-partitioned again based on other unused combination methods, and so on, until a target sample set is determined or there is no unused combination method corresponding to the first target sample set.
[0148] Continuing with the above example, if there is no second target sample set in the first sample subset A1 and the first sample subset A2 obtained by partitioning the first target sample set A based on the above first combination method 1, then other combination methods can be traversed and the unused combination method can be determined as the second combination method (such as the second combination method 2), and the first target sample set A can be re-partitioned based on the second combination method 2 to obtain the third sample subset A3 and the third sample subset A4.
[0149] If after training the preset model respectively based on the third sample subset A3 and the third sample subset A4, it is determined that the third sample subset A3 is the target sample set that causes the performance of the preset model to decline, then the dichotomy method continues to be used to partition the third sample subset A3. If after training the preset model respectively based on the third sample subset A3 and the third sample subset A4, neither of these two sample subsets causes the performance of the preset model to decline, then the first target sample set A can be re-partitioned again based on the unused combination method (such as the combination method 3).
[0150] Step 930, if there is no second combination method, then all the training samples in the first target sample set are determined as deteriorated samples.
[0151] In some embodiments, if there is no unused second combination method corresponding to the first target sample set, that is, all the combination methods corresponding to the first target sample set have been used to partition the first target sample set and the second target sample set has still not been determined, then it means that the first target sample set is a deteriorated sample group. Therefore, all the training samples in the first target sample set can be determined as deteriorated samples.
[0152] Such as Figure 11As shown, the last target sample set includes all the training samples in the deteriorated sample group. Therefore, after dividing the target sample set by all the corresponding combination methods, the obtained sample subsets cannot cause the performance degradation of the preset model. Thus, it can be determined that all the training samples in the target sample set are deteriorated samples.
[0153] Through the above solution, in the case where the training sample set includes deteriorated samples, the target sample set can be divided based on different combination methods to determine the deteriorated sample group that jointly affects the performance of the preset model in the training sample set. Thus, it can avoid the situation where the deteriorated sample group cannot be determined due to dividing the target sample set based on one combination method, and further improve the accuracy of determining deteriorated samples.
[0154] Figure 12 Another flowchart for determining deteriorated samples provided by an embodiment of this application is shown in Figure 12 As shown, the above method further includes steps 1210 to 1220.
[0155] Step 1210: Delete the deteriorated samples from the training sample set to obtain a target training sample set.
[0156] In some embodiments, after determining the deteriorated samples in the training sample set, the deteriorated samples are removed from the training sample set to obtain a target training sample set.
[0157] Step 1220: Train the preset model based on the target training sample set. After the training of the preset model is completed, if the performance of the preset model corresponding to the target training sample set is lower than the preset performance, continue to determine the deteriorated samples in the target training sample set until the performance of the preset model trained based on the training sample set after deleting the deteriorated samples is higher than or equal to the preset performance.
[0158] In some embodiments, after obtaining the above target training sample set, the preset model can be trained based on the target training sample set. If the performance of the preset model after training does not degrade, the preset model corresponding to the target training sample set can be used as the final preset model obtained in the current training round, and the preset model can be continuously trained in the next round based on other training sample sets.
[0159] However, if the training sample set contains both a deteriorated sample group and deteriorated samples that individually cause a decline in the performance of the preset model, the determined deteriorated samples may be the deteriorated samples that individually cause a decline in the performance of the preset model, and the deteriorated sample group may not be determined. Therefore, if the performance of the preset model still declines after training, the target training sample set after removing the deteriorated samples that individually cause a decline in the performance of the preset model can be screened for deteriorated samples again, so as to determine the deteriorated sample group in the target training sample set.
[0160] In some embodiments, the removed deteriorated samples can be stored in a database and can be used subsequently for periodic performance retesting of the preset model to verify whether the preset model can already correctly process these deteriorated samples without a decline in performance.
[0161] Through the above solution, it is possible to ensure that all deteriorated samples in the training sample set are determined, so that the preset model can be trained based on the training sample set after deleting all deteriorated samples, so as to improve or maintain the stability of the performance of the preset model.
[0162] In some embodiments, the above method for processing training samples can be applied to the processing of training samples of models such as email recognition models, SMS filtering models, malicious traffic recognition models, and malware detection models. The following takes the scenario of processing training samples of an email recognition model as an example for illustrative description.
[0163] Figure 13 It is a flowchart of a method for processing email data for model training provided by an embodiment of the present application. Figure 13 The method shown can be applied to a first server, such as Figure 13 shown, the method includes steps 1310 to step 1340.
[0164] Step 1310, obtain a first target data set that causes the performance of the email recognition model to be lower than the preset performance, and divide the first target data set into multiple first data subsets.
[0165] Refer to Figure 14 A scenario diagram of a method for processing email data for model training shown, such as Figure 14 shown, a mail security system is deployed on the first server, and the mail security system is used to process email data for model training. The second server and the third server are respectively communicatively connected to the first server. Among them, a mail recognition model is deployed on the second server to identify malicious emails in the email data; a mail gateway is deployed on the third server to obtain email data sent from an external network and send the email data to the first server.
[0166] In some embodiments, the first server may use the multiple email data sent by the third server as a training sample set (hereinafter referred to as the email data set) to train the email recognition model in the second server based on the email data set; and when the email data set causes a decline in the model performance of the email recognition model, recursively partition the email data set to identify the deteriorating email data therein.
[0167] In some embodiments, if the email data set (such as the first target data set) causes the model performance of the email recognition model to be lower than the preset performance, the deteriorating email data in the first target data set can be identified by recursively partitioning the first target data set, so as to remove the deteriorating email data from the first target data set. The first target data set includes multiple email data. Exemplarily, the first target data set can be divided into multiple first data subsets based on a preset partition value, a preset data quantity, or the minimum input amount of the training data corresponding to the email recognition model.
[0168] It can be understood that the implementation manner of step 1310 can refer to the description of step 110 and will not be elaborated here.
[0169] Step 1320: Send the multiple first data subsets to the email recognition model to train the email recognition model based on each of the first data subsets in the multiple first data subsets, and after the training is completed, determine the second target data set in the multiple first data subsets that causes the model performance of the email recognition model to be lower than the preset performance.
[0170] In some embodiments, after obtaining the multiple first data subsets, each first data subset is respectively input into the email recognition model to train the email recognition model based on each first data subset, so as to obtain the email recognition model trained based on each first data subset. When training the email recognition model based on each first sample data set, the email recognition model is the email recognition model before training based on the current email data set.
[0171] In some embodiments, after determining the email recognition models trained based on each first data subset, the model performance of the email recognition models corresponding to each first data subset can be tested to determine the model performance of the email recognition models corresponding to each first data subset. After determining the model performance of the email recognition models corresponding to each first data subset, it can be determined whether there is an email recognition model with a model performance lower than the preset performance. If so, the first data subset corresponding to the email recognition model is determined as the second target data set containing deteriorating email data. That is to say, the first data subset that causes the model performance of the email recognition model to be lower than the preset performance is determined as the second target data set.
[0172] It can be understood that the implementation of step 1320 can refer to the description of step 120 and will not be elaborated here.
[0173] Step 1330: When the number of email data in the second target dataset is less than or equal to a preset threshold, determine the email data in the second target dataset as deteriorated email data that causes the model performance of the email recognition model to be lower than the preset performance.
[0174] In some embodiments, to ensure the accuracy of determining deteriorated email data, the value of the preset threshold can be set to 1. Exemplarily, when the second target dataset includes one email data, it can be determined that the email data in the second target dataset is the deteriorated email data that causes the model performance of the email recognition model to decline.
[0175] It can be understood that the implementation of step 1330 can refer to the description of step 130 and will not be elaborated here.
[0176] In some embodiments, if the determined second target dataset includes multiple email data, then the second target dataset can continue to be divided based on a preset division value to obtain multiple data subsets, and then each data subset in the multiple data subsets is used to train the email recognition model, so as to determine the target dataset containing deteriorated email data based on the model performance of the email recognition model, and so on, until the determined target dataset includes one email data.
[0177] Through the above solution, when the email dataset causes the model performance of the preset model to decline, dividing the email dataset based on the email security system can continuously narrow the scope of the deteriorated email data to determine the deteriorated email data that causes the model performance of the email recognition model to decline, thereby improving the efficiency and accuracy of determining the deteriorated email data.
[0178] In some embodiments, when dividing the target dataset, the combination method of multiple email data in the target dataset can be determined based on a preset division value, a preset data quantity, or the minimum input amount of the training data corresponding to the email recognition model, and then the email data is divided into different data subsets based on the combination method. It can be understood that the method of dividing the target dataset based on different combination methods can refer to the description of steps 810 to 820 and will not be elaborated here.
[0179] In some embodiments, if, after partitioning the first target data set based on the first combination method to obtain multiple first data subsets and training the email recognition model, the model performance of the email recognition model does not degrade due to any of the multiple first sample subsets, that is, the second target data set does not exist in the multiple first data subsets, then in this case, the first target data set can be repartitioned based on another combination method (such as the second combination method) until the target data set is determined or all combination methods have been used to partition the first target data set. Herein, the second combination method is a combination method that has not been used.
[0180] In some embodiments, if there is no unused second combination method corresponding to the first target data set and the second target data set has still not been determined, then it means that the first target data set is a degraded data group. Therefore, all email data in the first target data set can be determined as degraded email data.
[0181] It can be understood that the method for determining the degraded data group as described above can refer to the descriptions in steps 910 to 930 and will not be elaborated herein.
[0182] Through the above solution, in the case where the email data set includes degraded email data, the target data set can be partitioned based on different combination methods to determine the degraded data group that jointly affects the model performance of the preset model in the email data. Thereby, it is possible to avoid the situation where the degraded data group cannot be determined due to partitioning the target data set based on one combination method, and further improve the accuracy of determining the degraded email data.
[0183] Figure 15 The following is a flowchart of a method for determining a first target data set provided by an embodiment of the present application. As Figure 15 shown, the method includes steps 1510 to 1530.
[0184] Step 1510: Obtain email data sent by a third server, and send multiple email data to an email recognition model to train the email recognition model based on the multiple email data.
[0185] In some embodiments, the email gateway in the third server can intercept emails entering the internal network (i.e., the network where the third server is located, and the first server, the second server, and the third server are in the same network), extract the email body, attachments, email headers, etc. from each email to obtain email data, and send the email data corresponding to each email to the first server. So that the first server can train the email recognition model in the second server based on the multiple email data.
[0186] In some embodiments, during the training of the email recognition model, the email recognition model can be trained based on multiple email datasets in sequence to gradually optimize or stabilize the model performance of the email recognition model, so that the trained email recognition model can obtain accurate email recognition results during the inference process.
[0187] Step 1520, after the training of the email recognition model is completed, if the model performance of the email recognition models corresponding to multiple email data is lower than the preset performance, then the multiple email data are determined as the first target dataset.
[0188] In some embodiments, after the training of the email recognition model based on the current email dataset is completed, the email security system can perform a performance test on the email recognition model to determine the model performance of the email recognition model corresponding to this email dataset. If the model performance of the email recognition model is lower than the preset performance, it means that the email dataset contains deteriorated email data. Among them, the model performance of the email recognition model can be the model performance of the email recognition model trained based on the previous email dataset.
[0189] In the case where the model performance of the email recognition model is lower than the preset performance, the multiple email data can be determined as the first target dataset containing deteriorated samples. Then, the above steps 1310 to 1330 are performed on the first target dataset to determine the deteriorated email data in the multiple email data.
[0190] Through the above solution, after the email recognition model is trained based on the email dataset, the model performance of the email recognition model can be tested to determine whether the email dataset contains deteriorated email data.
[0191] In some embodiments, after determining the first target dataset, the first target dataset can be divided into multiple second data subsets based on a preset division value; then, based on the minimum input amount of the email data corresponding to the email recognition model, the target data subset with the number of email data in the multiple second data subsets less than the minimum input amount is determined. The preset data ratio corresponding to the email recognition model is the ratio between the number of positive samples (i.e., malicious emails) and the number of negative samples (i.e., non-malicious emails) input into the email recognition model each time during training. It can be understood that the positive and negative samples in the email dataset can be determined according to the label information corresponding to each email data. Then, based on the minimum input amount of the email data corresponding to the email recognition model and the preset data ratio corresponding to the email recognition model, the number of email data to be added to the target data subset (i.e., the number of data to be added corresponding to the target data subset) is determined. It should be noted that the total number of email data in each second data subset can be equal or close.
[0192] After that, determine the number of malicious emails and the number of non-malicious emails in the email data to be added corresponding to the target data subset; finally, add the non-degraded malicious emails and non-degraded non-malicious emails to the target data subset according to the number of malicious emails and non-malicious emails to be added corresponding to the above target data subset, so as to obtain multiple second data subsets after addition. Among them, the non-degraded malicious emails and non-degraded non-malicious emails are email data that will not cause the model performance of the email recognition model to be lower than the preset performance.
[0193] It can be understood that the implementation method of training the preset model based on multiple second data subsets can refer to the descriptions in steps 410 to 430 and steps 510 to 530, which will not be elaborated here.
[0194] Through the above solution, while enabling the number of training data in each second data subset to meet the minimum input amount of the training data corresponding to the email recognition model, it can also meet the preset data ratio corresponding to the email recognition model, thereby avoiding the decline of the model performance of the email recognition model caused by the training data not meeting the minimum input amount or preset data ratio of the email data corresponding to the email recognition model, and further affecting the determination of degraded email data.
[0195] Figure 16 It is a flowchart of a method for determining a second target data set provided by an embodiment of the present application. As Figure 16 shown, the method includes steps 1610 to 1630.
[0196] Step 1610, obtain test email data, and send the test email data to the email recognition models corresponding to each first data subset respectively, so that the email recognition models corresponding to each first data subset identify whether the test email data is a malicious email and obtain the recognition results.
[0197] In some embodiments, after obtaining the corresponding trained email recognition model based on the above email data set, first data subset or second data subset, the model performance of the email recognition model can be determined by inputting test email data. Among them, the test email data may include multiple malicious emails and multiple non-malicious emails. After inputting the multiple email data into the trained malicious email recognition model, the malicious email recognition model identifies the malicious emails in the multiple email data, thereby obtaining the recognition results of the malicious emails.
[0198] Exemplarily, taking the determination of the model performance of the email recognition models corresponding to each first data subset as an example, the model performance of the email recognition models corresponding to each first data subset can be tested by test data respectively, so as to obtain the recognition results of the email recognition models corresponding to each first data subset for the test email data.
[0199] Step 1620: Obtain the recognition results of the email recognition models corresponding to each first data subset, and score the recognition results of the email recognition models corresponding to each first data subset to obtain the test scores of the email recognition models corresponding to each first data subset.
[0200] In some embodiments, after the trained email recognition model finishes processing the test email data, the recognition results of the trained email recognition model can be scored to obtain the test scores, so as to evaluate the model performance of the trained email recognition model based on the test scores. Exemplarily, taking the determination of the model performance of the email recognition models corresponding to each first data subset as an example, the recognition results of the email recognition models corresponding to each first data subset for the test email data can be obtained, and the recognition results of the email recognition models corresponding to each first data subset can be scored, so as to obtain the test scores of the email recognition models corresponding to each first data subset. Among them, the test scores are used to characterize the model performance of the email recognition models.
[0201] Step 1630: Based on the test scores of the email recognition models corresponding to each first data subset and the preset test scores of the email recognition models, determine the first data subsets in the multiple first data subsets whose test scores of the corresponding email recognition models are lower than the preset test scores as the second target data sets.
[0202] In some embodiments, after obtaining the test scores corresponding to the above-mentioned trained email recognition model, the test scores corresponding to the trained email recognition model can be compared with the preset test scores of the email recognition model, and it can be determined whether the test scores corresponding to the trained email recognition model are lower than the preset test scores. Among them, the preset test scores are used to characterize the preset performance of the email recognition model, that is, after training the email recognition model based on the previous email data set, the preset test scores obtained after testing the email recognition model corresponding to the previous email data set based on the test data; the method for determining the preset test scores can refer to the method for predicting scores above, and will not be elaborated here.
[0203] Exemplarily, after obtaining the test scores of the email recognition models corresponding to each of the above-mentioned first data subsets, based on the test scores of the email recognition models corresponding to each first data subset and the preset test scores of the email recognition model, the first data subsets in the multiple first data subsets whose test scores of the corresponding email recognition models are lower than the preset test scores can be determined as the second target data sets containing deteriorated email data.
[0204] It can be understood that the implementation manner of determining the second target data set can refer to the descriptions in steps 710 to 730, and will not be elaborated here.
[0205] Through the above solution, performance testing can be performed on the trained email recognition models corresponding to each data subset (such as the first data subset or the second data subset) based on the test data, so as to determine whether the model performance corresponding to the email recognition models trained based on each data subset has decreased, thereby enabling the determination of a target data set (such as the first target data set or the second target data set) containing deteriorated email data among multiple data subsets.
[0206] In some embodiments, after determining the above-mentioned deteriorated email data, the deteriorated email data can be deleted from the email data set and saved in the database. After obtaining the target email data set from which all deteriorated email data has been deleted, the email recognition model can be retrained based on the target email data set, thereby obtaining the final email recognition model obtained in the current training round. Additionally, the deteriorated email data in the database can be subsequently used for periodic performance retesting of the email recognition model to verify whether the email recognition model can correctly process this deteriorated data without a decrease in performance.
[0207] Figure 17 The flowchart of an email recognition method provided by an embodiment of the present application is as Figure 17 shown, and this method includes steps 1710 to 1740.
[0208] Step 1710, obtain a training data set (or a training sample set, etc.), and process the training data set to determine the deteriorated email data in the training data set.
[0209] In some embodiments, after the first server processes the training data set based on the processing method of the training samples provided in the above embodiments or the email data processing method for model training, the deteriorated email data in the training data set can be determined.
[0210] Step 1720, delete the deteriorated email data from the training data set to obtain a target training data set.
[0211] In some embodiments, after determining the deteriorated email data, the deteriorated email data can be deleted from the training data set, thereby obtaining a training data set that does not contain deteriorated email data, that is, the target training data set.
[0212] Step 1730, train the email recognition model to be trained based on the training data set to obtain a trained email recognition model.
[0213] In some embodiments, then, the first server can send the target training data set to the second server. After receiving the target training data set, the second server can train the email recognition model to be trained based on the target training data set and obtain a trained email recognition model.
[0214] Step 1740: Obtain the email data to be recognized, and recognize the email data to be recognized based on the trained email recognition model to determine whether the email data to be recognized is a malicious email or determine that the email data to be recognized is a non-malicious email.
[0215] In some embodiments, during the actual inference process, the email gateway in the third server can intercept the emails sent to the internal network, extract the email data in the emails, and then send it to the email security system in the first server; the email security system sends the received email data to the trained email recognition model deployed in the third server, so that the email recognition model can recognize the email data to determine whether the email corresponding to the email data is a malicious email; then, the email recognition model can send the email recognition result to the email security system.
[0216] If the email recognition result indicates that the email is a malicious email, the email security system filters the email (such as storing it in an isolation area), and can generate a malicious email processing log for the administrator to view; if the email recognition result indicates that the email is a non-malicious email, the email security system can send the email to the client for the user to receive and view the non-malicious email through the client, thereby avoiding the risk of the client being attacked by malicious emails.
[0217] Figure 18 It is a flowchart of another method for processing training samples provided by an embodiment of the present application. As Figure 18 shown, the method includes Step 1810 to Step 1830. The following combines Figure 18 to summarize the method for processing training samples provided by the present application.
[0218] Step 1801: Obtain a training sample set that causes the model performance of a preset model to decline.
[0219] It can be understood that the implementation manner of Step 1801 can refer to the description of Step 610, which will not be elaborated here.
[0220] Step 1802: After the training of the preset model is completed, if the model performance of the preset model corresponding to the training sample set is lower than the preset performance, determine the training sample set as the first target sample set.
[0221] It can be understood that the implementation manner of Step 1802 can refer to the description of Step 620, which will not be elaborated here.
[0222] Step 1803: Divide the first target sample set into multiple first sample subsets based on the target combination method, and train the preset model respectively based on each first sample subset in the multiple first sample subsets.
[0223] Exemplarily, the target combination modes include a first combination mode and a second combination mode. It can be understood that the implementation of step 1803 can refer to the descriptions of steps 110, 120, 810 to 820, which will not be elaborated here.
[0224] Step 1804: Determine whether there is a second target sample set in multiple first sample subsets that causes the performance of the preset model to decline.
[0225] Exemplarily, if there is a second target sample set in multiple first sample subsets that causes the performance of the preset model to decline, then execute step 1805; if there is no second target sample set in multiple first sample subsets that causes the performance of the preset model to decline, then execute step 1807.
[0226] It can be understood that the implementation of step 1804 can refer to the descriptions of steps 120, 710 to 730, which will not be elaborated here.
[0227] Step 1805: If there is a second target sample set in multiple first sample subsets that causes the performance of the preset model to decline, then determine whether the number of training samples in the second target sample set is less than or equal to a preset threshold.
[0228] Exemplarily, if the number of training samples in the second target sample set is less than or equal to the preset threshold, then execute step 1806; if the number of training samples in the second target sample set is greater than the preset threshold, then execute step 1803.
[0229] It can be understood that the implementation of step 1805 can refer to the description of step 130, which will not be elaborated here.
[0230] Step 1806: If the number of training samples in the second target sample set is less than or equal to the preset threshold, then determine the training samples in the second target sample set as deteriorated samples.
[0231] It can be understood that the implementation of step 1806 can refer to the description of step 130, which will not be elaborated here.
[0232] Step 1807: If there is no second target sample set in multiple first sample subsets that causes the performance of the preset model to decline, then determine whether all combination modes corresponding to the second target sample set have been used.
[0233] Exemplarily, if all combination modes corresponding to the second target sample set have been used to divide the second target sample set, then execute step 1808; if not all combination modes corresponding to the second target sample set have been used to divide the second target sample set, then execute step 1803.
[0234] It can be understood that the implementation of step 1807 can refer to the descriptions of step 910 and step 920, which will not be elaborated here.
[0235] Step 1808, if all combination methods corresponding to the second target sample set have been used to divide the second target sample set, then all training samples in the second target sample set are determined as deteriorated samples.
[0236] It can be understood that the implementation of step 1808 can refer to the description of step 930, which will not be elaborated here.
[0237] Step 1809, delete the deteriorated samples in the training sample set to obtain a target sample set, and retrain the preset model based on the target sample set.
[0238] Exemplarily, after executing step 1806 and step 1808, step 1809 can be executed. It can be understood that the implementation of step 1809 can refer to the descriptions of step 1210 to step 1220, which will not be elaborated here.
[0239] Step 1810, determine whether the model performance of the trained preset model has decreased.
[0240] It can be understood that the implementation of step 1810 can refer to the description of step 1220, which will not be elaborated here.
[0241] Step 1811, continue to train the preset model based on the next training sample set.
[0242] It can be understood that the implementation of step 1811 can refer to the description of step 1220, which will not be elaborated here.
[0243] Applying the technical solution provided by the present application, in the case where the training sample set causes the performance of the preset model to decrease, it is possible to determine the deteriorated samples that cause the model performance of the preset model to decrease by a method for the training sample set, so as to avoid the computational overhead of manual screening item by item and full retraining, and improve the efficiency and accuracy of determining deteriorated samples. Moreover, the method for processing training samples provided by the present application can be applied to the processing of model training samples in different scenarios, and has wide applicability.
[0244] Figure 19 It is a schematic diagram of a device for processing training samples provided by an embodiment of the present application. As Figure 19 shown, the device 1900 for processing training samples includes a division module 1901, a first training module 1902, and a first determination module 1903.
[0245] The partitioning module 1901 is configured to obtain a first target sample set that causes the model performance of a preset model to be lower than a preset performance, and partition the first target sample set into multiple first sample subsets.
[0246] The first training module 1902 is configured to train the preset model respectively based on each first sample subset in the multiple first sample subsets, and determine a second target sample set that causes the model performance of the preset model to be lower than the preset performance in the multiple first sample subsets after the training is completed.
[0247] The first determination module 1903 is configured to determine the training samples in the second target sample set as deterioration samples that cause the model performance of the preset model to be lower than the preset performance when the number of training samples in the second target sample set is less than or equal to a preset threshold.
[0248] In some embodiments, the partitioning module 1901 is further configured to: when the number of training samples in the second target sample set is greater than the preset threshold, partition the second target sample set into multiple second sample subsets; the first training module 1902 is further configured to: train the preset model based on the number of training samples in each second sample subset in the multiple second sample subsets and the minimum input amount of the training samples corresponding to the preset model, and determine a third target sample set that causes the model performance of the preset model to be lower than the preset performance in the multiple second sample subsets after the training is completed.
[0249] In some embodiments, the first training module 1902 is specifically configured to: if the number of training samples in each second sample subset is greater than or equal to the minimum input amount, train the preset model based on each second sample subset; if there is a target sample subset in the multiple second sample subsets where the number of training samples is less than the minimum input amount, determine the number of samples to be added corresponding to the target second sample subset based on the minimum input amount and the number of training samples in the target sample subset; obtain non-deterioration samples with the same number of samples to be added corresponding to the target sample subset, and add the non-deterioration samples to the target sample subset to obtain multiple added second sample subsets, and train the preset model based on the multiple added second sample subsets; wherein, the number of training samples in the target second sample subset is greater than or equal to the minimum input amount.
[0250] In some embodiments, the first training module 1902 is specifically configured to: determine the number of positive samples and the number of negative samples in the target sample subset; determine the number of positive samples to be added and the number of negative samples to be added in the samples to be added corresponding to the target sample subset according to the preset sample ratio corresponding to the preset model, the number of samples to be added corresponding to the target sample subset, the number of positive samples and the number of negative samples in the target sample subset; wherein the preset sample ratio is the ratio between the number of positive samples and the number of negative samples input into the preset model; obtain non-degraded positive samples with the same number as the number of positive samples to be added corresponding to the target sample subset, and non-degraded negative samples with the same number as the number of negative samples to be added corresponding to the target sample subset, and add the non-degraded positive samples and the non-degraded negative samples to the target sample subset to obtain multiple second sample subsets after addition.
[0251] As Figure 19 shown, the processing device 1900 for training samples further includes: a second determination module 1904.
[0252] In some embodiments, the first training module 1902 is further configured to: obtain a training sample set and train the preset model based on the training sample set; wherein the training sample set includes multiple training samples; the second determination module 1904 is configured to: after the training of the preset model is completed, if the model performance of the preset model corresponding to the training sample set is lower than the preset performance, determine the training sample set as the first target sample set; wherein, in the case that the test score of the preset model corresponding to the training sample set is lower than the preset test score of the preset model, the model performance of the preset model corresponding to the training sample set is lower than the preset performance; the test score and the preset test score are used to characterize the model performance of the preset model; the first training module 1902 is specifically configured to: obtain test samples and input the test samples into the preset models corresponding to the respective first sample subsets respectively, so that the preset models corresponding to the respective first sample subsets process the test samples and obtain processing results; obtain the processing results of the preset models corresponding to the respective first sample subsets, and score the processing results of the preset models corresponding to the respective first sample subsets to obtain the test scores of the preset models corresponding to the respective first sample subsets; based on the test scores of the preset models corresponding to the respective first sample subsets and the preset test score of the preset model, determine the first sample subsets among the multiple first sample subsets whose test scores of the corresponding preset models are lower than the preset test score as the second target sample set.
[0253] In some embodiments, the partitioning module 1901 is specifically configured to: determine the first combination mode corresponding to the multiple training samples in the first target sample set based on a preset partitioning value, a preset number of samples, or the minimum input amount of the training samples corresponding to the preset model, and the multiple training samples in the first target sample set; based on the first combination mode, partition the first target sample set into multiple first sample subsets.
[0254] As shown Figure 19 in FIG. Figure 19 , the processing apparatus 1900 for training samples further includes: a third determination module 1905.
[0255] In some embodiments, the third determination module 1905 is configured to: if there is no second target sample set in the multiple first sample subsets, determine whether there is a second combination method other than the first combination method corresponding to the first target sample set based on a preset division value and multiple training samples in the first target sample set; wherein, the second combination method is an unused combination method; the division module 1901 is further configured to: if there is a second combination method, re-divide the first target sample set into multiple third sample subsets based on the second combination method, and continue to train a preset model based on each third sample subset in the multiple third sample subsets until a target sample set is determined or there is no unused combination method corresponding to the first target sample set; the first determination module 1903 is further configured to: if there is no second combination method, determine all training samples in the first target sample set as degraded samples.
[0256] As shown Figure 19 in FIG. Figure 19 , the processing apparatus 1900 for training samples further includes: a first deletion module 1906.
[0257] In some embodiments, the first deletion module 1906 is configured to: delete the degraded samples from the training sample set to obtain a target training sample set; the first training module 1902 is further configured to: train a preset model based on the target training sample set, and after the training of the preset model is completed, if the model performance of the preset model corresponding to the target training sample set is lower than a preset performance, continue to determine the degraded samples in the target training sample set until the model performance of the preset model trained based on the training sample set after deleting the degraded samples is higher than or equal to the preset performance.
[0258] Figure 20 is a schematic diagram of a mail recognition apparatus provided by an embodiment of the present application. As shown Figure 20 in FIG. Figure 20 , the mail recognition apparatus 2000 includes a processing module 2001, a second deletion module 2002, a second training module 2003, and an identification module 2004.
[0259] The processing module 2001 is configured to obtain a training data set and process the training data set based on the processing method of training samples to determine the degraded mail data in the training data set.
[0260] Wherein, the training data set includes multiple mail data.
[0261] The second deletion module 2002 is configured to delete the degraded mail data from the training data set to obtain a target training data set.
[0262] The second training module 2003 is configured to train the mail recognition model to be trained based on a training data set, and obtain a trained mail recognition model.
[0263] The recognition module 2004 is configured to obtain mail data to be recognized, and recognize the mail data to be recognized based on the trained mail recognition model, so as to determine that the mail data to be recognized is a malicious mail or determine that the mail data to be recognized is a non-malicious mail.
[0264] Figure 21 A schematic diagram of a computing device provided by some embodiments of the present application. In some embodiments, the computing device may be a server, a terminal device, etc. The computing device includes one or more processors and a memory. The memory is configured to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the processing method of training samples or the mail recognition method in the above embodiments.
[0265] As Figure 21 shown, the computing device 2100 includes: a processor 2101 and a memory 2102. Exemplarily, the computing device 2100 may further include: a communication interface 2103 and a communication bus 2104.
[0266] Wherein, the processor 2101, the memory 2102 and the communication interface 2103 complete mutual communication through the communication bus 2104. The communication interface 2103 is used to communicate with network elements of other devices such as clients or other servers.
[0267] In some embodiments, the processor 2101 is used to execute the program 2105, and specifically may execute the relevant steps in the above embodiments of the processing method of training samples or the mail recognition method. Specifically, the program 2105 may include program code, and the program code includes computer-executable instructions.
[0268] Exemplarily, the processor 2101 may be a central processing unit CPU, or a specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits configured to implement some embodiments of the present application. The one or more processors that the computing device 2100 may include may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0269] In some embodiments, the memory 2102 is used to store the program 2105. The memory 2102 may include high-speed RAM memory and may also include non-volatile memory (NVM), such as at least one disk memory.
[0270] Specifically, the program 2105 can be called by the processor 2101 to cause the computing device 2100 to execute operations of the processing method for training samples or the mail recognition method.
[0271] Some embodiments of the present application provide a computer-readable storage medium that stores at least one executable instruction. When the executable instruction runs on the computing device 2100, it causes the computing device 2100 to execute the processing method for training samples or the mail recognition method in the above embodiments.
[0272] Specifically, the executable instruction can be used to cause the computing device 2100 to execute operations of the processing method for training samples or the mail recognition method.
[0273] For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.
[0274] The beneficial effects that can be achieved by the readable storage medium provided in some embodiments of the present application can refer to the beneficial effects in the corresponding processing method for training samples or the mail recognition method provided above, and will not be elaborated here.
[0275] It should be noted that in the application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0276] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description in the method embodiments.
[0277] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices.
[0278] For this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0279] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM).
[0280] In addition, a computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory. It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof.
[0281] In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0282] The above-described embodiments are only specific embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present application shall be included within the protection scope of the present application.
Claims
1. A method for processing training samples, characterized in that, Including: Obtain a first target sample set that causes the model performance of a preset model to be lower than a preset performance, and divide the first target sample set into multiple first sample subsets; Based on each of the first sample subsets in the multiple first sample subsets, train the preset model respectively, and after the training is completed, determine a second target sample set in the multiple first sample subsets that causes the model performance of the preset model to be lower than the preset performance; When the number of training samples in the second target sample set is less than or equal to a preset threshold, determine the training samples in the second target sample set as deterioration samples that cause the model performance of the preset model to be lower than the preset performance.
2. The method according to claim 1, characterized in that, The method further includes: When the number of training samples in the second target sample set is greater than the preset threshold, divide the second target sample set into multiple second sample subsets; Based on the number of training samples in each of the second sample subsets in the multiple second sample subsets and the minimum input amount of the training samples corresponding to the preset model, train the preset model, and after the training is completed, determine a third target sample set in the multiple second sample subsets that causes the model performance of the preset model to be lower than the preset performance.
3. The method according to claim 2, wherein The training the preset model based on the number of training samples in each of the second sample subsets in the multiple second sample subsets and the minimum input amount of the training samples corresponding to the preset model includes: If the number of training samples in each of the second sample subsets is greater than or equal to the minimum input amount, train the preset model based on each of the second sample subsets; If there is a target sample subset in the multiple second sample subsets where the number of training samples is less than the minimum input amount, determine the number of samples to be added corresponding to the target second sample subset based on the minimum input amount and the number of training samples in the target sample subset; Obtain non-deterioration samples with the same number of samples to be added corresponding to the target sample subset, and add the non-deterioration samples to the target sample subset to obtain multiple added second sample subsets, and train the preset model based on the multiple added second sample subsets; wherein, the number of training samples in the target second sample subset is greater than or equal to the minimum input amount.
4. The method according to claim 3, wherein The obtaining non-deterioration samples with the same number of samples to be added corresponding to the target sample subset and adding the non-deterioration samples to the target sample subset to obtain multiple added second sample subsets includes: Determine the number of positive samples and negative samples in the target sample subset; According to the preset sample ratio corresponding to the preset model, the number of samples to be added corresponding to the target sample subset, the number of positive samples and negative samples in the target sample subset, determine the number of positive samples to be added and the number of negative samples to be added in the samples to be added corresponding to the target sample subset; wherein, the preset sample ratio is the ratio between the number of positive samples and the number of negative samples input into the preset model. Obtain non-degraded positive samples with the same number as the positive samples to be added corresponding to the target sample subset, and non-degraded negative samples with the same number as the negative samples to be added corresponding to the target sample subset, and add the non-degraded positive samples and the non-degraded negative samples to the target sample subset to obtain the multiple second sample subsets after addition.
5. The method according to claim 1, wherein The method further includes: Obtain a training sample set and train the preset model based on the training sample set; wherein, the training sample set includes multiple training samples; After the training of the preset model is completed, if the model performance of the preset model corresponding to the training sample set is lower than the preset performance, then determine the training sample set as the first target sample set; wherein, when the test score of the preset model corresponding to the training sample set is lower than the preset test score of the preset model, the model performance of the preset model corresponding to the training sample set is lower than the preset performance; the test score and the preset test score are used to characterize the model performance of the preset model; The determining the second target sample set that causes the model performance of the preset model to be lower than the preset performance among the multiple first sample subsets after training includes: Obtain test samples and input the test samples into the preset models corresponding to each of the first sample subsets respectively, so that the preset models corresponding to each of the first sample subsets process the test samples and obtain processing results; Obtain the processing results of the preset models corresponding to each of the first sample subsets, and score the processing results of the preset models corresponding to each of the first sample subsets to obtain the test scores of the preset models corresponding to each of the first sample subsets; Based on the test scores of the preset models corresponding to each of the first sample subsets and the preset test score of the preset model, determine the first sample subsets among the multiple first sample subsets whose corresponding preset model test scores are lower than the preset test score as the second target sample set.
6. The method according to any one of claims 1-5, characterized in that, The dividing the first target sample set into multiple first sample subsets includes: Based on a preset division value, a preset sample quantity, or the minimum input quantity of the training samples corresponding to the preset model, and the multiple training samples in the first target sample set, determine the first combination mode corresponding to the multiple training samples in the first target sample set; Based on the first combination mode, divide the first target sample set into the multiple first sample subsets.
7. The method according to claim 6, wherein The method further includes: If the second target sample set does not exist in the multiple first sample subsets, then based on the preset division value and the multiple training samples in the first target sample set, determine whether there is a second combination mode corresponding to the first target sample set other than the first combination mode; wherein, the second combination mode is an unused combination mode; If the second combination mode exists, re-partition the first target sample set into multiple third sample subsets based on the second combination mode, and continue to train the preset model based on each of the third sample subsets in the multiple third sample subsets until the target sample set is determined or there is no unused combination mode corresponding to the first target sample set; If the second combination mode does not exist, determine all the training samples in the first target sample set as the deteriorated samples.
8. The method according to claim 5, wherein The method further includes: Delete the deteriorated samples from the training sample set to obtain a target training sample set; Train the preset model based on the target training sample set, and after the training of the preset model is completed, if the model performance of the preset model corresponding to the target training sample set is lower than the preset performance, continue to determine the deteriorated samples in the target training sample set until the model performance of the preset model trained based on the training sample set after deleting the deteriorated samples is higher than or equal to the preset performance.
9. A method for email recognition, characterized in that, It includes: Obtain a training data set, and process the training data set based on the processing method of the training samples according to any one of claims 1-8 to determine the deteriorated email data in the training data set; wherein, the training data set includes multiple email data; Delete the deteriorated email data from the training data set to obtain a target training data set; Train the email recognition model to be trained based on the target training data set to obtain a trained email recognition model; Obtain the email data to be recognized, and recognize the email data to be recognized based on the trained email recognition model to determine that the email data to be recognized is a malicious email or determine that the email data to be recognized is a non-malicious email.
10. A computing device, characterized in that, It includes: One or more processors; And A memory configured to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the processing method of the training samples according to any one of claims 1-8, or implement the email recognition method according to claim 9.