Classification model updating method and apparatus, and device and storage medium
By introducing the second classifier and Pearson correlation coefficient optimization feature group in the network flow classification model, the open current collection recognition problem is solved, the classification accuracy and speed are improved, and the effective identification of new categories is achieved.
Patent Information
- Application Number
- PCT/CN2024/135037
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-02
- Filing Date
- 2024-11-27
- Publication Date
- 2025-07-10
AI Technical Summary
The prior art has problems with open current collection identification in network flow classification, resulting in low classification accuracy and long time to generate negative samples, making it difficult to effectively identify new categories.
By introducing a second classifier into the first classification model, we can identify known classes and new class samples, optimize the statistical feature group using Pearson correlation coefficient, reduce the sample size and train the second classifier to achieve rapid identification of new classes.
It improves the accuracy and speed of the classification model, can effectively identify new categories, reduces computing resource consumption, and improves the efficiency of network flow data classification.
Smart Images

Figure CN2024135037_10072025_PF_FP_ABST
Abstract
Description
A classification model updating method, device, equipment and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on January 2, 2024, with application number 202410001999.5 and invention name "A classification model updating method, device, equipment and storage medium", the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of computer technology, and in particular to a classification model updating method, apparatus, device and storage medium. Background Art
[0004] Network Traffic Classification (NTC) classifies and identifies network flow data collected from various applications, playing a key role in ensuring Quality of Service (QoS), network security, and traffic trend analysis. NTC can achieve high classification accuracy when the test set consists of known classes. However, in reality, the network environment is dynamically changing, and classes not included in the training set may frequently appear during testing. This problem is known as the Open Set Flow Recognition (OSFR) problem.
[0005] Existing techniques, on the one hand, can generate negative samples of known classes by finding data close to training instances, and then locate the decision boundary based on the known classes and negative samples to distinguish known classes from new classes. However, the negative samples generated by this method are highly similar to the target samples, which may reduce the classification accuracy of the known classes. In addition, generating negative samples is time-consuming. On the other hand, Extreme Value Theory (EVT) is used to model positive training samples at the decision boundary. If the known class data is accurately modeled, the new class can be rejected. This allows convolutional neural networks and random forest models to be used to combine voting patterns with EVT to fit the Weibull distribution, and to reject new class samples using a threshold. However, experimental results show that the positive sample data does not fully fit the Weibull distribution, and the recall rate of the known class is lower than expected.
[0006] In summary, for the OSFR problem, the above methods all have the problem of low classification accuracy. Summary of the Invention
[0007] The embodiments of the present application provide a classification model updating method, apparatus, device and storage medium for improving the classification accuracy of the classification model.
[0008] In a first aspect, an embodiment of the present application provides a classification model updating method, the method comprising:
[0009] A first new class sample set is determined from a test sample set based on a first classifier in a first classification model; the first classifier is used to classify samples into first known class samples belonging to a recognizable first known class or first new class samples belonging to an unrecognizable first new class; a sample includes a value of at least one statistical feature corresponding to a piece of network flow data;
[0010] training a second classifier based on the first new class sample set, wherein the second classifier is configured to output types of samples other than the first new class samples and the first known class samples as an unrecognizable second new class;
[0011] The second classifier is added to the first classification model to obtain a second classification model, where the second classification model is used to classify the sample type into any one of the following: the first known class, the first new class, and the second new class.
[0012] In this scheme, the first classifier can identify that the sample belongs to the first known class sample, and based on the first classifier, the first new class sample set is determined from the test sample set, and the second classifier is trained based on the first new class sample set, so that the second classifier can identify that the type of the sample is the first new class, and the second classifier is added to the first classification model to obtain the second classification model. Since the second classification model contains the first classifier and the second classifier, the first classifier can identify that the type of the sample is the first known class, and the second classifier can identify that the type of the sample is the first new class. In this way, when the sample type is other than the first known class and the first new class (that is, the OSFR problem), the second classification model can classify it into the second new class. Compared with the first classifier outputting the sample type as the first known class or the first new class, the classification accuracy of the second classifier is higher.
[0013] Optionally, training a second classifier based on the first new class sample set includes: taking the first new class sample set and the first sample set as a second known class sample set, and training the second classifier based on the second known class sample set; the first sample set includes multiple first known class samples.
[0014] Through this method, the second classifier can identify the type of the sample as the second known class. In this way, when a sample type other than the second known class appears, the second classifier can quickly classify it into the second new class, thereby improving the classification speed of the second classification model.
[0015] Optionally, the method further includes: obtaining values of multiple statistical features corresponding to each of multiple network flow data to obtain multiple statistical feature sequences, wherein the values in a statistical feature sequence correspond to the same statistical feature; determining the Pearson correlation coefficient between any two statistical feature sequences in the multiple statistical feature sequences; determining a target statistical feature group from the multiple statistical features; wherein the target statistical feature group includes at least two statistical features, the Pearson correlation coefficient between two statistical feature sequences corresponding to any two statistical features in the at least two statistical features is less than a first threshold, and the time complexity of each statistical feature in the at least two statistical features meets a preset condition; and taking the value of each feature in the target statistical feature group corresponding to a network flow data as a sample.
[0016] Through this method, a piece of network flow data corresponds to multiple statistical features. If the values of all statistical features corresponding to a piece of network flow data are taken as a sample, the classification time of the classification model will be longer. According to the values of multiple statistical features corresponding to each piece of network flow data in the multiple network flow data, multiple statistical feature sequences are obtained. According to the Pearson correlation coefficient between any two statistical feature sequences, the target statistical feature group is determined from the multiple statistical features. In this way, the statistical features that make up the sample are fewer, the sample size is reduced, and the classification speed of the second classification model is improved.
[0017] Optionally, the first known class includes at least one subclass, and one network flow data corresponds to one subclass. The determining of the target statistical feature group from the multiple statistical features includes: obtaining a subclass sequence according to the subclass corresponding to each network flow data; determining the Pearson correlation coefficient between each statistical feature sequence and the subclass sequence; taking two statistical features whose Pearson correlation coefficient between the corresponding two statistical feature sequences is greater than or equal to the first threshold as a group of candidate statistical features, to obtain at least one group of candidate statistical features from the multiple statistical features; determining a statistical feature with a smaller Pearson correlation coefficient between the corresponding statistical feature sequence in each group of candidate statistical features and the subclass sequence, to obtain at least one statistical feature to be deleted; deleting the at least one statistical feature to be deleted from the multiple statistical features, and determining a statistical feature whose time complexity meets a preset condition from the remaining statistical features, to obtain the target statistical feature group.
[0018] Through this method, since the two statistical features whose Pearson correlation coefficient between the corresponding two statistical feature sequences is greater than or equal to the first threshold are relatively similar, for such a group of selected statistical features, only one of them can be retained, and the statistical features with the smaller Pearson correlation coefficient between the corresponding statistical feature sequence and the subclass sequence in each group of selected statistical features are determined to obtain at least one statistical feature to be deleted, and at least one statistical feature to be deleted is deleted from multiple statistical features, so that the remaining statistical features have a large influence on the sample category, and then the statistical features whose time complexity meets the preset conditions are determined from the remaining statistical features to obtain the target statistical feature group, so that a sample can well represent a network flow data, and the classification accuracy of the classification model trained by the sample is high.
[0019] Optionally, the determining of a first new class sample set from a test sample set based on the first classifier in the first classification model includes: modifying the type of the test sample determined by the first classifier to belong to the first known class and with a confidence level lower than a second threshold to the first new class.
[0020] Through this method, the confidence of some first known class samples identified by the first classifier is lower than the second threshold. If the classification accuracy of the first classifier is low, this first known class sample may actually not be a first known class sample, but a first new class sample. Therefore, the type of this first known class sample is modified to the first new class, so that the classification accuracy of the second classifier trained according to the first new class sample set is high.
[0021] Optionally, the method further includes: training the first classifier based on the first new class sample set and the first sample set to obtain an updated first classification model.
[0022] Through this method, if the classification accuracy of the first classifier is low, the first classifier can be trained using the first new class sample set and the first sample set. Since the first new class sample set also includes the first new class samples obtained by modifying the sample type, the classification accuracy of the first classifier trained in this way is high, so that the classification accuracy of the updated first classification model is high.
[0023] In a second aspect, an embodiment of the present application provides a classification model updating device, which includes a module / unit / technical means for executing the method in the above-mentioned first aspect or any optional implementation method of the first aspect.
[0024] Exemplarily, the device may include:
[0025] A processing module is configured to determine a first new class sample set from a test sample set based on a first classifier in a first classification model; the first classifier is configured to classify samples into first known class samples belonging to a recognizable first known class or first new class samples belonging to an unrecognizable first new class; a sample includes a value of at least one statistical feature corresponding to a piece of network flow data; a second classifier is trained based on the first new class sample set, the second classifier being configured to output the types of samples other than the first new class samples and the first known class samples as an unrecognizable second new class;
[0026] An adding module is used to add the second classifier to the first classification model to obtain a second classification model, and the second classification model is used to classify the type of samples into any one of the following: the first known class, the first new class, and the second new class.
[0027] Optionally, when training a second classifier based on the first new class sample set, the processing module is specifically used to: use the first new class sample set and the first sample set together as a second known class sample set, and train the second classifier based on the second known class sample set; the first sample set includes multiple first known class samples.
[0028] Optionally, the processing module is also used to: obtain the values of multiple statistical features corresponding to each network flow data in multiple network flow data, and obtain multiple statistical feature sequences, the values in a statistical feature sequence correspond to the same statistical feature; determine the Pearson correlation coefficient between any two statistical feature sequences in the multiple statistical feature sequences; determine a target statistical feature group from the multiple statistical features; wherein the target statistical feature group includes at least two statistical features, the Pearson correlation coefficient between two statistical feature sequences corresponding to any two statistical features in the at least two statistical features is less than a first threshold, and the time complexity of each statistical feature in the at least two statistical features meets a preset condition; and take the value of each feature in the target statistical feature group corresponding to a network flow data as a sample.
[0029] Optionally, the first known class includes at least one subclass, and one network flow data corresponds to one subclass. When the processing module determines the target statistical feature group from the multiple statistical features, it is specifically used to: obtain a subclass sequence according to the subclass corresponding to each network flow data; determine the Pearson correlation coefficient between each statistical feature sequence and the subclass sequence; take two statistical features whose Pearson correlation coefficient between the corresponding two statistical feature sequences is greater than or equal to the first threshold as a group of candidate statistical features, to obtain at least one group of candidate statistical features from the multiple statistical features; determine a statistical feature with a smaller Pearson correlation coefficient between the corresponding statistical feature sequence in each group of candidate statistical features and the subclass sequence, to obtain at least one statistical feature to be deleted; delete the at least one statistical feature to be deleted from the multiple statistical features, and determine a statistical feature whose time complexity meets a preset condition from the remaining statistical features, to obtain the target statistical feature group.
[0030] Optionally, when the processing module determines a first new class sample set from a test sample set based on the first classifier in the first classification model, it is specifically used to: modify the type of the test sample judged by the first classifier to belong to the first known class and with a confidence level lower than a second threshold to the first new class.
[0031] Optionally, the processing module is further configured to: train the first classifier based on the first new class sample set and the first sample set to obtain an updated first classification model.
[0032] In a third aspect, an embodiment of the present application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor executes the instructions stored in the memory, thereby enabling the at least one processor to perform the steps of the data processing method described in the first aspect above.
[0033] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the steps of the data processing method described in the first aspect above.
[0034] In addition, other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or may be understood by practicing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0036] FIG1 is a flow chart of a classification model updating method provided in an embodiment of the present application;
[0037] FIG2 is a schematic diagram of the structure of a classification model provided in an embodiment of the present application;
[0038] FIG3 is a schematic diagram of the structure of another classification model provided in an embodiment of the present application;
[0039] FIG4 is a structural diagram of a classification model updating device provided in an embodiment of the present application;
[0040] FIG5 is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. Unless there is a conflict, the embodiments in the present application and the features in the embodiments can be combined with each other in any way. In addition, although a logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than here.
[0042] The terms "first" and "second" in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any of its variations are intended to cover non-exclusive protection. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. "Multiple" in this application can mean at least two, for example, two, three or more, and the embodiments of this application are not limited thereto.
[0043] In addition, the term "and / or" in this article is simply a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the characters "three" in this article, unless otherwise specified, generally indicate that the related objects are in an "or" relationship.
[0044] NTC classifies and identifies network flow data collected from various applications. It can demonstrate high classification accuracy when the test set consists of known classes. However, in reality, the network environment changes dynamically, and classes not in the training set may constantly appear during testing. This problem is called Open Set Flow Recognition (OSFR).
[0045] Regarding the OSFR problem, existing technologies have the problem of low classification accuracy for network flow data.
[0046] In view of this, a technical solution of the embodiment of the present application is provided to improve the classification accuracy of the classification model.
[0047] In view of the above scenario, the classification model updating method provided by the present invention is described in detail below with reference to the accompanying drawings.
[0048] Referring to Figure 1, a flowchart of a classification model updating method is provided in accordance with an embodiment of the present application. This method can be executed by a computer device, such as a laptop computer, a desktop computer, or a server, and can also be applied to various devices with computing capabilities. The above devices are provided for illustration only and are not intended to be limiting in accordance with the present application.
[0049] The method is provided below by a computer device, and the method includes:
[0050] S101: Determine a first new class sample set from a test sample set based on a first classifier in a first classification model.
[0051] To simulate the continuous emergence of new classes in network flow data, the testing process is divided into multiple stages. The test sample set used in the next stage will contain new classes that were not present in the test sample set used in the previous stage. To distinguish test samples from different stages, the test sample set used by the first classifier to determine the first new class sample set is referred to as the test sample set at time t0. The test sample set at time t0 corresponds to the first testing stage.
[0052] The first classification model is introduced below. The test sample set at time t0 is input into the first classification model, and the first classification model outputs that the type of the test samples in the test sample set at time t0 is the first known class or the first new class.
[0053] The first classification model includes a first classifier, which is used to classify samples into first known class samples belonging to a recognizable first known class or first new class samples belonging to an unrecognizable first new class, that is, the first classifier is a binary classifier.
[0054] Optionally, the first known class includes at least one subclass, and the first classification model may also include at least one other classifier. The first known class sample determined based on the first classifier serves as the input of the at least one other classifier, and the at least one other classifier is used to output the type of the first known class sample as at least one subclass.
[0055] For example, a sample is input into the first classifier. If the first classifier outputs a label of +1, +1 indicates that the sample type belongs to the first known class, and the sample is a sample of the first known class. If the first classifier outputs a label of -1, -1 indicates that the sample type belongs to the first new class, and the sample is a sample of the first new class. The first known class includes subclasses 1 to 5. The first known class sample determined by the first classifier is used as input to at least one other classifier, and the output of the other classifier determines to which of subclasses 1 to 5 the sample type belongs.
[0056] It is understandable that any classifier in the at least one other classifier may be a binary classifier or a multi-classifier, and this embodiment of the present application does not limit this.
[0057] In one possible implementation, before determining the first new class sample set, a first classifier can be obtained by the following method. The first known class includes at least one subclass, and a classifier is trained to output samples of the type of any subclass of the at least one subclass. The unlabeled sample set is input into the classifier, and samples with a confidence level greater than or equal to a third threshold are used as samples in the first sample set, and samples with a confidence level less than the third threshold are used as samples in the second sample set. Positive samples are samples belonging to a certain class, and negative samples are samples that do not belong to a certain class. Samples in the first sample set are used as positive samples, and samples in the second sample set are used as negative samples. The first classifier is trained based on the first and second sample sets.
[0058] Confidence refers to the measure of the credibility or accuracy of a prediction result. In machine learning, confidence can also be defined as the estimated probability that a classifier classifies a sample as a certain category.
[0059] It is understandable that the third threshold can be specified according to actual needs, and the embodiment of the present application does not limit it.
[0060] After the first classifier is determined, a first new class sample set is determined from the test sample set at time t0 based on the first classifier.
[0061] Method 1: put the first new class samples judged by the first classifier into the buffer pool. When the number of the first new class samples in the buffer pool reaches a preset number, the multiple first new class samples in the buffer pool are used as the first new class sample set.
[0062] Through this method, the method of obtaining the first new class sample set is simple and easy to operate.
[0063] Method 2: In addition to placing the first new class samples judged by the first classifier into the cache pool, the type of the test samples judged by the first classifier to belong to the first known class and with a confidence level lower than the second threshold is modified to the first new class, and the first new class samples obtained by modifying the sample type are also placed into the cache pool. When the number of first new class samples in the cache pool reaches a preset number, multiple first new class samples in the cache pool are used as the first new class sample set.
[0064] It is understandable that the second threshold can be set according to actual needs and is not limited in the embodiment of the present application.
[0065] Through this method, the confidence of some first known class samples identified by the first classifier is lower than the second threshold. If the classification accuracy of the first classifier is low, this first known class sample may actually not be a first known class sample, but a first new class sample. Therefore, the type of this first known class sample is modified to the first new class, so that the classification accuracy of the second classifier trained according to the first new class sample set is high.
[0066] If the first new class sample set is obtained by the above method 2, the first classifier can be retrained based on the first new class sample set, so that the original first classifier in the first classification model is replaced by the retrained first classifier.
[0067] In a possible implementation, a first classifier is trained based on the first new class sample set and the first sample set to obtain an updated first classification model.
[0068] Through this method, if the classification accuracy of the first classifier is low, the first classifier can be trained using the first new class sample set and the first sample set. Since the first new class sample set also includes the first new class samples obtained by modifying the sample type, the classification accuracy of the first classifier trained in this way is high, so that the classification accuracy of the updated first classification model is high.
[0069] The first classification model is introduced above, and the samples are introduced below.
[0070] In the embodiment of the present application, a sample includes the value of at least one statistical feature corresponding to a piece of network flow data.
[0071] A piece of network flow data corresponds to multiple statistical features. In order to reduce the sample size and improve the processing speed of the classification model for the sample, some statistical features can be selected from the multiple statistics corresponding to a piece of network flow data, and the values of some statistical features can be used as a sample.
[0072] In one possible implementation, the values of multiple statistical features corresponding to each of the multiple network flow data are obtained to obtain multiple statistical feature sequences, and the values in a statistical feature sequence correspond to the same statistical feature; the Pearson correlation coefficient between any two statistical feature sequences in the multiple statistical feature sequences is determined; a target statistical feature group is determined from the multiple statistical features; wherein the target statistical feature group includes at least two statistical features, the Pearson correlation coefficient between the two statistical feature sequences corresponding to any two of the at least two statistical features is less than a first threshold, and the time complexity of each of the at least two statistical features meets a preset condition; the value of each feature in the target statistical feature group corresponding to a network flow data is taken as a sample.
[0073] Among them, the Pearson correlation coefficient refers to the sample correlation coefficient, which is used to measure the correlation between two variables; the time complexity qualitatively describes the running time of the algorithm. In the embodiment of the present application, the time complexity is used to describe the time required to calculate a certain feature.
[0074] It is understandable that the preset condition and the first threshold can be specified according to actual needs. For example, the preset condition includes that the time complexity of the statistical feature does not exceed the time complexity O(n).
[0075] For example, assuming that a piece of network flow data corresponds to four statistical features a to d, the values of the multiple statistical features corresponding to multiple pieces of network flow data are obtained. The values of the multiple statistical features corresponding to the first piece of network flow data are a1 to d1, the values of the multiple statistical features corresponding to the second piece of network flow data are a2 to d2, and so on, the values of the multiple statistical features corresponding to the nth piece of network flow data are an to dn, resulting in four statistical feature sequences a1 to an, b1 to bn, c1 to cn, and d1 to dn. Substituting the statistical feature sequences a1 to an and b1 to bn into the calculation formula for the Pearson correlation coefficient, the Pearson correlation coefficient between the statistical feature sequences a1 to an and the statistical feature sequences b1 to bn is obtained. The method for calculating the Pearson correlation coefficient between the remaining statistical feature sequences is similar to the method for calculating the Pearson correlation coefficient between the statistical feature sequences a1 to an and the statistical feature sequences b1 to bn, and will not be repeated here. If the Pearson correlation coefficient between the statistical feature sequence a1~an and the statistical feature sequence b1~bn is greater than or equal to the first threshold, and the Pearson correlation coefficient between any two statistical feature sequences in the remaining statistical features is less than the first threshold, then statistical feature a or statistical feature b is deleted from statistical features a~statistical features d. Assuming that statistical feature a is deleted, if the time complexity of statistical feature c does not meet the preset conditions, then the target statistical feature group includes statistical features b and statistical features d.
[0076] It can be understood that the above example only takes the number of all statistical features corresponding to a piece of network flow data as 4, and is not limited thereto.
[0077] Through this method, a piece of network flow data corresponds to multiple statistical features. If the values of all statistical features corresponding to a piece of network flow data are taken as a sample, the classification time of the classification model will be longer. According to the values of multiple statistical features corresponding to each piece of network flow data in the multiple network flow data, multiple statistical feature sequences are obtained. According to the Pearson correlation coefficient between any two statistical feature sequences, the target statistical feature group is determined from the multiple statistical features. In this way, the statistical features that make up the sample are fewer, the sample size is reduced, and the classification speed of the second classification model is improved.
[0078] Optionally, the first known class includes at least one subclass, and a network flow data corresponds to a subclass. When determining the target statistical feature from multiple statistical features, a subclass sequence is obtained according to the subclass corresponding to each network flow data; the Pearson correlation coefficient between each statistical feature sequence and the subclass sequence is determined; two statistical features whose Pearson correlation coefficient between the corresponding two statistical feature sequences is greater than or equal to the first threshold are taken as a group of candidate statistical features to obtain at least one group of candidate statistical features among multiple statistical features; statistical features with a smaller Pearson correlation coefficient between the corresponding statistical feature sequence and the subclass sequence in each group of candidate statistical features are determined to obtain at least one statistical feature to be deleted; at least one statistical feature to be deleted is deleted from the multiple statistical features, and a statistical feature whose time complexity meets a preset condition is determined from the remaining statistical features to obtain a target statistical feature group.
[0079] Using the above example where a network flow data item corresponds to four statistical features a to d, the subclasses corresponding to the first network flow data item to the subclasses corresponding to the nth network flow data item are X1 to Xn, that is, the subclass sequence is X1 to Xn; if the Pearson correlation coefficient between the statistical feature sequence a1 to an and the statistical feature sequence b1 to bn is greater than or equal to the first threshold, and the Pearson correlation coefficient between any two statistical feature sequences among the remaining statistical features is less than the first threshold, then the statistical feature a and the statistical feature b are a group of candidate statistical features, and the Pearson correlation coefficient 1 between X1 to Xn and a1 to an, and the Pearson correlation coefficient 2 between X1 to Xn and b1 to bn are determined. If the Pearson correlation coefficient 1 is less than the Pearson correlation coefficient 2, then the statistical feature a is used as the statistical feature to be deleted, and the statistical feature a is deleted from the statistical features a to d. Then, the statistical features whose time complexity meets the preset conditions are retained from the statistical features b to d to obtain the target statistical feature group.
[0080] It can be understood that X1~Xn are only used to distinguish the subclasses corresponding to different network flow data. The actual meanings of any two items in X1~Xn can be the same or different, and the embodiments of this application do not limit this. For example, the first known class includes subclasses 1 to 5, and X1 and X2 are both subclass 1.
[0081] Through this method, since the two statistical features whose Pearson correlation coefficient between the corresponding two statistical feature sequences is greater than or equal to the first threshold are relatively similar, for such a group of selected statistical features, only one of them can be retained, and the statistical features with the smaller Pearson correlation coefficient between the corresponding statistical feature sequence and the subclass sequence in each group of selected statistical features are determined to obtain at least one statistical feature to be deleted, and at least one statistical feature to be deleted is deleted from multiple statistical features, so that the remaining statistical features have a large influence on the sample category, and then the statistical features whose time complexity meets the preset conditions are determined from the remaining statistical features to obtain the target statistical feature group, so that a sample can well represent a network flow data, and the classification accuracy of the classification model trained by the sample is high.
[0082] S102: Train a second classifier based on the first new class sample set.
[0083] The second classifier is used to output the types of samples other than the first new class samples and the first known class samples as an unrecognizable second new class, that is, the second classifier is a binary classifier.
[0084] When training the second classifier, in order to improve the ability of the second classifier to identify the first new class sample set, the samples in the second sample set in step S101 can be used as negative samples to assist in training the second classifier.
[0085] In a possible implementation, samples in the first new class sample set are used as positive samples, and samples in the second sample set are used as negative samples to train a second classifier.
[0086] In this way, the second classifier only needs to identify the first new class and output the types of samples other than the first new class samples as the second new class. In this way, the time for training the second classifier is short and the computing resources consumed are small.
[0087] In another possible implementation, the first new class sample set and the first sample set in step S101 are used together as the second known class sample set, the samples in the second known class sample set are used as positive samples, and the samples in the second sample set are used as negative samples to train a second classifier.
[0088] In this way, since the second classifier can identify the second known class sample set, when a sample type other than the second known class appears, the second classifier can quickly classify it into the second new class, thereby improving the classification speed of the second classification model.
[0089] S103: Add a second classifier to the first classification model to obtain a second classification model.
[0090] The second classification model is used to classify the sample type into any one of the following: the first known class, the first new class, and the second new class.
[0091] After the first test phase, the second test phase begins. This phase uses the test sample set at time t1, which contains a new class compared to the test sample set at time t0. By inputting the test sample set at time t1 into the second classification model, the test sample type can be determined to be either the first known class, the first new class, or the second new class.
[0092] Furthermore, if in step S101, the first classification model also includes at least one other classifier for outputting the type of the first known class sample as at least one subclass, the second classification model is further used to classify the type of the first known class sample into at least one subclass.
[0093] The following describes how to add a second classifier to the first classification model.
[0094] In one possible implementation, if in step S102, the second classifier is trained based on the first new class sample set and the second sample set, then the second classifier is connected after the first classifier, and the samples that do not belong to the first known class determined by the first classifier are used as input to the second classifier.
[0095] Optionally, if the first classification model further includes at least one other classifier, the at least one other classifier is connected after the first classifier, and the first known class sample determined based on the first classifier is used as input of the at least one other classifier.
[0096] Taking the number of at least one other classifier as 1 as an example, refer to Figure 2, where (a) in Figure 2 and (b) in Figure 2 respectively provide schematic diagrams of a first classification model and a second classification model for an embodiment of the present application.
[0097] The first classification model includes classifier A1 and classifier A2. Classifier A1 is equivalent to the first classifier, and classifier A2 is equivalent to the other classifiers. In the first testing phase, test samples from the test sample set at time t0 are input into classifier A1. Test samples that classifier A1 can recognize are used as first known class samples. The first known class samples are then input into classifier A2, which determines to which of at least one subclass the first known class samples belong. Alternatively, the remaining test sample types that classifier A1 cannot recognize are output as the first new class.
[0098] Compared to the first classification model, the second classification model also includes classifier A3, which is equivalent to the second classifier. In the second test phase, the test samples in the test sample set at time t1 are input into classifier A1, and the test samples that can be recognized by classifier A1 are used as first known class samples, and the first known class samples are input into classifier A2. Classifier A2 determines which of the at least one subclass the first known class samples belong to; alternatively, the remaining test samples that cannot be recognized by classifier A1 are input into classifier A3, and classifier A3 outputs the type of the recognizable test samples as the first new class, and classifier A3 outputs the type of the unrecognizable test samples as the second new class.
[0099] In another possible implementation, if in step S102, the second classifier is trained based on the second known class sample set and the second sample set, then the second classifier is connected before the first classifier, and the second known class sample determined based on the second classifier is used as the input of the first classifier.
[0100] Optionally, if the first classification model further includes at least one other classifier, the at least one other classifier is connected after the first classifier, and the first known class sample determined based on the first classifier is used as input of the at least one other classifier.
[0101] Taking the number of at least one other classifier as 1 as an example, referring to FIG3 , FIG3 (a) and FIG3 (b) respectively provide schematic diagrams of another first classification model and a second classification model for an embodiment of the present application.
[0102] The first classification model includes classifier B1 and classifier B2. Classifier B1 is equivalent to the first classifier, and classifier B2 is equivalent to the other classifiers. In the first testing phase, test samples from the test sample set at time t0 are input into classifier B1. Classifier B1 outputs the type of unrecognizable test samples as the first new class. Alternatively, test samples recognizable by classifier B1 are input into classifier B2 as samples of the first known class. Classifier B2 determines to which of the at least one subclass the first known class sample belongs.
[0103] Compared to the first classification model, the second classification model also includes a classifier B3, which is equivalent to a second classifier. In the second test phase, the test samples in the test sample set at time t1 are input into classifier B3. Classifier B3 outputs the type of the unrecognizable test sample as a second new class; alternatively, classifier B3 inputs the recognizable test sample as a second known class sample into classifier B1. Classifier B1 outputs the type of the unrecognizable test sample as a first new class; alternatively, the test sample recognizable by classifier B1 is input into classifier B2 as a first known class sample, and classifier B2 determines to which class the first known class sample belongs of at least one subclass.
[0104] In the above schemes S101 to S103, the first classifier can identify that the sample belongs to the first known class sample, and based on the first classifier, the first new class sample set is determined from the test sample set, and the second classifier is trained based on the first new class sample set, so that the second classifier can identify that the type of the sample is the first new class, and the second classifier is added to the first classification model to obtain the second classification model. Since the second classification model contains the first classifier and the second classifier, the first classifier can identify that the type of the sample is the first known class, and the second classifier can identify that the type of the sample is the first new class. In this way, when the sample type is other than the first known class and the first new class (that is, the OSFR problem), the second classification model can classify it into the second new class. Compared with the first classifier outputting the sample type as the first known class or the first new class, the classification accuracy of the second classifier is higher.
[0105] In addition, in order to cope with new classes that continue to appear during the test process, the above solutions S101 to S103 can be repeatedly executed. For example, in the second test phase, the second classification model is used as the new first classification model, the second classifier in the second classification model is used as the new first classifier, and the test sample set at time t1 is used as the new test sample set. Based on the new first classification model, the new first classifier, and the new test sample set, the above solutions S101 to S103 are executed to obtain a new second classification model. The new second classification model can classify the test samples at time t2 into any one of the first known class, the first new class, the second new class, and the third new class, and so on.
[0106] Optionally, when the latest second classification model includes multiple binary classifiers, a multi-classifier can be trained based on the multiple binary classifiers, and the multiple binary classifiers can be replaced by the multi-classifier.
[0107] The above describes the method provided by the embodiment of the present application, and the following describes the device provided by the embodiment of the present application.
[0108] 4 , based on the same inventive concept, an embodiment of the present invention provides a classification model updating device.
[0109] Exemplarily, the apparatus 400 includes:
[0110] Processing module 401 is configured to determine a first new class sample set from a test sample set based on a first classifier in a first classification model; the first classifier is configured to classify samples into first known class samples belonging to a recognizable first known class or first new class samples belonging to an unrecognizable first new class; a sample includes a value of at least one statistical feature corresponding to a piece of network flow data; a second classifier is trained based on the first new class sample set, the second classifier being configured to output the type of samples other than the first new class samples and the first known class samples as an unrecognizable second new class;
[0111] The adding module 402 is used to add the second classifier to the first classification model to obtain a second classification model, where the second classification model is used to classify the sample type into any one of the following: the first known class, the first new class, and the second new class.
[0112] Optionally, when training a second classifier based on the first new class sample set, the processing module 401 is specifically used to: use the first new class sample set and the first sample set together as a second known class sample set, and train the second classifier based on the second known class sample set; the first sample set includes multiple first known class samples.
[0113] Optionally, the processing module 401 is also used to: obtain the values of multiple statistical features corresponding to each network flow data in multiple network flow data, and obtain multiple statistical feature sequences, the values in a statistical feature sequence correspond to the same statistical feature; determine the Pearson correlation coefficient between any two statistical feature sequences in the multiple statistical feature sequences; determine a target statistical feature group from the multiple statistical features; wherein the target statistical feature group includes at least two statistical features, the Pearson correlation coefficient between the two statistical feature sequences corresponding to any two statistical features in the at least two statistical features is less than a first threshold, and the time complexity of each statistical feature in the at least two statistical features meets a preset condition; and take the value of each feature in the target statistical feature group corresponding to a network flow data as a sample.
[0114] Optionally, the first known class includes at least one subclass, and one network flow data corresponds to one subclass. When the processing module 401 determines the target statistical feature group from the multiple statistical features, it is specifically used to: obtain a subclass sequence according to the subclass corresponding to each network flow data; determine the Pearson correlation coefficient between each statistical feature sequence and the subclass sequence; take two statistical features whose Pearson correlation coefficient between the corresponding two statistical feature sequences is greater than or equal to the first threshold as a group of candidate statistical features, to obtain at least one group of candidate statistical features from the multiple statistical features; determine a statistical feature with a smaller Pearson correlation coefficient between the corresponding statistical feature sequence in each group of candidate statistical features and the subclass sequence, to obtain at least one statistical feature to be deleted; delete the at least one statistical feature to be deleted from the multiple statistical features, and determine a statistical feature whose time complexity meets a preset condition from the remaining statistical features to obtain the target statistical feature group.
[0115] Optionally, when the processing module 401 determines a first new class sample set from a test sample set based on the first classifier in the first classification model, it is specifically used to: modify the type of the test sample judged by the first classifier to belong to the first known class and with a confidence level lower than a second threshold to the first new class.
[0116] Optionally, the processing module 401 is further configured to: train the first classifier based on the first new class sample set and the first sample set to obtain an updated first classification model.
[0117] As a possible product form of the above-mentioned device, referring to FIG5 , an embodiment of the present application further provides an electronic device 500, including:
[0118] At least one processor 501; and a communication interface 503 communicatively connected to the at least one processor 501; the at least one processor 501 executes instructions stored in the memory 502, so that the electronic device 500 executes the method in the embodiment shown in Figure 3 through the communication interface 503.
[0119] Optionally, the memory 502 is located outside the electronic device 500 .
[0120] Optionally, the electronic device 500 includes the memory 502, which is connected to the at least one processor 501 and stores instructions executable by the at least one processor 501. FIG5 shows with dotted lines that the memory 502 is optional for the electronic device 500.
[0121] The processor 501 and the memory 502 may be coupled via an interface circuit or may be integrated together, which is not limited here.
[0122] The specific connection medium between the processor 501, memory 502, and communication interface 503 is not limited in the embodiments of the present application. In Figure 5, the processor 501, memory 502, and communication interface 503 are connected via a bus 504. The bus is represented by a bold line in Figure 5. The connection between other components is only for illustrative purposes and is not intended to be limiting. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 5 only uses a single bold line, but this does not mean that there is only one bus or only one type of bus.
[0123] It should be understood that the processors mentioned in the embodiments of the present application can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented by software, the processor can be a general-purpose processor that is implemented by reading software code stored in a memory.
[0124] Exemplarily, the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0125] It should be understood that the memory mentioned in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DR RAM).
[0126] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) can be integrated into the processor.
[0127] It should be noted that the memory described herein is intended to include, but not be limited to, these and any other suitable types of memory.
[0128] As another possible product form, an embodiment of the present application also provides a computer-readable storage medium, which is used to store instructions. When the instructions are executed, the computer executes the method in the embodiment shown in Figure 1.
[0129] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0130] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or box in the flow chart and / or block diagram, as well as the combination of the flow chart and / or box in the flow chart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more flow charts and / or one or more boxes in the block diagram.
[0131] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0133] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include these modifications and variations.
Claims
1. A method for updating a classification model, characterized in that The method includes: Determining a first new class sample set from a test sample set based on a first classifier in a first classification model; the first classifier is used to classify a sample into a first known class sample belonging to a recognizable first known class or a first new class sample belonging to an unrecognizable first new class; a sample includes the values of at least one statistical feature corresponding to a network flow data; Training a second classifier based on the first new class sample set, the second classifier is used to output the type of a sample other than the first new class sample and the first known class sample as an unrecognizable second new class; Adding the second classifier to the first classification model to obtain a second classification model, the second classification model is used to classify the type of a sample into any one of the following: the first known class, the first new class, and the second new class.
2. The method according to claim 1, wherein The training the second classifier based on the first new class sample set includes: Using the first new class sample set and a first sample set together as a second known class sample set, and training the second classifier based on the second known class sample set; the first sample set includes a plurality of the first known class samples.
3. The method according to claim 1, characterized in that, The method further includes: Obtaining the values of a plurality of statistical features corresponding to each network flow data in a plurality of network flow data to obtain a plurality of statistical feature sequences, and the values in a statistical feature sequence correspond to the same statistical feature; Determining the Pearson correlation coefficient between any two statistical feature sequences among the plurality of statistical feature sequences; Determining a target statistical feature group from the plurality of statistical features; wherein, the target statistical feature group includes at least two statistical features, the Pearson correlation coefficient between any two statistical feature sequences corresponding to any two of the at least two statistical features is less than a first threshold, and the time complexity of each of the at least two statistical features meets a preset condition; Using the values of the features in the target statistical feature group corresponding to a network flow data as a sample.
4. The method according to claim 3, wherein The first known class includes at least one subclass, and a network flow data corresponds to a subclass. The determining the target statistical feature group from the plurality of statistical features includes: Obtaining a subclass sequence according to the subclass corresponding to each network flow data; Determining the Pearson correlation coefficient between each statistical feature sequence and the subclass sequence; Regarding two statistical features whose Pearson correlation coefficient between the corresponding two statistical feature sequences is greater than or equal to the first threshold as a group of candidate statistical features, and obtaining at least one group of candidate statistical features among the plurality of statistical features; Determining the statistical feature with a smaller Pearson correlation coefficient between the corresponding statistical feature sequence and the subclass sequence in each group of candidate statistical features to obtain at least one statistical feature to be deleted; Deleting the at least one statistical feature to be deleted from the plurality of statistical features, and determining the statistical features whose time complexity meets the preset condition from the remaining statistical features to obtain the target statistical feature group.
5. The method according to claim 1, wherein The determining the first new class sample set from the test sample set based on the first classifier in the first classification model includes: Modify the type of the test samples that are determined by the first classifier to belong to the first known class and have a confidence level lower than the second threshold to the first new class.
6. The method according to claim 5, characterized in that The method further includes: Training the first classifier based on the first new class sample set and the first sample set to obtain the updated first classification model; the first sample set includes a plurality of the first known class samples.
7. A classification model update device, characterized in that, The apparatus includes: A processing module, configured to determine a first new class sample set from a test sample set based on a first classifier in a first classification model; the first classifier is configured to classify a sample into a first known class sample belonging to an identifiable first known class or a first new class sample belonging to an unidentifiable first new class; a sample includes values of at least one statistical feature corresponding to a network flow data; training a second classifier based on the first new class sample set, the second classifier is configured to output the type of a sample other than the first new class sample and the first known class sample as an unidentifiable second new class; An adding module, configured to add the second classifier to the first classification model to obtain a second classification model, the second classification model is configured to classify the type of a sample into any one of the following: the first known class, the first new class, and the second new class.
8. The device according to claim 7, characterized in that, When training the second classifier based on the first new class sample set, the processing module is specifically configured to: Use the first new class sample set and the first sample set together as a second known class sample set, and train the second classifier based on the second known class sample set; the first sample set includes a plurality of the first known class samples.
9. An electronic device, characterized in that, Includes: At least one processor; And a memory and a communication interface communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the at least one processor, by executing the instructions stored in the memory, causes the electronic device to execute the method according to any one of claims 1 to 6 through the communication interface.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are run on a computer, the computer is caused to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Open set data labeling method and device, equipment, storage medium and program product
CN114330570A
Learning model for network flow online classification based on extreme learning machine algorithm
CN116522260A
Classification model updating method and device, equipment and storage medium
CN117997845A
Open set classification method based on classification utility
WO2021128704A1
Task agnostic open-set prototypes for few-shot open-set recognition
WO2023225426A1