Feature selection-based operating system identification training method and electronic equipment
Through the passive recognition method based on feature selection, the naive Bayesian model of feature selection and feature weights is used to achieve efficient identification of operating system types, solving the problems of resource occupation and false positives in traditional methods, and improving the recognition effect and adaptability.
Patent Information
- Application Number
- CN202510166758.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
During the identification process, traditional active operating system identification methods will cause false alarms of resource occupation and network protection mechanisms, resulting in poor detection effect and affecting the identification effect of operating system types.
A passive recognition method based on feature selection is adopted. By obtaining multiple traffic data of different operating system types, using correlation-based feature selection and mutual information feature scores, a more representative feature subset is selected, and a naive Bayesian model with feature weights is trained to achieve passive operating system type recognition.
It improves the recognition effect of operating system types, reduces dependence on predefined features, enhances the adaptability of complex and dynamic network environments, and overcomes the feature independence limitations of naive Bayes algorithms.
Smart Images

Figure CN120105094A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of network security technology, and in particular to a training method for operating system recognition based on feature selection and an electronic device. Background Art
[0002] With the rapid development of smart transportation, network security issues pose a major threat to the stability and reliability of transportation systems. Through the integration of technologies such as the Internet of Things (IoT) and artificial intelligence (AI), transportation systems have realized the intelligent collection, analysis and processing of traffic information, greatly improving the efficiency of traffic management. However, with the surge in the number of networked devices and the increase in network complexity, the risk of network attacks has also increased significantly. Botnets can quickly exhaust the computing resources or network bandwidth of the target host by controlling a large number of infected devices to launch attacks at the same time, making it unable to work normally. In the transportation system, botnet attacks may cause traffic signal control failure, loss of key sensor data, and interruption of communication links, posing a threat to the safety of the entire transportation system.
[0003] In the transportation system, the identification of operating system types is of great significance for asset management, network protection and security strategy formulation. Different operating system types have different security vulnerabilities, so quickly and accurately identifying the operating system type in the network environment becomes the key to ensuring the safety of the transportation system.
[0004] However, traditional active operating system identification methods require sending a large number of detection data packets to the target host, which may cause resource occupation and false alarms of network protection mechanisms (such as being blocked by firewalls and other measures), resulting in poor detection effects, thereby affecting the identification effect of the operating system type. Summary of the invention
[0005] The present disclosure is proposed in view of the above-mentioned situation, and its purpose is to provide a training method and electronic device for operating system identification based on feature selection that can support the identification of operating system types in a passive manner and can improve the identification effect of operating system types.
[0006] To this end, the first aspect of the present disclosure provides a training method for operating system identification based on feature selection, the training method comprising: obtaining multiple traffic data of different operating system types; obtaining the value of an initial feature set including a TCP feature set related to the TCP protocol and an IP feature set related to the IP protocol for each of the multiple traffic data, and using the multiple values of the initial feature sets as an initial training set; performing feature selection on the initial feature set using correlation-based feature selection and the initial training set to obtain at least one feature subset as a first feature set, and evaluating the features in the initial feature set using a mutual information-based evaluation function and the initial training set. Scoring is performed, a first preset number of features with the highest scores are selected from the initial feature set as the second feature set, and an intersection operation is performed based on the first feature set and the second feature set to obtain a target feature set; and multiple values of the target feature set are used as a target training set, and feature weights of the features in the target feature set are obtained based on the target training set, and a naive Bayes model combined with the feature weights of the features in the target feature set is trained using the target training set to obtain a trained model, and the trained model is used to receive the data to be identified corresponding to the target feature set in a network environment with passive traffic and output a classification result related to the operating system type.
[0007] In the first aspect of the present disclosure, the use of multiple feature selection can screen out more representative features (i.e., target feature sets), and improve the ability to capture features of complex and / or new botnets. In addition, through multiple feature selection, key features can be dynamically identified, reducing dependence on predefined features, and improving the recognition effect of operating system types (i.e., improving the ability to recognize operating systems). In addition, the complexity and dynamic changes of traffic in real network environments increase data noise, and multiple feature selection can effectively filter noise, extract more robust features, and improve the adaptability to complex and dynamic network environments. In addition, the naive Bayes model combined with feature weights can overcome the limitation that the features of the naive Bayes algorithm must be independent of each other to improve the performance and recognition accuracy of the naive Bayes model. In addition, the initial feature set is a feature set related to the network protocol, which can support the identification of operating system types in a passive manner.
[0008] In addition, in the training method involved in the first aspect of the present disclosure, optionally, the evaluation function based on mutual information satisfies the formula: MI (X t )=I(X t ; Y), where J MI (X t ) indicates X t The evaluation function, X t represents the set of values of the tth feature of the initial feature set in the initial training set, I(Xt ; Y) represents X expressed in the form of entropy t The mutual information between features and operating system types is shown in Figure 1. Y represents the set of operating system types. This can highlight the dependency between features and operating system types.
[0009] In addition, in the training method involved in the first aspect of the present disclosure, optionally, the at least one feature subset includes a first subset, a second subset, and a third subset whose scores decrease in sequence, and according to the scores of each feature in the second feature set, the second feature set is divided into a fourth subset, a fifth subset, and a sixth subset, the score of the feature in the fourth subset is greater than the score of the feature in the fifth subset and the score of the feature in the fifth subset is greater than the score of the feature in the sixth subset, and in the intersection operation, the first subset and the fourth subset are taken as a high-value subset, the second subset and the fifth subset are taken as a medium-value subset, and the high-value subset and the medium-value subset are taken as a union as the target feature set. In this case, high-value features can be retained, medium-value features can be comprehensively selected, and low-value features (i.e., features in the third and sixth subsets) can be eliminated, thereby improving the effectiveness of the target feature set.
[0010] In addition, in the training method involved in the first aspect of the present disclosure, optionally, the IP feature set includes an IP protocol version number, an IP message header length, a service type field, a total message length, an IP fragmentation flag, a fragmentation offset, a lifetime, a protocol flag, an IP message checksum, and an IP message option, and the TCP feature set includes a control flag, a reserved bit, a sliding window size, a TCP message checksum, an emergency pointer, a time sampling enable flag, a maximum message length, a window scaling factor, a selection confirmation message sequence number, a reception confirmation flag, an MD5 signature, a TCP header length, and a user timeout duration flag. In this case, the comprehensiveness of the initial feature set can be improved, and it is convenient to dynamically select more representative features from the initial feature set.
[0011] In addition, in the training method involved in the first aspect of the present disclosure, optionally, the first subset includes sliding window size, window scaling factor, control flag, and maximum message length, the second subset includes maximum message length, window scaling factor, control flag, and receive confirmation flag, the fourth subset includes window scaling factor, control flag, sliding window size, lifetime, maximum message length, and IP fragmentation flag, and the fifth subset includes receive confirmation flag, total message length, selection confirmation message sequence number, TCP message checksum, and user timeout duration identifier.
[0012] In addition, in the training method involved in the first aspect of the present disclosure, optionally, the feature weight of the feature in the target feature set is obtained using relative entropy, and the feature weight satisfies the formula: Among them, Weight represents the feature weight of M, M represents the feature in the target feature set, P(M j ) indicates that the value of M is M j The probability of C represents the type of operating system, P(C) represents the prior probability of C, P(C|M j ) indicates that the value of M is M j The conditional probability that the operating system type is C when J represents the number of samples in the target training set, M j It represents the value when the feature of the jth sample in the target training set is M. In this case, using relative entropy to obtain feature weights can improve the accuracy of measuring the correlation between features and operating system types, thereby improving the recognition effect of the naive Bayes model.
[0013] In addition, in the training method involved in the first aspect of the present disclosure, optionally, the naive Bayes model satisfies the formula: Among them, P post Indicates that each sample in the target training set belongs to Y k argMax represents the function used to determine the category with the largest posterior probability, Y k represents the kth operating system type, P(Y k ) indicates Y k The corresponding prior probability, I represents the number of features in the target feature set, x i represents the value of the i-th feature of the target feature set in the sample, P(x i |Y k ) indicates that when the classification result is Y k In the case of x i The corresponding conditional probability, Weight i represents the feature weight of the ith feature in the target feature set, and Π represents the product. Thus, combining the feature weights can improve the recognition effect of the naive Bayes model.
[0014] In addition, in the training method involved in the first aspect of the present disclosure, optionally, the target training set is grouped to obtain multiple groups of data, and in response to the number of types of samples of the corresponding operating system type in each group of the multiple groups of data being less than a preset number, the samples of the corresponding operating system type are copied to each group of the multiple groups of data to update the target training set. In this way, the samples can be expanded to solve the problem of uneven distribution of traffic data.
[0015] In addition, in the training method involved in the first aspect of the present disclosure, optionally, in response to the number of samples of the target operating system type in the target training set being less than a second preset number and the usage proportion of the target operating system type being greater than a preset proportion, the trained model is incrementally trained based on an incremental training set corresponding to the target operating system type, the target operating system type being one of the operating system types. Thus, the robustness of the trained model can be improved.
[0016] A second aspect of the present disclosure provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed, the training method involved in the first aspect of the present disclosure is implemented.
[0017] According to the present disclosure, a training method for operating system identification based on feature selection and an electronic device are provided, which can support the identification of operating system types in a passive manner and can improve the identification effect of operating system types. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The present disclosure will now be explained in further detail, by way of example only, with reference to the accompanying drawings.
[0019] Figure 1 is a schematic diagram showing a recognition environment involved in the examples of the present disclosure.
[0020] Figure 2 is an exemplary flow chart showing a training method involved in the examples of the present disclosure.
[0021] Figure 3 is a schematic diagram showing the acquisition of a target feature set involved in the example of the present disclosure.
[0022] Figure 4A is a diagram showing the distribution of TTL values involved in the examples of the present disclosure. Figure 4B is a diagram showing the value distribution of the control flag bits involved in the examples of the present disclosure. Figure 4C 2 is a diagram showing the distribution of IP message checksum values involved in the examples of the present disclosure. Figure 4D 2 is a diagram showing the distribution of TCP message checksum values involved in the examples of the present disclosure. Figure 4E is a diagram showing the distribution of values of the sliding window size involved in the examples of the present disclosure. Figure 4F is a diagram showing the distribution of values of the maximum message length involved in the examples of the present disclosure. Figure 4G is a diagram showing the value distribution of the window scaling factor involved in the examples of the present disclosure.
[0023] Figure 5 is a comparison chart showing the recognition performance of operating system types before and after the incremental training involved in the examples of the present disclosure. DETAILED DESCRIPTION
[0024] Hereinafter, with reference to the accompanying drawings, preferred embodiments of the present disclosure are described in detail. In the following description, the same symbols are given to the same components, and repeated descriptions are omitted. In addition, the accompanying drawings are only schematic diagrams, and the ratio of the dimensions of the components to each other or the shapes of the components, etc. may be different from the actual ones. It should be noted that the terms "including" and "having" in the present disclosure and any variations thereof, such as a process, method, system, product or device including or having a series of steps or units, are not necessarily limited to those steps or units clearly listed, but may include or have other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] First, relevant terms involved in the present disclosure are introduced.
[0026] "Passive traffic" can refer to traffic data obtained by monitoring the normal communication process of the target host. Passive traffic has the advantages of not interfering with the normal communication process and not being affected by intrusion detection mechanisms.
[0027] The disclosed example relates to a training method for operating system identification based on feature selection (hereinafter referred to as the training method). The trained model obtained by the training method can support the identification of operating system types in a passive manner and can improve the identification effect of operating system types. The passive method can utilize the existing traffic data in the network and can complete the identification of the operating system without interfering with normal network activities. That is, the trained model obtained by the training method involved in the disclosed example can be applied to the identification of operating system types in a network environment with passive traffic. In addition, the training method involved in the disclosed example can also be referred to as a learning method, a model generation method, etc.
[0028] In some examples, the training method involved in the examples of the present disclosure can be applied to scenarios with multiple operating systems in a transportation system. A transportation system usually contains a large number of heterogeneous devices, and efficient identification of their operating system types is of great significance for achieving lifecycle management of network security and improving network security protection capabilities.
[0029] Examples of the present disclosure will be described in detail below with reference to the accompanying drawings. Figure 1 is a schematic diagram showing a recognition environment involved in the examples of the present disclosure.
[0030] refer to Figure 1The identification environment may include various types of devices, such as FTP servers, office equipment, network management equipment, and DNS servers. The identification environment may also include network devices for interconnecting or connecting various types of devices to the Internet, such as routers, firewalls, and core switches. The identification environment may also include an identification system 100, which is used to receive flow data from network devices such as firewalls and identify the type of operating system corresponding to the flow data.
[0031] In some examples, the recognition system 100 may include a trained model, which may be obtained by training a machine learning model using the training method involved in the examples of the present disclosure. In some examples, the machine learning model may be a naive Bayes model. In other examples, the machine learning model may also be other models, such as a support vector machine, a random forest algorithm, or a neural network. The following describes an example of the present disclosure using a naive Bayes model as an example of the machine learning model, which does not limit the present disclosure.
[0032] Figure 2 is an exemplary flow chart showing a training method involved in the examples of the present disclosure.
[0033] Figure 3 is a schematic diagram showing the acquisition of a target feature set involved in the example of the present disclosure.
[0034] In some examples, reference Figure 2The training method may include obtaining multiple traffic data of different operating system types (step S101), obtaining the value of an initial feature set including a feature set related to a network protocol for each of the multiple traffic data, and using the values of the multiple initial feature sets as an initial training set (step S102), performing multiple feature selections on the initial feature set based on the initial training set and using feature selection based on correlation (CFS) and feature selection based on mutual information (MI) to obtain a target feature set (step S103), using the values of the multiple target feature sets as a target training set, obtaining feature weights of features in the target feature set based on the target feature set, and using the target training set to train a naive Bayes model combining the feature weights of features in the target feature set to obtain a trained model (step S104). In this case, the use of multiple feature selections can screen out more representative features (i.e., the target feature set) and improve the ability to capture features of complex and / or new botnets. In addition, multiple feature selection can dynamically identify key features, reduce dependence on predefined features, and improve the recognition effect of operating system types (i.e., improve the recognition ability of operating systems). In addition, the complexity and dynamic changes of traffic in real network environments increase data noise. Multiple feature selection can effectively filter noise, extract more robust features, and improve the adaptability to complex and dynamic network environments. In addition, the naive Bayesian model combined with feature weights can overcome the limitation that the features of the naive Bayesian algorithm must be independent of each other to improve the performance and recognition accuracy of the naive Bayesian model. In addition, the initial feature set is a feature set related to the network protocol, which can support the passive identification of operating system types.
[0035] In some examples, reference Figure 2 In step S101, in response to receiving multiple pieces of traffic data, the multiple pieces of traffic data may be divided according to the source IP address to obtain the operating system type of each piece of traffic data. Specifically, the multiple pieces of traffic data may be divided according to the source IP address to obtain the traffic data of the source IP address, and the source IP address and the operating system type may be matched to determine the operating system type of each piece of traffic data. This facilitates the subsequent addition of labels.
[0036] In some examples, the operating system type may be an operating system major version. For example, the operating system type may include at least one of Windows and Linux. In other examples, the operating system type may be an operating system detailed version. For example, the operating system type may include at least one of Windows XP, Windows 7, Windows 8, Windows 10, Linux 2, Linux 3, Mac, FreeBSD, OpenBSD, and Solaris.
[0037] In some examples, the flow data may be data transmitted via a network device. In this case, it is convenient to passively collect the flow data, thereby reducing interference with the target host.
[0038] In some examples, the plurality of flow data may be derived from existing data. The existing data may be existing historical data. Thus, the convenience of obtaining flow data may be improved and the difficulty of obtaining flow data may be reduced.
[0039] In some examples, the existing data may include an existing fingerprint library (eg, a fingerprint library from a p0f operating system identification tool) and historical traffic data in a real network environment (eg, pcap data packets).
[0040] In some examples, reference Figure 2 In step S102, the network protocol may include the TCP protocol and the IP protocol. Specifically, the value of the initial feature set of each of the multiple flow data may be obtained, and the initial feature set may include a TCP feature set related to the TCP protocol and an IP feature set related to the IP protocol.
[0041] In some examples, the IP feature set may include an IP protocol version number (Ver), an IP message header length (ihl), a service type field (Tos), a message total length (Len), an IP fragmentation flag bit (ip_flags), a fragmentation offset bit (frag), a lifetime (TTL), a protocol identifier bit (proto), an IP message checksum (ip_chksum), and an IP message option (ip_option). In this case, the comprehensiveness of the initial feature set can be improved, and it is convenient to dynamically select more representative features from the initial feature set.
[0042] In addition, the service type field can be used to specify the message priority. In some examples, the IP fragmentation flag can be divided into two cases: DF and MF. The survival time can indicate the survival time of the data packet in the traffic data in the network. The protocol identification bit can be used to indicate the protocol type of the carried data packet, such as TCP.
[0043] In some examples, the TCP feature set may include a control flag (tcp_flags), a reserved bit (Reserved), a sliding window size (Wsize), a TCP message checksum (tcp_chksum), an urgent pointer (urg), a time sampling enable flag (TSopt), a maximum message length (MSS), a window scaling factor (Scale), a selection confirmation message sequence number (SOK), a reception confirmation flag (ACK), an MD5 signature (MD), a TCP header length (Dataofs), and a user timeout duration identifier (UTO). In this case, the comprehensiveness of the initial feature set can be improved, and it is convenient to dynamically select more representative features from the initial feature set.
[0044] In addition, the control flag can be used to identify different control functions of TCP. The urgent pointer can be used to indicate the offset of the urgent message.
[0045] As described above, multiple values of the initial feature set are used as the initial training set. That is, the initial training set may include multiple groups of values of the initial feature set, and each group of values may include multiple values corresponding to multiple features in the initial feature set. In other words, the initial training set may include multiple samples, each sample may correspond to a group of values, and each group of values may include multiple values corresponding to multiple features in the initial feature set.
[0046] Since the feature dimensions of each feature are different, the initial training set is preprocessed to facilitate feature selection. Specifically, before using the initial training set to obtain the first feature set and the second feature set, the initial training set can be preprocessed, and the first feature set and the second feature set can be obtained using the preprocessed initial training set. This can facilitate feature selection.
[0047] In some examples, preprocessing may include at least one of the following:
[0048] The first method (hereinafter referred to as adding labels): labels are added to the samples in the initial training set using the operating system type of the host corresponding to the source IP address of the IP protocol.
[0049] The second type (hereinafter referred to as quantization processing): quantize the features whose values are characters in the initial training set according to the feature dimension. In some examples, the features can be quantized into values ranging from 1 to Z, where Z represents the feature dimension. For example: the quantization value of Ver can be 1 to 3; the quantization value of TTL can be 1 to 14; the quantization value of MSS can be 1 to 19; the quantization value of Wsize can be 1 to 118, etc.
[0050] The third type (hereinafter referred to as normalization): identify unexpected values in the initial feature set and normalize the unexpected values. Unexpected values can be values that are not within the consideration range of the feature dimension. In some examples, in normalization, unexpected values can be set to 0. That is, 0 can be used as a quantization value for other special values.
[0051] In the identification of operating system types, the first problem to be solved is which features to use as the basis for identification.
[0052] In some examples, reference Figure 2 In step S103, feature selection based on correlation and feature selection based on mutual information can be performed on the initial feature set to obtain a first feature set and a second feature set, and a target feature set is obtained based on the first feature set and the second feature set. In this case, feature selection based on correlation can characterize the degree of correlation between features, while feature selection based on mutual information can consider the dependency between features and operating system types. As described above, combining the feature sets of the two to obtain a target feature set can screen out more representative features.
[0053] In some examples, feature selection may be performed on the initial feature set using correlation-based feature selection and an initial training set to obtain at least one feature subset as the first feature set.
[0054] In addition, the features in each of the at least one feature subset may jointly determine the score of each feature subset, thereby highlighting the correlation between the features.
[0055] In some examples, in correlation-based feature selection, after obtaining the correlation matrix between features and operating system types and between features, the feature subset space can be searched using best-first search until at least one feature subset (e.g., three feature subsets) that maximizes the score (also referred to as value) of the feature subset is found. In some examples, the time complexity of correlation-based feature selection can be m*W(W-1) / 2, where m is the number of features in the feature subset and W is the number of features in the initial feature set.
[0056] In some examples, when searching the feature subset space, one may start with an empty set or a full set. Taking the empty set as an example, the feature that maximizes the score of the feature subset is continuously added to the empty set.
[0057] In some examples, in relevance-based feature selection, the scores of feature subsets can satisfy the formula:
[0058]
[0059] Among them, Merit srepresents the score of s containing m features, s represents a feature subset, represents the average value of the correlation between the features in s and the operating system type, represents the average value of the correlation between features in s. In some examples, symmetric uncertainty (SU) can be used to calculate and
[0060] In some examples, reference Figure 3 , at least one feature subset may include a first subset A1, a second subset B1, and a third subset C1 with decreasing scores. In some examples, feature subsets with adjacent scores in the first subset A1, the second subset B1, and the third subset C1 may have an intersection. Thus, the reliability of the features in the intersection can be improved.
[0061] Take the intersection of the first subset A1 and the second subset B1 as an example. In some examples, through correlation-based feature selection, the first subset A1 may include a sliding window size (Wsize), a window scaling factor (Scale), a control flag (tcp_flags), and a maximum message length (MSS), and the second subset may include a maximum message length (MSS), a window scaling factor (Scale), a control flag (tcp_flags), and a reception confirmation flag (ACK). The features and scores of the first subset A1 and the second subset B1 are shown in Table 1.
[0062]
[0063] Table 1 Scores of the first subset A1 and the second subset B1
[0064] In addition, mutual information is an effective indicator for measuring different variables, and its meaning is the common information between two variables. In some examples, feature selection based on mutual information and an initial training set can be used to perform feature selection on an initial feature set to obtain a second feature set. In some examples, in feature selection based on mutual information, an evaluation function based on mutual information and an initial training set can be used to score the features in the initial feature set, and a first preset number of features with the highest scores are selected from the initial feature set as the second feature set. In addition, the present disclosure does not limit the value of the first preset number, which can be determined, for example, by experiments.
[0065] In some examples, a forward search strategy can be used to select the first preset number of features with the highest scores from the initial feature set as the second feature set. Specifically, based on the initial training set, a forward search strategy can be used to score and sort the features in the initial feature set using an evaluation function, and the first preset number of features ranked first can be selected as the second feature set.
[0066] In addition, through feature selection based on mutual information, each feature in the second feature set can have a score. The features in the subsets divided by the second feature set each have their own score. In this way, the dependency between the feature and the operating system type can be highlighted. Figure 3 According to the scores of each feature in the second feature set, the second feature set can be divided into a fourth subset A2, a fifth subset B2 and a sixth subset C2, the scores of the features in the fourth subset A2 are greater than the scores of the features in the fifth subset B2 and the scores of the features in the fifth subset B2 are greater than the scores of the features in the sixth subset C2.
[0067] In addition, the set of values of the single features of the initial feature set in the initial training set and the set of operating system types can be regarded as two variables in the mutual information. In some examples, the evaluation function based on mutual information can satisfy the formula:
[0068] J MI (X t )=I(X t ; Y),
[0069] Among them, J MI (X t ) indicates X t The evaluation function, X t represents the set of values of the tth feature of the initial feature set in the initial training set (e.g. X t ={x 1 ,x 2 ,…,x W}), I(X t ; Y) represents X expressed in the form of entropy t The mutual information with Y, where Y represents the set of operating system types (e.g., Y = {y 1 ,y 2 ,…,y K}). Thus, the dependency between the features and the operating system types can be highlighted. In addition, W represents the number of samples in the initial training set, and K represents the number of types of operating system types.
[0070] In some examples, I(X t ; Y) can satisfy the formula:
[0071] I(X t ; Y) = H(Y) - H(Y|X t )=H(X t )-H(X t |Y),
[0072] Among them, H(X t ) indicates X t The entropy, H(X t|Y) is the conditional entropy, which means that X t degree of uncertainty.
[0073] In some examples, H(X t ) can satisfy the formula:
[0074]
[0075] Among them, x w For X t The value of the wth sample in (that is, X t The w-th value in w ) is x w The marginal probability distribution function of , W represents the number of samples in the initial training set.
[0076] In some examples, X t When the value in is a discrete value (i.e., X t is a discrete variable), p(x w ) can satisfy the formula:
[0077]
[0078] In some examples, H(X t |Y) can satisfy the formula:
[0079]
[0080] Among them, p(x w |y k ) is in y k Under the premise of x w The conditional probability distribution of p(x w ,y k ) is x w and k The joint probability distribution function, y k is the kth operating system type in Y, K represents the number of operating system types, W represents the number of samples in the initial training set, and Y represents the set of operating system types.
[0081] In some examples, the fourth subset A2, the fifth subset B2 and the sixth subset C2 may not have an intersection. Thus, the dependence on some features can be suppressed. In some examples, through feature selection based on mutual information, the fourth subset A2 may include a window scaling factor (Scale), a control flag (tcp_flags), a sliding window size (Wsize), a lifetime (TTL), a maximum message length (MSS), and an IP fragmentation flag (ip_flags), and the fifth subset B2 may include a reception confirmation flag (ACK), a total message length (Len), a selection confirmation message sequence number (SOK), a TCP message checksum (tcp_chksum), and a user timeout duration identifier (UTO).
[0082] The characteristics and scores of the fourth subset A2 and the fifth subset B2 are shown in Tables 2 and 3 respectively.
[0083]
[0084]
[0085] Table 2 Scores of the fourth subset A2
[0086]
[0087] Table 3 Scores of the fifth subset B2
[0088] As described above, the target feature set can be obtained based on the first feature set and the second feature set. In some examples, the target feature set can be obtained by performing an intersection operation based on the first feature set and the second feature set. Figure 3 In the intersection operation, the first subset A1 and the fourth subset A2 can be taken as the high-value subset A, the second subset B1 and the fifth subset B2 can be taken as the medium-value subset B (also called the common feature set), and the high-value subset A and the medium-value subset B can be taken as the target feature set. In this case, the high-value features can be retained, the medium-value features can be comprehensively selected, and the low-value features (i.e., the features in the third subset C1 and the sixth subset C2) can be eliminated, thereby improving the effectiveness of the target feature set.
[0089] For example, the first subset A1 may include a sliding window size, a window scaling factor, a control flag, and a maximum message length, the second subset B1 may include a maximum message length, a window scaling factor, a control flag, and a reception confirmation flag, the fourth subset A2 may include a window scaling factor, a control flag, a sliding window size, a survival time, a maximum message length, and an IP fragmentation flag, and the fifth subset B2 may include a reception confirmation flag, a total message length, a selection confirmation message sequence number, a TCP message checksum, and a user timeout duration identifier. In some examples, the high-value subset A may include a window scaling factor, a control flag, a sliding window size, a survival time, a maximum message length, and an IP fragmentation flag, the medium-value subset B may include a reception confirmation flag, and the target feature set may include a window scaling factor, a control flag, a sliding window size, a survival time, a maximum message length, an IP fragmentation flag, and a reception confirmation flag.
[0090] The following are the features in the target feature set:
[0091] Window scaling factor (Scale): When TCP is connected, it indicates the scaling multiple of the windows of both communicating parties, which is used to expand the window buffer during the transmission process.
[0092] Control flags (tcp_flags): represent different states of TCP data packets, which are distinguished by multiple identifiers, such as "S", "A", etc.
[0093] Sliding window size (Wsize): A field used by the TCP protocol for communication flow control. It is used to identify the amount of data that the system kernel's communication buffer can bear. Together with the window scaling factor, it constitutes an important part of TCP communication flow control.
[0094] Time to Live (TTL): defines the total number of routers that a data packet can pass through. Its main purpose is to prevent data packets from being transmitted repeatedly in the network, to clear redundant data packets in a timely manner, and to save network resources.
[0095] Maximum Message Size (MSS): A characteristic field in the TCP protocol options, which is defined as the length of the maximum data content in each data packet agreed upon by both parties of the TCP connection.
[0096] IP fragmentation flag (ip_flags): controls the fragmentation of IP packets during forwarding. The IP fragmentation flag can be divided into two types: DF and MF.
[0097] Receive confirmation flag (ACK): In the TCP / IP protocol stack, after either party in the communication confirms that the data sent by the other party is correct, it will reply with an ACK flag.
[0098] Figure 4Ais a diagram showing the distribution of TTL values involved in the examples of the present disclosure. Figure 4B is a diagram showing the value distribution of the control flag bits involved in the examples of the present disclosure. Figure 4C 2 is a diagram showing the distribution of IP message checksum values involved in the examples of the present disclosure. Figure 4D 2 is a diagram showing the distribution of TCP message checksum values involved in the examples of the present disclosure. Figure 4E is a diagram showing the distribution of values of the sliding window size involved in the examples of the present disclosure. Figure 4F is a diagram showing the distribution of values of the maximum message length involved in the examples of the present disclosure. Figure 4G is a diagram showing the value distribution of the window scaling factor involved in the examples of the present disclosure.
[0099] In addition, the present disclosure also uses the source IP address in the IP protocol field as a unique label to identify the operating system type, collects 261,432 data items by sampling and statistically obtains the value distribution of a part of the initial feature set, and displays it in the form of a bar chart.
[0100] Figure 4A and Figure 4B The figure is a distribution diagram of the values of TTL and control flag bits. In the sampled statistical data set, it can be seen that the value distribution of TTL is concentrated in 128, 127, 64 and 64-, among which 64- is the value when the value of TTL is less than 64.
[0101] In the distribution of control flag values, the "A" flag indicates the confirmation of the communication parties; "PA" is the PUSH flag, which notifies the receiver in the TCP connection to transmit the cached data to the receiving application; "S" is the TCP SYN flag, indicating the establishment of the connection; "FA" is the FIN flag, indicating the closing of the TCP connection; "R" is the RST flag, indicating the reset of the TCP connection. It can be seen from the distribution diagram that the "A" flag has the highest frequency of occurrence, far higher than the frequency of other TCP flags, followed by the "PA" and "S" flags.
[0102] By analyzing the frequency distribution of TTL and control flags, we can observe the differences between different communicating parties during the communication process. By combining these features, we can distinguish the differences in operating system types between different hosts, which serves as an important basis for identifying the operating system type.
[0103] Figure 4C and Figure 4D The distribution diagram of the IP packet checksum and TCP packet checksum, combined with the feature selection results (i.e., the target feature set) and Figure 4C and Figure 4DThe distribution of checksum values for different operating system types shown can lead to the following conclusion: In the checksum field, both the TCP message checksum and the IP message checksum have a wide range of values and little correlation with the specific operating system type.
[0104] Figure 4E is the distribution diagram of the sliding window size. Figure 4E Although the value range of the sliding window size field is large, it has an obvious value concentration area. The values of the sliding window size are concentrated in three intervals, which has a certain ability to distinguish different operating system types.
[0105] Figure 4F and Figure 4G is the distribution diagram of the maximum message length and window scaling factor. Figure 4F and Figure 4G In, presented with Figure 4A and Figure 4B Similar distribution, "*" indicates that the current statistical feature is not visible, that is, the maximum packet length and window scaling factor cannot be extracted from the traffic data. As can be seen from the figure, the maximum packet length and window scaling factor cannot be directly obtained in most packets. When this field has a specific value, the most common value of the maximum packet length is "1460", and the most common value of the window scaling factor is "7". From the statistical data in the above figure, it can be seen that the distribution of the values of the maximum packet length and the window scaling factor has a certain concentration trend. The maximum packet length values are concentrated on 1460 and "*", and the window scaling factor values are concentrated on "*" and "7".
[0106] In the actual operating system verification process, the amount of traffic data corresponding to different operating system types is also different. For example, some common Windows and Linux operating systems may correspond to multiple valid traffic data, while some uncommon operating systems have very few traffic data, and some even have only one.
[0107] In order to deal with the problem of uneven distribution of traffic data of different operating system types. In some examples, the target training set can be grouped to obtain multiple groups of data. In response to the number of types of samples of the corresponding operating system type in each group of the multiple groups of data being less than the preset types, the samples of the corresponding operating system type are copied to each group of the multiple groups of data to update the target training set. In this way, the samples can be expanded to solve the problem of uneven distribution of traffic data. In addition, the present disclosure does not limit the specific values of the preset types, for example, the preset types can be adjusted according to experimental results.
[0108] In some examples, the preset types may be 20. Taking 1,500,000 samples as an example, they may be divided into 10 groups, each group may have 150,000 data, and if there are less than 20 samples of a certain operating system type in each group of data, the samples of this operating system type may be copied to each group of data.
[0109] In some examples, the target feature set obtained by multiple feature selection can be verified. That is, the best feature subset can be selected from the initial feature set using multiple feature selection and verified. In some examples, in response to the validity of the target feature set meeting the preset requirements, the target training set can be determined based on the target feature set.
[0110] In some examples, multiple machine learning models can be used to verify the target feature set, and multiple machine learning models can be included in the naive Bayes model. Specifically, after comparing the recognition performance of multiple machine learning models, the effectiveness of the target feature set is verified, and it is also found that there is room for optimization in the application of the naive Bayes model. Therefore, it is proposed to solve the problem of independence between features of the naive Bayes model by feature weighting, and improve the recognition effect of the operating system type.
[0111] To this end, the present disclosure also provides some verification processes to verify the effectiveness of the target feature set determined by the multiple feature selection proposed in the present disclosure. Specifically, after the traffic data has been subjected to the above-mentioned preprocessing (such as quantization) and multiple feature selection, and the problem of uneven distribution of traffic data has been overcome, the target feature set is verified through a variety of traditional machine learning models, and the recognition performance (i.e., effectiveness) is compared from the three indicators of true positive rate, false positive rate, and precision rate.
[0112] The experiment uses three traditional machine learning models and Python's Scikit-learn to identify and verify the target feature set, and obtains the identification results of the operating system type corresponding to each host and the statistical results of the actual operating system type.
[0113] Combining the real statistical results with the recognition results obtained from the experiment, the statistical data of the recognition performance is calculated and compared, and the verification is carried out at two classification granularities: the main version of the operating system and the detailed version of the operating system. Among them, the main version classification granularity mainly distinguishes the types of operating systems at this level, such as Linux, Windows, VMware, etc., and the detailed version classification granularity distinguishes specific operating system types such as Windows7 and Windows 8 under Windows. The comparison of the recognition effect of the main version of the operating system and the comparison of the recognition effect of the detailed version of the operating system are shown in Table 4 and Table 5 respectively.
[0114]
[0115] Table 4 Comparison of recognition effects of major operating system versions
[0116]
[0117] Table 5 Comparison of operating system detailed version recognition results
[0118] It can be seen that based on the target feature set, SVM performs best, and the recognition effects of the naive Bayes model (NB) and the decision tree are not much different. Therefore, it can be considered that the target feature set obtained after multiple feature selection has the situation that the independence between features cannot be satisfied, which leads to the low recognition effect of the naive Bayes model; and the decision path obtained by decision tree training is easily affected by the missing fingerprint attributes in the test data set. The missing attributes easily lead to the decision path selection bias of the decision tree, which leads to the reduction of recognition effect; although SVM has better recognition performance than the previous two, in the problem of operating system type recognition, SVM handles the classification of multiple operating system types by constructing multiple binary classification SVMs. If K types of operating systems are classified, K(K-1) / 2 classifiers need to be constructed, which greatly occupies resources and is not conducive to practical use.
[0119] Therefore, the present disclosure considers implementing an optimized Naive Bayes model, thereby achieving a classification result with higher recognition performance.
[0120] In some examples, reference Figure 2 In step S104, the naive Bayes model may also be referred to as a model based on the naive Bayes algorithm. As described above, multiple values of the target feature set are used as the target training set. That is, the target training set may include multiple groups of values of the target feature set, and each group of values may include multiple values corresponding to multiple features in the target feature set. In other words, the target training set may include multiple samples, each sample may correspond to a group of values, and each group of values may include multiple values corresponding to multiple features in the target feature set.
[0121] As described above, the target training set can be used to train the naive Bayesian model that combines the feature weights of the features in the target feature set to obtain a trained model. The various features in the target feature set have different degrees of influence on the identification of the operating system type. For the target data set itself, by assigning different feature weights to different features, their importance can be better distinguished. For the naive Bayesian algorithm, the difference in feature weights is also related to the independence relationship between the various features. The naive Bayesian algorithm has such a premise assumption: the various features of the training data set are independent of each other, and the features in the target feature set of the present disclosure are all from the relevant fields in the network protocol, and there must be various subtle connections between these fields, making it difficult to make them completely independent of each other. Therefore, the recognition effect of the naive Bayesian algorithm in the target feature set of the operating system type is easily limited. Therefore, the naive Bayesian model that combines the feature weights of the features in the target feature set can improve the recognition effect of the operating system type.
[0122] In addition, the trained model (i.e., the naive Bayes model trained by the target training set) can be used to receive the data to be identified corresponding to the target feature set and output the classification result related to the operating system type. In some examples, the data to be identified can be traffic data in a network environment of passive traffic.
[0123] In some examples, relative entropy can be used to obtain the feature weights of features in the target feature set, and the feature weights can satisfy the formula:
[0124]
[0125] Among them, Weight represents the feature weight of M, M represents the feature in the target feature set, and P(M j ) indicates that the value of M is M j The probability of C represents the type of operating system, P(C) represents the prior probability of C, P(C|M j ) indicates that the value of M is M j The conditional probability that the operating system type is C when J represents the number of samples in the target training set, M j It represents the value when the feature of the jth sample in the target training set is M. In this way, the recognition effect of the naive Bayes model can be improved.
[0126] In some examples, a naive Bayes model combining the feature weights of the features in the target feature set can satisfy the formula:
[0127]
[0128] Among them, P post Indicates that each sample in the target training set belongs to Y kargMax represents the function used to determine the category with the largest posterior probability, Y k represents the kth operating system type (i.e., the kth category), P(Y k ) indicates Y k The corresponding prior probability, I represents the number of features in the target feature set, x i represents the value of the i-th feature of the target feature set in the sample, P(x i |Y k ) indicates that when the classification result is Y k In the case of x i The corresponding conditional probability, Weight i represents the feature weight of the i-th feature in the target feature set, and ∏ represents the product. Therefore, combining feature weights can improve the recognition effect of the naive Bayes model.
[0129] As mentioned above, Weight i Relative entropy can be used to obtain. In this case, using relative entropy to obtain feature weights can improve the accuracy of measuring the correlation between features and operating system types, thereby improving the recognition effect of the naive Bayes model.
[0130] As described above, the multiple flow data in step S101 may come from existing data. The existing data may not cover the characteristics of new operating systems (such as Windows 10 or Windows 11) and their updated versions, resulting in insufficient breadth and accuracy of identification.
[0131] In some examples, in response to the number of samples of the target operating system type in the target training set being less than the second preset number (i.e., the number of samples is insufficient or missing) and the usage ratio of the target operating system type being greater than the preset ratio (i.e., a large number of user groups are using it), the trained model is incrementally trained based on the incremental training set corresponding to the target operating system type. Thus, the robustness of the trained model can be improved. For example, the existing data (such as p0f, etc.) has the problem of missing fingerprints (i.e., missing samples) for Windows 10. In addition, the target operating system type can be one type of operating system type (e.g., Windows 10). In addition, the present disclosure does not limit the values of the second preset number and the preset ratio, and for example, they can be determined by statistical analysis of existing traffic data.
[0132] In some examples, the source of the incremental training set may be different from the source of the target training set. This helps to make up for the deficiencies of the target training set. In some examples, the target training set may be derived from existing data, and the incremental training set may be derived from experimental data. In this case, based on the target feature set obtained from the existing data and the trained model, incremental training is performed using the incremental training set obtained from the experimental data, which can improve the effectiveness and convenience of making up for the deficiencies of the target training set. In addition, the experimental data may be traffic data collected in an experimental environment. This enables the incremental training set to better match the target operating system type.
[0133] In some examples, the incremental training set can be determined by the HTTP protocol. In some examples, the incremental training set can be determined by the User_Agent field of the HTTP protocol. The shortcoming of the User-Agent field is that in a network environment with passive traffic, the field may be invisible or modified, so the use of this field alone has the problem of low recognition effect. The present disclosure uses this field in experimental data to construct an incremental training set, thereby expanding the number of samples of the target operating system type without affecting the recognition of the operating system type in a network environment with passive traffic in actual applications.
[0134] In some examples, the operating system type of the traffic data in the experimental data can be confirmed through the User_Agent field (that is, the label of the sample is determined), and the target feature set of the traffic data is determined based on the experimental data (see the relevant records of step S103 for details). In some examples, at least one set of values of the target feature set whose occurrence frequency is greater than the preset frequency can be selected from the experimental data as an incremental training set. In addition, the present disclosure does not limit the value of the preset frequency, which can be determined, for example, by experiments.
[0135] Taking Windows 10 as an example, the present disclosure determines the target feature set and multiple groups of values of the target feature set through the User_Agent field of the HTTP protocol and experimental data. According to the sorting of the frequency of occurrence of each group of values, the incremental training set determined by the principle of high frequency first is shown in Table 6.
[0136]
[0137] Table 6 Incremental training set
[0138] The examples disclosed herein also relate to an operating system identification method (also referred to as an operating system classification method). In some examples, the operating system identification method may include inputting the data to be identified in a network environment of passive traffic into a trained model trained by the training method involved in the examples disclosed herein to obtain a classification result related to the operating system type. In some examples, the values of a target feature set may be extracted from the traffic data to be identified, the values of the target feature set may be quantized and standardized to obtain the data to be identified, and the data to be identified may be input into the trained model to obtain a classification result.
[0139] The example of the present disclosure also relates to an electronic device, including a processor and a memory. The memory may store a computer program, and when the computer program is executed, one or more steps in the above-mentioned training method or operating system identification method are implemented.
[0140] The examples of the present disclosure also relate to a computer-readable storage medium, which can store at least one instruction, and when the at least one instruction is executed by a processor, one or more steps in the above-mentioned training method or operating system recognition method are implemented. The computer-readable storage medium can include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a micro drive, and a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nanosystem (including a molecular memory IC), or any type of medium or device suitable for storing instructions and / or data.
[0141] Figure 5 is a comparison chart showing the recognition performance of operating system types before and after the incremental training involved in the examples of the present disclosure.
[0142] In addition, in order to verify the effects of the training method and the scheme related to the training method (hereinafter referred to as the scheme) involved in the examples of this disclosure, this disclosure also describes some experimental results based on a real network environment, which does not limit this disclosure. The situation of the real network environment is as follows, and the Windows 10 operating system accounts for a high proportion, as shown in Table 7. Therefore, incremental training optimization of the Windows 10 series operating system is more important for real network environment applications.
[0143]
[0144] Table 7 Windows 10 system share
[0145] For the comparison of recognition accuracy of Windows series systems before and after Windows 10 incremental training optimization, since the operating system types of most samples in the real network environment of the experiment are Windows 7 and Windows 10, the recognition accuracy of these two system types is compared in five sample samples. The results are shown in Tables 8, 9 and Figure 5 As shown in Tables 8 and 9, after incremental training optimization, the recognition accuracy of Windows 10 operating system has been improved by about 20%, and the recognition accuracy of Windows 7 operating system, which also belongs to the Windows series, has no obvious impact. It proves that the incremental training optimization strategy proposed in the present disclosure effectively expands the sample of Windows 10, and has no obvious impact on the recognition performance of Windows 7, which is also in the Windows series, and avoids the impact on other system types belonging to the same Windows operating system.
[0146]
[0147]
[0148] Table 8 Before incremental training
[0149]
[0150] Table 9 After incremental training
[0151] In addition, the present disclosure also compares the present solution with the p0f identification tool in the same environment. In order to ensure that the test data covers enough operating system types and the number of samples of each operating system type is equivalent, the real network data collected by passive traffic is sampled, mainly covering common versions under Windows and Linux, where the number of samples of each operating system type is shown in Table 10.
[0152]
[0153] Table 10 Sample quantity distribution table
[0154] After completing the incremental training, in the experimental environment, the test indicators of the operating system identification method involved in the example of the present disclosure in Windows 7, Windows XP, Windows 8, Windows 10, Linux 2 and Linux 3 operating systems are shown in Table 11.
[0155]
[0156] Table 11 Test indicators after incremental training
[0157] The experimental results are shown in Table 12 and Table 13. The comparison of recall and precision in the table shows that the operating system identification method involved in the example of the present disclosure is generally much higher than the p0f identification tool in the two identification granularities of detailed version and main version. At the same time, the SVM, naive Bayes model and decision tree adopted in the target feature set verification process also have very obvious performance improvements. Among them, since p0f does not support the identification of Windows 10, the incremental training proposed in the present disclosure is also much better than the p0f identification tool in the two identification granularities of Windows 10.
[0158] It is worth noting that the recognition effects in Table 12 and Table 13 on Windows 8 and Windows XP are different from those on other operating systems. The reason is that in the corporate office environment where this experimental environment is located, the number of samples of Windows 8 and Windows XP operating systems is much lower than that of other types of operating systems. Therefore, when a target host is misidentified, the recall rate and precision rate are more affected.
[0159]
[0160] Table 12 Comparison of the recall rate of the operating system identification method of p0f and this solution
[0161]
[0162] Table 13 Comparison of the precision of the operating system identification method of p0f and this scheme
[0163] Based on the above experimental results, it can be clearly seen that this solution has great advantages in both recall rate and precision rate, whether it is the main version of the operating system or the detailed version of the operating system, and has good practical application effects in the real network environment under actual applications.
[0164] In various examples disclosed in the present invention, a multiple feature selection method is proposed, which combines two feature selection methods, namely, feature selection based on correlation and feature selection based on mutual information, and performs feature extraction on the basis of existing data to obtain a target feature set that supports passive identification, and performs analysis and experimental verification of the target feature set. On the basis of the obtained target feature set, in view of the problem that the naive Bayes model does not satisfy feature independence, it is proposed to optimize the naive Bayes model using feature weights to improve the recognition effect of the operating system type. In order to solve the problem of insufficient sample coverage of some operating system types, sample expansion is performed based on the HTTP protocol. In addition, comparative experiments are conducted using traffic data from a real network environment to verify the recognition capability of the solution involved in the examples disclosed in the present invention in terms of the target operating system type and its advantages in overall recognition performance.
[0165] Although the present disclosure is specifically described above in conjunction with the accompanying drawings and examples, it is to be understood that the above description does not limit the present disclosure in any form. Those skilled in the art may modify and change the present disclosure as needed without departing from the essential spirit and scope of the present disclosure, and these modifications and changes all fall within the scope of the present disclosure.
Claims
1. A training method for operating system recognition based on feature selection, characterized in that: The training method comprises: Get multiple traffic data of different operating system types; Obtaining the value of an initial feature set including a TCP feature set related to the TCP protocol and an IP feature set related to the IP protocol for each of the plurality of traffic data, and using the values of the plurality of initial feature sets as an initial training set; Performing feature selection on the initial feature set using correlation-based feature selection and the initial training set to obtain at least one feature subset as a first feature set, scoring the features in the initial feature set using a mutual information-based evaluation function and the initial training set, selecting a first preset number of features with the highest scores from the initial feature set as a second feature set, and performing an intersection and union operation based on the first feature set and the second feature set to obtain a target feature set; and Take multiple values of the target feature set as a target training set, obtain feature weights of the features in the target feature set based on the target training set, use the target training set to train a naive Bayesian model combined with the feature weights of the features in the target feature set to obtain a trained model, and the trained model is used to receive the data to be identified corresponding to the target feature set in a network environment with passive traffic and output a classification result related to the operating system type.
2. The training method according to claim 1, characterized in that: The evaluation function based on mutual information satisfies the formula: J MI (X t )=I(X t ;Y), Among them, J MI (X t ) indicates X t The evaluation function, X t represents the set of values of the tth feature of the initial feature set in the initial training set, I(X t ; Y) represents X expressed in the form of entropy t Mutual information with Y, where Y is the set of operating system types.
3. The training method according to claim 1, characterized in that: The at least one feature subset includes a first subset, a second subset and a third subset with decreasing scores. According to the scores of each feature in the second feature set, the second feature set is divided into a fourth subset, a fifth subset and a sixth subset. The score of the feature in the fourth subset is greater than the score of the feature in the fifth subset and the score of the feature in the fifth subset is greater than the score of the feature in the sixth subset. In the intersection operation, the first subset and the fourth subset are taken as a high-value subset, the second subset and the fifth subset are taken as an intersection as a medium-value subset, and the high-value subset and the medium-value subset are taken as a union as the target feature set.
4. The training method according to claim 1, characterized in that: The IP feature set includes the IP protocol version number, IP message header length, service type field, total message length, IP fragmentation flag, fragmentation offset, lifetime, protocol identification bit, IP message checksum, and IP message optional items; the TCP feature set includes the control flag, reserved bit, sliding window size, TCP message checksum, urgent pointer, time sampling enabled flag, maximum message length, window scaling factor, selection confirmation message sequence number, reception confirmation flag, MD5 signature, TCP header length, and user timeout duration flag.
5. The training method according to claim 4, characterized in that: The first subset includes sliding window size, window scaling factor, control flag, and maximum message length, the second subset includes maximum message length, window scaling factor, control flag, and receive confirmation flag, the fourth subset includes window scaling factor, control flag, sliding window size, lifetime, maximum message length, and IP fragmentation flag, and the fifth subset includes receive confirmation flag, total message length, selection confirmation message sequence number, TCP message checksum, and user timeout duration identifier.
6. The training method according to claim 1, characterized in that: The feature weights of the features in the target feature set are obtained using relative entropy, and the feature weights satisfy the formula: Among them, Weight represents the feature weight of M, M represents the feature in the target feature set, P(M j ) indicates that the value of M is M j The probability of C represents the type of operating system, P(C) represents the prior probability of C, P(C|M j ) indicates that the value of M is M j The conditional probability that the operating system type is C when J represents the number of samples in the target training set, M j Indicates the value when the feature of the j-th sample in the target training set is M.
7. The training method according to claim 6, characterized in that: The naive Bayes model satisfies the formula: Among them, P post Indicates that each sample in the target training set belongs to Y k argMax represents the function used to determine the category with the largest posterior probability, Y k represents the kth operating system type, P(Y k ) indicates Y k The corresponding prior probability, I represents the number of features in the target feature set, x i represents the value of the i-th feature of the target feature set in the sample, P(x i |Y k ) indicates that when the classification result is Y k In the case of x i The corresponding conditional probability, Weighti represents the feature weight of the i-th feature in the target feature set, and ∏ represents the product.
8. The training method according to claim 1, characterized in that: The target training set is grouped to obtain multiple groups of data. In response to the number of types of samples of the corresponding operating system type in each group of the multiple groups of data being less than a preset number of types, the samples of the corresponding operating system type are copied to each group of the multiple groups of data to update the target training set.
9. The training method according to claim 1, characterized in that: In response to the number of samples of the target operating system type in the target training set being less than a second preset number and the usage proportion of the target operating system type being greater than a preset proportion, incremental training is performed on the trained model based on an incremental training set corresponding to the target operating system type, and the target operating system type is one type of operating system type.
10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed, the training method according to any one of claims 1 to 9 is implemented.