Unknown protocol clustering method and system based on model interpretability analysis

By combining N-gram feature extraction and deep autoencoder dimensionality reduction with multiple clustering algorithms, the problems of feature vector repeatability and difficulty in estimating the number of clusters in unknown protocol clustering are solved, achieving more accurate clustering results and explainable display.

CN120692204APending Publication Date: 2025-09-23Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510794707.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-14
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the existing technology, cluster analysis of unlabeled protocol message data suffers from feature vector duplication and difficulty in estimating the number of clusters, resulting in inaccurate clustering results and affecting the classification and structural inference of unknown protocols.

Method used

N-gram feature extraction and deep autoencoder are used for feature dimensionality reduction, multiple clustering algorithms are combined to perform two clustering operations, and cross-category heat map analysis is used for feature visualization.

Benefits of technology

The accuracy and reliability of unknown protocol clustering are improved, the interpretability of the model is enhanced, and the feature differences and similarities of different categories are intuitively presented through the combination of multiple algorithms and feature visualization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120692204A_ABST
    Figure CN120692204A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and information security, and provides an unknown protocol clustering method and system based on model interpretability analysis. The method comprises the following steps: step 1, acquiring communication protocol data and performing feature extraction to obtain high-dimensional protocol features; step 2, inputting the high-dimensional protocol features into a depth auto-encoder for feature dimension reduction to obtain low-dimensional protocol features; 3, clustering the low-dimensional protocol features by adopting a plurality of clustering algorithms, and taking an optimal clustering result as a primary clustering result; step 4, clustering undefined label data in the primary clustering result by adopting multiple clustering algorithms, and taking an optimal clustering result as an unknown protocol clustering result; and step 5, combining the primary clustering result and the unknown protocol clustering result, calculating the features of each cluster in the results, performing cross-category heat map analysis to realize feature visualization, and generating a feature importance list.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and information security technology, and in particular to a method and system for clustering unknown protocols based on model interpretability analysis. Background Art

[0002] Protocol reverse engineering is a comprehensive technique that aims to delve deeply into a protocol's interaction patterns, grammatical structure, and semantics by analyzing network traffic, extracting data features, and inferring data frame specifications, without relying on protocol specifications. Its core goal is to fully understand the protocol's behavioral characteristics in order to effectively detect potential security threats and promote smooth interoperability between different systems. Clustering analysis is a key initial step in protocol reverse engineering. Faced with large amounts of unlabeled protocol message data, clustering techniques can automatically classify messages based on their inherent characteristics, providing a foundation for subsequent in-depth analysis. In practical applications, accurate clustering helps quickly identify anomalous or unknown groups of protocol messages. Clustering of private protocol message data enables classification and structural inference of unknown protocols by analyzing the format, semantics, and statistical characteristics of protocol messages. Existing research methods for this specific problem primarily focus on optimizing the combination of feature extraction and clustering algorithms.

[0003] In the feature extraction process, for message clustering based on bit-level definitions, the usual practice is to extract bit sequences of a specific length from the bit stream as the basic feature unit. For example, the message can be divided into multiple subsequences according to a fixed bit length, and the frequency of occurrence of each subsequence can be counted, and these frequencies can be used as components of the feature vector. However, this method has certain limitations in practical applications. Specifically, messages of multiple different protocols may show a high degree of similarity in the distribution of these bit sequences, resulting in the extracted feature vectors being repetitive. This repetitiveness will reduce the accuracy of subsequent clustering analysis, making it difficult to effectively distinguish messages of different protocol types, thereby affecting the reliability of the clustering results.

[0004] When it comes to optimizing clustering algorithms, traditional clustering methods, such as K-means and hierarchical clustering, are often used. For example, the K-means algorithm requires a pre-determined number of clusters, and the selection of initial cluster centers significantly impacts the final clustering results. In practical applications, especially when clustering proprietary protocol messages, accurately estimating the appropriate number of clusters is a challenge. Furthermore, messages from different protocols may be distributed in similar cluster centers, making it difficult to clearly assign clusters, leading to inaccurate clustering results. This inaccuracy impairs the ability to classify and infer the structure of unknown protocols, hindering in-depth analysis of proprietary protocols. Summary of the Invention

[0005] To address the above problems, the present invention proposes a method and system for clustering unknown protocols based on model interpretability analysis. The N-gram features are used to extract protocol features, and feature dimensionality reduction is performed through an autoencoder. Finally, a multi-clustering algorithm is used to perform two clusterings, and the clustering results are visualized.

[0006] In a first aspect, the present invention provides an unknown protocol clustering method based on model interpretability analysis, comprising:

[0007] Step 1: Obtain communication protocol data and perform feature extraction to obtain high-dimensional protocol features;

[0008] Step 2: Input the high-dimensional protocol features into a deep autoencoder for feature dimensionality reduction to obtain low-dimensional protocol features;

[0009] Step 3: clustering the low-dimensional protocol features respectively using multiple clustering algorithms, and taking the optimal clustering result as the initial clustering result; wherein the multiple clustering algorithms include K-means algorithm, DBSCAN, hierarchical clustering and Gaussian mixture model;

[0010] Step 4: clustering the undefined label data in the initial clustering result using the multiple clustering algorithms, and taking the optimal clustering result as the unknown protocol clustering result;

[0011] Step 5: Merge the initial clustering results and the unknown protocol clustering results and calculate the features of each cluster in the results, perform cross-category heat map analysis to achieve feature visualization, and generate a feature importance list.

[0012] Furthermore, in step 1, feature extraction is performed by adopting the N-gram algorithm.

[0013] Furthermore, the step 1 also includes: simultaneously expanding each byte into 8 bits through a binary feature extraction method, and then performing feature statistics to capture finer-grained protocol features.

[0014] Furthermore, the deep autoencoder includes three encoding layers, and the encoding layer structure includes a fully connected layer, a normalization processing layer and a regularization layer.

[0015] Furthermore, step 2 also includes using the SMOTE algorithm to balance data distribution of the encoded features.

[0016] Furthermore, in step 3, the optimal clustering result is determined based on the ARI / NMI index.

[0017] Furthermore, in step 4, the optimal clustering result is determined based on the silhouette coefficient and the DBI / CH index.

[0018] Furthermore, the elbow method is used to determine the number of clusters of the K-means algorithm, the hierarchical clustering algorithm and the Gaussian mixture model respectively.

[0019] In a second aspect, the present invention provides an unknown protocol clustering system based on model interpretability analysis, comprising:

[0020] A feature extraction unit is used to obtain communication protocol data and perform feature extraction to obtain high-dimensional protocol features;

[0021] A feature dimensionality reduction unit, configured to input the high-dimensional protocol features into a deep autoencoder for feature dimensionality reduction to obtain low-dimensional protocol features;

[0022] A first clustering unit is configured to cluster the low-dimensional protocol features respectively using a plurality of clustering algorithms, and use the optimal clustering result as the initial clustering result; wherein the plurality of clustering algorithms include K-means algorithm, DBSCAN, hierarchical clustering and Gaussian mixture model;

[0023] A second clustering unit is configured to cluster the undefined label data in the initial clustering result using the plurality of clustering algorithms, and use the optimal clustering result as the unknown protocol clustering result;

[0024] A feature visualization unit is used to merge the initial clustering results and the unknown protocol clustering results and calculate the features of each cluster in the results, perform cross-category heat map analysis to achieve feature visualization, and generate a feature importance list.

[0025] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described above when executing the program.

[0026] The beneficial effects of the present invention are:

[0027] The unknown protocol clustering method provided by the present invention extracts high-dimensional protocol features from bit-defined data frames, and uses deep autoencoders to perform feature dimensionality reduction, removes noise and redundant information in the data, highlights key features, and enables low-dimensional protocol features to better reflect the essential characteristics of the data. This provides higher-quality input for subsequent clustering algorithms, helps to improve the accuracy of clustering, and enables the clustering results to more truly reflect the inherent structure and category division of the protocol data. Different clustering algorithms have different advantages for different characteristics of data and clustering scenarios. Clustering through multiple clustering algorithms and selecting the optimal clustering results can make up for the shortcomings of a single algorithm, thereby improving the accuracy and reliability of the clustering results. The present invention uses cross-category heat map analysis to visualize the features of each cluster in the clustering results, intuitively presenting the feature differences and similarities between different categories, which helps analysts to have a deeper understanding of the decision-making basis of the clustering model, thereby enhancing the interpretability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A schematic diagram of a process for clustering unknown protocols based on model interpretability analysis provided by an embodiment of the present invention;

[0029] Figure 2 A heat map of the importance of features of each cluster provided by an embodiment of the present invention;

[0030] Figure 3 A cluster feature importance ranking diagram provided by an embodiment of the present invention;

[0031] Figure 4 A clustering result confusion diagram provided by an embodiment of the present invention;

[0032] Figure 5 A schematic diagram of the deep autoencoding structure provided by an embodiment of the present invention;

[0033] Figure 6 The training loss curve of the autoencoder provided in the embodiment of the present invention;

[0034] Figure 7 A comparison chart of data before and after sampling using the SMOTE algorithm provided by an embodiment of the present invention;

[0035] Figure 8 A schematic diagram of determining the optimal number of clusters using the elbow method provided in an embodiment of the present invention;

[0036] Figure 9 Schematic diagram of BIC scores for different GMM covariance types and cluster numbers provided by an embodiment of the present invention;

[0037] Figure 10 ARI and NMI value effect ranking diagram of each clustering algorithm provided in the embodiment of the present invention; Figure 10 -a is the ARI value effect sorting, Figure 10 -b is the NMI value effect sorting;

[0038] Figure 11 The silhouette coefficient, DBI and CH index graph corresponding to different cluster numbers provided in the embodiment of the present invention; Figure 11 -a is the silhouette coefficient graph, Figure 11 -b is the DBI index graph, Figure 11 -c is the CH index graph;

[0039] Figure 12 A structural framework diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0040] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0041] like Figure 1 As shown, an embodiment of the present invention provides an unknown protocol clustering method based on model interpretability analysis, including:

[0042] Step 1: Acquire communication protocol data and perform feature extraction to obtain high-dimensional protocol features. The communication protocol data is a bit-oriented data frame.

[0043] Step 2: Input the high-dimensional protocol features into the deep autoencoder for feature dimensionality reduction to obtain low-dimensional protocol features.

[0044] Step 3: Use multiple clustering algorithms to cluster the low-dimensional protocol features respectively, and use the optimal clustering result as the initial clustering result; the multiple clustering algorithms include K-means algorithm, DBSCAN, hierarchical clustering and Gaussian mixture model.

[0045] Step 4: Use multiple clustering algorithms to cluster the undefined label data in the initial clustering results, and use the optimal clustering result as the unknown protocol clustering result.

[0046] Step 5: Merge the initial clustering results and the unknown protocol clustering results and calculate the features of each cluster in the results. Perform cross-category heat map analysis to visualize the features and generate a feature importance list.

[0047] Specifically, we calculate the N-gram features of each cluster and identify the key n-gram patterns that distinguish the current protocol cluster from other clusters. We use Seaborn to perform cross-category heatmap analysis to visualize the features. Red indicates that the feature is significant in the current cluster, and blue indicates that it is more prominent in other clusters. We also generate a list of feature importance. Figure 2 As shown, the cross-category heat map of features provided by an embodiment of the present invention. Figure 3 This is a ranking diagram of the importance of each cluster feature provided by the embodiment of the present invention. Figure 4 As shown, t-SNE visualization is performed on different categories of data features.

[0048] The method provided by the embodiment of the present invention can improve the accuracy of segmentation of undefined categories and achieve more accurate recognition effect by extracting features of the communication protocol and using multiple algorithms for clustering.

[0049] On the basis of the above embodiment, the embodiment of the present invention provides a method for feature extraction in step 1: feature extraction is performed by adopting an N-gram algorithm.

[0050] Specifically, CountVectorizer / TfidfVectorizer is used to generate a character-level n-gram feature matrix. Character-level n-grams are used to capture local patterns in protocol data, enabling variable-length data input. An adjustable n-gram window algorithm automatically adjusts the feature extraction window size based on the length and structure of the input data, avoiding interference from high-frequency terms. Redundant bits can also be removed to extract valid fields for feature extraction.

[0051] Furthermore, through the binary feature extraction method, each byte is expanded into 8 bits, and then feature statistics are performed to capture more fine-grained protocol features.

[0052] Based on the above embodiment, this embodiment provides a method for constructing and training a deep autoencoder, including:

[0053] like Figure 5 As shown in the figure, a deep network consisting of an encoding layer (32 dimensions) and a decoding layer is constructed. Through multi-layer nonlinear transformations, high-dimensional features are encoded in low dimensions, solving the curse of dimensionality. The number of encoding layer nodes can be adaptively set according to the input feature dimension, balancing information preservation and overfitting risks.

[0054] The improved binary cross entropy function is used to train the deep network, in which the improved binary cross loss entropy function adds two steps when calculating the loss: clipping the prediction value and taking the maximum value. The purpose of clipping the prediction value is to prevent infinite loss values ​​caused by extreme values ​​(such as 0 or 1). Taking the maximum value is to ensure that the loss value is always non-negative and avoid negative values ​​from interfering with the training process. After the deep network training is completed, the encoding layer is extracted as a deep autoencoder for feature dimensionality reduction. In addition, the early stopping mechanism is used to prevent the model training from overfitting, thereby saving the best encoder model. Figure 6 As shown in Figure 2, a schematic diagram of the deep network training loss curve verifies the convergence of the model.

[0055] The cropping operation is expressed as follows:

[0056] y pred =clip(y,1e-7,1-1e-7)

[0057] Among them, y pred Represents the eigenvalues ​​of the clipped deep network output, clip represents clipping, y represents the eigenvalues ​​of the deep network output, and clips the predicted eigenvalues ​​to the range of [1e-7, 1-1e-7] to avoid infinite values ​​in logarithmic calculations.

[0058] The binary cross entropy loss function is expressed as follows:

[0059] L=-(y true log(y pred )+(1-y true )log(1-y pred ))

[0060] Among them, y pred represents the feature value of the deep network output after clipping, y true represents the true feature value of the training data,

[0061] The non-negativity guarantee is expressed as follows:

[0062] L final =max(L,0)

[0063] Among them, L final Represents the final loss value, L represents the calculated binary cross entropy loss value, which represents the difference between the model prediction value and the true value.

[0064] Therefore, the improved binary cross entropy loss function can be expressed as:

[0065] L final =max(-(y true ·log(clip(y,1e-7,1-1e-7))+(1-ytrue )·log(1-clip(y,1e-7,1-1e-7))),0).

[0066] Based on the above embodiment, step 2 further includes using the SMOTE algorithm to balance data distribution of the encoded features.

[0067] Specifically, in the coding space, the SMOTE algorithm is applied to the coding features of the minority class samples to generate the coding features of the new minority class samples. The distance between the minority class samples is calculated, the nearest neighbor of each minority class sample is found, and then new sample codes are generated between these nearest neighbors through interpolation and other methods. Figure 7 As shown in Figure 2, the data samples before and after applying the SMOTE algorithm can help improve the balanced distribution of data by operating the encoded features with the SMOTE algorithm.

[0068] As a nonlinear dimensionality reduction space, deep autoencoders can learn the intrinsic representation of data, map the original high-dimensional data to a low-dimensional encoding space, extract the intrinsic characteristics of the data, and perform SMOTE oversampling in the encoding space to generate samples that are more consistent with the actual distribution of the data. Compared with performing SMOTE directly in the original space, the generated minority class samples are of higher quality.

[0069] Based on the above embodiment, the elbow method is used to determine the number of clusters of the K-means algorithm, the hierarchical clustering algorithm and the Gaussian mixture model. Figure 8 As shown in FIG, by analyzing the curve of the inertia of the clustering results changing with the number of clusters, the inflection point (ie, the “elbow point”) is found as the optimal number of clusters.

[0070] Based on the above embodiment, in step 3 and step 4, the covariance matrix of the Gaussian mixture model is first determined. Specifically, the BIC scores of different covariance matrices are calculated and the BIC scores are used as statistical indicators to select the Gaussian mixture model. Figure 9 As shown, a schematic diagram of BIC scores for different Gaussian mixture model covariance types and cluster numbers.

[0071] Based on the above embodiment, in step 3 of this embodiment, the optimal clustering result is determined according to the ARI / NMI index.

[0072] ARI is an adjusted Rand index used to evaluate the similarity between clustering results and true labels, taking into account random guessing. The value range is between [-1,1]. The larger the value, the more similar the clustering result is to the true label. When the clustering result is completely consistent with the true label, the ARI index is 1; when the clustering result is completely unrelated to the true label, the ARI index is close to 0; when the correlation between the clustering result and the true label is lower than random guessing, the ARI index may be negative. The NMI index is an information theory indicator used to measure the degree of mutual information between the clustering result and the true label. The value range is between [0,1]. The larger the value, the more mutual information there is between the clustering result and the true label, that is, the more similar the clustering result is to the true label. The larger the value of the NMI index, the more mutual information there is between the clustering result and the true label, and the better the clustering effect. Figure 10 As shown in the figure, the effect ranking diagram of ARI and NMI index of each clustering algorithm is shown.

[0073] Based on the above embodiment, in step 4 of this embodiment, the optimal clustering result is determined according to the silhouette coefficient and the DBI / CH index.

[0074] The silhouette coefficient is used to measure the quality of clustering results. It combines the cohesion and separation of clusters and can reflect the closeness of data points in the cluster to which they belong and the degree of separation from other clusters. The value range of the silhouette coefficient is between [-1,1]. The larger the value, the better the clustering effect. DBI is used to evaluate the quality of clustering algorithms. The clustering effect is measured by calculating the ratio of the sum of intra-class distances to the inter-class distances. The smaller the DBI value, the tighter the samples within the class, the higher the inter-class separation, and the better the clustering effect. The CH index evaluates the clustering quality by the ratio of the intra-class squared error to the sum of the inter-class squared errors. The larger the value, the better the clustering effect. The larger the value of the CH index, the higher the inter-class separation, the better the intra-class closeness, and the more ideal the clustering effect. Figure 11 As shown, a schematic diagram of the silhouette coefficient, DBI and CH index corresponding to different cluster numbers.

[0075] Based on the above embodiment, this embodiment further provides an unknown protocol clustering system based on model interpretability analysis, including:

[0076] A feature extraction unit is used to obtain communication protocol data and perform feature extraction to obtain high-dimensional protocol features;

[0077] The feature dimensionality reduction unit is used to input the high-dimensional protocol features into the deep autoencoder for feature dimensionality reduction to obtain low-dimensional protocol features;

[0078] The first clustering unit is used to cluster the low-dimensional protocol features using multiple clustering algorithms, and use the optimal clustering result as the initial clustering result; the multiple clustering algorithms include K-means algorithm, DBSCAN, hierarchical clustering and Gaussian mixture model;

[0079] The second clustering unit is used to cluster the undefined label data in the initial clustering result using multiple clustering algorithms, and use the optimal clustering result as the unknown protocol clustering result;

[0080] The feature visualization unit is used to calculate the features of each cluster in the unknown protocol clustering results, perform cross-category heat map analysis to achieve feature visualization, and generate a feature importance list.

[0081] Based on the above embodiments, Figure 12 As shown, this embodiment further provides an electronic device, which may include: a processor (processor) 1201, a communication interface (Communications Interface) 1202, a memory (memory) 1203 and a communication bus 1204, wherein the processor 1201, the communication interface 1202, and the memory 1203 communicate with each other via the communication bus 1204. The processor 1201 may call the logic instructions in the memory 1203 to execute the method provided in the above embodiment, for example, including:

[0082] Step 1: Acquire communication protocol data and perform feature extraction to obtain high-dimensional protocol features; Step 2: Input the high-dimensional protocol features into the deep autoencoder for feature dimensionality reduction to obtain low-dimensional protocol features; Step 3: Use multiple clustering algorithms to cluster the low-dimensional protocol features separately, and use the optimal clustering result as the initial clustering result; the multiple clustering algorithms include K-means algorithm, DBSCAN, hierarchical clustering and Gaussian mixture model; Step 4: Use multiple clustering algorithms to cluster the undefined label data in the initial clustering results, and use the optimal clustering result as the unknown protocol clustering result; Step 5: Merge the initial clustering results and the unknown protocol clustering results and calculate the features of each cluster in the results, perform cross-category heat map analysis to achieve feature visualization, and generate a feature importance list.

[0083] In addition, when the logic instructions in the above-mentioned memory 1203 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for clustering unknown protocols based on model interpretability analysis, characterized in that: include: Step 1: Obtain communication protocol data and perform feature extraction to obtain high-dimensional protocol features; Step 2: Input the high-dimensional protocol features into a deep autoencoder for feature dimensionality reduction to obtain low-dimensional protocol features; Step 3: clustering the low-dimensional protocol features respectively using multiple clustering algorithms, and taking the optimal clustering result as the initial clustering result; wherein the multiple clustering algorithms include K-means algorithm, DBSCAN, hierarchical clustering and Gaussian mixture model; Step 4: clustering the undefined label data in the initial clustering result using the multiple clustering algorithms, and taking the optimal clustering result as the unknown protocol clustering result; Step 5: Merge the initial clustering results and the unknown protocol clustering results and calculate the features of each cluster in the results, perform cross-category heat map analysis to achieve feature visualization, and generate a feature importance list.

2. The unknown protocol clustering method based on model interpretability analysis according to claim 1, characterized in that: In step 1, feature extraction is performed by adopting the N-gram algorithm.

3. The unknown protocol clustering method based on model interpretability analysis according to claim 1, characterized in that: The step 1 also includes: simultaneously expanding each byte into 8 bits through a binary feature extraction method, and then performing feature statistics to capture finer-grained protocol features.

4. The unknown protocol clustering method based on model interpretability analysis according to claim 1, characterized in that: The deep autoencoder includes three encoding layers, and the encoding layer structure includes a fully connected layer, a normalization processing layer and a regularization layer.

5. The unknown protocol clustering method based on model interpretability analysis according to claim 1, characterized in that: The step 2 also includes using the SMOTE algorithm to balance the data distribution of the encoded features.

6. The unknown protocol clustering method based on model interpretability analysis according to claim 1, characterized in that: In step 3, the optimal clustering result is determined based on the ARI / NMI index.

7. The unknown protocol clustering method based on model interpretability analysis according to claim 1, characterized in that: In step 4, the optimal clustering result is determined based on the silhouette coefficient and DBI / CH index.

8. The unknown protocol clustering method based on model interpretability analysis according to claim 1, characterized in that: The elbow method is used to determine the number of clusters of the K-means algorithm, the hierarchical clustering algorithm and the Gaussian mixture model respectively.

9. An unknown protocol clustering system based on model interpretability analysis, characterized in that: include: A feature extraction unit is used to obtain communication protocol data and perform feature extraction to obtain high-dimensional protocol features; A feature dimensionality reduction unit, configured to input the high-dimensional protocol features into a deep autoencoder for feature dimensionality reduction to obtain low-dimensional protocol features; A first clustering unit is used to cluster the low-dimensional protocol features using multiple clustering algorithms, and use the optimal clustering result as the initial clustering result; The multiple clustering algorithms include K-means algorithm, DBSCAN, hierarchical clustering and Gaussian mixture model; A second clustering unit is configured to cluster the undefined label data in the initial clustering result using the plurality of clustering algorithms, and use the optimal clustering result as the unknown protocol clustering result; A feature visualization unit is used to merge the initial clustering results and the unknown protocol clustering results and calculate the features of each cluster in the results, perform cross-category heat map analysis to achieve feature visualization, and generate a feature importance list.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.