Block cipher algorithm identification method based on distance measurement and ensemble learning

By combining distance measurement and ensemble learning methods, ciphertext features are extracted and integrated learning classifiers are built, and the problems of accuracy and robustness of the existing technology in complex ciphertext environments are solved, and more efficient cipher algorithm recognition and better generalization capabilities are achieved.

CN119939349AActive Publication Date: 2025-05-06GUILIN UNIV OF ELECTRONIC TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510033956.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

In a complex hybrid ciphertext environment, existing machine learning-based cipher algorithm recognition methods are difficult to adapt to the diversity and complexity of ciphertext, resulting in limited accuracy and effectiveness of feature extraction. In addition, a single classifier is difficult to fully utilize the advantages of the model when facing ciphertext data of different types and complexities, resulting in insufficient robustness and generalization capabilities of the recognition results.

Method used

The block cryptographic algorithm recognition method based on distance measurement and integrated learning is adopted. By calculating the distance measurement value between plain text and cipher text and the information entropy value of cipher text, cipher text features are extracted, and an integrated learning classifier is built using a stacking method for model training and testing, improving the recognition accuracy of the encryption algorithm used for unknown cipher text data.

Benefits of technology

It improves the accuracy and robustness of cryptographic algorithm recognition, enhances the generalization ability of the model, is suitable for various encryption scenarios and data distribution environments, and has good classification effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939349A_ABST
    Figure CN119939349A_ABST
Patent Text Reader

Abstract

The invention discloses a block cipher algorithm identification method based on distance measurement and ensemble learning, and the method comprises the steps: firstly selecting a public data set, carrying out the division, building a plaintext data set, encrypting the plaintext data set through a set block cipher algorithm, and obtaining a ciphertext data set; calculating a distance metric value between the plaintext and the ciphertext and an information entropy value of ciphertext data, extracting ciphertext features to obtain a feature data set, and establishing a corresponding label set; constructing an integrated learning classifier by using a stacking method, carrying out model training by using a feature data set and a label, and finding an optimal hyper-parameter classifier model through experimental optimization; and finally, for unknown ciphertext data, using the trained ensemble learning classifier to complete the identification of the cryptographic algorithm. According to the method, stacking integration is adopted as an algorithm of a classifier model, feature extraction is carried out based on distance measurement and information entropy, the feature extraction capability and generalization of the model can be enhanced, and therefore the accuracy of model recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of information security, and in particular to a block cipher algorithm identification method based on distance measurement and ensemble learning. Background Art

[0002] Cryptographic algorithm identification is a method of extracting features from an encryption system to identify and classify the specific encryption algorithm type used, given only the ciphertext. This method extracts key indicators that reflect the unique behaviors and patterns of different encryption algorithms by analyzing the statistical features, structural features, and other implicit characteristics of the ciphertext.

[0003] At present, in the research of cryptographic algorithms based on machine learning, the feature extraction method often uses NIST randomness test, entropy and other methods to analyze the randomness of ciphertext, and uses this as the feature value for identification and classification. In order to improve the efficiency of cryptographic algorithm recognition, machine learning methods are used as classification algorithms to construct classification identifiers, extract ciphertext features and train models. Commonly used machine learning methods include logistic regression, decision tree, SVM, random forest, AdaBoost, Bagging, etc. However, in a complex mixed ciphertext environment, the randomness test index is difficult to adapt to the diversity and complexity of ciphertext, resulting in the accuracy and effectiveness of feature extraction being limited. When a single classifier faces ciphertext data of different types and complexity, it may not be able to take into account the recognition requirements of all features, and it is difficult to give full play to the advantages of the model, resulting in insufficient robustness and generalization ability of the recognition results, and it is difficult to effectively improve the classification results.

[0004] In view of the above problems, the present invention proposes a block cipher algorithm identification method based on distance measurement and ensemble learning. The distance measurement value between ciphertext and plaintext is calculated by multiple parties, and combined with the ciphertext information entropy value to form a feature vector. The stacking method is used to construct an ensemble learning classifier to train and test the model, thereby improving the recognition accuracy of the encryption algorithm used for unknown ciphertext data. Summary of the invention

[0005] The present invention proposes a block cipher algorithm identification method based on distance metric and ensemble learning. The method first selects a public data set for division to establish a plaintext data set, and encrypts the plaintext data set with a set block cipher algorithm to obtain a ciphertext data set; calculates the distance metric value between the plaintext and the ciphertext and the information entropy value of the ciphertext data, extracts the ciphertext features, obtains the feature data set, and establishes a corresponding label set; uses a stacking method to construct an ensemble learning classifier, trains the model with the feature data set and labels, and finds the optimal hyperparameter classifier model through experimental tuning; finally, for unknown ciphertext data, uses the trained ensemble learning classifier to complete the identification of the cipher algorithm.

[0006] The technical solution for achieving the purpose of the present invention is:

[0007] A block cipher algorithm identification method based on distance measurement and ensemble learning specifically includes the following steps:

[0008] (1) Data preparation and encryption;

[0009] Select a public data set and divide it into different sizes to obtain a plaintext data set, and use different block cipher algorithms to encrypt all the plaintext data in the plaintext data set to generate ciphertext data, thereby forming a ciphertext data set;

[0010] (2) Feature extraction based on distance measurement and information entropy;

[0011] Calculate the distance measurement value between the plaintext vector and the ciphertext vector and the information entropy value of the ciphertext vector, extract the ciphertext features, obtain the feature data set, and set the corresponding cryptographic algorithm label for each feature data in the feature data set;

[0012] (3) Construction and training of ensemble learning classifiers;

[0013] Build an ensemble learning classifier based on the stacking method, select the gradient boosting decision tree as the meta-classifier, random forest, SVM, and decision tree as the base classifier, train the model with feature data sets and labels, and find the optimal hyperparameter classifier model through experimental tuning;

[0014] (4) Identify cryptographic algorithms based on trained ensemble learning classifiers;

[0015] For unknown ciphertext data, the trained ensemble learning classifier is used to identify the block cipher algorithm.

[0016] In the block cipher algorithm identification method based on distance metric and ensemble learning of the present invention, the data preparation and encryption in step (1) are specifically performed as follows:

[0017] (1.1) Select a public dataset as a plaintext dataset. The public dataset can be a text file, an image file, etc.

[0018] (1.2) Select or divide pn files with file sizes of 1KB, 16KB, and 32KB from the public data set to form three plaintext data sets of different sizes. Finally, the plaintext data set Plaintext = {pl 1,1 ,pl 1,2 ,…,pl i,j ,…pl 3,pn}, where pl i,j Indicates the j-th file with file size i, 1≤i≤3, 1≤j≤pn;

[0019] (1.3) Convert all files in the plaintext dataset Plaintext into binary files;

[0020] (1.4) For all file data of different sizes in the plaintext dataset Plaintext, the open source cryptographic libraries Crypto and GmSSL are used to encrypt the plaintext data with the selected k block cipher algorithms, using random keys and CBC encryption mode to form a ciphertext dataset Cipher={cipher=3×pn×k}. 1,1,1 ,ciph 1,1,2 ,…,ciph i,j,q ,…ciph 3,pn,k}, where ciph i,j,q It means that the jth file pl with file size i in the plaintext data set Plaintext is ciphered by the qth block cipher algorithm. i,j Encrypted ciphertext file, 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0021] (1.5) Convert all ciphertexts in the ciphertext dataset Cipher into binary files.

[0022] In the block cipher algorithm identification method based on distance metric and ensemble learning of the present invention, the feature extraction based on distance metric and information entropy in step (2) is specifically performed as follows:

[0023] (2.1) To quantify the similarity or difference between plaintext and ciphertext, the distance metric between plaintext and ciphertext is calculated to reflect the degree of change imposed on the plaintext data by the encryption algorithm during the encryption process and its characteristic differences, thereby revealing the unique patterns and behavioral characteristics generated by different encryption algorithms when processing the same plaintext; here, the ciphertext features will be extracted by calculating the distance between plaintext and ciphertext. The specific method is as follows:

[0024] (2.1.1) In order to effectively extract the ciphertext features, each binary plaintext file and ciphertext file in the plaintext dataset Plaintext and the ciphertext dataset Cipher are first vectorized to better capture the information transformation and data distribution differences in the encryption process and form the feature vector of the ciphertext data; for the plaintext file pl i,j , forming the plaintext vector pll i,j , whose length is len(pll i,j ), then pll i,j (t) represents the tth byte of the plaintext vector, 1≤t≤len(pll i,j );For the ciphertext file ciph i,j,q , forming the ciphertext vector ciphh i,j,q , whose length is len(ciphhi,j,q )=len(pll i,j ), then ciphh i,j,q (t) represents the tth byte of the ciphertext vector, 1≤t≤len(ciphh i,j,q );

[0025] (2.1.2) Euclidean distance is the straight-line distance between two points in Euclidean space. In the field of machine learning, Euclidean distance is often used to evaluate the similarity or difference between data. Here, the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The Euclidean distance of is used as the characteristic value of the distance metric, expressed as Fod(q, pll i,j ,ciphh i,j,q ), the calculation formula is:

[0026]

[0027] For the ciphertext data set Cipher, the Euclidean distance between all ciphertext files and the corresponding plaintext files is calculated, and the Euclidean distance feature set FOD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fod 1,1,2 ,…,Fod i,j,q ,…Fod 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0028] For the eigenvalue Fod in FOD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lod i,j,q , Lod i,j,q =q; corresponding to FOD, forming the Euclidean distance feature set label LOD = {Lod 1,1,1 ,Lod 1,1,2 ,…,Lod i,j,q ,…Lod 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0029] (2.1.3) The Hamming distance is a method to measure the number of different characters between two strings of equal length. The calculation method is to perform an XOR operation on the two strings of equal length and count the number of characters that are 1, which is the Hamming distance between the two variables. The Hamming distance reflects the distribution of the bit-level differences between the plaintext vector and the ciphertext vector after encryption by the algorithm. It is expressed as Fhd(q, pll i,j ,ciphh i,j,q ) represents the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithmi,j,q The Hamming distance is calculated as:

[0030]

[0031] For the ciphertext data set Cipher, the Hamming distance between all ciphertext files and the corresponding plaintext files is calculated, and the Hamming distance feature set FHD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fhd 1,1,2 ,…,Fhd i,j,q ,…Fhd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0032] For the characteristic value Fhd in FHD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lhd i,j,q , Lhd i,j,q =q; corresponding to FHD, forming the Hamming distance feature set label LHD = {Lhd 1,1,1 ,Lhd 1,1,2 ,…,Lhd i,j,q ,…Lhd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0033] (2.1.4) Manhattan distance is a method for calculating the distance between two points in a regular grid. It represents the sum of the absolute distances between the two points in the standard coordinate system in two-dimensional coordinates. Manhattan distance reveals the cumulative effect of the difference between plaintext and ciphertext after encryption. i,j ,ciphh i,j,q ) represents the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The Manhattan distance is calculated as:

[0034]

[0035] For the ciphertext data set Cipher, the Manhattan distance between all ciphertext files and the corresponding plaintext files is calculated, and the Manhattan distance feature set FMD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fmd 1,1,2 ,…,Fmd i,j,q ,…Fmd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0036] For the eigenvalue Fmd in FMD i,j,q, according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lmd i,j,q , Lmd i,j,q =q; corresponding to FMD, forming the Manhattan distance feature set label LMD = {Lmd 1,1,1 ,Lmd 1,1,2 ,…,Lmd i,j,q ,…Lmd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0037] (2.1.5) Cosine distance is a method of measuring the difference between two vectors in a vector space by using the cosine of the angle between two vectors. When the cosine is close to 1, the angle is close to 0 degrees, indicating that the two vectors are more similar. i,j ,ciphh i,j,q ) represents the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The cosine distance is calculated as:

[0038]

[0039] For the ciphertext data set Cipher, the cosine distance between all ciphertext files and the corresponding plaintext files is calculated, and the cosine distance feature set FCD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fcd 1,1,2 ,…,Fcd i,j,q ,…Fcd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0040] For the eigenvalue Fcd in FCD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lcd i,j,q ,Lcd i,j,q =q; corresponding to FCD, forming a cosine distance feature set label LCD = {Lcd 1,1,1 ,Lcd 1,1,2 ,…,Lcd i,j,q ,…Lcd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0041] (2.2) Information entropy is usually used as an indicator to measure data uncertainty. In classification tasks, information entropy indicates the uniformity of data distribution in each category. Information entropy is often used to evaluate the randomness of ciphertext data. The higher the entropy value, the more dispersed the data distribution. Fed(q,ciphh i,j,q) represents the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The information entropy is calculated as follows:

[0042]

[0043] Among them, T represents the size of the symbol set and is the ciphertext vector ciphh i,j,q The number of different symbols in ciphh, f represents the specific symbol value. For binary ciphertext vectors, the value of f is 0 or 1, so T = 2; ciphh i,j,q (f) represents the ciphertext vector ciphh i,j,q The relative frequency or probability of the occurrence of the symbol with value f in ;

[0044] Calculate the information entropy of all ciphertext files for the ciphertext data set Cipher, and obtain the information entropy feature set FED of the ciphertext data set Cipher = {Fed 1,1,1 ,Fed 1,1,2 ,…,Fed i,j,q ,…Fed 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0045] For the eigenvalue Fed in FED i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Led i,j,q ,Led i,j,q =q; corresponding to FED, forming the information entropy feature set label LED = {Led 1,1,1 ,Led 1,1,2 ,…,Led i,j,q ,…Led 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0046] (2.3) After calculating the distance metric eigenvalues ​​and information entropy eigenvalues ​​of all ciphertext files and corresponding plaintext files in the ciphertext data set Cipher, the feature data set Fea is obtained, where Fea = {FOD, FHD, FMD, FCD, FED};

[0047] (2.4) Relative to the feature data set Fea, the corresponding label set is Lab = {LOD, LHD, LMD, LCD, LED}. The label set Lab and the feature data set Fea together constitute the data set (Fea, Lab) for training and testing the integrated learning classifier model.

[0048] In the block cipher algorithm identification method based on distance metric and ensemble learning of the present invention, the construction and training of the ensemble learning classifier described in step (3) are specifically performed as follows:

[0049] Ensemble learning is a machine learning method that improves the overall model performance by combining multiple base learners. Ensemble learning algorithms include voting, bagging, boosting, stacking, etc. Here, stacking is selected to build the classifier model. Stacking is a more complex ensemble method. By building a multi-level model structure, the prediction results of multiple base classifiers are used as new features and input into a meta-classifier for final prediction. Stacking can effectively reduce overfitting and the deviation of a single model, and can capture more complex feature relationships. The specific method is as follows:

[0050] (3.1) Data standardization;

[0051] For the prepared data sets (Fea, Lab), they are first standardized. The main purpose of standardization is to adjust the different feature scales of the data to a similar range to eliminate the deviation caused by different dimensions between features; the standardized data has a mean of 0 and a variance of 1; after standardization, the data set is split into a training set (Fea) according to a ratio of 8:2 using the train_test_split function of the open source library sklearn. train ,Lab train ) and the test set (Fea test ,Lab test );

[0052] (3.2) Define the base classifier;

[0053] Random forest, support vector machine, and decision tree are selected as base classifiers. Random forest itself is an integrated learning model composed of multiple decision trees, which has strong generalization ability and anti-overfitting ability, but its interpretability is poor and the model is too complex; support vector machine performs well in processing high-dimensional or nonlinear data and focuses on boundary samples, but it has high requirements for parameter optimization and slow training speed; decision tree is a simple and easy-to-interpret model that can capture complex patterns in data, but single-tree models often have weak generalization ability and are prone to overfitting; for the three classifier models, through experimental tuning, the hyperparameters are set as follows:

[0054] (3.2.1) Random forest generated 100 decision trees, and the random generator seed was 42 to ensure the repeatability of the results;

[0055] (3.2.2) Support vector machine probability is set to True to enable probability estimation to output the probability of each category. Kernel is set to RBF, which means using radial basis function as kernel function. RBF kernel can process nonlinear data and achieve better classification effect by mapping original features to high-dimensional space.

[0056] (3.2.3) The random seed of the decision tree is set to 42 to ensure the repeatability of the results;

[0057] (3.3) Select meta-classifier;

[0058] Gradient boosted decision tree (GBDT) is selected as the meta-classifier of stacking ensemble. GBDT can model complex feature interactions and nonlinear relationships, improve the overall performance of the model, and gradually correct the errors of the previous model through the gradient descent method to improve the classification accuracy. Through experimental tuning, the number of decision trees n_estimators is set to 100, the learning efficiency learning_rate is set to 0.3, and the random generator seed is set to 42;

[0059] (3.4) Build and train the classifier model;

[0060] Based on the selected base classifier and meta-classifier, a stacking ensemble method is used to input the training set (Fea train ,Lab train ), train the stacked ensemble learning classifier; the stacking ensemble parameters are set to: Passthrough = True, pass the original features to the meta-classifier together with the base classifier output; Cv = 5, use the five-fold cross-validation method to solve the problem that GBDT as a meta-classifier may overfit the prediction results of the base classifier; Stack_method = auto, set the stacking method to automatic stacking, and automatically select the most appropriate prediction output method according to the available methods of the base classifier; finally, a trained ensemble learning classifier is obtained.

[0061] In the block cipher algorithm identification method based on distance measurement and ensemble learning of the present invention, step (4) performs cipher algorithm identification based on a trained ensemble learning classifier, and the specific steps are as follows:

[0062] (4.1) Test set of input ciphertext data (Fea test ,Lab test ) into the trained ensemble learning classifier;

[0063] (4.2) Based on the trained ensemble learning classifier, use the prediction method of the classifier model to predict the test set and generate the predicted label vector y pred, predict the label value corresponding to each feature data in the test set, and then determine the cryptographic algorithm used when the ciphertext data is encrypted according to the label value.

[0064] The beneficial effects of the present invention are:

[0065] (1) The present invention discloses a block cipher algorithm identification method based on distance metric and ensemble learning, which uses a variety of different distance metric calculation methods to quantify the distribution and relationship between plaintext data and ciphertext data in feature space, mines the unique features generated by different cryptographic algorithms during the encryption process, captures the subtle differences and patterns reflected by various cryptographic algorithms when generating ciphertext, and improves the feature extraction capability;

[0066] (2) The method of the present invention adopts stacked ensemble learning to construct a multi-level classifier model. By combining the prediction results of multiple base classifiers and using a high-level meta-classifier to make the final decision, it is beneficial to optimize and integrate the prediction information of different models and improve the overall classification accuracy;

[0067] (3) The method of the present invention has good versatility, is applicable to various encryption scenarios and data distribution environments, and has good classification effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 This is a flow chart of a block cipher algorithm identification method based on distance measurement and ensemble learning of the present invention;

[0069] Figure 2 This is a flow chart of feature extraction based on distance measurement and information entropy in the present invention.

[0070] Specific implementation examples

[0071] The present invention will be further described in detail below in conjunction with the embodiments and drawings, but the present invention is not limited thereto. Example

[0072] A block cipher algorithm identification method based on distance measurement and ensemble learning, referring to Figure 1 , including the following steps:

[0073] (1) Data preparation and encryption;

[0074] (2) Feature extraction based on distance measurement and information entropy;

[0075] (3) Construction and training of ensemble learning classifiers;

[0076] (4) Identify cryptographic algorithms based on the trained ensemble learning classifier.

[0077] In the block cipher algorithm identification method based on distance metric and ensemble learning of the present invention, the data preparation and encryption in step (1) are specifically performed as follows:

[0078] (1.1) Select a public dataset as a plaintext dataset. The public dataset can be a text file, an image file, etc.

[0079] (1.2) Select or divide pn files with file sizes of 1KB, 16KB, and 32KB from the public data set to form three plaintext data sets of different sizes. Finally, the plaintext data set Plaintext = {pl 1,1 ,pl 1,2 ,…,pl i,j ,…pl 3,pn}, where pl i,j Indicates the j-th file with file size i, 1≤i≤3, 1≤j≤pn;

[0080] (1.3) Convert all files in the plaintext dataset Plaintext into binary files;

[0081] (1.4) For all file data of different sizes in the plaintext dataset Plaintext, the open source cryptographic libraries Crypto and GmSSL are used to encrypt the plaintext data with the selected k block cipher algorithms, using random keys and CBC encryption mode to form a ciphertext dataset Cipher={cipher=3×pn×k}. 1,1,1 ,ciph 1,1,2 ,…,ciph i,j,q ,…ciph 3,pn,k}, where ciph i,j,q It means that the jth file pl with file size i in the plaintext data set Plaintext is ciphered by the qth block cipher algorithm. i,j Encrypted ciphertext file, 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0082] (1.5) Convert all ciphertexts in the ciphertext dataset Cipher into binary files.

[0083] In the block cipher algorithm identification method based on distance metric and ensemble learning of the present invention, the feature extraction based on distance metric and information entropy in step (2) is referred to as Figure 2 , the specific steps are as follows:

[0084] (2.1) To quantify the similarity or difference between plaintext and ciphertext, the distance metric between plaintext and ciphertext is calculated to reflect the degree of change imposed on the plaintext data by the encryption algorithm during the encryption process and its characteristic differences, thereby revealing the unique patterns and behavioral characteristics generated by different encryption algorithms when processing the same plaintext; here, the ciphertext features will be extracted by calculating the distance between plaintext and ciphertext. The specific method is as follows:

[0085] (2.1.1) In order to effectively extract the ciphertext features, each binary plaintext file and ciphertext file in the plaintext dataset Plaintext and the ciphertext dataset Cipher are first vectorized to better capture the information transformation and data distribution differences in the encryption process and form the feature vector of the ciphertext data; for the plaintext file pl i,j , forming the plaintext vector pll i,j , whose length is len(pll i,j ), then pll i,j (t) represents the tth byte of the plaintext vector, 1≤t≤len(pll i,j );For the ciphertext file ciph i,j,q , forming the ciphertext vector ciphh i,j,q , whose length is len(ciphh i,j,q )=len(pll i,j ), then ciphh i,j,q (t) represents the tth byte of the ciphertext vector, 1≤t≤len(ciphh i,j,q );

[0086] (2.1.2) Euclidean distance is the straight-line distance between two points in Euclidean space. In the field of machine learning, Euclidean distance is often used to evaluate the similarity or difference between data. Here, the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The Euclidean distance of is used as the characteristic value of the distance metric, expressed as Fod(q, pll i,j ,ciphh i,j,q ), the calculation formula is:

[0087]

[0088] For the ciphertext data set Cipher, the Euclidean distance between all ciphertext files and the corresponding plaintext files is calculated, and the Euclidean distance feature set FOD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fod 1,1,2 ,…,Fod i,j,q ,…Fod 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0089] For the eigenvalue Fod in FOD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lod i,j,q , Lod i,j,q =q; corresponding to FOD, forming the Euclidean distance feature set label LOD = {Lod 1,1,1 ,Lod 1,1,2 ,…,Lod i,j,q ,…Lod 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0090] (2.1.3) The Hamming distance is a method to measure the number of different characters between two strings of equal length. The calculation method is to perform an XOR operation on the two strings of equal length and count the number of characters that are 1, which is the Hamming distance between the two variables. The Hamming distance reflects the distribution of the bit-level differences between the plaintext vector and the ciphertext vector after encryption by the algorithm. It is expressed as Fhd(q, pll i,j ,ciphh i,j,q ) represents the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The Hamming distance is calculated as:

[0091]

[0092] For the ciphertext data set Cipher, the Hamming distance between all ciphertext files and the corresponding plaintext files is calculated, and the Hamming distance feature set FHD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fhd 1,1,2 ,…,Fhd i,j,q ,…Fhd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0093] For the characteristic value Fhd in FHD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lhd i,j,q , Lhd i,j,q =q; corresponding to FHD, forming the Hamming distance feature set label LHD = {Lhd 1,1,1 ,Lhd 1,1,2 ,…,Lhd i,j,q ,…Lhd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0094] (2.1.4) Manhattan distance is a method for calculating the distance between two points in a regular grid. It represents the sum of the absolute distances between the two points in the standard coordinate system in two-dimensional coordinates. Manhattan distance reveals the cumulative effect of the difference between plaintext and ciphertext after encryption. i,j ,ciphh i,j,q ) represents the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The Manhattan distance is calculated as:

[0095]

[0096] For the ciphertext data set Cipher, the Manhattan distance between all ciphertext files and the corresponding plaintext files is calculated, and the Manhattan distance feature set FMD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fmd 1,1,2 ,…,Fmd i,j,q ,…Fmd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0097] For the eigenvalue Fmd in FMD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lmd i,j,q , Lmd i,j,q =q; corresponding to FMD, forming the Manhattan distance feature set label LMD = {Lmd 1,1,1 ,Lmd 1,1,2 ,…,Lmd i,j,q ,…Lmd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0098] (2.1.5) Cosine distance is a method of measuring the difference between two vectors in a vector space by using the cosine of the angle between two vectors. When the cosine is close to 1, the angle is close to 0 degrees, indicating that the two vectors are more similar. i,j ,ciphh i,j,q ) represents the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The cosine distance is calculated as:

[0099]

[0100] For the ciphertext data set Cipher, the cosine distance between all ciphertext files and the corresponding plaintext files is calculated, and the cosine distance feature set FCD of the ciphertext data set Cipher is obtained.1,1,1 ,Fcd 1,1,2 ,…,Fcd i,j,q ,…Fcd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0101] For the eigenvalue Fcd in FCD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lcd i,j,q ,Lcd i,j,q =q; corresponding to FCD, forming a cosine distance feature set label LCD = {Lcd 1,1,1 ,Lcd 1,1,2 ,…,Lcd i,j,q ,…Lcd 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0102] (2.2) Information entropy is usually used as an indicator to measure data uncertainty. In classification tasks, information entropy indicates the uniformity of data distribution in each category. Information entropy is often used to evaluate the randomness of ciphertext data. The higher the entropy value, the more dispersed the data distribution. Fed(q,ciphh i,j,q ) represents the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The information entropy is calculated as follows:

[0103]

[0104] Among them, T represents the size of the symbol set and is the ciphertext vector ciphh i,j,q The number of different symbols in ciphh, f represents the specific symbol value. For binary ciphertext vectors, the value of f is 0 or 1, so T = 2; ciphh i,j,q (f) represents the ciphertext vector ciphh i,j,q The relative frequency or probability of the occurrence of the symbol with value f in ;

[0105] Calculate the information entropy of all ciphertext files for the ciphertext data set Cipher, and obtain the information entropy feature set FED of the ciphertext data set Cipher = {Fed 1,1,1 ,Fed 1,1,2 ,…,Fed i,j,q ,…Fed 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0106] For the eigenvalue Fed in FED i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Led i,j,q ,Led i,j,q=q; corresponding to FED, forming the information entropy feature set label LED = {Led 1,1,1 ,Led 1,1,2 ,…,Led i,j,q ,…Led 3,pn,k}, where 1≤i≤3, 1≤j≤pn, 1≤q≤k;

[0107] (2.3) After calculating the distance metric eigenvalues ​​and information entropy eigenvalues ​​of all ciphertext files and corresponding plaintext files in the ciphertext data set Cipher, the feature data set Fea is obtained, where Fea = {FOD, FHD, FMD, FCD, FED};

[0108] (2.4) Relative to the feature data set Fea, the corresponding label set is Lab = {LOD, LHD, LMD, LCD, LED}. The label set Lab and the feature data set Fea together constitute the data set (Fea, Lab) for training and testing the integrated learning classifier model.

[0109] In the block cipher algorithm identification method based on distance metric and ensemble learning of the present invention, the construction and training of the ensemble learning classifier described in step (3) are specifically performed as follows:

[0110] Ensemble learning is a machine learning method that improves the overall model performance by combining multiple base learners. Ensemble learning algorithms include voting, bagging, boosting, stacking, etc. Here, stacking is selected to build the classifier model. Stacking is a more complex ensemble method. By building a multi-level model structure, the prediction results of multiple base classifiers are used as new features and input into a meta-classifier for final prediction. Stacking can effectively reduce overfitting and the deviation of a single model, and can capture more complex feature relationships. The specific method is as follows:

[0111] (3.1) Data standardization;

[0112] For the prepared data sets (Fea, Lab), they are first standardized. The main purpose of standardization is to adjust the different feature scales of the data to a similar range to eliminate the deviation caused by different dimensions between features; the standardized data has a mean of 0 and a variance of 1; after standardization, the data set is split into a training set (Fea) according to a ratio of 8:2 using the train_test_split function of the open source library sklearn. train ,Lab train ) and the test set (Fea test ,Lab test);

[0113] (3.2) Define the base classifier;

[0114] Random forest, support vector machine, and decision tree are selected as base classifiers. Random forest itself is an integrated learning model composed of multiple decision trees, which has strong generalization ability and anti-overfitting ability, but its interpretability is poor and the model is too complex; support vector machine performs well in processing high-dimensional or nonlinear data and focuses on boundary samples, but it has high requirements for parameter optimization and slow training speed; decision tree is a simple and easy-to-interpret model that can capture complex patterns in data, but single-tree models often have weak generalization ability and are prone to overfitting; for the three classifier models, through experimental tuning, the hyperparameters are set as follows:

[0115] (3.2.1) Random forest generated 100 decision trees, and the random generator seed was 42 to ensure the repeatability of the results;

[0116] (3.2.2) Support vector machine probability is set to True to enable probability estimation to output the probability of each category. Kernel is set to RBF, which means using radial basis function as kernel function. RBF kernel can process nonlinear data and achieve better classification effect by mapping original features to high-dimensional space.

[0117] (3.2.3) The random seed of the decision tree is set to 42 to ensure the repeatability of the results;

[0118] (3.3) Select meta-classifier;

[0119] Gradient boosted decision tree (GBDT) is selected as the meta-classifier of stacking ensemble. GBDT can model complex feature interactions and nonlinear relationships, improve the overall performance of the model, and gradually correct the errors of the previous model through the gradient descent method to improve the classification accuracy. Through experimental tuning, the number of decision trees n_estimators is set to 100, the learning efficiency learning_rate is set to 0.3, and the random generator seed is set to 42;

[0120] (3.4) Build and train the classifier model;

[0121] Based on the selected base classifier and meta-classifier, a stacking ensemble method is used to input the training set (Fea train ,Lab train), train the stacked ensemble learning classifier; the stacking ensemble parameters are set to: Passthrough = True, pass the original features to the meta-classifier together with the base classifier output; Cv = 5, use the five-fold cross-validation method to solve the problem that GBDT as a meta-classifier may overfit the prediction results of the base classifier; Stack_method = auto, set the stacking method to automatic stacking, and automatically select the most appropriate prediction output method according to the available methods of the base classifier; finally, a trained ensemble learning classifier is obtained.

[0122] In the block cipher algorithm identification method based on distance measurement and ensemble learning of the present invention, step (4) performs cipher algorithm identification based on a trained ensemble learning classifier, and the specific steps are as follows:

[0123] (4.1) Test set of input ciphertext data (Fea test ,Lab test ) into the trained ensemble learning classifier;

[0124] (4.2) Based on the trained ensemble learning classifier, use the prediction method of the classifier model to predict the test set and generate the predicted label vector y pred , predict the label value corresponding to each feature data in the test set, and then determine the cryptographic algorithm used when the ciphertext data is encrypted according to the label value.

Claims

1. A block cipher algorithm identification method based on distance metric and ensemble learning, characterized in that: The following steps are involved: (1) Data preparation and encryption; Select a public data set and divide it into different sizes to obtain a plaintext data set, and use different block cipher algorithms to encrypt all the plaintext data in the plaintext data set to generate ciphertext data, thereby forming a ciphertext data set; (2) Feature extraction based on distance measurement and information entropy; Calculate the distance measurement value between the plaintext vector and the ciphertext vector and the information entropy value of the ciphertext vector, extract the ciphertext features, obtain the feature data set, and set the corresponding cryptographic algorithm label for each feature data in the feature data set; (3) Construction and training of ensemble learning classifiers; Build an ensemble learning classifier based on the stacking method, select the gradient boosting decision tree as the meta-classifier, random forest, SVM, and decision tree as the base classifier, train the model with feature data sets and labels, and find the optimal hyperparameter classifier model through experimental tuning; (4) Identify cryptographic algorithms based on trained ensemble learning classifiers; For unknown ciphertext data, the trained ensemble learning classifier is used to identify the block cipher algorithm.

2. A block cipher algorithm identification method based on distance metric and ensemble learning according to claim 1, characterized in that: The data preparation and encryption described in step (1) are specifically performed as follows: (1.1) Select a public dataset as a plaintext dataset. The public dataset can be a text file, an image file, etc. (1.2) Select or divide pn files with file sizes of 1KB, 16KB, and 32KB from the public data set to form three plaintext data sets of different sizes. Finally, the plaintext data set Plaintext = {pl 1,1 ,pl 1,2 ,…,pl i,j ,…pl 3,pn }, where pl i,j Indicates the j-th file with file size i, 1≤i≤3, 1≤j≤pn; (1.3) Convert all files in the plaintext dataset Plaintext into binary files; (1.4) For all file data of different sizes in the plaintext dataset Plaintext, the open source cryptographic libraries Crypto and GmSSL are used to encrypt the plaintext data with the selected k block cipher algorithms, using random keys and CBC encryption mode to form a ciphertext dataset Cipher={cipher=3×pn×k}. 1,1,1 ,ciph 1,1,2 ,…,ciph i,j,q ,…ciph 3,pn,k }, where ciph i,j,q It means that the jth file pl with file size i in the plaintext data set Plaintext is ciphered by the qth block cipher algorithm. i,j Encrypted ciphertext file, 1≤i≤3, 1≤j≤pn, 1≤q≤k; (1.5) Convert all ciphertexts in the ciphertext dataset Cipher into binary files.

3. A block cipher algorithm identification method based on distance metric and ensemble learning according to claim 1, characterized in that: The specific method of feature extraction based on distance measurement and information entropy in step (2) is as follows: (2.1) To quantify the similarity or difference between plaintext and ciphertext, the distance metric between plaintext and ciphertext is calculated to reflect the degree of change imposed on the plaintext data by the encryption algorithm during the encryption process and its characteristic differences, thereby revealing the unique patterns and behavioral characteristics generated by different encryption algorithms when processing the same plaintext; here, the ciphertext features will be extracted by calculating the distance between plaintext and ciphertext. The specific method is as follows: (2.1.1) In order to effectively extract the ciphertext features, each binary plaintext file and ciphertext file in the plaintext dataset Plaintext and the ciphertext dataset Cipher are first vectorized to better capture the information transformation and data distribution differences in the encryption process and form the feature vector of the ciphertext data; for the plaintext file pl i,j , forming the plaintext vector pll i,j , whose length is len(pll i,j ), then pll i,j (t) represents the tth byte of the plaintext vector, 1≤t≤len(pll i,j );For the ciphertext file ciph i,j,q , forming the ciphertext vector ciphh i,j,q , whose length is len(ciphh i,j,q )=len(pll i,j ), then ciphh i,j,q (t) represents the tth byte of the ciphertext vector, 1≤t≤len(ciphh i,j,q ); (2.1.2) Euclidean distance is the straight-line distance between two points in Euclidean space. In the field of machine learning, Euclidean distance is often used to evaluate the similarity or difference between data. Here, the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The Euclidean distance of is used as the characteristic value of the distance metric, expressed as Fod(q, pll i,j ,ciphh i,j,q ), the calculation formula is: For the ciphertext data set Cipher, the Euclidean distance between all ciphertext files and the corresponding plaintext files is calculated, and the Euclidean distance feature set FOD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fod 1,1,2 ,…,Fod i,j,q ,…Fod 3,pn,k }, where 1≤i≤3, 1≤j≤pn, 1≤q≤k; For the eigenvalue Fod in FOD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lod i,j,q , Lod i,j,q =q; corresponding to FOD, forming the Euclidean distance feature set label LOD = {Lod 1,1,1 ,Lod 1,1,2 ,…,Lod i,j,q ,…Lod 3,pn,k }, where 1≤i≤3, 1≤j≤pn, 1≤q≤k; (2.1.3) The Hamming distance is a method to measure the number of different characters between two strings of equal length. The calculation method is to perform an XOR operation on the two strings of equal length and count the number of characters that are 1, which is the Hamming distance between the two variables. The Hamming distance reflects the distribution of the bit-level differences between the plaintext vector and the ciphertext vector after encryption by the algorithm. It is expressed as Fhd(q, pll i,j ,ciphh i,j,q ) represents the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The Hamming distance is calculated as: For the ciphertext data set Cipher, the Hamming distance between all ciphertext files and the corresponding plaintext files is calculated, and the Hamming distance feature set FHD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fhd 1,1,2 ,…,Fhd i,j,q ,…Fhd 3,pn,k }, where 1≤i≤3, 1≤j≤pn, 1≤q≤k; For the characteristic value Fhd in FHD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lhd i,j,q , Lhd i,j,q =q; corresponding to FHD, forming the Hamming distance feature set label LHD = {Lhd 1,1,1 ,Lhd 1,1,2 ,…,Lhd i,j,q ,…Lhd 3,pn,k }, where 1≤i≤3, 1≤j≤pn, 1≤q≤k; (2.1.4) Manhattan distance is a method for calculating the distance between two points in a regular grid. It represents the sum of the absolute distances between the two points in the standard coordinate system in two-dimensional coordinates. Manhattan distance reveals the cumulative effect of the difference between plaintext and ciphertext after encryption. i,j ,ciphh i,j,q ) represents the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The Manhattan distance is calculated as: For the ciphertext data set Cipher, the Manhattan distance between all ciphertext files and the corresponding plaintext files is calculated, and the Manhattan distance feature set FMD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fmd 1,1,2 ,…,Fmd i,j,q ,…Fmd 3,pn,k }, where 1≤i≤3, 1≤j≤pn, 1≤q≤k; For the eigenvalue Fmd in FMD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lmd i,j,q , Lmd i,j,q =q; corresponding to FMD, forming the Manhattan distance feature set label LMD = {Lmd 1,1,1 ,Lmd 1,1,2 ,…,Lmd i,j,q ,…Lmd 3,pn,k }, where 1≤i≤3, 1≤j≤pn, 1≤q≤k; (2.1.5) Cosine distance is a method of measuring the difference between two vectors in a vector space by using the cosine of the angle between two vectors. When the cosine is close to 1, the angle is close to 0 degrees, indicating that the two vectors are more similar. i,j ,ciphh i,j,q ) represents the plaintext vector pll i,j and the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The cosine distance is calculated as: For the ciphertext data set Cipher, the cosine distance between all ciphertext files and the corresponding plaintext files is calculated, and the cosine distance feature set FCD of the ciphertext data set Cipher is obtained. 1,1,1 ,Fcd 1,1,2 ,…,Fcd i,j,q ,…Fcd 3,pn,k }, where 1≤i≤3, 1≤j≤pn, 1≤q≤k; For the eigenvalue Fcd in FCD i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Lcd i,j,q ,Lcd i,j,q =q; corresponding to FCD, forming a cosine distance feature set label LCD = {Lcd 1,1,1 ,Lcd 1,1,2 ,…,Lcd i,j,q ,…Lcd 3,pn,k }, where 1≤i≤3, 1≤j≤pn, 1≤q≤k; (2.2) Information entropy is usually used as an indicator to measure data uncertainty. In classification tasks, information entropy indicates the uniformity of data distribution in each category. Information entropy is often used to evaluate the randomness of ciphertext data. The higher the entropy value, the more dispersed the data distribution. Fed(q,ciphh i,j,q ) represents the ciphertext vector ciphh encrypted by the qth cryptographic algorithm i,j,q The information entropy is calculated as follows: Among them, T represents the size of the symbol set and is the ciphertext vector ciphh i,j,q The number of different symbols in ciphh, f represents the specific symbol value. For binary ciphertext vectors, the value of f is 0 or 1, so T = 2; ciphh i,j,q (f) represents the ciphertext vector ciphh i,j,q The relative frequency or probability of the occurrence of the symbol with value f in ; Calculate the information entropy of all ciphertext files for the ciphertext data set Cipher, and obtain the information entropy feature set FED of the ciphertext data set Cipher = {Fed 1,1,1 ,Fed 1,1,2 ,…,Fed i,j,q ,…Fed 3,pn,k }, where 1≤i≤3, 1≤j≤pn, 1≤q≤k; For the eigenvalue Fed in FED i,j,q , according to the qth cryptographic algorithm for encrypting the plaintext file, set the corresponding label value to Led i,j,q ,Led i,j,q =q; corresponding to FED, forming the information entropy feature set label LED = {Led 1,1,1 ,Led 1,1,2 ,…,Led i,j,q ,…Led 3,pn,k }, where 1≤i≤3, 1≤j≤pn, 1≤q≤k; (2.3) After calculating the distance metric eigenvalues ​​and information entropy eigenvalues ​​of all ciphertext files and corresponding plaintext files in the ciphertext data set Cipher, the feature data set Fea is obtained, where Fea = {FOD, FHD, FMD, FCD, FED}; (2.4) Relative to the feature data set Fea, the corresponding label set is Lab = {LOD, LHD, LMD, LCD, LED}. The label set Lab and the feature data set Fea together constitute the data set (Fea, Lab) for training and testing the integrated learning classifier model.

4. The method for identifying a block cipher algorithm based on distance metric and ensemble learning according to claim 1, characterized in that: The construction and training of the ensemble learning classifier described in step (3) are specifically performed as follows: (3.1) Data standardization; For the prepared data set (Fea, Lab), it is first standardized. The main purpose of standardization is to adjust the different feature scales of the data to a similar range to eliminate the deviation caused by different dimensions between features; the standardized data has a mean of 0 and a variance of 1. After the data set is standardized, the train_test_split function of the open source library sklearn is used to split it into a training set (Fea) in a ratio of 8:

2. train ,Lab train ) and the test set (Fea test ,Lab test ); (3.2) Define the base classifier; In order to solve the problem that single tree models often have weak generalization ability and are prone to overfitting, three classifiers, random forest, support vector machine, and decision tree, are selected as base classifiers. Through experimental tuning, the hyperparameters are set as follows: (3.2.1) Random forest generated 100 decision trees, and the random generator seed was 42 to ensure the repeatability of the results; (3.2.2) Support vector machine probability is set to True to enable probability estimation to output the probability of each category. Kernel is set to RBF, which means using radial basis function as kernel function. RBF kernel can process nonlinear data and achieve better classification effect by mapping original features to high-dimensional space. (3.2.3) The random seed of the decision tree is set to 42 to ensure the repeatability of the results; (3.3) Select meta-classifier; Gradient boosted decision tree (GBDT) is selected as the meta-classifier of stacking ensemble. GBDT can model complex feature interactions and nonlinear relationships, improve the overall performance of the model, and gradually correct the errors of the previous model through the gradient descent method to improve the classification accuracy. Through experimental tuning, the number of decision trees n_estimators is set to 100, the learning efficiency learning_rate is set to 0.3, and the random generator seed is set to 42; (3.4) Build and train the classifier model; Based on the selected base classifier and meta-classifier, a stacking ensemble method is used to input the training set (Fea train ,Lab train ), train the stacked ensemble learning classifier; the stacking ensemble parameters are set to: Passthrough = True, pass the original features to the meta-classifier together with the base classifier output; Cv = 5, use the five-fold cross-validation method to solve the problem that GBDT as a meta-classifier may overfit the prediction results of the base classifier; Stack_method = auto, set the stacking method to automatic stacking, and automatically select the most appropriate prediction output method according to the available methods of the base classifier; finally, a trained ensemble learning classifier is obtained.

5. The method for identifying a block cipher algorithm based on distance metric and ensemble learning according to claim 1, characterized in that: The cryptographic algorithm recognition based on the trained ensemble learning classifier described in step (4) is specifically performed as follows: (4.1) Test set of input ciphertext data (Fea test ,Lab test ) into the trained ensemble learning classifier; (4.2) Based on the trained ensemble learning classifier, use the prediction method of the classifier model to predict the test set and generate the predicted label vector y pred , predict the label value corresponding to each feature data in the test set, and then determine the cryptographic algorithm used when the ciphertext data is encrypted according to the label value.

Citation Information

Patent Citations

  • Encrypted network flow monitoring method

    CN114465786A

  • Cryptographic algorithm identification method and related device

    CN115048638A

  • Classifying data with deep learning neural records incrementally refined through expert input

    US20150254555A1