Block cipher algorithm identification method based on multi-modal feature fusion

By converting ciphertext data into text, image, and voice modalities and performing deep learning feature fusion, the problem of low recognition accuracy of block cipher algorithms is solved, and higher recognition accuracy and robustness are achieved.

CN120744807APending Publication Date: 2025-10-03GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510795853.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing block cipher algorithm identification methods are difficult to effectively distinguish different block cipher algorithms when the encryption algorithm is unknown, and the identification accuracy is not high, especially the effects of deep learning and manually designed feature extraction methods are limited.

Method used

A multimodal feature fusion method is adopted to convert the ciphertext data into three modalities: text, image, and voice. The features are extracted separately using deep learning models, and mid-term fusion is performed through the transformer model. The deep learning models suitable for each modality are integrated for weighted fusion, and finally the encryption algorithm recognition of unknown ciphertext data is realized.

Benefits of technology

It improves the accuracy and robustness of block cipher algorithm recognition, has stronger generalization ability, is suitable for various encryption scenarios and data distributions, and has good classification effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744807A_ABST
    Figure CN120744807A_ABST
Patent Text Reader

Abstract

The invention discloses a block cipher algorithm identification method based on multi-modal feature fusion. The method comprises the following steps: firstly, converting ciphertext data into three modal representations of a text, an image and voice; then integrating a deep learning model suitable for each mode through a multi-mode feature fusion strategy, and performing feature extraction on ciphertext data of different modes; on this basis, based on a transform model and by adopting a medium-term fusion strategy, performing weighted fusion on the features of different modals, and training the model; and finally, based on the trained model, realizing identification of an encryption algorithm adopted by an unknown ciphertext. According to the method, feature extraction is completed through different encoders, a feature fusion thought is adopted, and a transform model containing a self-attention mechanism is introduced to perform multi-modal feature fusion, so that the robustness and generalization ability of the model are effectively improved, and the recognition accuracy of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security, and in particular to a block cipher algorithm recognition method based on multimodal feature fusion. Background Art

[0002] Cryptographic algorithm identification is a prerequisite for cryptanalysis. Currently, commonly used feature-based cryptographic algorithm identification research methods include statistical methods and machine learning methods. Statistical methods extract various statistical indicators of ciphertext, such as randomness detection indicators, as ciphertext features for identification. Machine learning methods extract ciphertext features, construct a classifier, and train the classification and recognition model on a training dataset to achieve classification predictions for the ciphertext data in the test set.

[0003] Because it's difficult to ensure the effectiveness of ciphertext feature extraction methods, deep learning has been introduced into cryptographic algorithm recognition technology to automatically learn and extract features, reducing reliance on manual feature design and improving recognition accuracy and robustness. However, for block ciphers, while different block ciphers have their own unique characteristics, distinguishing the encryption algorithm used by different block ciphers based on the ciphertext remains challenging when the encryption algorithm is unknown. Currently, both manually designed features and deep learning-based automatic feature learning and extraction offer low recognition accuracy.

[0004] In order to better identify cryptographic algorithms, the present invention proposes a block cipher algorithm identification method based on multimodal feature fusion. Multimodal feature fusion refers to the effective integration of data from different modalities, such as text, images, audio, etc., to give full play to the advantages of each modality in information expression, thereby improving the overall performance of the model in tasks such as classification and recognition. From the perspective of fusion methods, multimodal feature fusion is divided into early fusion, late fusion, hybrid fusion, intermediate fusion, etc. The method of the present invention first converts the ciphertext data into different modal representations, and then makes full use of the advantages of deep learning in feature expression, integrates deep learning models suitable for each modality for feature extraction, and adopts a mid-term fusion strategy to achieve feature fusion, and completes the identification of the encryption algorithm used for unknown ciphertext data. Summary of the Invention

[0005] The present invention proposes a block cipher algorithm recognition method based on multimodal feature fusion. The method first converts ciphertext data into three modal representations: text, image, and speech. Secondly, through a multimodal feature fusion strategy, a deep learning model suitable for each modality is integrated to extract features of the ciphertext data of different modalities respectively. On this basis, based on the transformer model and adopting a mid-term fusion strategy, the features of different modalities are weightedly fused and the model is trained. Finally, the encryption algorithm used for unknown ciphertext is recognized based on the trained model.

[0006] The technical solution for achieving the purpose of the present invention is:

[0007] A block cipher algorithm identification method based on multimodal feature fusion specifically includes the following steps:

[0008] (1) Data preparation and preprocessing;

[0009] Random plaintext is selected as the plaintext dataset, and all the plaintext data in the plaintext dataset are encrypted using different block cipher algorithms to generate the original ciphertext dataset; the ciphertext data is preprocessed and converted into ciphertext datasets in three modalities: text, image, and voice;

[0010] (2) Text modality feature extraction;

[0011] The convolutional neural network-based text classification model TextCNN is used as the text modality encoder model. After operations such as text input, word embedding, convolution feature extraction, maximum pooling, and feature splicing, the text modality ciphertext dataset is subjected to feature extraction to obtain the feature set of the text modality data.

[0012] (3) Image modality feature extraction;

[0013] A convolutional neural network (CNN) is used as the image modality encoder model. After image input, local feature extraction by the convolution layer, dimensionality reduction by the pooling layer, flattening layer, and fully connected layer operations, the image modality ciphertext dataset is subjected to feature extraction to obtain the feature set of the image modality data.

[0014] (4) Speech modal feature extraction;

[0015] A self-supervised learning framework Wav2Vec2 is used as the speech modality encoder model. After voice input, model encoding, feature aggregation and other operations, feature extraction is performed on the speech modality ciphertext dataset to obtain the feature set of speech modality data.

[0016] (5) Feature fusion and model training of different modal data;

[0017] Based on the transformer model and using a mid-term fusion strategy, feature data from different modal feature sets are aligned, feature weighted fusion is completed, and model training is performed.

[0018] (6) Model prediction and encryption algorithm identification of unknown ciphertext data;

[0019] For unknown ciphertext data, the trained model is used to identify the cryptographic algorithm.

[0020] In the block cipher algorithm identification method based on multimodal feature fusion of the present invention, the data preparation and preprocessing in step (1) are specifically performed as follows:

[0021] (1.1) Select n random plaintext data as the plaintext data set Plaintext={p1,p2,…,p i ,…,p n}, p i Represents the i-th plaintext data, 1≤i≤n;

[0022] (1.2) For all the data in the plaintext dataset Plaintext, the open source cryptographic libraries Crypto and GmSSL are used to encrypt the plaintext data with the selected k block cipher algorithms, using random keys and ECB encryption mode, forming an original ciphertext dataset of size n×k saved in binary format Cipher={c 1,1 ,c 1,2 ,…,c q,i ,…,c k,n}, where c q,i Indicates that the qth block cipher algorithm is used to cipher the i-th data p in the plaintext dataset Plaintext i The encrypted ciphertext data, 1≤q≤k, 1≤i≤n; and the maximum length of all ciphertext data in Cipher is calculated as LenC;

[0023] (1.3) Preprocess all ciphertext data in the original ciphertext dataset Cipher to remove redundant spaces and invalid characters; and for each ciphertext data, if its length is less than LenC, fill it with zero value to obtain the text modal ciphertext dataset Cipher text ={CT 1,1 ,CT 1,2 ,…,CT q,i ,…,CT k,n}, where CT q,i Indicates the ciphertext data c q,i The processed text modal ciphertext data; and the Cipher text Convert all binary ciphertext data in to hexadecimal form;

[0024] (1.4) Call the Python image conversion function to convert all the ciphertext data in the original ciphertext dataset Cipher into grayscale images, and set the image size ratio to 3×4; for the ciphertext data c q,i , according to the grayscale image size, each byte value is mapped to the pixel value in the image in turn, and the image is generated using the grayscale mode. If the ciphertext data c q,i If the length of is not enough to convert to a grayscale image, zero padding is performed to complete the grayscale image conversion; all ciphertext data are converted into grayscale images using the same method to obtain the image modal ciphertext dataset Cipher image ={CI 1,1 ,CI 1,2 ,…,CI q,i ,…,CI k,n}, where CI q,i Indicates the ciphertext data c q,i The converted image modal ciphertext data;

[0025] (1.5) Use the speech synthesis engine tool TTS to perform speech encoding on all the ciphertext data in the original ciphertext dataset Cipher. Set the speech synthesis language to English, the speech rate to 1, and the unified sampling frequency to 16kHz to obtain the speech modal ciphertext dataset Cipher. audio ={CA 1,1 ,CA 1,2 ,…,CA q,i ,…,CA k,n}, where CA q,i Indicates the ciphertext data c q,i The converted voice modal ciphertext data.

[0026] In the block cipher algorithm recognition method based on multimodal feature fusion of the present invention, the text modal feature extraction in step (2) is specifically performed as follows:

[0027] (2.1) Build a TextCNN encoder model based on convolutional neural network;

[0028] (2.1.1) The TextCNN encoder model consists of a word embedding layer, three one-dimensional convolutional layers, a global max pooling layer, and a dropout layer;

[0029] (2.1.2) The word embedding layer is used to map the input word index sequence into a dense word vector, thereby outputting a two-dimensional matrix to form a set of word vectors for the input text data. Setting the parameter embed_dim to 128 indicates that the dimension of each word embedding vector is 128.

[0030] (2.1.3) The convolutional layer is used to extract local semantic features from word vector data and output feature maps corresponding to the activation distribution of the corresponding local pattern. Three one-dimensional convolutional layers are connected in sequence, with kernel_size of 3, 4, and 5, respectively, to capture the corresponding local features and output feature maps of different granularities, such as triplets, quadruplets, and quintuples.

[0031] (2.1.4) Perform a global max pooling operation on the output feature map of each convolution kernel, converting the convolution feature map into a uniform dimensional representation. Extract the most significant activation value from each feature map to obtain a scalar vector. Concatenate the pooling results corresponding to the three convolution kernels to form a 384-dimensional overall feature vector.

[0032] (2.1.5) Add the Dropout layer to enhance the generalization ability of the model. Set the parameter to 0.5 and randomly drop 50% of the neurons to prevent overfitting of the model.

[0033] (2.2) The text modality dataset Cipher text Each ciphertext data in is segmented with characters as the basic unit to build a vocabulary. Each word corresponds to an index number to form a word index sequence with uniform length.

[0034] (2.3) The TextCNN encoder model constructed according to step (2.1) completes the feature extraction of a text modal ciphertext data each time and outputs a text feature vector with a dimension of 384;

[0035] (2.4) Complete the feature extraction of all text modal ciphertext data and obtain the text modal feature dataset Fea text .

[0036] In the block cipher algorithm recognition method based on multimodal feature fusion of the present invention, the image modal feature extraction in step (3) is specifically performed as follows:

[0037] (3.1) Construct a convolutional neural network (CNN) encoder model;

[0038] (3.1.1) The Convolutional Neural Network (CNN) encoder model consists of two convolutional blocks, a flattening layer, and a fully connected layer;

[0039] (3.1.2) Each convolution block consists of one convolution layer and one maximum pooling layer. The convolution layer is used to extract local visual features and output feature maps, and the maximum pooling layer is used to reduce the feature dimension of the feature maps. The first convolution block sets the number of input channels to 1, the number of output channels to 32, the convolution kernel size to 3×3, the padding to 1, and the activation function to ReLU. The maximum pooling layer sets the window size to 2×2. The second convolution block sets the number of input channels to 32, the number of output channels to 64, the convolution kernel size to 3×3, the padding to 1, and the activation function to ReLU. The maximum pooling layer sets the window size to 2×2.

[0040] (3.1.3) The flattening layer is used to flatten the obtained multi-dimensional feature map into a one-dimensional vector for input into the fully connected layer;

[0041] (3.1.4) The fully connected layer maps the flattened feature vector to a 128-dimensional space and adds nonlinearity through the ReLU activation function, ultimately outputting a fixed-length 128-dimensional image feature vector.

[0042] (3.2) Image modal ciphertext dataset Cipher image All grayscale images in are resized to 32×32;

[0043] (3.3) The convolutional neural network (CNN) encoder model constructed according to step (3.1) extracts features from a preprocessed grayscale image each time and outputs a fixed-length 128-dimensional image feature vector;

[0044] (3.4) Complete the feature extraction of all grayscale images and obtain the feature dataset Fea of ​​the image modality image .

[0045] In the block cipher algorithm recognition method based on multimodal feature fusion of the present invention, the speech modal feature extraction in step (4) is specifically performed as follows:

[0046] (4.1) Construct a Wav2Vec2 encoder model for feature extraction of speech modality data;

[0047] (4.1.1) The Wav2Vec2 encoder model, a self-supervised learning framework developed by Facebook AI, is used to extract features from speech modal data. For the input ciphertext data, the model first performs low-level modeling using convolutional layers. It then uses a multi-layer Transformer to perform global context modeling on the audio, outputting a time series feature matrix of size T × B, where T is the number of time frames and B is the feature dimension.

[0048] (4.1.2) Add a max pooling layer to aggregate the variable-length time series features to generate a 768-dimensional fixed-length vector;

[0049] (4.2) The speech modal ciphertext dataset Cipher audio All the data in the are sequentially input into the constructed Wav2Vec2 encoder model for feature extraction, and the speech modal features of each speech modal ciphertext data are obtained, and finally the speech modal feature dataset Fea of ​​all speech modal ciphertext data is obtained. audio .

[0050] In the block cipher algorithm recognition method based on multimodal feature fusion of the present invention, the feature fusion and model training of the different modal data in step (5) are specifically performed as follows:

[0051] (5.1) The characteristic data set Fea of ​​the three modes is divided into text 、Fea image 、Fea audio Perform segmentation to obtain training sets Fea containing three different feature subsets with algorithm labels. train And the test set Fea containing 3 different feature subsets without algorithm labels test , the algorithm tag is the serial number of the set block cipher algorithm;

[0052] (5.2) To ensure that the feature dimensions of the three modalities in the training set are the same, a fully connected mapping is called to map the feature datasets of the three modalities in the training set to the same dimension D, D = 521;

[0053] (5.3) Call a transformer-based self-attention mechanism model and set num_heads to 4, indicating the use of 4 attention heads. Take the feature datasets of the three modalities in the training set as input data and stack them into a sequence H with a dimension of 3×D=1563. Use the self-attention mechanism to calculate the sequence H and obtain the weighted trimodal feature representation H. att =Attention(H);

[0054] (5.4) Perform weighted fusion on the output trimodal feature representation to obtain the vector FTtrain fusion , FTtrain fusion =H att [0]+H att [1]+H att [2];

[0055] (5.5) The fused features are represented as FTtrain fusionThe input is sent to the fully connected layer, calculated using the nonlinear activation function ReLU, and then outputs its classification label Label through a Softmax layer; the accuracy evaluation index is used to evaluate the model classification effect, and the model is iterated to complete the model training and obtain a trained multimodal feature fusion model.

[0056] In the block cipher algorithm identification method based on multimodal feature fusion of the present invention, the model prediction and encryption algorithm identification of unknown ciphertext data in step (6) are specifically performed as follows:

[0057] (6.1) To ensure that the feature dimensions of the three modalities in the test set are the same, a fully connected mapping is called to map the feature datasets of the three modalities in the test set to the same dimension D, D = 521;

[0058] (6.2) The processed test set Fea test The data is used as the input data of the trained multimodal feature fusion model. First, follow steps (5.3) and (5.4) to obtain the feature fusion representation FTest of the test set during the model processing. fusion ;

[0059] (6.3) In the trained model, FTest fusion The input is sent to the fully connected layer, and after calculation using the nonlinear activation function ReLU, it is output through a Softmax layer with the probability of different classification labels. The label with the highest probability is the prediction result, thereby determining the cryptographic algorithm used when the unknown ciphertext data is encrypted.

[0060] The beneficial effects of the present invention are:

[0061] (1) The present invention proposes a block cipher algorithm recognition method based on multimodal feature fusion. The method selects three modal data types, namely, ciphertext, image, and speech, and extracts features through different encoders. The method adopts the idea of ​​feature fusion and introduces a self-attention mechanism to fuse multimodal features, thereby effectively improving recognition accuracy and having stronger robustness and generalization ability than a single modal input model.

[0062] (2) The method of the present invention has good versatility and is applicable to the encryption algorithm identification of unknown ciphertext data in various encryption scenarios and data distribution environments, and has good classification effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 This is a flow chart of a block cipher algorithm recognition method based on multimodal feature fusion according to the present invention.

[0064] Specific implementation examples

[0065] The present invention will be further described in detail below with reference to the embodiments and drawings, but the present invention is not limited thereto. Example

[0066] A block cipher algorithm recognition method based on multimodal feature fusion, referring to Figure 1 , including the following steps:

[0067] (1) Data preparation and preprocessing;

[0068] (2) Text modality feature extraction;

[0069] (3) Image modality feature extraction;

[0070] (4) Speech modal feature extraction;

[0071] (5) Feature fusion and model training of different modal data;

[0072] (6) Model prediction and encryption algorithm identification of unknown ciphertext data.

[0073] In the block cipher algorithm identification method based on multimodal feature fusion of the present invention, the data preparation and preprocessing in step (1) are specifically performed as follows:

[0074] (1.1) Select n random plaintext data as the plaintext data set Plaintext={p1,p2,…,p i ,…,p n}, p i represents the i-th plaintext data, where n = 1000, 1≤i≤n;

[0075] (1.2) For all the data in the plaintext dataset Plaintext, the open source cryptographic libraries Crypto and GmSSL are used to encrypt the plaintext data with the selected k block cipher algorithms, using random keys and ECB encryption mode, forming an original ciphertext dataset of size n×k saved in binary format Cipher={c 1,1 ,c 1,2 ,…,c q,i ,…,c k,n}, where k = 5, c q,i Indicates that the qth block cipher algorithm is used to cipher the i-th data p in the plaintext dataset Plaintext i The encrypted ciphertext data, 1≤q≤k, 1≤i≤n; and the maximum length of all ciphertext data in Cipher is calculated as LenC;

[0076] (1.3) Preprocess all ciphertext data in the original ciphertext dataset Cipher to remove redundant spaces and invalid characters; and for each ciphertext data, if its length is less than LenC, fill it with zero value to obtain the text modal ciphertext dataset Cipher text ={CT 1,1 ,CT 1,2 ,…,CT q,i ,…,CT k,n}, where CT q,i Indicates the ciphertext data c q,i The processed text modal ciphertext data; and the Cipher text Convert all binary ciphertext data in to hexadecimal form;

[0077] (1.4) Call the Python image conversion function to convert all the ciphertext data in the original ciphertext dataset Cipher into grayscale images, and set the image size ratio to 3×4; for the ciphertext data c q,i , according to the grayscale image size, each byte value is mapped to the pixel value in the image in turn, and the image is generated using the grayscale mode. If the ciphertext data c q,i If the length of is not enough to convert to a grayscale image, zero padding is performed to complete the grayscale image conversion; all ciphertext data are converted into grayscale images using the same method to obtain the image modal ciphertext dataset Cipher image ={CI 1,1 ,CI 1,2 ,…,CI q,i ,…,CI k,n}, where CI q,i Indicates the ciphertext data c q,i The converted image modal ciphertext data;

[0078] (1.5) Use the speech synthesis engine tool TTS to perform speech encoding on all the ciphertext data in the original ciphertext dataset Cipher. Set the speech synthesis language to English, the speech rate to 1, and the unified sampling frequency to 16kHz to obtain the speech modal ciphertext dataset Cipher. audio ={CA 1,1 ,CA 1,2 ,…,CA q,i ,…,CA k,n}, where CA q,i Indicates the ciphertext data c q,i The converted voice modal ciphertext data.

[0079] In the block cipher algorithm recognition method based on multimodal feature fusion of the present invention, the text modal feature extraction in step (2) is specifically performed as follows:

[0080] (2.1) Build a TextCNN encoder model based on convolutional neural network;

[0081] (2.1.1) The TextCNN encoder model consists of a word embedding layer, three one-dimensional convolutional layers, a global max pooling layer, and a dropout layer;

[0082] (2.1.2) The word embedding layer is used to map the input word index sequence into a dense word vector, thereby outputting a two-dimensional matrix to form a set of word vectors for the input text data. Setting the parameter embed_dim to 128 indicates that the dimension of each word embedding vector is 128.

[0083] (2.1.3) The convolutional layer is used to extract local semantic features from word vector data and output feature maps corresponding to the activation distribution of the corresponding local pattern. Three one-dimensional convolutional layers are connected in sequence, with kernel_size of 3, 4, and 5, respectively, to capture the corresponding local features and output feature maps of different granularities, such as triplets, quadruplets, and quintuples.

[0084] (2.1.4) Perform a global max pooling operation on the output feature map of each convolution kernel, converting the convolution feature map into a uniform dimensional representation. Extract the most significant activation value from each feature map to obtain a scalar vector. Concatenate the pooling results corresponding to the three convolution kernels to form a 384-dimensional overall feature vector.

[0085] (2.1.5) Add the Dropout layer to enhance the generalization ability of the model. Set the parameter to 0.5 and randomly drop 50% of the neurons to prevent overfitting of the model.

[0086] (2.2) The text modality dataset Cipher text Each ciphertext data in is segmented with characters as the basic unit to build a vocabulary. Each word corresponds to an index number to form a word index sequence with uniform length.

[0087] (2.3) The TextCNN encoder model constructed according to step (2.1) completes the feature extraction of a text modal ciphertext data each time and outputs a text feature vector with a dimension of 384;

[0088] (2.4) Complete the feature extraction of all text modal ciphertext data and obtain the text modal feature dataset Fea text .

[0089] In the block cipher algorithm recognition method based on multimodal feature fusion of the present invention, the image modal feature extraction in step (3) is specifically performed as follows:

[0090] (3.1) Construct a convolutional neural network (CNN) encoder model;

[0091] (3.1.1) The Convolutional Neural Network (CNN) encoder model consists of two convolutional blocks, a flattening layer, and a fully connected layer;

[0092] (3.1.2) Each convolution block consists of one convolution layer and one maximum pooling layer. The convolution layer is used to extract local visual features and output feature maps, and the maximum pooling layer is used to reduce the feature dimension of the feature maps. The first convolution block sets the number of input channels to 1, the number of output channels to 32, the convolution kernel size to 3×3, the padding to 1, and the activation function to ReLU. The maximum pooling layer sets the window size to 2×2. The second convolution block sets the number of input channels to 32, the number of output channels to 64, the convolution kernel size to 3×3, the padding to 1, and the activation function to ReLU. The maximum pooling layer sets the window size to 2×2.

[0093] (3.1.3) The flattening layer is used to flatten the obtained multi-dimensional feature map into a one-dimensional vector for input into the fully connected layer;

[0094] (3.1.4) The fully connected layer maps the flattened feature vector to a 128-dimensional space and adds nonlinearity through the ReLU activation function, ultimately outputting a fixed-length 128-dimensional image feature vector.

[0095] (3.2) Image modal ciphertext dataset Cipher image All grayscale images in are resized to 32×32;

[0096] (3.3) The convolutional neural network (CNN) encoder model constructed according to step (3.1) extracts features from a preprocessed grayscale image each time and outputs a fixed-length 128-dimensional image feature vector;

[0097] (3.4) Complete the feature extraction of all grayscale images and obtain the feature dataset Fea of ​​the image modality image .

[0098] In the block cipher algorithm recognition method based on multimodal feature fusion of the present invention, the speech modal feature extraction in step (4) is specifically performed as follows:

[0099] (4.1) Construct a Wav2Vec2 encoder model for feature extraction of speech modality data;

[0100] (4.1.1) The Wav2Vec2 encoder model, a self-supervised learning framework developed by Facebook AI, is used to extract features from speech modal data. For the input ciphertext data, the model first performs low-level modeling using convolutional layers. It then uses a multi-layer Transformer to perform global context modeling on the audio, outputting a time series feature matrix of size T × B, where T is the number of time frames and B is the feature dimension.

[0101] (4.1.2) Add a max pooling layer to aggregate the variable-length time series features to generate a 768-dimensional fixed-length vector;

[0102] (4.2) The speech modal ciphertext dataset Cipher audio All the data in the are sequentially input into the constructed Wav2Vec2 encoder model for feature extraction, and the speech modal features of each speech modal ciphertext data are obtained, and finally the speech modal feature dataset Fea of ​​all speech modal ciphertext data is obtained. audio .

[0103] In the block cipher algorithm recognition method based on multimodal feature fusion of the present invention, the feature fusion and model training of the different modal data in step (5) are specifically performed as follows:

[0104] (5.1) The characteristic data set Fea of ​​the three modes is divided into text 、Fea image 、Fea audio Perform segmentation to obtain training sets Fea containing three different feature subsets with algorithm labels. train And the test set Fea containing 3 different feature subsets without algorithm labels test , the algorithm tag is the serial number of the set block cipher algorithm;

[0105] (5.2) To ensure that the feature dimensions of the three modalities in the training set are the same, a fully connected mapping is called to map the feature datasets of the three modalities in the training set to the same dimension D, D = 521;

[0106] (5.3) Call a transformer-based self-attention mechanism model and set num_heads to 4, indicating the use of 4 attention heads. Take the feature datasets of the three modalities in the training set as input data and stack them into a sequence H with a dimension of 3×D=1563. Use the self-attention mechanism to calculate the sequence H and obtain the weighted trimodal feature representation H. att =Attention(H);

[0107] (5.4) Perform weighted fusion on the output trimodal feature representation to obtain the vector FTtrain fusion , FTtrain fusion =H att [0]+H att [1]+H att [2];

[0108] (5.5) The fused features are represented as FTtrain fusion The input is sent to the fully connected layer, calculated using the nonlinear activation function ReLU, and then outputs its classification label Label through a Softmax layer; the accuracy evaluation index is used to evaluate the model classification effect, and the model is iterated to complete the model training and obtain a trained multimodal feature fusion model.

[0109] In the block cipher algorithm identification method based on multimodal feature fusion of the present invention, the model prediction and encryption algorithm identification of unknown ciphertext data in step (6) are specifically performed as follows:

[0110] (6.1) To ensure that the feature dimensions of the three modalities in the test set are the same, a fully connected mapping is called to map the feature datasets of the three modalities in the test set to the same dimension D, D = 521;

[0111] (6.2) The processed test set Fea test The data is used as the input data of the trained multimodal feature fusion model. First, follow steps (5.3) and (5.4) to obtain the feature fusion representation FTest of the test set during the model processing. fusion ;

[0112] (6.3) In the trained model, FTest fusion The input is sent to the fully connected layer, and after calculation using the nonlinear activation function ReLU, it is output through a Softmax layer with the probability of different classification labels. The label with the highest probability is the prediction result, thereby determining the cryptographic algorithm used when the unknown ciphertext data is encrypted.

Claims

1. A block cipher algorithm identification method based on multimodal feature fusion, characterized in that: The following steps are involved: (1) Data preparation and preprocessing; Select random plaintext as the plaintext data set, and encrypt all the plaintext data in the plaintext data set with different block cipher algorithms to generate the original ciphertext text data set; Preprocess the ciphertext data and convert it into ciphertext datasets in three modalities: text, image, and voice. (2) Text modality feature extraction; The convolutional neural network-based text classification model TextCNN is used as the text modality encoder model. After operations such as text input, word embedding, convolution feature extraction, maximum pooling, and feature splicing, the text modality ciphertext dataset is subjected to feature extraction to obtain the feature set of the text modality data. (3) Image modality feature extraction; A convolutional neural network (CNN) is used as the image modality encoder model. After image input, local feature extraction by the convolution layer, dimensionality reduction by the pooling layer, flattening layer, and fully connected layer operations, the image modality ciphertext dataset is subjected to feature extraction to obtain the feature set of the image modality data. (4) Speech modal feature extraction; A self-supervised learning framework Wav2Vec2 is used as the speech modality encoder model. After voice input, model encoding, feature aggregation and other operations, feature extraction is performed on the speech modality ciphertext dataset to obtain the feature set of speech modality data. (5) Feature fusion and model training of different modal data; Based on the transformer model and using a mid-term fusion strategy, feature data from different modal feature sets are aligned, feature weighted fusion is completed, and model training is performed. (6) Model prediction and encryption algorithm identification of unknown ciphertext data; For unknown ciphertext data, the trained model is used to identify the cryptographic algorithm.

2. The method for identifying a block cipher algorithm based on multimodal feature fusion according to claim 1, characterized in that: The data preparation and preprocessing described in step (1) are as follows: (1.1) Select n random plaintext data as the plaintext data set Plaintext={p1,p2,…,p i ,…,p n }, p i Represents the i-th plaintext data, 1≤i≤n; (1.2) For all the data in the plaintext dataset Plaintext, the open source cryptographic libraries Crypto and GmSSL are used to encrypt the plaintext data with the selected k block cipher algorithms, using random keys and ECB encryption mode, forming an original ciphertext dataset of size n×k saved in binary format Cipher={c 1,1 ,c 1,2 ,…,c q,i ,…,c k,n }, where c q,i Indicates that the qth block cipher algorithm is used to cipher the i-th data p in the plaintext dataset Plaintext i The encrypted ciphertext data, 1≤q≤k, 1≤i≤n; and the maximum length of all ciphertext data in Cipher is calculated as LenC; (1.3) Preprocess all ciphertext data in the original ciphertext dataset Cipher to remove redundant spaces and invalid characters; and for each ciphertext data, if its length is less than LenC, fill it with zero value to obtain the text modal ciphertext dataset Cipher text ={CT 1,1 ,CT 1,2 ,…,CT q,i ,…,CT k,n }, where CT q,i Indicates the ciphertext data c q,i The processed text modal ciphertext data; and the Cipher text Convert all binary ciphertext data in to hexadecimal form; (1.4) Call the Python image conversion function to convert all the ciphertext data in the original ciphertext dataset Cipher into grayscale images, and set the image size ratio to 3×4; for the ciphertext data c q,i , according to the grayscale image size, each byte value is mapped to the pixel value in the image in turn, and the image is generated using the grayscale mode. If the ciphertext data c q,i If the length of is not enough to convert to a grayscale image, zero padding is performed to complete the grayscale image conversion; all ciphertext data are converted into grayscale images using the same method to obtain the image modal ciphertext dataset Cipher image ={CI 1,1 ,CI 1,2 ,…,CI q,i ,…,CI k,n }, where CI q,i Indicates the ciphertext data c q,i The converted image modal ciphertext data; (1.5) Use the speech synthesis engine tool TTS to perform speech encoding on all the ciphertext data in the original ciphertext dataset Cipher. Set the speech synthesis language to English, the speech rate to 1, and the unified sampling frequency to 16kHz to obtain the speech modal ciphertext dataset Cipher. audio ={CA 1,1 ,CA 1,2 ,…,CA q,i ,…,CA k,n }, where CA q,i Indicates the ciphertext data c q,i The converted voice modal ciphertext data.

3. The method for identifying a block cipher algorithm based on multimodal feature fusion according to claim 1, characterized in that: The text modality feature extraction described in step (2) is specifically performed as follows: (2.1) Build a TextCNN encoder model based on convolutional neural network; (2.1.1) The TextCNN encoder model consists of a word embedding layer, three one-dimensional convolutional layers, a global max pooling layer, and a dropout layer; (2.1.2) The word embedding layer is used to map the input word index sequence into a dense word vector, thereby outputting a two-dimensional matrix to form a set of word vectors for the input text data. Setting the parameter embed_dim to 128 indicates that the dimension of each word embedding vector is 128. (2.1.3) The convolutional layer is used to extract local semantic features from word vector data and output feature maps corresponding to the activation distribution of the corresponding local pattern. Three one-dimensional convolutional layers are connected in sequence, with kernel_size of 3, 4, and 5, respectively, to capture the corresponding local features and output feature maps of different granularities, such as triplets, quadruplets, and quintuples. (2.1.4) Perform a global max pooling operation on the output feature map of each convolution kernel, converting the convolution feature map into a uniform dimensional representation. Extract the most significant activation value from each feature map to obtain a scalar vector. Concatenate the pooling results corresponding to the three convolution kernels to form a 384-dimensional overall feature vector. (2.1.5) Add the Dropout layer to enhance the generalization ability of the model. Set the parameter to 0.5 and randomly drop 50% of the neurons to prevent overfitting of the model. (2.2) The text modality dataset Cipher text Each ciphertext data in is segmented with characters as the basic unit to build a vocabulary. Each word corresponds to an index number to form a word index sequence with uniform length. (2.3) The TextCNN encoder model constructed according to step (2.1) completes the feature extraction of a text modal ciphertext data each time and outputs a text feature vector with a dimension of 384; (2.4) Complete the feature extraction of all text modal ciphertext data and obtain the text modal feature dataset Fea text .

4. The method for identifying a block cipher algorithm based on multimodal feature fusion according to claim 1, characterized in that: The image modality feature extraction described in step (3) is specifically performed as follows: (3.1) Construct a convolutional neural network (CNN) encoder model; (3.1.1) The convolutional neural network (CNN) encoder model consists of two convolutional blocks, a flattening layer, and a fully connected layer. (3.1.2) Each convolutional block consists of one convolutional layer and one maximum pooling layer. The convolutional layer is used to extract local visual features and output a feature map, and the maximum pooling layer is used to reduce the feature dimension of the feature map. The first convolutional block has 1 input channel, 32 output channels, a 3×3 convolution kernel size, 1 padding, and a ReLU activation function. The maximum pooling layer has a 2×2 window size. The second convolutional block has 32 input channels, 64 output channels, a 3×3 convolution kernel size, 1 padding, and a ReLU activation function. The maximum pooling layer has a 2×2 window size. (3.1.3) The flattening layer flattens the obtained multi-dimensional feature map into a one-dimensional vector for input into the fully connected layer; (3.1.4) The fully connected layer maps the flattened feature vector to a 128-dimensional space and adds nonlinearity through the ReLU activation function, ultimately outputting a fixed-length 128-dimensional image feature vector. (3.2) Image modal ciphertext dataset Cipher image All grayscale images in are resized to 32×32; (3.3) The convolutional neural network (CNN) encoder model constructed according to step (3.1) extracts features from a preprocessed grayscale image each time and outputs a fixed-length 128-dimensional image feature vector; (3.4) Complete the feature extraction of all grayscale images and obtain the feature dataset Fea of ​​the image modality image .

5. The method for identifying a block cipher algorithm based on multimodal feature fusion according to claim 1, characterized in that: The speech modality feature extraction described in step (4) is specifically performed as follows: (4.1) Construct a Wav2Vec2 encoder model for feature extraction of speech modality data; (4.1.1) The Wav2Vec2 encoder model, a self-supervised learning framework developed by Facebook AI, is used to extract features from speech modal data. For the input ciphertext data, the model first performs low-level modeling using convolutional layers. It then uses a multi-layer Transformer to perform global context modeling on the audio, outputting a time series feature matrix of size T × B, where T is the number of time frames and B is the feature dimension. (4.1.2) Add a max pooling layer to aggregate the variable-length time series features to generate a 768-dimensional fixed-length vector; (4.2) The speech modal ciphertext dataset Cipher audio All the data in the are sequentially input into the constructed Wav2Vec2 encoder model for feature extraction, and the speech modal features of each speech modal ciphertext data are obtained, and finally the speech modal feature dataset Fea of ​​all speech modal ciphertext data is obtained. audio .

6. The method for identifying a block cipher algorithm based on multimodal feature fusion according to claim 1, characterized in that: The feature fusion and model training of the different modal data in step (5) are as follows: (5.1) The characteristic data set Fea of ​​the three modes is divided into text 、Fea image 、Fea audio Perform segmentation to obtain training sets Fea containing three different feature subsets with algorithm labels. train And the test set Fea containing 3 different feature subsets without algorithm labels test , the algorithm tag is the serial number of the set block cipher algorithm; (5.2) To ensure that the feature dimensions of the three modalities in the training set are the same, a fully connected mapping is called to map the feature datasets of the three modalities in the training set to the same dimension D, D = 521; (5.3) Call a transformer-based self-attention mechanism model and set num_heads to 4, indicating the use of 4 attention heads. Take the feature datasets of the three modalities in the training set as input data and stack them into a sequence H with a dimension of 3×D=1563. Use the self-attention mechanism to calculate the sequence H and obtain the weighted trimodal feature representation H. att =Attention(H); (5.4) Perform weighted fusion on the output trimodal feature representation to obtain the vector FTtrain fusion , FTtrain fusion =H att [0]+H att [1]+H att [2]; (5.5) The fused features are represented as FTtrain fusion The input is sent to the fully connected layer, calculated using the nonlinear activation function ReLU, and then outputs its classification label Label through a Softmax layer; the accuracy evaluation index is used to evaluate the model classification effect, and the model is iterated to complete the model training and obtain a trained multimodal feature fusion model.

7. The method for identifying a block cipher algorithm based on multimodal feature fusion according to claim 1, characterized in that: The model prediction described in step (6) and the encryption algorithm identification of unknown ciphertext data are specifically performed as follows: (6.1) To ensure that the feature dimensions of the three modalities in the test set are the same, a fully connected mapping is called to map the feature datasets of the three modalities in the test set to the same dimension D, D = 521; (6.2) The processed test set Fea test The data is used as the input data of the trained multimodal feature fusion model. First, follow steps (5.3) and (5.4) to obtain the feature fusion representation FTest of the test set during the model processing. fusion ; (6.3) In the trained model, FTest fusion The input is sent to the fully connected layer, and after calculation using the nonlinear activation function ReLU, it is output through a Softmax layer with the probability of different classification labels. The label with the highest probability is the prediction result, thereby determining the cryptographic algorithm used when the unknown ciphertext data is encrypted.

Citation Information

Cited By

  • Block cipher identification method based on quantum self-organizing fuzzy neural network

    CN121396429A