Method and device for constructing neural network difference partition device based on multi-scale convolution
By adopting multi-scale convolution and residual tower reconstruction methods in differential dividers, the neural network structure is optimized to adapt to the ciphertext characteristics, and the problem of lack of universality and generalization of differential dividers in the prior art is solved, and the high accuracy distinction and key recovery attacks of lightweight symmetric encryption algorithms are achieved.
Patent Information
- Application Number
- CN202510348795.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The prior art lacks universality and generalization when building differential dividers, and fails to effectively consider the characteristics of the ciphertext structure, resulting in poor distinction effect.
Using a solution based on multi-scale convolution and residual tower reconstruction, the neural network differential divider is constructed and optimized, the number of residual towers and convolution layers is adjusted, and the multi-scale idea of pyramid convolution is combined with the convolution kernel size to adapt to the ciphertext structure.
High accuracy distinction between lightweight symmetric encryption algorithms with reduced rounds to 5-7 rounds was achieved, and 11 rounds of key recovery attacks of symmetric cryptography algorithms were successfully carried out, which significantly improved the accuracy of ciphertext distinction and key recovery and deciphering success rate.
Smart Images

Figure CN120146111A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of the combination of cryptography and artificial intelligence, and more specifically, to a method and device for constructing a neural network differential distinguisher based on multi-scale convolution. Background Art
[0002] Nowadays, with the popularization of big data systems, Internet of Things (IoT) systems, and various intelligent application systems, cryptography is facing new challenges while developing rapidly. For example, the emergence of quantum computers has given rise to a large number of studies on quantum and post-quantum cryptography, and researchers need to develop new cryptographic algorithms and protocols, such as lattice-based and code-based cryptography. And in this era of everything being interconnected, while a large number of IoT devices are being put into use, their own storage and computing capabilities are limited and they cannot use conventional cryptographic algorithms. Therefore, block cipher algorithms that are simple in design, high in operation efficiency, and suitable for resource-constrained environments have been proposed and received extensive attention and research. The design and analysis of block ciphers are two complementary research fields. On the one hand, using existing cryptographic analysis means, cipher designers expect to design cryptographic algorithms that can resist all existing known attacks. On the other hand, for existing cryptographic algorithms, cryptanalysts also expect to be able to attack the cryptographic algorithms by finding some security flaws in the algorithms. These two complementary aspects are constantly promoting the development of block cipher theory.
[0003] In cryptography, the meanings of the two terms "analysis" and "attack" are equivalent. Cryptanalysis also studies the security strength of a cipher when using different attack means to break the cipher, which is divided into practical security and theoretical security. Evaluating the practical security of an encryption algorithm is done by considering the actual computational amount required in the process of breaking the cipher, without considering the attacker's computing power and computing time. This evaluation usually involves key recovery attacks, that is, using known information (including partial plaintext and its corresponding ciphertext) to recover the encryption key of the cryptographic algorithm or some bits of the key. Generally, the security analysis of block ciphers mainly focuses on key recovery attacks. In cryptanalysis, the well-known Kerckhoff's assumption is usually followed, that is, "the security of a cryptographic system should be based on the secrecy of the key, rather than on the secrecy of algorithm details". Under Kerckhoff's assumption, the analysis of cipher security is essentially an evaluation of the secrecy of the key, rather than an evaluation of the encryption algorithm itself. Therefore, the security analysis of encryption algorithms mainly focuses on the secrecy of the key, rather than the specific details of the encryption algorithm. This analysis method believes that even if an attacker understands the working principle of the encryption algorithm, as long as the key is kept secret, the encryption system can still remain secure. Therefore, the focus of cryptanalysis is to determine the strength and secrecy of the encryption key.
[0004] For symmetric cryptographic algorithms, there are various cryptanalysis methods, among which the most important ones include differential cryptanalysis, linear cryptanalysis, integral attacks, and related attack variants. The analysis and attack of lightweight symmetric cryptography can generally be divided into two stages: distinguisher construction and key recovery. In the distinguisher construction stage, the attacker looks for non-random characteristics in the cryptographic algorithm, such as linear correlations in the internal state, or output differences that produce abnormal distributions given specific input differences. The key recovery stage targets the round functions before and after the constructed distinguisher and uses these non-random characteristics to (partially) recover the key bits. Essentially, the attacker makes guesses about the key information and checks for the existence of non-random characteristics after encrypting or decrypting several round functions. If the probability distribution of the statistical data meets the expectation, the guessed key is considered likely to be correct. Optimizing key recovery attacks is an important aspect of cryptanalysis. Similarly, the analysis of symmetric cryptography is carried out by extracting and applying non-random statistical characteristics in the cryptographic algorithm. In the early 1990s, Eli Biham and Adi Shamir proposed the differential analysis method at the Crypto conference. Differential cryptanalysis belongs to the chosen-plaintext attack method. It separates block ciphers from random permutations through the probability propagation characteristics of specific plaintext difference values during the encryption process and conducts key recovery attacks on this basis. Let the plaintext input pair be (x, x * ), then the difference value of x, x * is defined as Δ x = x ⊕ x * . After the first round of iteration, the intermediate ciphertext has a difference value of Δ x1 . After n rounds of iteration, a sequence of difference values Ω = (Δ x , Δ x1 ...... Δ xn ) is obtained. This sequence of difference values is called the n-round differential path of the block cipher, and the differential characteristic characterizes the differential propagation characteristics during the encryption process.
[0005] In the prior art, CN118114229A discloses a differential distinguisher for symmetric cryptography based on a residual neural network, including an initial convolution module, a residual module, a prediction module, etc. It mainly modifies the prediction module, and the modification method is simple hierarchical stacking. CN119519940A discloses a design method for a differential distinguisher of symmetric cryptography based on deep learning. A preliminary symmetric cryptography differential distinguisher is established based on an input module, an initial convolution module, a residual module, and a prediction module, and the preliminary symmetric cryptography differential distinguisher is optimized to generate several different residual structure distinguishers. It improves the residual module and the activation function, and combines different numbers of one-dimensional convolutional layers with batch normalization and the Hardswish activation function in the residual module. However, the existing methods only make specific improvements for the current task, lack universality and generalization, and do not consider the characteristics of the ciphertext structure itself, resulting in still poor performance of the constructed differential distinguisher. Summary of the Invention
[0006] The present invention provides an optimization method for constructing a differential distinguisher structure based on a multi-scale convolution optimization scheme, and uses the differential distinguisher constructed based on this method for related password breaking. The neural network differential distinguisher is constructed and optimized by adopting schemes such as multi-scale convolution and residual tower reconstruction, realizing high-accuracy differentiation of lightweight symmetric encryption algorithms with a reduction of rounds to 5-7 rounds. On this basis, a key recovery attack is successfully implemented for an 11-round symmetric cryptography algorithm, significantly improving the ciphertext differentiation accuracy and the success rate of key recovery and password breaking.
[0007] To achieve the above object, the first aspect of the present invention provides a method for constructing a neural network differential distinguisher based on multi-scale convolution, including:
[0008] A neural network differential distinguisher suitable for the ciphertext structure is initially constructed by adjusting the residual neural network. The initially constructed neural network differential distinguisher includes an input module, an initial convolution module, a residual module, and a prediction module. The input module is used to receive input data from the ciphertext pair, the initial convolution module is used to extract the features of the input data of the ciphertext pair, the residual module is used to extract the deep features of the input data of the ciphertext pair, and the prediction module is used to map the input features to the output label to obtain the final result;
[0009] The number of residual towers and the number of convolutional layers in the residual tower in the neural network differential distinguisher are adjusted to initially optimize the neural network differential distinguisher, and in combination with the multi-scale convolution idea of pyramid convolution, the composition of the convolutional kernel sizes of the neural network differential distinguisher is adjusted to obtain an optimized neural network differential distinguisher.
[0010] In one implementation, the number of residual towers and the number of convolutional layers in the residual towers within the neural network differential differentiator are adjusted to preliminarily optimize the neural network differential differentiator. Combining the multi-scale convolution idea of pyramid convolution, the composition of the convolutional kernel sizes of the neural network differential differentiator is adjusted to obtain an optimized neural network differential differentiator, including:
[0011] Referring to the main network construction scheme of Res2Net, the residual module is modified. After feature extraction in the initial convolutional module, the obtained feature tensor matrix is input into the residual module. Multiple copies of the feature tensor matrix are jointly used as learning data and input into the neural network. The way of splitting the channel dimension of the feature tensor at the connection of the residuals for learning and then splicing again is modified to adding the results. The convolutional kernel sizes in a single residual tower and the initial convolutional kernel sizes between the residual towers increase one by one, and an optimized first neural network differential differentiator ND1 is constructed;
[0012] And / or referring to the main network construction scheme of DensneNet, the residual module is modified. The features learned by each convolutional layer are connected to subsequent convolutional layers using a residual structure as part of the input, and the features are learned and fused multiple times. Four feature fusion layers are set to save and fuse features between different convolutional layers, increasing the number of convolutional layers in a single residual tower and increasing the convolutional kernel sizes within and directly between the residual towers, and an optimized second neural network differential differentiator ND2 is constructed;
[0013] And / or referring to and applying the idea of parallel convolution, the residual module is modified. It is composed of 5 parallel double-layer convolutional groups. The large residual tower has 5 double-layer convolutional layers, and the small residual tower has 2 convolutional layers. The convolutional kernel sizes between the small residual tower and the large residual tower increase one by one, and an optimized third neural network differential differentiator ND3 is constructed;
[0014] And / or referring to and applying the ideas of parallel convolution and multiple feature learning, the initial convolutional module is modified. A single-layer convolutional operation with a convolutional kernel size of 1 is used for feature extraction. The single-layer convolutional operation with a convolutional kernel size of 1 is performed three times, and after the second and third convolutional operations, convolutional operations with larger convolutional kernels are performed 1 time and 2 times respectively. Finally, the results of the three convolutional operations are added to obtain the final feature matrix and input into the residual module, and an optimized fourth neural network differential differentiator ND4 is constructed.
[0015] In one implementation, the number of convolutional layers in a single residual tower is increased from 2 to 5. The convolutional kernel size in a single residual tower starts from 1 and increases one by one, with an increment of 2. The initial convolutional kernel sizes between the residual towers increase one by one, with an increment of 2.
[0016] In one implementation, the method further includes:
[0017] Generating a ciphertext dataset and training the optimized neural network differential distinguisher.
[0018] In one implementation, generating a ciphertext dataset and training the optimized neural network differential distinguisher includes:
[0019] Simulating the Speck symmetric cipher algorithm to generate a dataset for training and validation;
[0020] Setting the learning rate, Epoch, and Batch Size, and using the Adam algorithm with default parameters in Keras to optimize for the cross-entropy loss function and a small penalty of L2 weight regularization. The learning rate uses a cyclic learning rate. Among them, the obtained network is stored at the end of each epoch, and the best network obtained is evaluated according to the test set, which is not used for training;
[0021] Training the neural network to obtain a differential distinguisher model file, which contains the discrimination accuracy of each neural network differential distinguisher for 5 - 8 rounds.
[0022] In one implementation, the method further includes: performing a key recovery attack based on the optimized neural network differential distinguisher to break the corresponding lightweight symmetric cipher.
[0023] In one implementation, performing a key recovery attack based on the optimized neural network differential distinguisher to break the corresponding lightweight symmetric cipher includes:
[0024] Based on neutral bits, Bayesian optimization, and the UCB problem, creating a key recovery attack algorithm. First, use the obtained ND2 neural network differential distinguisher for 6 and 7 rounds to extend the recovery rounds to 11 rounds;
[0025] Using the existing ND2 neural network differential distinguisher to break the Speck encryption with reduced rounds to 11 rounds.
[0026] Based on the same inventive concept, the second aspect of the present invention provides a construction device for a neural network differential distinguisher based on multi-scale convolution, including:
[0027] An initial construction module for initially constructing a neural network differential distinguisher suitable for the ciphertext structure by adjusting the residual neural network. The initially constructed neural network differential distinguisher includes an input module, an initial convolution module, a residual module, and a prediction module. The input module is used to receive input data from the ciphertext pair, the initial convolution module is used to extract the features of the ciphertext pair input data, the residual module is used to extract the deep features of the ciphertext pair input data, and the prediction module is used to map the input features to output labels to obtain the final result;
[0028] An optimization module, which is used to adjust the number of residual towers and the number of convolutional layers in the residual towers within the neural network differential differentiator, preliminarily optimize the neural network differential differentiator, and combine the multi-scale convolution idea of pyramid convolution to adjust the composition of the convolutional kernel sizes of the neural network differential differentiator, so as to obtain an optimized neural network differential differentiator.
[0029] Based on the same inventive concept, the third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for constructing a neural network differential differentiator based on multi-scale convolution described in the first aspect.
[0030] Based on the same inventive concept, the fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for constructing a neural network differential differentiator based on multi-scale convolution described in the first aspect.
[0031] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:
[0032] The present invention proposes an optimization method for constructing a differential differentiator structure based on a multi-scale convolution optimization scheme. By adjusting the residual neural network, a neural network differential differentiator suitable for the ciphertext structure is initially constructed. Then, the number of residual towers and the number of convolutional layers in the residual towers within the neural network differential differentiator are adjusted to preliminarily optimize the neural network differential differentiator. Combining the multi-scale convolution idea of pyramid convolution, the composition of the convolutional kernel sizes of the neural network differential differentiator is adjusted to obtain an optimized neural network differential differentiator. The optimized differential differentiator achieves high-precision differentiation of lightweight symmetric encryption algorithms reduced to 5-7 rounds.
[0033] On this basis, for the 11-round symmetric cipher algorithm, a key recovery attack is successfully carried out to achieve the goal of breaking. This achievement effectively improves the analysis level of lightweight symmetric encryption algorithms, bringing new improvement ways and optimization ideas to cryptography research. In the field of information security, this technology can be used to evaluate and improve the security of encryption algorithms, especially in resource-limited environments such as the Internet of Things and mobile communications, where it has extremely crucial application value. Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 This is a flowchart of the construction method of the neural network differential distinguisher based on multi-scale convolution in the embodiments of the present invention;
[0036] Figure 2 This is a framework diagram of initially constructing a neural network differential distinguisher in the embodiments of the present invention;
[0037] Figure 3 This is a schematic structural diagram of the residual module of the optimized first neural network differential distinguisher ND1 in the embodiments of the present invention;
[0038] Figure 4 This is a schematic structural diagram of the residual module of the optimized second neural network differential distinguisher ND2 in the embodiments of the present invention;
[0039] Figure 5 This is a schematic structural diagram of the residual module of the optimized third neural network differential distinguisher ND3 in the embodiments of the present invention;
[0040] Figure 6 This is a schematic structural diagram of the residual module of the optimized fourth neural network differential distinguisher ND4 in the embodiments of the present invention. Detailed implementation manners
[0041] The present invention discloses a method for constructing and optimizing a neural network differential distinguisher by using optimization schemes such as multi-scale convolution to achieve high-accuracy breaking of the reduced-round symmetric cipher algorithm, which includes the following steps: Step 1, propose a generalizable scheme for optimizing the residual tower architecture of the differential distinguisher based on the idea of multi-scale convolution. Step 2, construct a neural network differential distinguisher with stronger distinguishing effect for the Speck encryption algorithm with reduced rounds from 5 to 7 based on the scheme in Step 1. Further, it also includes Step 3, generate a ciphertext data set and train to obtain a differential distinguisher model. Step 4, perform a key recovery attack based on the optimized neural network differential distinguisher to break the corresponding lightweight symmetric cipher.
[0042] The present invention constructs and optimizes a neural network differential distinguisher through schemes such as multi-scale convolution and residual tower reconstruction, achieving a high-accuracy distinction for lightweight symmetric encryption algorithms reduced to 5-7 rounds. On this basis, a key recovery attack was successfully implemented against the 11-round symmetric cipher algorithm, thus achieving the purpose of breaking. This achievement significantly improves the analysis ability of lightweight symmetric encryption algorithms, providing new improvement methods and optimization ideas for cryptography research. In the field of information security, this technology can be used to evaluate and improve the security of encryption algorithms, especially in resource-constrained environments such as the Internet of Things and mobile communications, and has important application value. Through in-depth analysis of lightweight symmetric encryption algorithms, the present invention provides a powerful tool for cryptography research and practice, promoting technological progress in related fields.
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] Embodiment 1
[0045] This embodiment discloses a method for constructing a neural network differential distinguisher based on multi-scale convolution, including:
[0046] S1: Initially construct a neural network differential distinguisher suitable for ciphertext structures by adjusting a residual neural network. The initially constructed neural network differential distinguisher includes an input module, an initial convolution module, a residual module, and a prediction module. The input module is used to receive input data from ciphertext pairs. The initial convolution module is used to extract features of the input data of the ciphertext pairs. The residual module is used to extract deep features of the input data of the ciphertext pairs. The prediction module is used to map the input features to output labels to obtain the final result;
[0047] S2: Adjust the number of residual towers and the number of convolutional layers in the residual towers within the neural network differential distinguisher to initially optimize the neural network differential distinguisher, and combine the multi-scale convolution idea of pyramid convolution to adjust the composition of the convolutional kernel sizes of the neural network differential distinguisher to obtain an optimized neural network differential distinguisher.
[0048] The present invention defines the number of plaintext pairs that satisfy the input difference as Δ x and the last difference, i.e., the output difference as Δ xn as N D (Δ x , Δ xn ) The input difference Δx Transform to the output difference Δ xn with probability R P is defined as shown in formula (1):
[0049]
[0050] If the corresponding probabilities R of all differential paths can be calculated P (Δ x , Δ xn ), the differential distribution table (DDT) of the corresponding encryption algorithm is obtained, where the occurrence probability is much higher than 1 / 2 m The differential path of the difference is called a high-probability differential path. Finding the high-probability differential path with the highest probability is a prerequisite for differential analysis. The subsequent steps are as follows:
[0051] (1) Let the length of the sub-key of the r-th round to be recovered be L, and set a counter v for each key to be guessed i , which is used as the score of the candidate key;
[0052] (2) Uniformly select and generate random plaintexts p 1 , p 2 ...... p n , and let p i be XORed with the Δ x value in the high-probability differential path to obtain the plaintext pair (p i , p i ′ ). After encrypting r + 1 rounds, the ciphertext pair (c i , c i ′ ) is obtained;
[0053] (3) Use the differential distinguisher to filter all the ciphertext pairs, and then use the random key to decrypt the ciphertext pair (c i , c i ′ ). If the difference value of the decrypted ciphertext pair is Δ xn , the value of the counter v i is incremented by one, and the key with the largest resulting value is considered the correct key.
[0054] As a commonly used cryptographic attack technique, differential cryptanalysis is widely used in the analysis of cryptographic algorithms due to its high efficiency, universality, and the ability to reveal the internal characteristics of cryptographic algorithms. From the most basic DES encryption algorithm to Blowfish, Speck, LBlock, etc. encryption, differential analysis has now become one of the security indicators that must be considered in the design and analysis of symmetric cryptography.
[0055] Specifically, step S1 is the initial construction step. By adjusting the residual neural network, a neural network differential distinguisher suitable for the ciphertext structure is initially constructed. The specific structure of the network is as follows Figure 2 It can be seen that the neural network differential distinguisher consists of an input module (Module1), an initial convolution module (Module2), a residual module (Module3), and a prediction module (Module4). The input module is used to receive the data input from the ciphertext pair. A pair of (C0, C1) ciphertexts of Speck32 / 64 can be written as a sequence of four sixteen-bit words (w0, w1, w2, w3), which reflects the word-oriented structure of the ciphertext. In this network, wi is a row vector of a 4×16 matrix, and the input layer consists of 64 units, also arranged in a 4×16 array. In the initial convolution layer, a single-layer convolution operation with a convolution kernel size of 1 is used to extract the features of the ciphertext matrix in the input layer. The purpose of this step is to imitate the exclusive-or operation in the cryptographic operation. Since the exclusive-or operation cannot be performed in the neural network, this convolution layer performs convolution learning on the four bits that are exclusive-or'ed with each other in the cryptographic operation to extract the features therein. A BN layer and a Relu activation function are added after the convolution layer. Next, the residual module is used to improve the feature extraction ability of the model. In this module, 10 residual towers composed of two convolutional neural network layers are used to extract deeper features, that is, each residual tower contains 2 one-dimensional convolutions with a convolution kernel size of 3. A BN layer and a Relu activation function are also added after each convolution layer. Information loss is reduced through residual connections. This module is used for main feature extraction. Finally, 2 fully connected layers with 64 neurons are added as the prediction module. This module maps the input features to the output labels, predicts the real pair and the random pair, and outputs the final result to obtain the accuracy rate. The principle of the adopted and improved neural network differential distinguisher is approximately to construct an approximate cryptographic differential distribution table (DDT) during the learning stage and use this information to directly classify the ciphertext pair.
[0056] Step S2 is the optimization based on the model constructed in S1, including adjusting the number of residual towers in the neural network differential distinguisher and the number of convolution layers in the residual tower, and combining the multi-scale convolution idea of pyramid convolution to adjust the composition of the convolution kernel sizes of the neural network differential distinguisher.
[0057] In one implementation, by adjusting the number of residual towers in the neural network differential distinguisher and the number of convolution layers in the residual tower, the neural network differential distinguisher is initially optimized, and combining the multi-scale convolution idea of pyramid convolution, the composition of the convolution kernel sizes of the neural network differential distinguisher is adjusted to obtain the optimized neural network differential distinguisher, including:
[0058] Referring to the main network construction scheme of Res2Net, the residual module is modified. After feature extraction in the initial convolution module, the obtained feature tensor matrix is input into the residual module. Multiple copies of the feature tensor matrix are copied and used as learning data to input into the neural network. The method of splitting the channel dimension of the feature tensor at the connection of the residual and then learning and splicing again is modified to adding the results. The convolutional kernel sizes of the convolutional layers within a single residual tower and the initial convolutional kernel sizes between residual towers increase one by one, and the optimized first neural network differential discriminator ND1 is constructed;
[0059] And / or referring to the main network construction scheme of DensneNet, the residual module is modified. The features learned by each convolutional layer are connected to subsequent convolutional layers using a residual structure as part of the input, and the features are learned and fused multiple times. Among them, 4 feature fusion layers are set to save and fuse features between different convolutional layers, the number of convolutional layers in a single residual tower is increased, and the convolutional kernel sizes within and directly between residual towers are increased, and the optimized second neural network differential discriminator ND2 is constructed;
[0060] And / or referring to and applying the idea of parallel convolution, the residual module is modified. It is composed of 5 parallel double-layer convolutional groups. The large residual tower has 5 double-layer convolutional layers, and the small residual tower has 2 convolutional layers. The convolutional kernel sizes between the small residual tower and the large residual tower increase one by one, and the optimized third neural network differential discriminator ND3 is constructed;
[0061] And / or referring to and applying the idea of parallel convolution and multiple feature learning, the initial convolution module is modified. A single-layer convolution operation with a convolutional kernel size of 1 is used for feature extraction. The single-layer convolution operation with a convolutional kernel size of 1 is performed three times, and after the second and third convolution operations, 1 and 2 convolution operations with larger convolutional kernels are performed respectively. Finally, the results of the three convolutions are added to obtain the final feature matrix and input into the residual module, and the optimized fourth neural network differential discriminator ND4 is constructed.
[0062] Specifically, referring to the main network construction scheme of Res2Net, the first optimized differential discriminator ND1 is constructed. This differential discriminator has the ability of multi-scale feature representation and effectively improves the discrimination performance. Specifically: after feature extraction in the initial convolution module, the 16x32 feature tensor matrix is input into the residual module. Multiple copies of the feature tensor matrix are copied and used as learning data to input into the neural network, and the scheme of splicing again after splitting channels in Res2Net is modified to adding results suitable for the ciphertext structure, and the first optimized differential discriminator ND1 is constructed. The specific structure of the modified residual module is as Figure 3 shown. (FeatureTensorMatrix represents the feature tensor matrix)
[0063] Since Res2Net is an image neural network itself, in the residual block, the channel dimension of the feature tensor is segmented for learning and then concatenated again. In the invention, to make the neural network structure more suitable for learning the ciphertext structure, the results are directly added at the connection of the residuals, thus maintaining the integrity of the ciphertext structure. ND1 first inputs the 16x32 feature tensor matrix into the residual module after feature extraction in the initial convolution module. This embodiment changes the behavior of segmenting the matrix in the traditional Res2net network to preserve the integrity of the ciphertext structure. The 5 convolutional layers in the convolutional module are designed in a stepped shape, and the features learned by the previous convolutional layer and the subsequent replicated ciphertext data are used as the input data for the subsequent convolutional layer. The convolutional kernel size of the convolutional layers within a single residual tower and the initial convolutional kernel size between the residual towers increase by 2 one by one.
[0064] Referring to the main network construction scheme of DensneNet, the second optimized differential discriminator ND2 is constructed. The differential discriminator has a similar ciphertext discrimination ability to ND1. The specific structure of its modified residual module is as Figure 4 shown. This discriminator is the best discriminator for balancing accuracy, training loss, and time among them. The innovative idea of the ND2 residual tower module is similar to that of DenseNet, that is, the features learned by each convolutional layer are connected to the subsequent convolutional layer using the residual structure as part of the input, and the features are learned and fused multiple times. Four feature fusion layers (i.e., Figure 4 the circular plus sign part in) are set for feature preservation and fusion between different convolutional layers. In addition, the number of convolutional layers in a single residual tower is increased from 2 to 5, and the increasing relationship of the convolutional kernel size within and directly between the residual towers is also applied.
[0065] Referring to and applying the idea of parallel convolution, the third optimized differential discriminator ND3 is constructed. The differential discriminator has a similar ciphertext discrimination ability to ND1. The specific structure of its modified residual module is as Figure 5 shown. The residual module of the ND3 differential discriminator consists of 5 parallel double-layer convolutional groups. The small residual tower ( Figure 5 the structure formed by the two convolutional layers in each vertical row in) has two convolutional layers, but the entire residual tower contains five small residual towers. The convolutional kernel size between the small and large residual towers ( Figure 5 the residual module shown) still increases by 2 one by one. Such grouped convolution operations can make the convolution operations in the network more flexible and increase the expressive ability of the model. By increasing the number of groups of grouped convolution, the network can increase the width of the network, thereby increasing the capacity and representation ability of the model.
[0066] Similarly referring to and applying the ideas of parallel convolution and multiple feature learning, the fourth optimized differential distinguisher ND4 is constructed. The ciphertext discrimination ability of this differential distinguisher is slightly inferior to that of ND1-ND3, but the training loss and time are also significantly reduced, making it a lightweight differential distinguisher. The specific structure of its modified initial convolution module is as follows Figure 6 shown. For the initial convolution module, ND4 uses a single-layer convolution operation with a kernel size of 1 to extract features, which can be closer to the bit operation mode in the encryption algorithm and obtain the correlation information between the bits for exclusive OR operation. The single-layer convolution operation with a kernel size of 1 is performed three times (i.e., Figure 6 the three convolution operations in the first row in
[0067] ). After the second and third times, convolution operations with larger kernel sizes are performed 1 time and 2 times respectively, and finally the results of the three convolutions (i.e., the first convolution containing one convolution operation, the second convolution containing two convolutions, and the third convolution containing three convolutions) are added to obtain the final feature matrix as the input to the residual module.
[0068] Comparative document 2 - CN119519940A discloses a design method for symmetric cipher differential distinguishers based on deep learning. Its improvement is to add a simple attention mechanism inside the residual tower, which is different from the innovative ideas of multi-scale convolution and feature multiple learning in the present invention. In addition, comparative document 2 only modifies the residual module, while the improvement of the present invention targets both the residual module and the initial convolution module, which is more comprehensive compared. Similar to the disadvantages of comparative document 1, this improvement in comparative document 2 does not have high universality. What the present invention proposes is not just four excellent network structures, but an improvement scheme that can be widely applied in symmetric cipher discrimination experiments. The proposal of the four network structures is to prove the advantages of this improvement scheme. In addition, another defect of comparative document 2 is the low discrimination accuracy. Compared with the accuracy in the present invention, it is significantly lower, with only 93.04% in 5 rounds. Generally speaking, compared with the prior art, in the discrimination experiment and key recovery attack against Speck32 / 64, the various distinguishers proposed in the present invention have achieved much better discrimination accuracy and key deciphering success rate than those in comparative document 2, and also have certain advantages compared with comparative document 1. What the present invention proposes is a neural network differential distinguisher construction scheme with extremely strong generalization, which has great reference value for research in this field, while comparative documents 1 and 2 only optimize simple neural network structures.
[0069] In one embodiment, the number of convolutional layers in a single residual tower is increased from 2 to 5. The size of the convolutional kernel in a single residual tower starts from 1 and increases incrementally by 2 one by one. The initial size of the convolutional kernel between residual towers increases incrementally by 2 one by one.
[0070] Combining the multi-scale convolution idea of pyramid convolution, use the Speck ciphertext reduced to 5 rounds to test the advantages and disadvantages of different composition schemes, and determine the best convolution composition scheme. The comparison table is shown in Table 1.
[0071]
[0072]
[0073] Table 1
[0074] In one embodiment, the method further includes:
[0075] Generate a ciphertext data set and train the optimized neural network differential distinguisher (the discrimination accuracy and improvement amplitude of each distinguisher for the Speck encryption algorithm are shown in Table 2).
[0076]
[0077] Table 2
[0078] The training process specifically includes:
[0079] Simulate the Speck symmetric cipher algorithm to generate the data sets for training and validation;
[0080] Set the learning rate, Epoch, and Batch Size, and use the Adam algorithm with default parameters in Keras to optimize for the cross-entropy loss function and a small penalty of L2 weight regularization. The learning rate adopts a cyclic learning rate. Among them, the obtained network is stored at the end of each epoch, and the best network obtained is evaluated according to the test set, and this test set is not used for training;
[0081] Train the neural network to obtain a differential distinguisher model file, which contains the discrimination accuracies of each neural network differential distinguisher for 5 - 8 rounds.
[0082] In the specific implementation process, the training set is 1 million ciphertext data, with real pairs and random pairs each accounting for half, and the test set is 100,000 ciphertext data.
[0083] The network is trained for a total of 200 epochs. The batch size is set to 5000. Use the Adam algorithm with default parameters in Keras to optimize for the cross-entropy loss function and a small penalty of L2 weight regularization (regularization parameter). The learning rate adopts a cyclic learning rate, as shown in Formula 1, where. The obtained network is stored at the end of each epoch, and the best network obtained is evaluated according to the test set, and this test set is not used for training.
[0084]
[0085] α = 10 -4 、β = 2 * 10 -3 、n = 9、l i is the current learning rate. Here, the parameters have no practical significance because it is only to achieve a cyclic learning rate every 10 epochs.
[0086] Training the neural network to obtain a differential distinguisher model file contains the discrimination accuracies of each distinguisher for 5 - 8 rounds. This is an important indicator for evaluating the distinguisher. Secondly, it is the weight file (h5 file) of the distinguisher.
[0087] In one implementation, the method further includes: performing a key recovery attack based on the optimized neural network differential distinguisher to break the corresponding lightweight symmetric cipher.
[0088] Please refer to Figure 1 , which is the flowchart of the method for constructing a neural network differential distinguisher based on multi-scale convolution in the embodiments of the present invention.
[0089] The steps of the key recovery attack specifically include:
[0090] Based on neutral bit positions, Bayesian optimization, and the UCB problem, a key recovery attack algorithm is created. First, a 6- or 7-round ND2 neural network differential distinguisher is obtained and used to extend the recovery rounds to 11 rounds.
[0091] Use the existing ND2 neural network differential distinguisher to break the Speck encryption with reduced rounds to 11 rounds.
[0092] In the specific implementation process, based on neutral bit positions, Bayesian optimization, and the UCB problem, a key recovery attack algorithm is created. First, a 6- or 7-round ND2 differential distinguisher is obtained and used to extend the recovery rounds to 11 rounds. The specific steps are as follows:
[0093] (1) Set two rounds of preset differentials (0x211, 0xa04) → (0x40, 0), that is, select plaintext pairs with a difference of (0x211, 0xa04). After two rounds of encryption, a high-probability differential (0x40, 0) is obtained, and its probability is 2 -6 , extend 7 rounds to 9 rounds;
[0094] (2) According to the property that the Speck encryption can be extended by one round without consumption, the 9-round differential distinguisher can be extended to 10 rounds without consumption;
[0095] (3) In the key recovery attack, the key is used to decrypt N-round ciphertext pairs to obtain the (N - 1)-round ciphertext, and then the differential distinguisher is used for judgment. Therefore, after 10 rounds, it can be extended by one more round to reach 11 rounds.
[0096] The key recovery attack algorithm for the 11-round Speck encryption algorithm is as follows:
[0097] Set the number of recovery attacks to 100. In each recovery attack, 100 ciphertext structures of 11 rounds are generated; each ciphertext structure is obtained by flipping neutral bit positions. A single ciphertext pair can be flipped by 6 bit positions to obtain a ciphertext structure containing 64 ciphertext pairs. Generate 100 ciphertext structures;
[0098] Search the ciphertext structures. Set the maximum number of iterations to 500. In each iteration, the UCB algorithm is used to select a ciphertext structure with the highest current priority. The first one selected is the first ciphertext structure.
[0099] For the selected ciphertext structure, decrypt and score it. Use the Bayesian optimization algorithm and iterate five times. In each Bayesian optimization iteration, 32 selected candidate keys are used to decrypt all ciphertext pairs. The first key is randomly selected. Decrypt using the selected key, score the decrypted ciphertext with the distinguisher, and calculate the mean value of the key score and the distinguisher score through formula 2;
[0100]
[0101] S k represents the key score, Z i is the output signal of the i-th input sample given by the neural differentiator, k is the block size set to 1 here, i is just a counter with no meaning, and N is the total number of ciphertext pairs
[0102] Calculate the Euclidean distances scores of each candidate key with a 32-bit key difference in the range of [0, 2^16), and select the 32 keys with the smallest distances as the keys to be used in the next iteration; in this way, only 160 candidate last-round sub-keys need to be tried for each ciphertext structure, and output the used keys and the corresponding key scores;
[0103] For candidate keys with scores exceeding C 11 = 10, perform the next step, use the Bayesian optimization algorithm to recommend 160 10-round candidate keys again, and calculate the corresponding scores. If the score exceeds C 10 = 10, then consider the two selected keys as key guesses;
[0104] Calculate the bit difference between the guessed key and the true key to determine whether the key recovery is successful.
[0105] Using the existing ND2 neural network differential differentiator, break the Speck encryption from 12 rounds to 11 rounds. In 1000 key recovery attacks, the number of successful key recovery attacks using the ND2 differential differentiator reached 618 times, and the success rate was approximately 61.8%, which was 9.7% higher than the existing mainstream algorithms, achieving significant results.
[0106] Embodiment 2
[0107] Based on the same inventive concept, this embodiment discloses a construction device of a neural network differential differentiator based on multi-scale convolution, including:
[0108] An initial construction module for initially constructing a neural network differential differentiator suitable for the ciphertext structure by adjusting the residual neural network. The initially constructed neural network differential differentiator includes an input module, an initial convolution module, a residual module, and a prediction module. The input module is used to receive input data from ciphertext pairs, the initial convolution module is used to extract features of the input data of ciphertext pairs, the residual module is used to extract deep features of the input data of ciphertext pairs, and the prediction module is used to map the input features to output labels to obtain the final result;
[0109] An optimization module, which is used to adjust the number of residual towers and the number of convolutional layers in the residual towers within the neural network differential discriminator, preliminarily optimize the neural network differential discriminator, and combine the multi-scale convolution idea of pyramid convolution to adjust the composition of the convolutional kernel sizes of the neural network differential discriminator, so as to obtain an optimized neural network differential discriminator.
[0110] Since the system introduced in the second embodiment of the present invention is the system adopted for constructing the neural network differential discriminator based on multi-scale convolution in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of this system, so it will not be elaborated here. Any system adopted by the method in the first embodiment of the present invention falls within the scope of protection of the present invention.
[0111] Embodiment Three
[0112] Based on the same inventive concept, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in Embodiment One.
[0113] Since the computer-readable storage medium introduced in the third embodiment of the present invention is the computer-readable storage medium adopted for constructing the neural network differential discriminator based on multi-scale convolution in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of this computer-readable storage medium, so it will not be elaborated here. Any computer-readable storage medium adopted by the method in the first embodiment of the present invention falls within the scope of protection of the present invention.
[0114] Embodiment Four
[0115] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method described in Embodiment One.
[0116] Since the computer device introduced in the fourth embodiment of the present invention is the computer device adopted for constructing the neural network differential discriminator based on multi-scale convolution in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of this computer device, so it will not be elaborated here. Any computer device adopted by the method in the first embodiment of the present invention falls within the scope of protection of the present invention.
[0117] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0118] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0119] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and variations.
Claims
1. A method for constructing a neural network differential discriminator based on multi-scale convolution, characterized in that: include: A neural network differential distinguisher suitable for the ciphertext structure is preliminarily constructed by adjusting the residual neural network, wherein the preliminarily constructed neural network differential distinguisher includes an input module, an initial convolution module, a residual module and a prediction module, the input module is used to receive input data from the ciphertext pair, the initial convolution module is used to extract the features of the ciphertext pair input data, the residual module is used to extract the deep features of the ciphertext pair input data, and the prediction module is used to map the input features to the output labels to obtain the final result; The number of residual towers and the number of convolutional layers in the neural network differential discriminator are adjusted to perform preliminary optimization of the neural network differential discriminator. Combined with the multi-scale convolution idea of pyramid convolution, the convolution kernel size composition of the neural network differential discriminator is adjusted to obtain the optimized neural network differential discriminator.
2. The method for constructing a neural network differential classifier based on multi-scale convolution as claimed in claim 1, characterized in that: The number of residual towers and the number of convolutional layers in the neural network differential discriminator are adjusted to preliminarily optimize the neural network differential discriminator. In addition, the convolution kernel size composition of the neural network differential discriminator is adjusted in combination with the multi-scale convolution idea of pyramid convolution to obtain the optimized neural network differential discriminator, including: Referring to the main network construction scheme of Res2Net, the residual module is modified. After the initial convolution module performs feature extraction, the obtained feature tensor matrix is input into the residual module. Multiple copies of the feature tensor matrix are copied and input into the neural network as learning data. The residual connection is divided for the channel dimension of the feature tensor, and then the learning and then splicing are modified to add the results. The convolution kernel size of the convolution layer in a single residual tower and the initial convolution kernel size between the residual towers are increased one by one, and the optimized first neural network differential distinguisher ND1 is constructed. And / or refer to the main network construction scheme of DensneNet, modify the residual module, connect the features learned by each convolution layer to the subsequent convolution layer with a residual structure, and learn and fuse the features multiple times as part of the input, wherein 4 feature fusion layers are set to preserve and fuse the features between different convolution layers, increase the number of convolution layers in a single residual tower, increase the size of the convolution kernels in the residual tower and the residual tower, and construct an optimized second neural network differential distinguisher ND2; And / or refer to and apply the idea of parallel convolution, modify the residual module, adopt 5 parallel double-layer convolution groups, the large residual tower has 5 double-layer convolution layers, the small residual tower has two convolution layers, the convolution kernel size between the small residual tower and the large residual tower increases one by one, and construct the optimized third neural network difference distinguisher ND3; And / or refer to and apply the ideas of parallel convolution and multiple feature learning, modify the initial convolution module, use a single-layer convolution operation with a convolution kernel size of 1 to extract features, perform the single-layer convolution operation with a convolution kernel size of 1 three times, and after the second and third convolution operations, perform convolution operations with a larger convolution kernel once and twice respectively, finally add the results of the three convolutions to obtain the final feature matrix input into the residual module, and construct the optimized fourth neural network differential distinguisher ND4.
3. The method for constructing a neural network differential classifier based on multi-scale convolution as claimed in claim 2, characterized in that: The number of convolutional layers in a single residual tower is increased from 2 to 5. The size of the convolution kernel in a single residual tower starts from 1 and increases one by one in increments of 2. The initial size of the convolution kernel between residual towers increases one by one in increments of 2.
4. The method for constructing a neural network differential classifier based on multi-scale convolution as claimed in claim 1, characterized in that: The method further comprises: Generate a ciphertext dataset and train the optimized neural network differential discriminator.
5. The method for constructing a neural network differential classifier based on multi-scale convolution as claimed in claim 4, characterized in that: Generate a ciphertext dataset and train the optimized neural network differential discriminator, including: Simulate the Speck symmetric encryption algorithm to generate data sets for training and verification; Set the learning rate, Epoch, and Batch Size, and use the Adam algorithm with default parameters in Keras to optimize the cross entropy loss function and the small penalty of L2 weight regularization. The learning rate uses a cyclic learning rate. At the end of each epoch, the obtained network is stored and the best network is evaluated based on the test set, which is not used for training; The neural network is trained to obtain the differential discriminator model file, which contains the discrimination accuracy of each neural network differential discriminator for 5-8 rounds.
6. The method for constructing a neural network differential classifier based on multi-scale convolution as claimed in claim 1, characterized in that: The method also includes: performing a key recovery attack based on the optimized neural network differential distinguisher to decipher the corresponding lightweight symmetric cipher.
7. The method for constructing a neural network differential classifier based on multi-scale convolution as claimed in claim 6, characterized in that: Based on the optimized neural network differential distinguisher, key recovery attack is carried out to decipher the corresponding lightweight symmetric cipher, including: Based on neutral bits, Bayesian optimization and UCB problem, a key recovery attack algorithm is created. First, the ND2 neural network differential distinguisher obtained in 6 and 7 rounds is used to expand the recovery rounds to 11 rounds. The existing ND2 neural network differential distinguisher is used to decipher the Speck encryption with 11 rounds reduced.
8. A device for constructing a neural network differential classifier based on multi-scale convolution, characterized in that: include: An initial construction module, used to preliminarily construct a neural network differential distinguisher suitable for the ciphertext structure by adjusting the residual neural network, wherein the preliminarily constructed neural network differential distinguisher includes an input module, an initial convolution module, a residual module and a prediction module, the input module is used to receive input data from the ciphertext pair, the initial convolution module is used to extract features of the ciphertext pair input data, the residual module is used to extract deep features of the ciphertext pair input data, and the prediction module is used to map the input features to output labels to obtain the final result; The optimization module is used to adjust the number of residual towers and the number of convolution layers in the neural network differential discriminator, perform preliminary optimization on the neural network differential discriminator, and adjust the convolution kernel size composition of the neural network differential discriminator in combination with the multi-scale convolution idea of pyramid convolution to obtain the optimized neural network differential discriminator.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for constructing a neural network differential discriminator based on multi-scale convolution as described in any one of claims 1 to 7 is implemented.
10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the method for constructing a neural network differential discriminator based on multi-scale convolution as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Bypass analysis model construction method and device for protection strategy data
CN115189872A
Data anomaly detection method based on multi-scale residual classifier
CN115733673A
Symmetric cipher difference partition device based on residual neural network
CN118114229A
Distributed privacy-preserving computing on protected data
US20200311300A1
Operating system and method of a fully homomorphic encryption neural network model
US20250086437A1