Method and device for constructing bypass analysis model for protection strategy data

By building a bypass analysis model for protection policy data, the CNN model is optimized to adapt to data sets with first-order masks and random delay protection policies, the problem of difficult to crack the encryption implementation with protection policy in the prior art is solved, and efficient key acquisition performance is achieved.

CN115189872BActive Publication Date: 2025-05-09ARMY ENG UNIV OF PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210811536.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-11
Publication Date
2025-05-09
Estimated Expiration
2042-07-11

AI Technical Summary

Technical Problem

When using deep learning technology for bypass cryptographic analysis, it is difficult to effectively crack the encryption implementation with protection strategies, and the key acquisition effect is not ideal.

Method used

A bypass analysis model for protection policy data is constructed. By building a simple CNN model specifically for SCA scenarios, and selecting a data set with a first-order mask and a random delay protection policy for optimization, the optimized model SESCAnew is generated to improve the performance of key acquisition.

Benefits of technology

Through the optimized model SESCAnew, effective encryption of encryption with protection policy is achieved, significantly improving the key acquisition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115189872B_ABST
    Figure CN115189872B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for constructing a bypass analysis model for protection strategy data, wherein the method includes: constructing a simple CNN model specifically for SCA scenarios, and selecting a specified data set of the bypass leakage public database ASCAD database as the experimental object, and the specified data set has a first-order mask and a random delay protection strategy; generating a SESCA basic model SESCAbase according to the simple CNN model; through the selected specified data set, based on the preset structural parameter optimization rules, the other structural parameters and training parameters in SESCAbase except for the specifically set structural parameters are gradually optimized to obtain the optimized model SESCAnew to achieve key acquisition performance. The basic model is applied to the ASCAD data set with a first-order mask and a random delay protection strategy, and the various hyperparameters of the model are optimized within the most targeted experimental parameter range, so as to improve the decryption performance of the SESCA model, and a new optimized architecture SESCAnew for attack with a protection strategy is obtained, thereby achieving an excellent decryption effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of bypass analysis, and in particular to a method and device for constructing a bypass analysis model for protection strategy data. Background Art

[0002] Side Channel Analysis (SCA) is a technique that bypasses the tedious analysis of encryption algorithms and uses information related to computational data (such as execution time, power consumption, electromagnetic radiation, etc.) leaked from encryption algorithms in hardware encryption implementations to crack encryption systems in combination with statistical analysis methods.

[0003] Nowadays, with the development of machine learning, deep learning technology with excellent performance in image classification and target recognition has become popular. Studies have shown that the application of convolutional neural network algorithms under deep learning in side-channel analysis can produce better key acquisition performance.

[0004] At present, there are two main CNN structures that have successfully achieved key acquisition using CNNSCA at home and abroad, namely variants based on the two network structures of Alexnet and VGGnet. Among them, Alexnet, the champion structure of the 2012 ILSVRC, was applied to the SCA scenario in the literature. Its Alex-CNNSCA network model achieved attack-unprotected encryption, but in fact, the key acquisition effect was not ideal when the experiment was with protected targets. The 2013 ILSVRC champion network ZFNet has not changed much compared with the first ILSVRC champion network Alexnet in 2012. The 2014 ILSVRC runner-up structure VGGnet also achieved key acquisition for unprotected targets in SCA applications. VGG-CNNSCA models with different parameters have been proposed in the literature. Among them, the VGG-CNNSCA proposed in the literature has the best key acquisition effect. When designing and optimizing this model, the author also experimented with the bypass leakage data set with protection (first-order mask and random delay), and its key acquisition results were not ideal. The 2014 ILSVRC champion network GoogLeNet and the 2015 ILSVRC champion network ResNet were also used in SCA, but the results were average, and this conclusion has been confirmed in the literature. The last ILSVRC champion network in 2017 was SENet proposed by Momenta and Oxford University. Currently, there are few literatures that apply this network to SCA scenarios. Although the evolving CNNSCA has overcome the shortcomings of previous modeling methods and has improved key acquisition performance when implementing unprotected encryption, there are currently few CNNSCA experiments on protected encryption. Summary of the invention

[0005] The purpose of the present invention is to provide a method and device for constructing a bypass analysis model for protection strategy data, aiming to solve the above-mentioned problems.

[0006] One or more embodiments of this specification provide a method for constructing a bypass analysis model for protection strategy data, including:

[0007] S1. Build a simple CNN model specifically for SCA scenarios, and select a specified data set of the bypass leakage public database ASCAD database as the experimental object, the specified data set with a first-order mask and a random delay protection strategy;

[0008] S2. Generate a SESCA base model SESCAbase according to the simple CNN model;

[0009] S3. By selecting the designated data set, based on the preset structural parameter optimization rules, the other structural parameters and training parameters in the SESCAbase except for the specifically set structural parameters are gradually optimized to obtain the optimized model SESCAnew to achieve key acquisition performance.

[0010] One or more embodiments of the present specification provide a bypass analysis model construction device for protection strategy data, characterized by comprising:

[0011] A model building module builds a simple CNN model specifically for SCA scenarios, and selects a specified data set of the bypass leakage public database ASCAD database as the experimental object, wherein the specified data set has a first-order mask and a random delay protection strategy;

[0012] A model generation module generates a SESCA base model SESCAbase according to the simple CNN model;

[0013] The parameter optimization module gradually optimizes the other structural parameters and training parameters in the SESCAbase except for the specifically set structural parameters based on the preset structural parameter optimization rules by selecting the specified data set, and obtains the optimized model SESCAnew to achieve key acquisition performance.

[0014] One or more embodiments of the present specification provide an electronic device, comprising: a processor; and a memory arranged to store computer-executable instructions, wherein when the computer-executable instructions are executed, the processor implements the steps of the bypass analysis model construction method for protection strategy data.

[0015] One or more embodiments of the present specification provide a storage medium for storing computer-executable instructions, which, when executed, implement the steps of the above-mentioned bypass analysis model construction method for protection strategy data.

[0016] By adopting the embodiment of the present invention, the basic model SESCAbase is applied to a data set of the Bypass Channel Leakage Common Database (ASCAD) with a first-order mask and a random delay protection strategy. In this specific application scenario, unnecessary experimental parameters are excluded, and various hyperparameters of the model are optimized within the most targeted experimental parameter range to improve the decryption performance of the SESCA model. A new optimization architecture SESCAnew for attack band protection strategy is obtained, thereby achieving excellent decryption effect.

[0017] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0019] Figure 1 A flow chart of a bypass analysis model construction method for protection strategy data provided by one or more embodiments of this specification;

[0020] Figure 2 A structural diagram of SESCAbase provided for one or more embodiments of this specification;

[0021] Figure 3 to Figure 13 Schematic diagrams of experimental results of Experiments 1 to 11 provided in one or more embodiments of this specification;

[0022] Fig.14 A schematic diagram of the module composition of a bypass analysis model building device for protection strategy data provided by one or more embodiments of this specification;

[0023] Fig.15 A schematic diagram of the structure of an electronic device provided for one or more embodiments of this specification. DETAILED DESCRIPTION

[0024] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0026] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0027] Method Embodiment

[0028] According to an embodiment of the present invention, a method for constructing a bypass analysis model for protection strategy data is provided. Figure 1 FIG. 1 is a flow chart of a method for constructing a bypass analysis model for protection strategy data according to an embodiment of the present invention. Figure 1 As shown, the bypass analysis model construction method for protection strategy data according to an embodiment of the present invention specifically includes:

[0029] S1. Build a simple CNN model specifically for SCA scenarios, and select a specified data set (which can be a third data set) of the bypass leakage public database ASCAD database as the experimental object, wherein the specified data set has a first-order mask and a random delay protection strategy.

[0030] S2. Generate a SESCA base model SESCAbase according to the simple CNN model;

[0031] S3. By selecting the designated data set, based on the preset structural parameter optimization rules, the other structural parameters and training parameters in the SESCAbase except for the specifically set structural parameters are gradually optimized to obtain the optimized model SESCAnew to achieve key acquisition performance.

[0032] By adopting the embodiment of the present invention, the basic model SESCAbase is applied to a data set of the Bypass Channel Leakage Common Database (ASCAD) with a first-order mask and a random delay protection strategy. In this specific application scenario, unnecessary experimental parameters are excluded, and various hyperparameters of the model are optimized within the most targeted experimental parameter range to improve the decryption performance of the SESCA model. A new optimization architecture SESCAnew for attack band protection strategy is obtained, thereby achieving excellent decryption effect.

[0033] The following is a detailed description of the bypass analysis model construction method for protection strategy data:

[0034] Generate the SESCA base model SESCAbase according to the simple CNN model, specifically including:

[0035] S201. Determine that the convolution block of the simple CNN model is composed of a CONV layer, a BN layer, and an ACT layer, and add a POOL layer after the convolution block to reduce the feature dimension to form a new convolution block, and repeat the new convolution block in the network model for a preset number of times until an output of a preset reasonable size is obtained;

[0036] S202. Introduce a preset number of FC layers, and use a softmax function in the last FC layer, wherein the softmax function outputs a classification prediction result;

[0037] S203. A SEnet module is embedded between the convolution layer and the pooling layer of the simple CNN model to reduce the gradient diffusion problem in error back propagation, thereby obtaining a SESCA base model SESCAbase.

[0038] Here we build a new CNN simple model specifically for SCA scenarios, and select the third data set of the ASCAD database as the experimental object. This data set is collected from the encryption implementation with two protection strategies: mask and delay. The convolution block of the simple CNN model consists of a CONV layer, a BN layer, and an ACT layer. A POOL layer is added after this block to reduce the feature dimension. The new convolution block is repeated n times in the network model until an output of a reasonable size is obtained. Then, n FC layers are introduced, and a softmax function is used in the last FC layer, which outputs the classification prediction result. In order to improve the classification and recognition performance of the CNNSCA model, the SEnet module is deliberately embedded in the simple model. Its main function is to reduce the gradient diffusion problem in the error back propagation. This is the first time it has been used in judging the security of the CNNSCA model with a protection strategy. The SEnet module will be embedded between the convolution layer and the pooling layer of the simple model convolution block, and the simple model containing the SEnet module will be renamed the SESCA base model-SESCAbase. The structure of this newly designed SESCAbase is as follows Figure 2 shown.

[0039] The structural parameter optimization rule includes at least one of the following rules:

[0040] Rule 1: The convolutional layers in the same convolutional block set the same parameters to keep the amount of data generated by different layers unchanged;

[0041] Rule 2: The dimension of each pooling window is 2, and the window sliding step is also 2, which reduces the dimension of the input data by a factor of 2;

[0042] Rule 3: In the convolutional layer of the i-th block (starting from i=1), the number of convolution kernels is n:n i =n1×2 i-1 , i ≥ 2. This rule keeps the amount of data processed by different convolution blocks as constant as possible.

[0043] It should be noted here that the structural parameter optimization rules cited in this section are also cited in some scenarios in traditional technologies. Therefore, this application selects this structural parameter optimization rule and innovates on its basis.

[0044] The other structural parameters include at least one of the number of convolutional layers in each convolutional block, the number of convolutional kernels in the convolutional layer, the size of the convolutional kernel, the pooling method of the pooling layer, the channel ratio of the Excitation layer of the SE module, the number of cycles of the SE module, and the number of channels of the fully connected layer.

[0045] By selecting the designated data set, based on the preset structural parameter optimization rules, the other structural parameters and training parameters in the SESCAbase except the structural parameters with specific settings are gradually optimized to obtain the optimized model SESCAnew, so as to realize the optimization process of the structural parameters in the key acquisition performance, specifically including:

[0046] S301. In the SESCAbase, the number of convolutional layers of each convolutional block is set to 1, and the upper limit of the number of convolutional layers of the convolutional blocks in the SESCAbase is set to 3, the baseline is the convolutional structure corresponding to the number of convolutional layers of each convolutional block is set to 1, and the convolutional layer parameter configuration is determined through corresponding experiments;

[0047] In the initial SESCAbase structure, the number of convolutional layers in each convolutional block is 1, and the convolutional structure is named Cnov1. Referring to the number of convolutional layers of different convolutional blocks in the Alexnet and VGGnet16 prototypes, it is found that the minimum number is 1 and the maximum is 3, and the small number is distributed in the front convolutional blocks, and the large number is in the back. Therefore, the upper limit of the number of convolutional layers of the SESCAbase convolutional block is set to 3, and the baseline is Cnov1. Through two sets of necessary experiments, a certain convolutional layer parameter configuration can be obtained. When training the SESCAbase model, a training batch of 200 is temporarily used (this parameter does not affect the adjustment of the structural parameters. Unless otherwise specified in the experiments in this article, this batch parameter is used for experiments). In addition, the basis for stopping training is that the model prediction success rate is close to 1.

[0048] Experiment 1: Set the model with 2 convolution layers in all 5 convolution blocks, and keep other parameters consistent with SESCAbase. This structure is named Cnov2. Then set the number of convolution layers in the first 4 convolution blocks to 2, the number of convolution layers in the last convolution block to 3, and keep other parameters consistent with SESCAbase. This structure is named Cnov3. The specific settings of the number of convolution layers in each convolution block of Cnov1 to 3 are shown in Table 1 below:

[0049] Table1 SESCAbase.Conv1-7 Configuration

[0050]

[0051] Table 1

[0052] The three constructed structures were trained and tested. The results of Experiment 1 are as follows: Figure 3 shown.

[0053] from Figure 3It is found that the conjecture entropy of the Cnov3 structure with an upper limit of 3 convolution layers cannot converge. Further analysis shows that when the number of convolution blocks with 3 convolution layers exceeds 2 or more, the amount of calculation and parameters of model training increases exponentially. The 8G GPU memory used in the experiment is directly exhausted and the code cannot be run. This parameter setting method has no practical significance. Therefore, the upper limit of the number of convolution layers for each convolution block is determined to be 2.

[0054] Experiment 2: Figure 3 As shown in Figure 1, the guessed entropy of the Cnov2 and Cnov1 structures also has no downward convergence trend. Therefore, it is necessary to further refine the convolution layer parameters of each convolution block based on the conclusion of Experiment 1. Therefore, four structures Cnov4 to 7 are set, and each structure reduces the number of convolution layers in each convolution block of Cnov2 by 1. The specific settings of the number of convolution layers of each convolution block of Cnov4 to 7 are shown in Table 1. These constructed structures are trained and tested, and then compared with the test results of the Cnov1 to 2 structures. The results of Experiment 2 are shown in Figure 1. Figure 4 shown.

[0055] from Figure 4 It can be seen that the curve of the corresponding example Cnov7 converges well. Finally, the convolutional layer parameters of the Cnov7 structure are determined as the convolutional layer parameter benchmark of SESCAnew.

[0056] S302. Testing the number of convolution kernels of the first convolution layer through corresponding experiments to determine a benchmark corresponding to the number of convolution kernels;

[0057] Increasing the number of convolution kernels means extracting features from more dimensional channels of the input data. Increasing the amount of feature data can make the convolutional network model better trained, but it will inevitably increase the amount of computation and storage of the experimental equipment, and will inevitably increase the training time of the model. Therefore, while ensuring that the model efficiency loss is not large, the model training time can be reduced by reducing the number of convolution kernels. As can be seen from Rule 3, to determine the benchmark of the number of convolution kernels, it is only necessary to test the number of convolution kernels of the first convolution layer. At the same time, referring to the CNNSCA model, the upper limit of the number of convolution kernels in its convolution layer is 512, which can achieve the key acquisition effect. Therefore, the paper also adjusts the upper limit of the number of convolution kernels of SESCAnew to 512.

[0058] Experiment 3: The four network structures tested are named filter1, filter2, filter3, and filter4. The convolution kernel values ​​of the first convolution layer are 8, 16, 32, and 64 respectively. The number of convolution kernels of the remaining four convolution blocks also increases by 2 times respectively. The upper limit of the number of convolution kernels is always 512. The other structural parameters are the parameters of the current SESCAnew. The filter1~4 structures are trained and tested. The results of Experiment 3 are shown in Figure 3. Figure 5 shown.

[0059] exist Figure 5 As shown in the figure, the curve of filter3 has good convergence. Therefore, the filter3 structure is used as the basis for selecting the convolution kernel parameters, and the convolution kernel parameters of SESCAnew are updated synchronously.

[0060] S303. When the pooling mode of the pooling layer is set to average pooling or maximum pooling, the pooling window and the pooling step size are set to 2;

[0061] The pooling method of the SESCAbase pooling layer is initially set to AveragePool. Another commonly used pooling method is MaxPooling. According to Rule 2, no matter which pooling method is used, the pooling window and pooling step size are still set to 2.

[0062] Experiment 4: We will test the impact of two pooling methods, AveragePool and MaxPooling, on the current SESCAnew structure. The results of Experiment 4 are as follows: Figure 6 shown.

[0063] exist Figure 6 Obviously, the entropy convergence of SESCAnew using average pooling is better than that of maximum pooling, so the pooling method of SESCAnew chooses average pooling.

[0064] S304. Initially set the convolution kernel size to 1*3;

[0065] In SESCAbase, the convolution kernel size of each convolution layer is initially set to 1x3, or 3 for short. When building a deep learning network, the size of the convolution kernel is often reduced by increasing the depth of the network to reduce the computational complexity of the network. In the VGG-CNNSCA structure, the convolution kernel uses a larger size of 11, and in the Alex-CNNSCA structure, a smaller size of 3 is used. Considering these factors comprehensively, Experiment 5 is designed.

[0066] Experiment 5: Here we name five structures kernel3, kernel5, kernel7, kernel9, and kernel11, corresponding to the five convolution kernel sizes of 3, 5, 7, 9, and 11. The other parameters of these structures are the same as the current SESCAnew. Experiment 5 will test the experimental effects of these five structures. The results are as follows: Figure 7 shown.

[0067] exist Figure 7 In the structure kernel5 with a convolution kernel size of 5, the guessed entropy continues to converge downward with small fluctuations, so the convolution kernel size of SESCAnew is set to 5.

[0068] S305. Initialize and add a SE module to each of the last four convolution blocks of the SESCAbase, and set the initial value of the channel ratio rate of the Excitation layer of the SE module to 1 / 16, and set the rate value of the SESCAnew to 1 / 8 through corresponding experiments;

[0069] In essence, the attention mechanism of deep learning is similar to the selective visual attention mechanism of humans. Its core function is to select the more critical information for the current task goal from a large amount of information. In the last four convolution blocks of SESCAbase, an SE fixed module is initialized and added respectively. The channel ratio rate of the Excitation layer of the SE module is initially set to 1 / 16, but whether this conventional setting is suitable in the SCA scenario needs further verification.

[0070] Experiment 6: Focusing only on the SE module of the current SESCAnew model, we test the experimental effects of SESCAnew when the channel ratio rate of the Excitation layer of the SE module is 1 / 4, 1 / 8, 1 / 16, and 1 / 32. The experimental results are shown in the figure. Figure 8 shown.

[0071] from Figure 8 It is found that the curve corresponding to the SErate8 in the figure has the best convergence effect, and the channel ratio of the Excitation layer of the corresponding SE module is 1 / 8, so the parameter of SESCAnew is set to 1 / 8.

[0072] S306. Set the SE module cycle parameter of the SESCAnew to 3;

[0073] Experiment 7: Based on Experiment 6, the experimental effect of using 1, 2, and 3 SE modules in the last four convolution blocks in the current SESCAnew structure is tested. The results of Experiment 7 are as follows Fig. 9 shown.

[0074] from Fig. 9 As shown in the figure, the guessed entropy curve of SEtime3 converges the fastest, and SEtime3 means that the last four convolution blocks of SESCAnew use the SE module three times respectively. Therefore, the SE module loop parameter of SESCAnew is set to 3.

[0075] S307 . Set the number of all-connected channels of the SESCAnew to 1024.

[0076] In the CNNSCA model, the FC layer uses 4096 channels, which is the same as the number of channels in the fully connected layer of the original VGGnet and Alexnet models. Considering that the classification task is 1000 categories, and only 256 categories are needed in the SCA scenario, this parameter can be appropriately adjusted to reduce the training complexity of the model.

[0077] Experiment 8: We will test the experimental effects of the current SESCAnew model using four different FC layer channel numbers: 4096, 3072, 2048, and 1024. The reason for not setting the channel number lower than 1024 is that if the vector dimension changes sharply from the convolution layer to the FC layer, it will increase the computational complexity. The results of Experiment 8 are shown in Figure 8. Fig.10 shown.

[0078] from Fig.10 It is found that when the number of channels in the FC layer is 1024, the guessed entropy of the SESCAnew model converges better. Therefore, 1024 is selected as the number of channels in the FC layer of the SESCAnew model.

[0079] In summary, the structural parameters of SESCAnew have been optimized, and the new structural parameters are shown in Table 2 below:

[0080] Table2 SESCAnew Configuration

[0081]

[0082] Table 2

[0083] By selecting the designated data set, based on the preset structural parameter optimization rules, the other structural parameters and training parameters in the SESCAbase except the structural parameters with specific settings are gradually optimized to obtain the optimized model SESCAnew, so as to achieve the optimization process of the training parameters in the key acquisition performance, specifically including:

[0084] S308. Setting training parameters, the training parameters include learning rate, batch learning size, and number of iterations;

[0085] S309. Optimize the training parameters through corresponding experiments.

[0086] Optimizing the training parameters through corresponding experiments specifically includes:

[0087] The learning rate is set to 0.001, the batch size is set to 64, and the number of iterations is set to 300.

[0088] Among them, all experiments were conducted using the two parameters of batch learning size 200 and learning rate 1x10-4. In addition, the number of iterations was not strictly controlled. When the prediction accuracy was close to 1 during the model training phase, the training was stopped. Although these training parameters do not affect the tuning experiments of various structural parameters, they are not the best settings. The convolutional network of deep learning is applied to the bypass experiment, and these training parameters should also be tuned according to the bypass signal data actually processed. The order of training parameter tuning is usually learning rate, batch learning size, and number of iterations. In the three sets of experiments for training parameter optimization, the current SESCAnew structure is used.

[0089] The learning rate is a hyperparameter set manually. The learning rate is used to adjust the size of the weight change, thereby controlling the training speed of the model. The learning rate is generally between 0 and 1. If the learning rate is too large, it will accelerate learning in the early stage of model training, making the model easier to approach the local or global optimal solution, but there will be large fluctuations in the later stage of training, and even the value of the loss function will hover around the minimum value, and it will always be difficult to reach the optimal solution; if the learning rate is too small, the model weight adjustment will be too slow and the training time will be too long.

[0090] Experiment 9: We will test the experimental effects of five commonly used learning rates on the current SESCAnew model, namely lrate1 = 1x10-2, lrate2 = 1x10-3, lrate3 = 1x10-4, lrate4 = 1x10-5, lrate5 = 1x10-6. The experimental results are as follows Fig.11 shown.

[0091] Fig.11 It can be seen that the curve corresponding to the legend lrate2 has a better convergence effect, and its corresponding learning rate is 1x10-3, so 1x10-3 is selected as the learning rate of the SESCAnew structure.

[0092] The appropriate batch size is important for model optimization. This parameter does not need to be fine-tuned, just take a rough number, usually 2 to the power of n (GPU can perform better with batches of power 2). Too large a batch size will be limited by the GPU memory, affecting the computing speed, and cannot be increased indefinitely; nor can it be too small, which may cause the network model to fail to fit.

[0093] Experiment 10: Based on the size of the ASCAD dataset (50,000 training data), we selected batch size values ​​of 32, 64, 128, and 256 for the experiment. The results of Experiment 10 are shown in Figure 10. Fig.12 shown.

[0094] from Fig.12It can be seen that when the batch size is 64, the guess entropy of SESCAnew converges fastest and is most stable. Therefore, 64 is selected as the batch size of SESCAnew.

[0095] Usually, the number of iterations (epochs) is related to the prediction performance of the CNN model. If the model is overfitted (the accuracy reaches 1), there is no need to continue training; on the contrary, if all epochs have completed the calculation, and the model's loss value is still decreasing, the prediction accuracy is very low, then this epoch is too small and should be increased. In the SCA scenario, the number of iterations of CNN model training mainly refers to the actual key acquisition effect of the model, which is measured by the guess entropy.

[0096] Experiment 11: The current SESCAnew model is trained under 7 iteration parameters, namely 100, 150, 200, 250, 300, 350, and 400. Then the 7 trained models are used for key acquisition experiments. The results are as follows: Fig.13 shown.

[0097] from Fig.13 It shows that the guessed entropy curve corresponding to epoch 300 has the best convergence effect. Therefore, 300 is selected as the training iteration number of SESCAnew.

[0098] According to the 11 groups of experiments mentioned above, the optimal settings of the SESCAnew structural parameters and training parameters are demonstrated. The SESCAnew model contains 5 convolutional blocks, 6 convolutional layers (the first four convolutional blocks each have 1 convolutional layer, and the fifth convolutional block contains 2 convolutional layers), and 3 fully connected layers. The convolution kernel size of each convolutional layer is 5, the activation function is ReLU, and the padding is Same. Each convolutional block is equipped with a pooling layer, and the pooling layer selects the average pooling mode, and the pooling window is (2,2). The number of output channels of the convolutional layers in convolutional blocks 1-5 starts from 32 and increases by 2 in turn. In convolutional blocks 2-4, 3 SE modules are added after the convolutional layer of each convolutional block, and the channel ratio of the Excitation layer of the SE module is set to 1 / 8. The number of output channels in the first two fully connected layers is set to 1024, and the activation function is ReLU. The number of output channels of the third fully connected layer is the number of target categories, 256, and the classification function is Softmax. The global configuration loss function is crossentropy, the optimization method is RMSprop, the training learning rate is 1x10^-3, the batch learning size is 64, and the number of iterations is 300. All the parameters of the newly obtained SESCAnew are shown in Table 3:

[0099] Table3 SESCAnew Configuration

[0100]

[0101] Table 3

[0102] Device Example 1

[0103] The embodiment of the present invention provides a bypass analysis model construction device for protection strategy data, such as Fig.14 As shown, including:

[0104] The model building module 1410 builds a simple CNN model specifically for the SCA scenario, and selects a specified data set of the bypass leakage public database ASCAD database as an experimental object, wherein the specified data set has a first-order mask and a random delay protection strategy;

[0105] A model generation module 1420 generates a SESCA base model SESCAbase according to the CNN simple model;

[0106] The parameter optimization module 1430 gradually optimizes the other structural parameters and training parameters in the SESCAbase except for the specifically set structural parameters based on the preset structural parameter optimization rules by selecting the specified data set, and obtains the optimized model SESCAnew to achieve key acquisition performance.

[0107] The embodiment of the present invention is a device embodiment corresponding to the above method embodiment. The specific operations of each module can be understood by referring to the description of the method embodiment, which will not be repeated here.

[0108] Device Example 2

[0109] The embodiment of the present invention provides a bypass analysis model construction device for protection strategy data, such as Fig.15 As shown, it includes: a memory 1510, a processor 1520, and a computer program stored in the memory 1510 and executable on the processor 1520, and when the computer program is executed by the processor 1520, the following method steps are implemented:

[0110] S1. Build a simple CNN model specifically for SCA scenarios, and select a specified data set of the bypass leakage public database ASCAD database as the experimental object, the specified data set with a first-order mask and a random delay protection strategy;

[0111] S2. Generate a SESCA base model SESCAbase according to the simple CNN model;

[0112] S3. By selecting the designated data set, based on the preset structural parameter optimization rules, the other structural parameters and training parameters in the SESCAbase except for the specifically set structural parameters are gradually optimized to obtain the optimized model SESCAnew to achieve key acquisition performance.

[0113] Device Example 3

[0114] An embodiment of the present invention provides a computer-readable storage medium, on which a program for implementing information transmission is stored. When the program is executed by the processor 1520, the following method steps are implemented:

[0115] S1. Build a simple CNN model specifically for SCA scenarios, and select a specified data set of the bypass leakage public database ASCAD database as the experimental object, the specified data set with a first-order mask and a random delay protection strategy;

[0116] S2. Generate a SESCA base model SESCAbase according to the simple CNN model;

[0117] S3. By selecting the designated data set, based on the preset structural parameter optimization rules, the other structural parameters and training parameters in the SESCAbase except for the specifically set structural parameters are gradually optimized to obtain the optimized model SESCAnew to achieve key acquisition performance.

[0118] The computer-readable storage medium in this embodiment includes, but is not limited to, ROM, RAM, magnetic disk or optical disk, etc.

[0119] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a bypass analysis model for protection strategy data, characterized in that: include: S1. Build a simple CNN model specifically for SCA scenarios, and select a specified data set of the bypass leakage public database ASCAD database as the experimental object, the specified data set with a first-order mask and a random delay protection strategy; S2. Generate a SESCA base model SESCAbase according to the simple CNN model; S2 specifically includes: S201. Determine that the convolution block of the simple CNN model is composed of a CONV layer, a BN layer, and an ACT layer, and add a POOL layer after the convolution block to reduce the feature dimension to form a new convolution block, and repeat the new convolution block in the network model for a preset number of times until an output of a preset reasonable size is obtained; S202. Introduce a preset number of FC layers, and use a softmax function in the last FC layer, wherein the softmax function outputs a classification prediction result; S203. A SEnet module is embedded between the convolution layer and the pooling layer of the simple CNN model to reduce the gradient diffusion problem in error back propagation, thereby obtaining a SESCA base model SESCAbase; S3. By selecting the designated data set, based on the preset structural parameter optimization rules, the other structural parameters and training parameters in the SESCAbase except for the specifically set structural parameters are gradually optimized to obtain the optimized model SESCAnew to achieve key acquisition performance.

2. The method according to claim 1, characterized in that The structural parameter optimization rule includes at least one of the following rules: Rule 1: The convolutional layers in the same convolutional block set the same parameters to keep the amount of data generated by different layers unchanged; Rule 2: The dimension of each pooling window is 2, and the window sliding step is also 2, which reduces the dimension of the input data by a factor of 2; Rule 3: In the convolutional layer of the i-th block, the number of convolution kernels is n: n i =n 1 ×2i-1 , i ≥ 2, this rule makes the amount of data processed by different convolution blocks remain as constant as possible; Rule 4: The convolution kernel size of all convolutional layers is the same.

3. The method according to claim 2, characterized in that The other structural parameters include at least one of the number of convolutional layers in each convolutional block, the number of convolutional kernels in the convolutional layer, the size of the convolutional kernel, the pooling method of the pooling layer, the channel ratio of the Excitation layer of the SE module, the number of cycles of the SE module, and the number of channels of the fully connected layer.

4. The method according to claim 3, characterized in that The S3 specifically includes: S301. In the SESCAbase, the number of convolutional layers of each convolutional block is set to 1, and the upper limit of the number of convolutional layers of the convolutional blocks in the SESCAbase is set to 3, the baseline is the convolutional structure corresponding to the number of convolutional layers of each convolutional block is set to 1, and the convolutional layer parameter configuration is determined through corresponding experiments; S302. Testing the number of convolution kernels of the first convolution layer through corresponding experiments to determine a benchmark corresponding to the number of convolution kernels; S303. Set the pooling mode of the pooling layer to average pooling, and set the pooling window and pooling step size to 2; S304. Set the convolution kernel size to 5; S305. Initialize and add a SE module to each of the last four convolution blocks of the SESCAbase, and set the initial value of the channel ratio rate of the Excitation layer of the SE module to 1 / 16, and set the rate value of the SESCAnew to 1 / 8 through corresponding experiments; S306. Set the SE module cycle parameter of the SESCAnew to 3; S307 . Set the number of all-connected channels of the SESCAnew to 1024.

5. The method according to claim 1, characterized in that The S3 specifically includes: S308. Setting training parameters, the training parameters include learning rate, batch learning size, and number of iterations; S309. Optimize the training parameters through corresponding experiments.

6. The method according to claim 5, characterized in that The S308 specifically includes: The learning rate is set to 0.001, the batch size is set to 64, and the number of iterations is set to 300.

7. A bypass analysis model construction device for protection strategy data, characterized in that: include: A model building module builds a simple CNN model specifically for SCA scenarios, and selects a specified data set of the bypass leakage public database ASCAD database as the experimental object, wherein the specified data set has a first-order mask and a random delay protection strategy; A model generation module generates a SESCA base model SESCAbase according to the simple CNN model; The model generation module is specifically used for: Determine that the convolution block of the simple CNN model consists of a CONV layer, a BN layer, and an ACT layer, and add a POOL layer after the convolution block to reduce the feature dimension to form a new convolution block, and repeat the new convolution block for a preset number of times in the network model until an output of a preset reasonable size is obtained; A preset number of FC layers are introduced, and a softmax function is used in the last FC layer, wherein the softmax function outputs a classification prediction result; The SEnet module is embedded between the convolution layer and the pooling layer of the simple CNN model to reduce the gradient diffusion problem in the error back propagation, thereby obtaining the SESCA base model SESCAbase; The parameter optimization module gradually optimizes the other structural parameters and training parameters in the SESCAbase except for the specifically set structural parameters based on the preset structural parameter optimization rules by selecting the specified data set, and obtains the optimized model SESCAnew to achieve key acquisition performance.

8. An electronic device, characterized in that: include: processor; as well as, A memory arranged to store computer executable instructions, wherein when the computer executable instructions are executed, the processor is caused to implement the steps of the bypass analysis model building method for protection strategy data according to any one of claims 1 to 6.

9. A storage medium, characterized in that: Used to store computer executable instructions, which, when executed by a processor, implement the steps of the bypass analysis model construction method for protection strategy data as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Side channel analysis method based on deep learning

    CN111565189A

  • Modular design method for deep learning bypass analysis

    CN112231773A