Application layer data encryption method and system based on SM4 algorithm and abstract verification
By analyzing the information characteristics of plaintext blocks, dynamically adjusting the control parameters of the Logistic chaotic mapping algorithm, combining the SM4 algorithm to generate dynamic S-box and wheel keys, the traditional SM4 algorithm solves the problem of differential attacks and data leakage of high-repetitive plaintext blocks, and realizes more efficient application-level data encryption and integrity verification.
Patent Information
- Application Number
- CN202510846940.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Traditional SM4 algorithms have the risk of data leakage when facing differential attacks, linear attacks and side channel analysis. Especially in the ECB encryption mode, high-repetitive plaintext blocks generate the same ciphertext blocks, and lack dynamic adaptability to data correlation and sensitive information in the plaintext blocks, resulting in poor encryption effects.
By analyzing the information repetition, semantic correlation and information sensitivity of the plaintext block, the control parameters of the Logistic chaotic mapping algorithm are dynamically adjusted, and the dynamic S-box and wheel key are generated in combination with the SM4 algorithm, each plaintext block is encrypted, and the ciphertext data integrity is verified through the digital digest.
It enhances the anti-attack capability of plaintext blocks, reduces the risk of data leakage, improves the encryption effect of application-level data, and ensures data integrity and security during communication.
Smart Images

Figure CN120378087A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data encryption technology, and particularly to an application layer data encryption method and system based on the SM4 algorithm and digest verification. Background Art
[0002] With the rapid development of information technology, the importance of data encryption in ensuring information security has become increasingly prominent. As a block cipher algorithm, the SM4 algorithm has been widely used in the field of application layer data encryption. However, the traditional SM4 algorithm has certain limitations. Among them, the S-box and round key generation are based on a fixed mathematical structure, which makes the traditional SM4 algorithm vulnerable to attacks such as differential attacks, linear attacks, and side-channel analysis, and there is a risk of data leakage. Especially in the ECB encryption mode, when there is high repetition in the application layer data, the ciphertext blocks generated by the same plaintext block are highly similar, and attackers can infer the plaintext content through the statistical characteristics of the ciphertext, seriously threatening data security; moreover, the traditional SM4 algorithm lacks the dynamic adaptation ability to the data correlation and sensitive information within the plaintext block. Since the data within different plaintext blocks has different characteristics, and the traditional encryption method often uses a unified encryption strategy and strength for processing, it is difficult to adjust the encryption complexity intensity according to the data characteristics within different plaintext blocks, resulting in a poor encryption effect on the application layer data. Summary of the Invention
[0003] In order to solve the above technical problems, an application layer data encryption method and system based on the SM4 algorithm and digest verification are provided to solve the existing problems.
[0004] The solution of this application to solve the technical problems is to provide an application layer data encryption method and system based on the SM4 algorithm and digest verification, including the following steps: In the first aspect, an embodiment of this application provides an application layer data encryption method based on the SM4 algorithm and digest verification, and the method includes the following steps: Obtain the text data to be encrypted in the application layer; divide the text data into multiple plaintext blocks, perform word segmentation processing on each plaintext block, and obtain each vocabulary; Analyze the number of characters in the interval when each vocabulary in each plaintext block appears twice adjacent to each other in the text data, and the relevant situation with adjacent vocabulary when they appear twice adjacent to each other, and calculate the information repetition degree of each plaintext block in combination with the number of times each vocabulary appears in the text data; Obtain the semantic correlation degree of each plaintext block through the co-occurrence situation of each vocabulary and its adjacent vocabulary in each plaintext block in the text data, and the number of sentences included in each plaintext block; Determine the information sensitivity of each plaintext block according to the digital situation included in each plaintext block and the relevant situation between each vocabulary and the preset sensitive vocabulary; Determine the adjustment coefficient for each plaintext block based on the information repetition degree, semantic association degree, and information sensitivity; based on the adjustment coefficient, adjust the control parameters of the chaotic mapping algorithm to obtain the adjusted control parameters corresponding to each plaintext block, combine the chaotic mapping algorithm and the SM4 algorithm to encrypt each plaintext block to obtain ciphertext data, and verify the integrity of the ciphertext data during communication transmission through digital digest technology.
[0005] Preferably, the dividing the text data into multiple plaintext blocks includes: encoding the text data to obtain the byte sequence corresponding to the text data, and dividing the text data into multiple plaintext blocks according to the byte sequence.
[0006] Preferably, the calculating the information repetition degree of each plaintext block includes: Calculate the interval distance between each occurrence of each vocabulary in each plaintext block in the text data and the next occurrence; Convert the phrase formed by each vocabulary in each plaintext block and the adjacent vocabulary each time it appears in the text data into a word vector; Calculate the correlation degree of the word vectors between each occurrence and the next occurrence of each vocabulary in each plaintext block in the text data, denoted as the first correlation degree; Count the frequency of each vocabulary in each plaintext block in the text data; The information repetition degree is the sum value of the product of the frequency, the interval distance, and the first correlation degree of all vocabularies in each plaintext block.
[0007] Preferably, the obtaining the semantic association degree of each plaintext block includes: Denote the vocabulary adjacent to the position of each vocabulary in each plaintext block in the text data as the adjacent vocabulary; Select the maximum value between the frequency of each vocabulary in each plaintext block and the frequency of the adjacent vocabulary, denoted as the highest frequency; Count the number of times the phrase formed by each vocabulary and the adjacent vocabulary in each plaintext block appears together in the text data, denoted as the co-occurrence times; Calculate the cumulative sum of the ratio of the co-occurrence times to the highest frequency of all vocabularies in each plaintext block; count the number of sentences included in each plaintext block; The semantic association degree is the ratio of the cumulative sum to the number of sentences.
[0008] Preferably, the determining the information sensitivity of each plaintext block includes: Extract all digital groups in the text data; pre-construct a preset digital library with a specific format and a preset sensitive word library according to prior knowledge; Calculate the minimum value of the difference between the number of digits of each digit group and the number of digits of all specific types of digits in the specific format digit library; Calculate the maximum value of the correlation degree between the word vectors of each word in each plaintext block and the word vectors of all sensitive words in the sensitive word library, and denote it as the second correlation degree; Denote the sum of the second correlation degrees of all words in each plaintext block as the first sensitivity value; If there are digits in each plaintext block, obtain the digit group to which the digit belongs, and use the result of negative mapping of the minimum value of the corresponding digit group as the second sensitivity value of each plaintext block; if there is no digit group in each plaintext block, the second sensitivity value of each plaintext block is 0; Use the sum of the first sensitivity value and the second sensitivity value as the information sensitivity of each plaintext block.
[0009] Preferably, the adjustment coefficient is the normalized result of the product of the information redundancy, the semantic association degree, and the information sensitivity.
[0010] Preferably, the th plaintext block corresponds to the adjusted control parameter The calculation formula is: , where is the preset initial control parameter, is the preset limit factor, is the th adjustment coefficient of the plaintext block, is the preset first value, is the preset second value, where the preset first value is less than the preset second value.
[0011] Preferably, the obtaining of the ciphertext data includes: based on the adjusted control parameter corresponding to each plaintext block, encrypt each plaintext block by combining the chaotic mapping algorithm and the SM4 algorithm, output the ciphertext block corresponding to each plaintext block, and splice the ciphertext blocks of all plaintext blocks corresponding to the text data to obtain the ciphertext data.
[0012] Preferably, the verification of the integrity of the ciphertext data during the communication transmission process includes: through the digital digest technology, generate a digital digest of the ciphertext data, the sending end transmits the generated digital digest together with the ciphertext data, after the receiving end receives the ciphertext data, decrypts the ciphertext data and regenerates the digital digest, and compares the newly generated digital digest with the original digital digest to judge the integrity of the ciphertext data.
[0013] In a second aspect, an embodiment of the present application further provides an application layer data encryption system based on the SM4 algorithm and digest verification, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the application layer data encryption method based on the SM4 algorithm and digest verification described in any one of the above are implemented.
[0014] The present application has at least the following beneficial effects: By analyzing the repetition of words in each plaintext block in the text data and calculating the information repetition degree of each plaintext block, the beneficial effect is that it considers the same situation of the text information contained in each plaintext block and the text information contained in other plaintext blocks, so as to reflect the possibility that multiple plaintext blocks may generate the same ciphertext block or the similarity of the generated ciphertext blocks, in order to evaluate the risk of leakage of the plaintext block, and thus the more likely it is to strengthen the encryption complexity; secondly, by analyzing the association situation of adjacent words in each plaintext block each time they appear and calculating the semantic association degree of each plaintext block, the beneficial effect is that it considers the relevance between the words in the plaintext block and reflects the risk that the plaintext block can be inferred and cracked; further, by analyzing the digital information or sensitive words contained in each plaintext block and calculating the information sensitivity of each plaintext block, the beneficial effect is that it considers the sensitivity of the information in the plaintext block and reflects the importance of the information contained in the plaintext block, and evaluates the degree of protection required for the plaintext block; determining the adjustment coefficient of each plaintext block, the beneficial effect is that by comprehensively considering the repeatability, word relevance, and sensitivity of each plaintext block, it evaluates the adjustment of the control parameters of the introduced Logistic chaotic mapping algorithm to improve the encryption complexity of the plaintext block, avoid generating similar ciphertexts for multiple plaintext blocks with high repeatability, reduce the risk of mode leakage in the ECB mode, weaken the possibility of the attacker cracking the key using context association, and can enhance the encryption intensity for the plaintext block with sensitive information; obtaining the adjusted control parameters corresponding to each plaintext block, combining the Logistic chaotic mapping algorithm and the SM4 algorithm, and encrypting each plaintext block to obtain ciphertext data, the beneficial effect is that through the control parameters corresponding to different plaintext blocks, the S-box and round keys are dynamically generated using the Logistic chaotic mapping algorithm, breaking the static mode of the fixed S-box and key of the traditional SM4 algorithm, and then the dynamically generated S-box and round keys are used to encrypt each plaintext block using the SM4 algorithm, enhancing the anti-attack ability of the plaintext block, reducing the risk of leakage of the plaintext block, and improving the encryption effect of the application layer data; verifying the integrity of the ciphertext data during communication transmission through the digital digest technology, the beneficial effect is that it verifies the integrity of the text data in the application layer during communication transmission, evaluates the possibility of the text data being tampered with, and ensures the security and integrity of data protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The following further elaborates in detail on the application layer data encryption method based on the SM4 algorithm and digest verification of the present application in conjunction with the accompanying drawings.
[0016] Figure 1 It is a flowchart of the steps of the application layer data encryption method based on the SM4 algorithm and digest verification provided by an embodiment of the present application; Figure 2 It is a flowchart of the steps of the method for obtaining the information repetition degree of each plaintext block provided by an embodiment of the present application; Figure 3 It is a flowchart of the steps of the method for obtaining ciphertext data provided by an embodiment of the present application. Specific Embodiments
[0017] In order to make the objectives, technical solutions, and advantages of the present application clearer, the following further elaborates in detail on the application layer data encryption method and system based on the SM4 algorithm and digest verification proposed in the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0019] Please refer to Figure 1 , which shows a flowchart of the steps of the application layer data encryption method based on the SM4 algorithm and digest verification provided by an embodiment of the present application. The method includes the following steps: Step 1, obtain the text data to be encrypted in the application layer.
[0020] The application layer is located at the topmost layer of the network protocol stack and is the interface between users or applications and the network. It realizes data formatting, encryption, and transmission through standardized protocols. During computer communication, the data in the application layer is the business data processed at the highest layer in network communication and directly faces users or applications. It is the effective information content that users or applications actually need to transmit, such as web page content, file content, message text, etc. Therefore, obtain the text data to be encrypted in the application layer.
[0021] Thus, the text data to be encrypted in the application layer is obtained.
[0022] Step 2, divide the text data into multiple plaintext blocks, perform word segmentation on each plaintext block to obtain each vocabulary; analyze the number of characters in the interval when each vocabulary in each plaintext block appears twice adjacent to each other in the text data, and the relevant situation with adjacent vocabulary when they appear twice adjacent to each other. Combine the number of times each vocabulary appears in the text data to calculate the information repetition degree of each plaintext block.
[0023] As a block cipher algorithm, the security of the SM4 algorithm highly depends on the non - linear characteristics of the S - box and the strength of the key expansion algorithm. Traditional S - boxes and key generations are usually based on fixed mathematical structures, which are at risk of targeted attacks, such as differential attacks and linear attacks. Moreover, the S - box and round keys of the traditional SM4 algorithm are static or semi - static, and attackers may crack them through long - term observation or side - channel analysis. The Logistic chaotic mapping algorithm is a chaotic search method. Since the generated chaotic sequence can exhibit higher randomness and has the characteristics of initial - value sensitivity, pseudo - randomness, and ergodicity, it has received extremely high attention in the field of secure communication. Dynamically generating the S - box or round keys through chaotic mapping can significantly increase the non - linear degree of the SM4 algorithm and reduce the possibility of data being attacked and cracked.
[0024] The control parameters of the Logistic chaotic mapping algorithm affect the encryption security of the SM4 algorithm. If the control parameters are not selected reasonably, the ciphertext after the SM4 algorithm encrypts the text data may leak statistical characteristics, increasing the risk of differential or linear attacks. Therefore, reasonable control parameters need to be selected to dynamically generate the S - box or round keys, and then the text is encrypted through the SM4 algorithm. Among them, the SM4 algorithm is a block cipher algorithm. Therefore, before encrypting the text data, the text data to be encrypted needs to be block - processed. Specifically: Encode the text data to obtain the byte sequence corresponding to the text data, and divide the text data into multiple plaintext blocks according to the byte sequence; In this embodiment, the GBK encoding method is used to encode the text data. Among them, the GBK encoding is a well - known technology and will not be elaborated here. As other implementation methods, implementers can use other methods of existing technologies, such as Unicode encoding, ASCII encoding, etc. This embodiment does not make special restrictions on this.
[0025] It should be noted that when using the SM4 algorithm for encryption, the size of a plaintext block is 128 bits, that is, 16 bytes. After GBK encoding, a Chinese character and Chinese punctuation each occupy 2 bytes, and English punctuation, English letters, and numbers each occupy 1 byte. Therefore, the text data is divided into multiple plaintext blocks according to the number of bytes that each plaintext block can accommodate; if the number of bytes of the characters contained in the plaintext block is less than 16 bytes, the PKCS7 padding method is used to pad the plaintext block. Among them, the PKCS7 padding method is a well - known technology and will not be elaborated here.
[0026] Perform word - segmentation processing on the text segments contained in each plaintext block to obtain each vocabulary; In this embodiment, the N-gram model is used for word segmentation. The N-gram model is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of the prior art. For example, jieba word segmentation, etc. This embodiment does not make special restrictions on this.
[0027] Secondly, the text data to be encrypted contains multiple sentences. After the plaintext is segmented into blocks, the text fragments included in each plaintext block may exist in different sentences, and the characters within the same sentence will also exist in different plaintext blocks. By analyzing the distribution of the words included in each plaintext block in the original text data, the information redundancy is calculated. Among them, the step flow chart of the method for obtaining the information redundancy of each plaintext block provided in the embodiment of the present application is as Figure 2 shown, specifically as follows: Count the frequency of each word in each plaintext block in the text data; Calculate the interval distance between each occurrence of each word in each plaintext block in the text data and the next occurrence; Convert the word group formed by each word in each plaintext block and the adjacent word each time it appears in the text data into a word vector; In this embodiment, the Word2Vec model is used to obtain word vectors. The Word2Vec model is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of the prior art. For example, the BERT algorithm, etc. This embodiment does not make special restrictions on this.
[0028] It should be noted that for the sake of easy understanding, assume that for a piece of text data: "I love ancient Chinese poems, and ancient Chinese poems are very artistic. I love programming, and programming is very interesting", there are 2 sentences. Assume that the first plaintext block contains the text fragment "I love ancient Chinese poems,", which has 7 characters and 1 Chinese punctuation mark, a total of 16 bytes; Secondly, after word segmentation, it is: ["I", "love", "history", "ancient Chinese poems"]. For the word "love", it first appears in "I love history" in the text data and the second time in "I love programming". Then the number of characters between the two appearances of the word "love" in the text data is 14, and the punctuation marks are 2. Therefore, the interval distance is 16; The word group formed by the word "love" and the adjacent word when it first appears is "I love history", and the word group formed by the word "love" and the adjacent word when it appears the second time is "I love programming". The formed word groups are converted into word vectors through the Word2Vec model.
[0029] Calculate the correlation degree of the word vectors between each occurrence and the next occurrence of each word in each plaintext block in the text data, denoted as the first correlation degree; In this embodiment, the degree of correlation is calculated by computing the cosine similarity of the word vectors between each occurrence and the next occurrence of each word within each plaintext block in the text data. The calculation of the cosine similarity is a well-known technique and will not be elaborated here. As other embodiments, implementers may adopt other methods of the prior art, such as the Pearson correlation coefficient, etc. This embodiment places no special restrictions on this.
[0030] The sum value of the product of the frequency, the interval distance, and the first correlation degree of all words within each plaintext block is used as the information redundancy degree of each plaintext block. It should be noted that the larger the frequency, the more frequently the word appears in the text data, and the greater the possibility that the word is the key information in the text data; the larger the interval distance, the more likely the word is divided into different plaintext blocks, and then different plaintext blocks may all contain the word information. The larger the first correlation degree, the greater the consistency of the meaning contained when the word appears each time in the text data, that is, the greater the possibility of the consistency of the information contained in different plaintext blocks. The greater the obtained information redundancy degree, the higher the frequency of the repeated occurrence of the text segment information contained in the plaintext block in the text data, and the more likely it is that there are other plaintext blocks that also contain the same text information. When using the SM4 algorithm for encryption subsequently, multiple plaintext blocks are more likely to generate the same ciphertext block or the generated ciphertext blocks have a large similarity, resulting in a poor encryption effect on the text data and being more easily cracked by attackers. Therefore, it is necessary to increase the control parameter of the Logistic chaotic mapping algorithm to generate a more complex chaotic sequence to improve the subsequent encryption effect and thus enhance the anti-attack ability.
[0031] Thus, the information redundancy degree of each plaintext block is obtained.
[0032] Step 3: Obtain the semantic association degree of each plaintext block based on the co-occurrence situation of each word and its adjacent words within each plaintext block in the text data, and the number of sentences contained in each plaintext block.
[0033] Furthermore, since each plaintext block contains less information, and the number of bytes corresponding to an entire sentence in text data is generally greater than the number of bytes that a single plaintext block can contain. For example, for a piece of text data "User ID 456789 should be properly kept. The password is qwer1234. Do not disclose it. The password of User ID 123456 has been reset", there are two sentences. The first plaintext block is "User ID 456789 should", the second plaintext block is "be properly kept. The password is", the third plaintext block is "qwer1234. Do not disclose", the fourth plaintext block is "it. User ID 1234", and the fifth plaintext block is "56's password has been reset". Therefore, for the first plaintext block, it contains partial information of the sentence "User ID 456789 should be properly kept. The password is qwer1234. Do not disclose it", and the fourth plaintext block contains partial content of two sentences. Therefore, for the content in the same sentence, since the entire sentence has the characteristic of semantic coherence, there is a correlation between the lexical information at different positions in the sentence. When an attacker conducts key cracking, they can use the correlation of the information contained between the words for assisted cracking.
[0034] Based on the above analysis, by analyzing the correlation between different words within each plaintext block and calculating the semantic correlation degree, specifically: The words adjacent to the position of each word within each plaintext block in the text data are recorded as adjacent words; The maximum value between the frequency of each word within each plaintext block and the frequency of the adjacent word is recorded as the highest frequency; The number of times the phrases formed by each word within each plaintext block and the adjacent word appear together in the text data is counted and recorded as the co-occurrence frequency; It should be noted that for the convenience of understanding, for the word "user", the adjacent word of the word "user" is "ID", and the phrase "user ID" appears 2 times in the text data.
[0035] Count the number of sentences contained in each plaintext block; It should be noted that a sentence usually ends with a period, a question mark, or an exclamation mark. Therefore, the positions of the period, question mark, and exclamation mark within the plaintext block can be counted to obtain the number of sentences contained in each plaintext block.
[0036] Calculate the sum of the ratios of the co-occurrence frequency to the highest frequency of all words within each plaintext block, and take the ratio of the sum to the number of sentences as the semantic correlation degree of each plaintext block; It should be noted that the larger the ratio of the co-occurrence times to the highest times, the more times the two words co-occur, indicating a stronger correlation between the two words. The larger the number of sentences, the more sentences are included in the plaintext block, the more complex the semantic structure, the more dispersed the semantic correlation may be, and the relatively lower the correlation between words. The greater the obtained semantic correlation degree, the stronger the semantic correlation between the words included in the plaintext block, and the easier it is for the attacker to use these correlations to crack the key. Therefore, it is more necessary to increase the control parameter of the Logistic chaotic mapping algorithm to generate a more complex chaotic sequence to improve the subsequent encryption effect and enhance the anti-attack ability.
[0037] Thus, the semantic correlation degree of each plaintext block is obtained.
[0038] Step 4: Determine the information sensitivity of each plaintext block according to the digital situation included in each plaintext block and the correlation between each word and the preset sensitive words.
[0039] Secondly, if the text data to be encrypted contains digital information, compared with pure log text, the text with digital information is more sensitive; secondly, there will also be sensitive words in the text data, such as words like number, account voucher, password, etc. These sensitive information requires a higher protection level; therefore, by analyzing the sensitivity of the word information and the situation of containing digital information in each plaintext block, calculate the information sensitivity, specifically: Extract all digital groups in the text data; In this embodiment, a regular expression is used to extract digital groups in the text data. Among them, the regular expression is a well-known technology and will not be elaborated here.
[0040] According to prior knowledge, a preset specific format digital library and a preset sensitive word library are pre-constructed; It should be noted that according to prior knowledge, the structural characteristics of sensitive numbers are defined. Each specific type of digital structure is defined in the specific format digital library. For example, the fixed length of a password is 8 digits, the fixed length of a user number is 6 digits, the fixed length of an ID card number is 18 digits, the fixed length of a mobile phone number is 11 digits, the fixed length of a verification code is 6 digits, etc.; secondly, collect multiple sensitive words, such as password, number, leakage, payment, card number, amount, address, etc., so as to construct a sensitive word library.
[0041] Calculate the minimum value of the difference between the number of digits of each digital group and the number of digits of all specific types of digital structures in the specific format digital library; In this embodiment, calculate the minimum value of the absolute value of the difference between the number of digits of each digital group and the number of digits of all specific types of digital structures in the digital format structure library.
[0042] It should be noted that, for example, for a text data "User ID 456789 please keep it properly, the password is qwer1234, do not disclose it. The password of User ID 123456 has been reset", the number of digits of a digital group "qwer1234" is 8, and the fixed length of the password in the specific format digital library is 8 digits. The difference in the number of digits between the two is the smallest, and the minimum value is 0.
[0043] Calculate the maximum value of the correlation degree between the word vectors of each vocabulary in each plaintext block and the word vectors of all sensitive vocabularies in the sensitive vocabulary library, and denote it as the second correlation degree; In this embodiment, the Word2Vec model is used to obtain word vectors. Among them, the Word2Vec model is a well-known technology and will not be elaborated here. Secondly, the correlation degree is calculated by calculating the maximum value of the cosine similarity between the word vectors of each vocabulary in each plaintext block and the word vectors of all vocabularies in the sensitive vocabulary library. Among them, the calculation of the cosine similarity is a well-known technology and will not be elaborated here.
[0044] Denote the sum of the second correlation degrees of all vocabularies in each plaintext block as the first sensitive value; If there are numbers in each plaintext block, obtain the digital group to which the number belongs, and use the result of the negative mapping of the minimum value of the corresponding digital group as the second sensitive value of each plaintext block; if there is no digital group in each plaintext block, the second sensitive value of each plaintext block is 0; In this embodiment, the specific process of negative mapping is: perform negative mapping through the exponential function. Assume that the minimum value is denoted as , then use the result of as the second sensitive value, where is the exponential function with the natural constant as the base.
[0045] It should be noted that, for example, for a text data "User ID 456789 please keep it properly, the password is qwer1234, do not disclose it. The password of User ID 123456 has been reset", the fifth plaintext block is "The password of 56 has been reset". There are numbers 5 and 6 in this plaintext block, and the digital group to which these two numbers belong is 123456.
[0046] Take the sum of the first sensitive value and the second sensitive value as the information sensitivity of each plaintext block; It should be noted that the greater the second relevance, the more likely the word belongs to the sensitive words. The greater the first sensitivity value, the more sensitive the word information contained in the plaintext block. The greater the second sensitivity value, the more sensitive the digital information contained in the plaintext block. The greater the obtained information sensitivity, the more important the information contained in the plaintext block, and the greater its protection degree should be, that is, the encryption of the plaintext block by the SM4 algorithm should be more complex, and the control parameter of the Logistic chaotic mapping algorithm should be increased.
[0047] Thus, the information sensitivity of each plaintext block is obtained.
[0048] Step 5: Based on the information repetition degree, semantic association degree, and information sensitivity, determine the adjustment coefficient for each plaintext block; based on the adjustment coefficient, adjust the control parameter of the chaotic mapping algorithm to obtain the adjusted control parameter corresponding to each plaintext block, and combine the chaotic mapping algorithm and the SM4 algorithm to encrypt each plaintext block to obtain ciphertext data, and verify the integrity of the ciphertext data during communication transmission through digital digest technology.
[0049] Furthermore, based on the information repetition degree, the semantic association degree, and the information sensitivity, determine the adjustment coefficient, specifically: Take the normalized result of the product of the information repetition degree, the semantic association degree, and the information sensitivity as the adjustment coefficient for each plaintext block; In this embodiment, the sigmoid function is used for normalization processing. The sigmoid function is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of the existing technology, such as the softmax function, the tanh function, etc. This embodiment does not make special restrictions on this.
[0050] It should be noted that the greater the adjustment coefficient, the more complex the encryption effect required for encrypting the plaintext block. Therefore, the control parameter of the Logistic chaotic mapping algorithm should be increased more. On the contrary, it indicates that the importance of the information in the plaintext block is relatively low, and the control parameter of the Logistic chaotic mapping algorithm can be reduced.
[0051] Secondly, based on the adjustment coefficient, adjust the control parameter of the chaotic mapping algorithm, specifically: Among them, is the adjusted control parameter corresponding to the th plaintext block, is the preset initial control parameter, is the preset limit factor, is the adjustment coefficient of the th plaintext block, is a preset first value, is a preset second value, where the preset first value is less than the preset second value; In this embodiment, the preset initial control parameter takes the value of 3.8, and the preset limit factor takes the value of 0.2, and its function is to control the adjustment range to avoid excessive adjustment of the control parameter; the preset first value takes the value of 0.4, and the preset second value takes the value of 0.7. As other implementation manners, the implementer can set them according to the actual situation.
[0052] It should be noted that by setting different control parameters of the chaotic mapping algorithm according to the inconsistency of the importance degree of the information contained in each plaintext block, the security of the SM4 algorithm is improved, and the risk of text data leakage is reduced.
[0053] Based on the adjusted control parameter corresponding to each plaintext block, a chaotic sequence is generated by combining with the Logistic chaotic mapping algorithm. By using the generated chaotic sequence through a certain mapping relationship, an S-box is dynamically generated. The round keys are generated by using the chaotic sequence. Through the dynamically generated S-box and round keys, each plaintext block is encrypted by using the SM4 algorithm, and the ciphertext block corresponding to each plaintext block is output; the ciphertext blocks of all plaintext blocks corresponding to the text data are spliced to obtain the ciphertext data; It should be noted that the process of chaotic encryption by using the Logistic chaotic mapping algorithm and the SM4 algorithm is a well-known technology and will not be elaborated here.
[0054] Among them, the step flow chart of the method for obtaining the ciphertext data provided by the embodiment of the present application is as Figure 3 shown.
[0055] Secondly, in order to ensure the integrity and consistency of the text data, through the digital digest technology, a digital digest of the ciphertext data is generated. The sending end transmits the generated digital digest and the ciphertext data together. After receiving the ciphertext data, the receiving end decrypts the ciphertext data and regenerates the digital digest, and compares the newly generated digital digest with the original digital digest to determine whether they are consistent, so as to evaluate whether the ciphertext data is tampered with during the communication transmission process, so as to ensure the security and integrity of the text data in the application layer.
[0056] It should be noted that the digital digest technology is a well-known technology and will not be elaborated here.
[0057] Based on the same inventive concept as the above method, an embodiment of the present application further provides an application layer data encryption system based on the SM4 algorithm and digest verification, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above application layer data encryption methods based on the SM4 algorithm and digest verification are implemented.
[0058] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0059] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0060] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made. Therefore, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application all belong to the protection scope of the technical solution of the present application.
Claims
1. An application layer data encryption method based on the SM4 algorithm and digest verification, characterized in that The method includes the following steps: Obtain the text data to be encrypted in the application layer; divide the text data into multiple plaintext blocks, perform word segmentation processing on each plaintext block to obtain each vocabulary; Analyze the number of characters in the interval when each vocabulary in each plaintext block appears twice adjacent to each other in the text data, and the relevant situation with adjacent vocabulary when appearing twice adjacent to each other, and combine the number of times each vocabulary appears in the text data to calculate the information repetition degree of each plaintext block; Obtain the semantic correlation degree of each plaintext block through the co-occurrence situation of each vocabulary and its adjacent vocabulary in each plaintext block in the text data, and the number of sentences included in each plaintext block; Determine the information sensitivity of each plaintext block according to the digital situation included in each plaintext block and the relevant situation between each vocabulary and the preset sensitive vocabulary; Based on the information repetition degree, semantic correlation degree and information sensitivity, determine the adjustment coefficient of each plaintext block; based on the adjustment coefficient, adjust the control parameters of the chaotic mapping algorithm to obtain the adjusted control parameters corresponding to each plaintext block, combine the chaotic mapping algorithm and the SM4 algorithm, encrypt each plaintext block to obtain ciphertext data, and verify the integrity of the ciphertext data during the communication transmission process through digital digest technology.
2. The application layer data encryption method based on the SM4 algorithm and digest verification according to claim 1, characterized in that, The dividing the text data into multiple plaintext blocks includes: encoding the text data to obtain the byte sequence corresponding to the text data, and dividing the text data into multiple plaintext blocks according to the byte sequence.
3. The application layer data encryption method based on the SM4 algorithm and digest verification according to claim 1, wherein The calculating the information repetition degree of each plaintext block includes: Calculate the interval distance between each time each vocabulary in each plaintext block appears and the next appearance in the text data; Convert the phrase formed by each vocabulary and its adjacent vocabulary each time each vocabulary in each plaintext block appears in the text data into a word vector; Calculate the correlation degree of the word vectors between each time each vocabulary in each plaintext block appears and the next appearance, denoted as the first correlation degree; Count the frequency of each vocabulary in each plaintext block appearing in the text data; The information repetition degree is the sum value of the product of the frequency, the interval distance and the first correlation degree of all vocabularies in each plaintext block.
4. The application layer data encryption method based on the SM4 algorithm and digest verification according to claim 3, characterized in that, The obtaining the semantic correlation degree of each plaintext block includes: Denote the vocabulary adjacent to the position where each vocabulary in each plaintext block is located in the text data as the adjacent vocabulary; Select the maximum value between the frequency of each vocabulary in each plaintext block and the frequency of the adjacent vocabulary, denoted as the highest frequency; Count the number of times the phrase formed by each vocabulary and the adjacent vocabulary in each plaintext block appears together in the text data, denoted as the co-occurrence times; Calculate the cumulative sum of the ratio of the co-occurrence times of all vocabularies in each plaintext block to the highest frequency; count the number of sentences included in each plaintext block; The semantic correlation degree is the ratio of the cumulative sum to the number of sentences.
5. The application layer data encryption method based on the SM4 algorithm and digest verification according to claim 1, characterized in that, The determining the information sensitivity of each plaintext block includes: Extract all digital groups in the text data; according to prior knowledge, pre-construct a preset specific format digital library and a preset sensitive word library; Calculate the minimum value of the difference between the number of digits of each digital group and the number of digits of all specific types of numbers in the specific format digital library; Calculate the maximum value of the correlation degree between the word vectors of each vocabulary in each plaintext block and the word vectors of all sensitive vocabularies in the sensitive vocabulary library, and denote it as the second correlation degree; Denote the sum of the second correlation degrees of all vocabularies in each plaintext block as the first sensitivity value; If there are numbers in each plaintext block, obtain the number group to which the number belongs, and use the result of negative mapping of the minimum value of the corresponding number group as the second sensitivity value of each plaintext block; if there is no number group in each plaintext block, the second sensitivity value of each plaintext block is 0; Use the sum of the first sensitivity value and the second sensitivity value as the information sensitivity of each plaintext block.
6. The application layer data encryption method based on the SM4 algorithm and digest verification according to claim 1, wherein The adjustment coefficient is the normalized result of the product of the information repetition degree, the semantic association degree, and the information sensitivity.
7. The application layer data encryption method based on the SM4 algorithm and digest verification according to claim 1, characterized in that, The formula for calculating the adjusted control parameter corresponding to the th plaintext block is as follows: , where is the preset initial control parameter, is the preset limit factor, is the adjustment coefficient of the th plaintext block, is the preset first value, is the preset second value, where the preset first value is less than the preset second value.
8. The application layer data encryption method based on the SM4 algorithm and digest verification according to claim 1, characterized in that, The obtaining of the ciphertext data includes: based on the adjusted control parameter corresponding to each plaintext block, encrypt each plaintext block by combining the chaotic mapping algorithm and the SM4 algorithm, output the ciphertext block corresponding to each plaintext block, and splice the ciphertext blocks of all plaintext blocks corresponding to the text data to obtain the ciphertext data.
9. The application layer data encryption method based on the SM4 algorithm and digest verification according to claim 1, characterized in that, The verification of the integrity of the ciphertext data during communication transmission includes: generating a digital digest of the ciphertext data through the digital digest technology, the sending end transmits the generated digital digest together with the ciphertext data, after the receiving end receives the ciphertext data, decrypts the ciphertext data and regenerates the digital digest, and compares the newly generated digital digest with the original digital digest to judge the integrity of the ciphertext data.
10. An application layer data encryption system based on the SM4 algorithm and digest verification, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the application layer data encryption method based on the SM4 algorithm and digest verification as described in any one of claims 1-9.
Citation Information
Patent Citations
Data encryption method and system based on chaotic block cipher
CN115348101A
DATA ENCRYPTION METHOD WITH CHAOTIC CHANGES OF THE ROUND KEY BASED ON DYNAMIC CHAOS
EA201600099A1
Method of semantic transposition of text into an unrelated semantic domain for secure, deniable, stealth encryption
WO2023141715A1
Cited By
Financial data storage method, device, system and equipment and medium
CN120951360A