Hash abstract generation method, system and device and storage medium

Through information redistribution uniformization processing and field-type sequence generation hash digest solves the security and performance problems of existing hash functions and realizes efficient and secure hash digest generation.

CN120337267APending Publication Date: 2025-07-18DONGGUAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510565769.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing hash functions such as MD5 and SHA-1 have collision attack problems. SHA-256 has slow computing speed and increased storage overhead. SHA-3 has high implementation complexity and is difficult to perform well in hardware.

Method used

Through information redistribution and uniformization processing, character strings are generated, base numbers and true numbers are determined, and the output summary of the hash function is generated using field-type sequences to reduce the collision rate and improve safety.

Benefits of technology

Improves the security of hash processing, reduces the collision rate, and simplifies the implementation process, making it easier to promote and apply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337267A_ABST
    Figure CN120337267A_ABST
Patent Text Reader

Abstract

The invention discloses a Hash abstract generation method, system and device and a storage medium. The method comprises the following steps: acquiring input information, and encoding the input information to generate a character string; performing information redistribution homogenization processing on the character string to generate a first sequence; determining a base number and a true number according to the first sequence, and determining a field pattern sequence based on the base number and the true number; and extracting a sequence with a first preset length from the field pattern sequence, and generating an output abstract of the hash function. According to the embodiment of the invention, through information redistribution uniformization, the change of each character affects abstract output, and the security of hash processing is improved; on the other hand, an output abstract is generated through a field pattern sequence, and the collision rate is reduced; the scheme is simple and easy to implement and convenient to popularize. The method can be widely applied to the technical field of computers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method, system, device and storage medium for generating a hash digest. Background Art

[0002] A hash function is a function that maps data of any size to a data value of a fixed size. It has a wide range of applications in computer science and cryptography. Hash functions in related technologies, such as MD5, have problems with collision attacks and need to improve security; SHA-3 has problems with complex implementation. Summary of the Invention

[0003] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.

[0004] To this end, an object of the present invention is to provide an efficient method, system, device and storage medium for generating a hash digest.

[0005] In order to achieve the above technical object, on the one hand, an embodiment of the present invention provides a method for generating a hash digest, including the following steps: obtaining input information, and performing encoding processing on the input information to generate a string; performing information redistribution and homogenization processing on the string to generate a first sequence; determining a base number and a mantissa according to the first sequence, and determining a field type sequence based on the base number and the mantissa; extracting a sequence with a first preset length from the field type sequence to generate an output digest of the hash function. In the embodiment of the present application, through information redistribution and homogenization, the change of each character affects the output digest, improving the security of hash processing; on the other hand, generating the output digest through the field type sequence reduces the collision rate; this solution is simple and easy to implement and is convenient for popularization.

[0006] In some embodiments, the performing information redistribution and homogenization processing on the string to generate a first sequence includes:

[0007] Adding the forward arrangement sequence and the reverse arrangement sequence of the string to obtain a first-round sequence;

[0008] Dividing the first-round sequence into two segments, and adding the forward arrangement sequence and the reverse arrangement sequence of the corresponding string of each segment to obtain a second-round sequence; the second-round sequence includes two segments of second sequences;

[0009] Sequentially performing splitting and forward and reverse addition processing on the strings of each segment of the second sequence, updating the quantity and character length of the second sequence until the character length of each segment of the obtained second sequence is less than a second preset length, generating a uniform sequence;

[0010] Perform an expansion process on the uniform sequence to obtain a first sequence; the length of the first sequence is an integer multiple of the first preset length.

[0011] In some embodiments, the adding the forward arrangement sequence and the reverse arrangement sequence of the string to obtain a first-round sequence includes:

[0012] If the length of the sequence after addition is greater than the length of the string, discard the highest-order character to obtain the first-round sequence.

[0013] In some embodiments, the determining a base number and a true number according to the first sequence and determining a field type sequence based on the base number and the true number includes:

[0014] Split the first sequence according to the first preset length, and add the several split sequences to obtain a third sequence; wherein, the length of the third sequence is the first preset length;

[0015] Split the third sequence into a fourth sequence and a fifth sequence, determine the base number according to the fourth sequence, and determine the true number according to the fifth sequence;

[0016] Determine a field type sequence based on the base number and the true number.

[0017] In some embodiments, the determining a base number and a true number according to the first sequence and determining a field type sequence based on the base number and the true number includes:

[0018] Split the first sequence according to the first preset length, and add the several split sequences to obtain a third sequence;

[0019] Split the third sequence into several segments, split each segment, and determine the base number and the true number of the corresponding segment according to the sequence after splitting each segment to obtain several pairs of base numbers and true numbers;

[0020] Determine several field type sequences based on the several pairs of base numbers and true numbers.

[0021] In some embodiments, the method further includes:

[0022] If the splitting of the second sequence is an even split and the second preset length is 2, determine the number of rounds of splitting and forward and reverse addition processing according to the length of the string.

[0023] On the other hand, an embodiment of the present invention provides a method for encrypting and transmitting data, including:

[0024] Obtain target data and a digital signature; the digital signature is obtained by performing a hash operation and private key encryption on the target data;

[0025] Perform a hashing operation on the target data through the hashing digest generation method as described above to obtain a target output digest;

[0026] Decrypt the target output digest with the public key to obtain a target value;

[0027] If the target value is the same as the digital signature, receive the target data.

[0028] On the other hand, an embodiment of the present invention provides a hashing digest generation system, including:

[0029] A first module, configured to obtain input information and perform encoding processing on the input information to generate a string;

[0030] A second module, configured to perform information redistribution and homogenization processing on the string to generate a first sequence;

[0031] A third module, configured to determine a base number and a true number according to the first sequence, and determine a field type sequence based on the base number and the true number;

[0032] A fourth module, configured to extract a sequence with a first preset length from the field type sequence to generate an output digest of the hash function.

[0033] On the other hand, an embodiment of the present invention provides a hashing digest generation device, including:

[0034] At least one processor;

[0035] At least one memory, configured to store at least one program;

[0036] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned hashing digest generation method.

[0037] On the other hand, an embodiment of the present invention provides a storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the above-mentioned hashing digest generation method when executed by the processor.

[0038] The embodiments of the present application at least include the following beneficial effects: The method provided by the embodiments of the present invention includes: obtaining input information, encoding the input information to generate a string; performing information redistribution and homogenization processing on the string to generate a first sequence; determining a base number and a true number according to the first sequence, and determining a field type sequence based on the base number and the true number; extracting a sequence with a first preset length from the field type sequence to generate an output digest of a hash function. Through information redistribution and homogenization, each character change affects the output digest in the embodiments of the present application, improving the security of hash processing; on the other hand, generating the output digest through the field type sequence reduces the collision rate; this solution is simple and easy to implement and is convenient for popularization. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the relevant technical solution drawings in the embodiments of the present invention or the prior art. It should be understood that the drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 FIG. is a schematic flowchart of an embodiment of the hash digest generation method provided by the present invention;

[0041] Figure 2 FIG. is a schematic flowchart of an embodiment of data encryption in the data transmission method provided by the present invention;

[0042] Figure 3 FIG. is a schematic flowchart of an embodiment of data decryption in the data transmission method provided by the present invention;

[0043] Figure 4 FIG. is a schematic structural diagram of an embodiment of the hash digest generation system provided by the present invention;

[0044] Figure 5 FIG. is a schematic structural diagram of an embodiment of the hash digest generation device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adjusted adaptively according to the understanding of those skilled in the art.

[0046] Exponential field type: It refers to the specific integer value after actually expanding the integer exponent x e in accordance with multiplication, where x and e represent the base and mantissa of the exponential integer. For example, the exponent 5 4 = 5 × 5 × 5 × 5 = 625, so 625 is the field type of the exponent 5 4 and 7 108 =(18646,11341,71613,14493,26161,68039,56698,97144,64158,58248,58341,59291,20740,30215,43182,06422,51722,18248,01), the exponential integer 7 108 has 92 digits in length after expansion.

[0047] This application can regard the above exponential field type as a sequence of length N, and use s = (d N d N-1 …d1) to represent this sequence of length N, where d1 is the first element starting from the right, and so on, where {d1, d2, …, d N-1 , d N} ∈ {0, 1, 2, …, 8, 9}.

[0048] A hash function is a function that maps data of any size to a data value of a fixed size (usually called a "hash value" or "digest"). It has a wide range of applications in computer science and cryptography. The following are the functions and characteristics of the hash function:

[0049] a. Data integrity verification: The hash function can be used to verify whether the data has been tampered with during transmission or storage. By comparing the hash values of the original data and the received data, the integrity of the data can be judged.

[0050] b. Fast data search: In data structures (such as hash tables), the hash function can quickly map the data to a specific location, thus accelerating search, insertion, and deletion operations.

[0051] c. Digital signature and authentication: In digital signatures, the hash function is used to generate a message digest to ensure that the message has not been modified during transmission and to provide authentication.

[0052] d. Password storage: When storing user passwords, the plaintext password is usually not stored directly, but its hash value is stored to enhance security.

[0053] e. Data deduplication: Hash functions can be used to detect duplicate data. For example, in a file storage system, the hash values of files are compared to determine whether their contents are the same.

[0054] Hash function characteristics: a. Deterministic: For the same input, a hash function always generates the same output. b. Fast computation: A hash function should be able to compute the output quickly to support efficient data processing and querying. c. Collision resistance: An ideal hash function should make it difficult to find two different input data that generate the same hash value (i.e., a collision). Although collisions may occur theoretically, a good hash function can make the probability extremely low. d. One-wayness: A hash function should be one-way, that is, the original input cannot be deduced from the hash value. This characteristic is particularly important in cryptography. e. Uniformity: A hash function should be able to distribute the input data evenly into the output space to reduce the possibility of collisions. f. Sensitivity to minor changes: Minor changes in the input data (such as changing a single character) will result in a large change in the output hash value. This characteristic is called the "avalanche effect".

[0055] Common hash functions:

[0056] a. MD5: Although it is fast, it is no longer recommended for security-related applications due to security vulnerabilities.

[0057] b. SHA-1: It is more secure than MD5, but it also has some known weaknesses and is gradually being phased out.

[0058] c. SHA-256: Belonging to the SHA-2 family, it is widely used in security applications and is currently considered secure.

[0059] d. SHA-3: The latest hash standard, providing higher security and flexibility.

[0060] In summary, hash functions play an important role in data processing, storage, and security. Their unique characteristics enable them to be widely used in multiple fields.

[0061] The following is a specific analysis of the four hash functions MD5, SHA-1, SHA-256, and SHA-3, including their operation processes, characteristics, and disadvantages.

[0062] a. MD5: Operational Process: Input Chunking: The input data is divided into chunks of 512 bits (64 bytes). Padding: Padding bits are added before the last chunk to make its length 448 bits (i.e., 512 - 64), and then 64 bits representing the original data length are appended at the end. Initializing Variables: Four 32-bit initial values (A, B, C, D) are used for calculation. Processing Each Chunk: Each 512-bit chunk is processed through four loops (each loop using different operations and constants), updating the values of A, B, C, D. Output: The final A, B, C, D are combined into a 128-bit hash value. Disadvantages: Collision Vulnerability: There are known effective collision attacks, and attackers can construct two different inputs to produce the same MD5 hash value. Insufficient Security: Not suitable for security-sensitive applications such as digital signatures or certificates.

[0063] b. SHA-1: Operational Process: Input Chunking: The input data is divided into chunks of 512 bits. Padding: Similar to MD5, the data is padded to 448 bits and 64 bits of the original length are appended. Initializing Variables: Five 32-bit initial values (A, B, C, D, E) are used. Processing Each Chunk: Each chunk is processed through 80 loop operations (divided into four stages), gradually updating the values of A - E. Output: The final A, B, C, D, E are combined into a 160-bit hash value. Disadvantages: Collision Attack: SHA-1 also has known collision attacks, and practical collision attacks have been made public. Security Issues: Although SHA-1 is more secure than MD5, it is no longer recommended for new security applications.

[0064] c. SHA-256: Operational Process: Input Chunking: The input data is divided into chunks of 512 bits. Padding: The same as the previous two, padded to 448 bits and 64 bits representing the length are appended. Initializing Variables: Eight 32-bit initial values are used. Processing Each Chunk: Each chunk is processed through 64 loop operations (based on logical functions and constants), gradually updating the eight variables. Output: The final eight variables are combined into a 256-bit hash value. Disadvantages: Speed: Compared with MD5 and SHA-1, the calculation speed is slower, especially in resource-constrained environments. Larger Output: Although it provides higher security, the 256-bit hash value may cause an increase in storage overhead in some applications.

[0065] d. SHA-3: Operational Process: Input Chunking: The data is divided into chunks of arbitrary size and the "sponge" structure is used. Padding: Specific padding rules are used. Initializing the State: An array of states (usually 1600 bits) is initialized. Absorbing Process: The input data is absorbed chunk by chunk to update the state. Compressing Process: The state is compressed through the internal Keccak function to generate the output. Output: Hash values of arbitrary size can be generated, usually 256 bits, 512 bits, etc. Disadvantages: Complexity: Compared with the SHA-2 series, the structure and implementation of SHA-3 are more complex. Performance Issues: On some hardware, the performance of SHA-3 may be inferior to that of the SHA-2 series.

[0066] The summary is shown in Table 1 below.

[0067] Hash function Output length Security Main drawbacks MD5 128 bits Low Collision attack, insufficient security SHA-1 160 bits Medium Collision attack, no longer recommended SHA-256 256 bits High Slow speed, increased storage overhead SHA-3 Variable High Complex implementation, poor performance on some hardware

[0068] Table 1

[0069] The hash digest generation method and system according to the embodiments of the present invention will be described in detail below with reference to the accompanying drawings. First, the hash digest generation method according to the embodiments of the present invention will be described with reference to the accompanying drawings.

[0070] Refer to Figure 1 , a hash digest generation method is provided in the embodiments of the present invention. The hash digest generation method in the embodiments of the present invention can be applied to a terminal, a server, or software running on a terminal or a server, etc. The terminal can be a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The hash digest generation method in the embodiments of the present invention mainly includes the following steps:

[0071] S100: Obtain the input information and perform encoding processing on the input information to generate a string;

[0072] S200: Perform information redistribution and homogenization processing on the string to generate a first sequence;

[0073] S300: Determine the base number and the true number according to the first sequence, and determine the field type sequence based on the base number and the true number;

[0074] S400: Extract a sequence of a first preset length from the field type sequence to generate the output digest of the hash function.

[0075] In some possible embodiments, the present application can generate a decimal string from the input information through encoding processing. The information redistribution homogenization process is used to scramble the string so that each character in the scrambled string has the same influence on the output digest of the hash function. The first preset length in the present application is the fixed length of the string generated by the hash function, which can be adjusted according to actual needs, and the present application does not limit the specific value of the first preset length. In some embodiments, the length of the string in the present application is greater than or equal to the first preset length. If the length of the string obtained after preprocessing the input information is less than the first preset length, some characters of the preprocessed string are filled in itself to generate a string with a length of the first preset length.

[0076] Expanding the integer exponent to form a field type sequence has unpredictability and uniformity of character distribution. The present invention proposes an invention technique based on the integer exponent field type as a hash function. In the process, first, the input information is encoded into a string composed of decimal letters. In the second step, this string is subjected to a series of uniformly distributed scrambling processes, so that each character can have an equivalent influence on the subsequent generated output. In the third step, the result of the second step is used to generate the base and true number of the integer exponent, and this integer exponent is expanded to form a field type sequence. In the fourth step, 80 fixed-length strings are taken from this field type sequence as the output digest of the hash function. The invention technique of the present invention can meet many characteristics required by an ideal hash function, such as certainty, fast calculation, collision resistance, one-wayness, uniformity, and sensitivity to small changes.

[0077] In some embodiments, the information redistribution homogenization process for the string to generate a first sequence includes:

[0078] Adding the forward arrangement sequence and the reverse arrangement sequence of the string to obtain a first-round sequence;

[0079] Dividing the first-round sequence into two segments, and adding the forward arrangement sequence and the reverse arrangement sequence of the corresponding string in each segment to obtain a second-round sequence; the second-round sequence includes two segments of second sequences;

[0080] Successively performing splitting and forward-backward addition processing on the strings in each segment of the second sequence, updating the quantity and character length of the second sequence until the character length in each segment of the obtained second sequence is less than the second preset length, to generate a uniform sequence;

[0081] Performing an expansion process on the uniform sequence to obtain a first sequence; the length of the first sequence is an integer multiple of the first preset length.

[0082] Dividing the first-round sequence into two segments can be an even split. If the character length of the first-round sequence is odd, it is determined that the length of the first segment after splitting is 1 greater than the length of the second segment. In some embodiments, the first preset length is 2, and the character length of each obtained second sequence is 1. The length of the uniform sequence is the same as the length of the string. If the length of the uniform sequence is greater than the first preset length, for the length of the subsequent field pattern sequence, it is necessary to set the length of the uniform sequence to an integer multiple of the first preset length here. Specifically, the length of the first sequence after filling can be made an integer multiple of the first preset length by filling some of its own characters into itself. Specifically, the uniform sequence is processed for expansion to obtain the first sequence, including:

[0083] If the length of the uniform sequence is greater than the first preset length, some characters in the uniform sequence are filled into the uniform sequence to obtain the first sequence;

[0084] If the length of the uniform sequence is equal to the first preset length, the uniform sequence is used as the first sequence.

[0085] The first quantity of some characters to be filled is determined through the following steps:

[0086] Obtain that the length of the string is the first value, and determine that the specific value of the integer multiple is the second value;

[0087] Determine that the first quantity is the product of the first preset length and the second value minus the first value.

[0088] In some embodiments, the step of adding the forward arrangement sequence and the reverse arrangement sequence of the string to obtain the first-round sequence includes:

[0089] If the length of the sequence after addition is greater than the length of the string, discard the highest-order character to obtain the first-round sequence.

[0090] It can be understood that for the positive and negative addition processing in subsequent rounds, the method of discarding some characters is also adopted to make the character length after addition the same as the character length of each character before addition, without changing the length of the uniform sequence.

[0091] In some embodiments, the step of determining the base number and the true number according to the first sequence and determining the field pattern sequence based on the base number and the true number includes:

[0092] The first sequence is split according to the first preset length, and several sequences after splitting are added to obtain a third sequence; wherein, the length of the third sequence is the first preset length;

[0093] The third sequence is split into a fourth sequence and a fifth sequence, and the base number is determined according to the fourth sequence, and the true number is determined according to the fifth sequence;

[0094] Determine a field pattern sequence based on the base number and the mantissa.

[0095] Exemplarily, the fourth sequence can be directly used as the base number, and the sum of the characters of the fifth sequence is divided by a preset value to obtain the mantissa. The preset value can be set according to requirements and can be 2, 4, etc., and the present application does not make specific limitations. Similarly, when adding several split sequences, if a carry occurs, some characters are discarded to make the length of the obtained third sequence the same as the first preset length.

[0096] In some embodiments, the determining a base number and a mantissa according to the first sequence and determining a field pattern sequence based on the base number and the mantissa includes:

[0097] Split the first sequence according to the first preset length, and add the several split sequences to obtain a third sequence;

[0098] Split the third sequence into several segments, and split each segment. Determine the base number and mantissa corresponding to each segment according to the sequences after splitting each segment, and obtain several pairs of base numbers and mantissas;

[0099] Determine several field pattern sequences based on the several pairs of base numbers and mantissas.

[0100] It can be understood that the present application can directly obtain a field pattern sequence according to the first sequence, or split the first sequence to obtain multiple field pattern sequences. Those skilled in the art can select a specific implementation manner from the perspective of collision probability and operation complexity. Generally, the method of generating multiple field pattern sequences has a lower collision probability.

[0101] In some embodiments, the method further includes:

[0102] If the splitting of the second sequence is uniform splitting and the second preset length is 2, determine the number of rounds of splitting and forward and reverse addition processing according to the length of the string.

[0103] On the other hand, an embodiment of the present invention provides a method for encrypted transmission of data, including:

[0104] Obtain target data and a digital signature; the digital signature is obtained by performing a hash operation and private key encryption on the target data;

[0105] Perform a hash operation on the target data through the hash digest generation method as described above to obtain a target output digest;

[0106] Decrypt the target output digest with a public key to obtain a target value;

[0107] If the target value is the same as the digital signature, receive the target data.

[0108] It can be understood that the data transmission process of this application is as follows:

[0109] Obtain target data and a digital signature; the digital signature is obtained by performing a hash operation and private key encryption on the target data;

[0110] Perform encoding processing on the target data to generate a string; perform information redistribution and homogenization processing on the string to generate a first sequence; determine the base number and the true number according to the first sequence, and determine a field type sequence based on the base number and the true number; extract a sequence with a first preset length from the field type sequence to generate a target output digest of the target data based on a hash function;

[0111] Decrypt the target output digest with the public key to obtain a target value;

[0112] If the target value is the same as the digital signature, receive the target data; if the target value is not the same as the digital signature, reject receiving the target data.

[0113] Of course, the hash digest generation method provided by this application can also be used in other fields. In any field where a hash function can be applied, it can be processed through the hash digest generation method provided by this application. This application does not limit the specific application fields.

[0114] The following uses a specific example to introduce the method provided by this application in detail:

[0115] A hash function is a function that maps data of any size to a data value of a fixed size. It has a wide range of applications in computer science and cryptography. Hash functions can be used for digital signatures, message integrity detection, authentication detection of message origin, etc., and also constitute the security guarantee of various cryptographic systems and protocols. Since expanding integer exponents by multiplication operations to form a field sequence has unpredictability and uniformity of character distribution, and the characteristic that reversing the exponent integer from the field is a one-way function. We propose an invention technology based on integer exponent fields as hash functions, which can generate hash digests with extremely low collision rates after verification. On the other hand, this application also newly creates a rule that can evenly distribute input information. The input information is used to generate the base and mantissa of the integer exponent based on the result of this operation. This processing rule for evenly distributing input information can cause all input characters to affect the base and mantissa, and thus achieve the ability to control the output of the hash digest, which fully realizes the goal of being sensitive to minor changes. In addition, the multiplication operations and the processing rules for uniform distribution listed in this invention technology both have low complexity. Therefore, it can be concluded that this invention technology can meet the many characteristics required for an ideal hash function, such as determinism, fast calculation, collision resistance, one-wayness, uniformity, and sensitivity to minor changes.

[0116] A hash function is a function H that maps a message M of any length to a hash value h of a fixed length (assume the length is m). Its mathematical expression is as follows:

[0117] h = H(M)

[0118] For a hash function to have one-wayness, it must satisfy the following characteristics:

[0119] Given M, it is easy to calculate h;

[0120] Given h, it is difficult to reverse M according to H(M) = h;

[0121] Given M, it is difficult to find another message M' such that H(M) = H(M').

[0122] Hash functions are a class of mathematical functions with the following three characteristics:

[0123] a. The input of a hash function can be a string of any length;

[0124] b. A hash function produces an output of a fixed size (such as a 256-bit output);

[0125] c. A hash function can perform effective calculations and the calculation time is reasonable. For an n-bit string, the complexity of calculating its hash function is O(n).

[0126] According to the above conditions, the technical solution of the present invention encodes the input information through four steps: ① encoding, ② reallocating and scrambling, ③ generating the base and the mantissa of the exponent integer and expanding the exponent field type, and ④ generating a fixed-length hash value, so as to use the exponent field type as a new hash function of the function.

[0127] We use the following five parts to elaborate the technical solution of the present invention in detail.

[0128] I. Carry is a complex non-linear problem

[0129] Multiplying two decimal numbers results in complex carry relationships, which vary depending on the specific numbers. There are a large number of states that need to be verified. Therefore, it is only possible to deduce the relationships for individual cases. It belongs to a non-linear problem and it is impossible to sort out carry relationships applicable to any multiplication of two decimal numbers. The following takes the multiplication operation of binary numbers as an example to illustrate the complexity of decimal number multiplication. The multiplication operation of binary numbers can be based on the addition operation of two binary numbers. The addition result of 1100100101 and 0101101100 is as follows:

[0130]

[0131] The above addition operation of two n-bit binary numbers requires the following five steps or operations for a total of n times. If the kth bit of one number has fewer than n bits, n-k 0s can be added in front of it to meet the same bit length.

[0132] The addition operation process of these five steps is organized as follows:

[0133] 1. Starting from the right, observe the two bits above and below, and whether there is a carry;

[0134] 2. If both are 0 and there is no carry, write 0 and then perform the operation on the next bit to the left;

[0135] 3. If there are the following two situations: (a). Both are 0 and there is a carry; (b). Only one of them is 0 and there is no carry; then write 1 and then perform the operation on the next bit to the left;

[0136] 4. If there are the following two situations: (a). Both are 1 and there is no carry; (b). Only one of them is 0 and there is a carry; then write 0 and then perform the operation on the next bit to the left;

[0137] 5. If both are 1 and there is a carry, write 1 and mark the carry information for the next field to the left, and then perform the operation on the next bit to the left.

[0138] The multiplication result of the aforementioned 1100100101 and 0101101100 is as follows:

[0139]

[0140] Since 0101101100 has five 1s, a total of five addition operations on the multiplicand 1100100101 with different shifts (different weight values) are required to obtain the result of 1100100101×0101101100 = 1000111100010011100. For each addition operation, the above five steps a to e need to be checked for each field. The result will change according to the specific numbers. It is necessary to calculate the addition result of the last two digits one field at a time from right to left in sequence. We cannot directly infer the value of the addition operation result of several fields skipped on the left from the results of several fields with lower weight values on the right. In other words, the result of 1100100101×0101101100 is very difficult to obtain through the "estimation method". Here, the so-called "estimation method" means that it cannot be calculated by other rules other than the addition procedure of the above five steps a - e. Generally speaking, binary multiplication can be realized from the binary addition operation logic. When performing the above five steps a to e for checking, it depends on whether the two numbers are 0 or 1 and whether there is a carry in the previous column. The complexity can be summarized into the following six states in total. Therefore, during the process of adding two n-bit binary numbers, a total of 6n state judgments are required to obtain the result. So it is not a simple process. As shown in Table 2.

[0141]

[0142] Table 2

[0143] If we try to perform decimal addition from the perspective of operation efficiency, considering that the values of the numbers range from 0 to 9 and the carry status also has ten states from 0 to 9, the total complexity can be summarized into 550 states. The following only lists the 55 states with a carry value of 2 from the previous field on the right as follows. Since the carry value of the previous field can be divided into ten values from 0 to 9, the total complexity can be summarized into 10 * 55 = 550 states. This can be summarized as a concept of using a decision tree to realize decimal addition operation. However, dealing with such a large state structure will affect the difficulty and challenge of hardware implementation.

[0144] Furthermore, we need to generalize the addition operation in decimal system to the multiplication of two decimal numbers, which can be explored from the basic concept of "multiple". Multiplying the integer a by b is equivalent to adding the integer a to itself b times. If a has n decimal characters, then a + a requires 550n total state judgments. To complete the multiplication of a by b, there are a total of 550nb such states. Is there a way to systematically summarize these states into rules? If the answer is yes, then we have reason to say that the exponential field type can be deduced according to certain rules, which has predictability or estimability, and the values of some fields on the left can be directly inferred from the results of some fields on the right. The problem is whether we can estimate the huge number of 550nb states using rules. The 55 states with a carry value of 2 in the previous field are listed in Table 3.

[0145]

[0146]

[0147] Table 3

[0148] The following equation shows the result of implementing multiplication 78456321 * 122 = 9571671162 through the logic of addition. First, split 122 into three parts: 100 + 20 + 2, which are the units digit, tens digit, and hundreds digit. So, the top two terms 78456321 + 78456321 in the equation correspond to the result of 78456321 * 2 for the units digit, the middle two terms 0784563210 + 784563210 correspond to the result of 78456321 * 20, and the bottom term 07845632100 is equal to 078456321 * 100. Adding these five terms gives 9571671162:

[0149]

[0150] II. Analysis of the exponential field type sequence as a hash function:

[0151] In this subsection, we use the entropy H that measures the amount of information in a message to measure the degree of uniform distribution of the exponential field type. Suppose a discrete information source has n elements {v0, v1,..., v n-1}, and their respective probabilities are {p0, p1,..., p n-1}, then the value of its entropy is:

[0152]

[0153] If we regard the exponential field type (d0d1...d N-1 ) as a random sequence of length N, d k-1∈ {0, 1, 2, …, 9} represents the k-th element. Therefore, this discrete information source has n = 10 elements. When the value of N is large enough and the random sequence has a high degree of randomness, then p0 ≈ p1 ≈ … ≈ p9 = 0.1. We use:

[0154]

[0155] to represent the information entropy value of the ideal uniform distribution, and use as the exponent field type (d0d1…d N-1 ) the entropy obtained according to the actual p k value, and then define as the level of uniformity of the exponent field type (d0d1…d N-1 ).

[0156] It can be understood that:

[0157] a. The uniformity level of the exponent field type depends on the length of the exponent field type and is independent of the selected base.

[0158] b. The simulation calculation results show that regardless of the base, when the length of the exponent field type sequence is in the order of 90 - 100 numbers, a 95% uniformity level can be achieved.

[0159] c. Considering the difficulty of calculating large exponent field types based on requirements, a high uniformity level goal can also be achieved by concatenating several exponent field types with low lengths at the head and tail to form a field type sequence with sufficient length.

[0160] d. When the length of the exponent field type is insufficient, the uniformity level can be moderately increased to 99% by converting decimal digits to 0 or 1 digits.

[0161] Take three relevant field type sequences as examples for verification and explanation:

[0162] 56 123 =(1064455037 9089249868 0125463186 5312565699 3627459243 9518875633 9061727975 7086406900 9327000832 3409649232 6986378194 7424181748 0553288470 3089431014 7953621115 0168750494 0974916797 3105784921 9488883031 8388869149 2469730197 897216)216 bits

[0163] 56 124 =(5960948212 2899799260 8702593844 5750367916 43137717661305703549 8745676663 9683878645 2231204661 1094035703 1123717890 55754177891098415433 7300813682 8540278244 0945002766 9459534064 9392395562 91377449782977667235 7830489108 2244096)217bits

[0167] 55 124 =(6382245375 0697873386 7459180905 1801824757 15181147926211908816 4514608803 5463943844 3779354156 0769820434 1783094728 77170465586023628663 5168373084 8780377213 2330698387 1588961126 6266802369 61774994597362820059 0610504150 390625)216bits

[0171] A detailed analysis and comparison of the above three field sequences can confirm that there is no pattern in how these decimal elements are distributed in their fields. It is impossible to deduce the regularity. Only when the base and the true number are known, can we use the multiplication relationship, such as from 56 124 =56×56 123 The relationship between the first field and the second field can be obtained from the first field. Although the two have such a relationship, due to the complex nonlinear problem of carry mentioned above, the distribution of decimal elements in these two fields is still very different, and it is impossible to estimate the field from the fact that the two have the same base. However, if you want to deduce the third field from the first two fields, there is no connection between them, although the latter two have the same power. In addition, even if the first field is known, according to If we want to deduce the base and true number from the first field, we must first analyze the two factors 2 and 7, and then find the individual powers. This is a complicated problem. In particular, the solution may not be unique. For example, because 56 123 =(2 3 ×7) 123 =2369 ×7 123 =(2 9 ×7 3 ) 41 =175616 41 。And the calculation of the field pattern is easy.

[0172] When the base number is composed of multiple relatively large factors, it becomes more difficult to reverse-deduce the base number and the true number from the expanded field pattern. In addition, if the head and tail of the field pattern are removed to form a new sequence, or two field patterns are added and merged into a new sequence, or multiple short field patterns are connected end to end to form a long field pattern sequence, it is even more challenging to reverse-deduce the base number and the true number from them.

[0173] Combining the two conclusions:

[0174] 1. Based on the non-linear property of dealing with multiplication carry, the expansion of the exponential field pattern sequence has the characteristics of unpredictability and uniform distribution of the field pattern. A slight change in the base number or the true number can generate a completely different field pattern sequence.

[0175] 2. Expanding the exponential field pattern sequence has the property of a one-way function. It is extremely challenging to reverse-deduce the base number and the true number from the exponential field pattern sequence.

[0176] Based on these two conclusions, it is suitable to be used as a hash function to generate hash digests.

[0177] III. Satisfying the information uniform redistribution rule for input sensitivity

[0178] During the process of hash function processing, the variable-length input will be converted into a fixed-length hash value as the output. During this process of generating the hash digest, any minor change in the input will cause a large change in the output. This characteristic is called input sensitivity. This depends on whether any individual character in the input can affect the output during the operation of the hash function. To achieve this goal, the technology of the present invention proposes a processing mechanism that can disrupt the input information and evenly redistribute the information.

[0179] In the first round, the input information is first encoded into a string composed of decimal letters. For example, according to the base64 encoding method, an English letter or a decimal letter can be compiled into one or two decimal characters respectively. Then, based on the generated string, a new character string of reverse arrangement is formed into a new character 3||2 string. After adding the large integers represented by these two strings, a new string integer is generated. If the addition operation causes an overflow due to carry, the highest bit is discarded, so the generated new string still maintains the same length. In this processing process, if any character in the original information changes, regardless of whether it occurs in the left half or any position in the right half string, it will cause changes in the character content of two positions in the new string. In other words, a change in one character will become a change in two characters. The following formula is used as an example for illustration:

[0180]

[0181] In the above formula, N is the number of digits of the string. The field type sequence of length N is represented by the formula s=(d N d N-1 …d1). After the first-round operation on the d N-i th element in its right half, it will affect the left and right half sequences of the newly generated string, and it will occur at the elements at the c i th and c N-i th positions.

[0182]

[0183] In the above formula, the || symbol means concatenating the two sequences on its left and right. At this time, the length of s=(d N d N-1 …d1) is the same as the length of .

[0184] The second-round operation is to divide the newly generated string in the first round into two segments, the front and the back. As shown in the above formula, if the length of the string is odd at this time, the left segment in the front is allocated one more character than the right segment in the back as a principle. For the two segments in the front and the back, in other words, c1 = c 11 ||c 12 and c2 = c 21 ||c 22 . After adding the two strings in the middle to their reverse arrangement sequences in sequence, a new string is generated respectively. If the addition operation generates a carry and an overflow, the highest bit is discarded, so the length is still maintained. In this processing process, if any character in the original information changes, it will cause changes in the character content of four positions in the newly generated string. Just take the processing of the left half c1 = c 11 ||c 12Take the part of "c" as an example and illustrate as follows, where c is set i is located at c1 = c 11 ||c 12 in the "c" of 11 the end, specifically as follows:

[0185]

[0186] In the third round, the string generated in the second round will be divided into four segments from left to right (i.e., four second sequences are obtained). If the length of the string is odd at this time, it is still based on the principle that the left segment has one more character than the right segment in the back. Similarly, according to the above principle, the reverse arrangement is generated separately within each of the four segments, and then in the four independent segments, after adding the string and its reverse arrangement sequence in order, a new string is generated respectively. If there is a carry and overflow in the addition operation, the highest bit will be discarded, so the four independent segments still maintain the same length. In this processing process, at this time, any change in the original character, no matter where it occurs, will cause changes in the characters at eight positions in the newly generated string.

[0187] For example of how to segment, let N = 66. At this stage, the operation process of the first round will be applied to c1 = c 11 ||c 12 and c2 = c 21 ||c 22 , at this time c = c1||c2 = c 11 ||c 12 ||c 21 ||c 22 . Since N = 66, after the first round, it is divided into two segments of 33||33. At this time, this 33 refers to the length of |c1| = 33 mentioned above. Since 33 is odd, in the next round, it will be cut into two segments of 17||16, and the previous segment is one more in length than the latter segment. Therefore, N = 66 will be divided into four segments of 17||16||17||16 after two rounds (n = 2). At this time 2 n = 2 2 = 4, n = 2. Next, 17 is divided into two segments of 9||8, 16 is divided into 8||8, then when dividing, 9 is divided into two segments of 5||4, and then the part of 5 will become two segments of 3||2. In the next step, the part of 3||2 is cut into 2||1||1||1, and finally the part of 3||2 becomes 1||1||1||1||1.

[0188] Summarize as follows:

[0189] 66 → 33 || 33 (the result of n = 1 has 2 segments of the second sequence) → 17 || 16 || 17 || 16 (the result of n = 2 has 4 segments of the second sequence) → 9 || 8 || 8 || 8 || 9 || 8 || 8 || 8 (the result of n = 3 has 8 segments of the second sequence) → 5 || 4 || 4 || 4 || 4 || 4 || 4 || 4 || 5 || 4 || 4 || 4 || 4 || 4 || 4 || 4 (the result of n = 4 has 16 segments of the second sequence) → 3 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 3 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2 || 2

[0190] (After the fifth round of processing, the result of n = 5 has 32 segmented second sequences)

[0191] → 2 || 1 || 1 || 1 || 1 || 1 || … || 1 || 1 || 1 || 1 || 1 || … || 2 || 1 || 1 || 1 || 1 || … || 1 || 1 || 1 || 1 || 1 || 1 (the result of n = 6 has 64 segments of the second sequence)

[0192] (The final result after the seventh round of processing).

[0193] This process of successively halving the length of the previous round in each round and then adding and merging it with the reverse arrangement sequence continues until the length of each segment is less than 2 in the end. If the number of such recursions is n (i.e., the number of rounds in this application), then if any character in the original information changes, it will cause 2 n changes in the characters at 2 positions in the newly generated string. In other words, this way of scrambling the input information can distribute the influence of any character in the input to 2 n different positions in the entire sequence, so it can meet the requirement of the sensitivity to the input during the operation of the hash function. According to the example of N = 66 above, select Therefore, if seven rounds of information averaging and redistribution processing are performed, 64 = 2 6 < N(= 66) < 2 7 = 128, ensuring that the values at all individual positions in the sequence can affect the values at all positions. IV. The process of generating the output summary of this technology

[0194] The first step: Information encoding, padding, and homogenization processing. Encode the input information into a string composed of decimal letters. For example, one English letter or decimal letter can be compiled into two decimal characters according to the base64 encoding method, with s = (d N d N-1…d1) indicates that this is a sequence of length N. It refers to the length of the encoded information of the input information for generating the hash digest. If its length is less than 80, the leading character part is directly concatenated at its end for padding. It should be noted that 80 here is an exemplary example, and this application does not limit the specific length threshold. Those skilled in the art can adjust it according to actual needs. The character padding process can be as follows:

[0195]

[0196] This formula indicates that the previous 80 - N characters need to be supplemented. Or because the information is too short, it needs to be concatenated multiple times to meet the condition of a length of 80. For example:

[0197]

[0198] To avoid confusion caused by using too many symbols, in the following, regardless of the situation of s, s1, s2, they are all uniformly represented as s=(d N d N-1 …d1) indicates that this is an input sequence of length N≥80.

[0199] Second step: Information redistribution and homogenization processing flow. In this step, the multi - round information scrambling process described above is executed, which is called the information homogenization processing flow. Its purpose is to evenly distribute all input characters so that they all have the same influence on the subsequent execution of the hash function. For s=(d N d N-1 …d1) generated in the first step, the result after passing through the information homogenization processing flow is represented by the following formula:

[0200] s=(d N d N-1 …d1)→e=(e N e N-1 …e1)

[0201] Next, it is packaged into a sequence with a final length of 80. Let N be the length of the encoded information of the input information for generating the hash digest (that is, the length of the string in this application, the first value). Taking N = 80k - d as an example, the previous d characters e N e N-1 …e N-d+1 can be inserted and concatenated at the end to generate k (the second value) sequences of length 80 {e1, e2, …, e k}, which is expressed as follows:

[0202]

[0203] Third step: Generate the base and the mantissa of the exponential integer. Perform the addition and combination of k sequences {e1, e2, …, e k} with each sequence having a length of 80. If carry and overflow occur during the addition operation, discard the highest bit, so the length remains 80 characters. The result is shown as follows:

[0204]

[0205] Considering that the output digest is 80 characters long, these 80 characters are extracted from the field type sequence generated by performing the expansion of the exponential integer x e . We need to examine how to generate a pair of integers (x, e) for the base and the mantissa, or multiple pairs of integers (x1, e1), (x2, e2), …, (x 80 h 79 …h1) to complete this task. We can make a trade-off from two perspectives: the probability of collision and the computational complexity. k ,e k )

[0206] One way is to divide (h 80 h 79 …h1) into two equal parts. Derive the base from the first part and the mantissa from the second part. Extract the first 80 characters from the field type sequence of an x e as the output digest O 80 , as shown in the following formula:

[0207]

[0208] This formula shows that this part is the field type sequence generated by expanding the exponential integer x e . Assume its length is N, and then sequentially extract the first 80 decimal digits from the beginning as the output of the hash digest.

[0209] According to the field type sequence generated by expanding the exponential integer x e , estimate its length to be

[0210] For example

[0211] Next, further analyze the feasibility of the above rule. From the perspective of uniform character distribution, the probability of any character in (h 80 h 79 …h1) appearing as any value in the set {0, 1, 2, …, 9} is the same, which is one-tenth. The expected value of randomly selecting a character is 4.5.

[0212]

[0213] Therefore, the following estimation results can be obtained. This rule can obtain a field type sequence with an average character length of 102, which is sufficient to extract 80 characters from it as the output summary.

[0214] E(x) = E(h 80 +h 79 +…+h 41 ) = 40×E(h i ) = 40×4.5 = 180,

[0215]

[0216] This method of dividing (h 80 h 79 …h1) into two equal parts, deriving the base number from the previous segment and the mantissa from the latter segment. A relatively large problem is that it requires multiplying a base number with a hundreds digit by a power of a tens digit, which requires a considerable amount of hardware equipment or computing power.

[0217] Second, divide (h 80 h 79 …h1) into five grouped sequences of 16 lengths in sequence from front to back to generate {h1, h2, h3, h4, h5}. The process is as follows.

[0218] h = (h 80 h 79 …h1) → (h 80 h 79 …h 65 )||(h 64 h 63 …h 49 )||(h 48 h 47 …h 33 )||(h 32 h 31 …h 17 )||(h 16 h 15 …h1) = h1||h2||h3||h4||h5

[0219] These five grouped sequences of 16 lengths independently generate five output summaries of 16 lengths in sequence, and then concatenate them to form the final 80 characters as the output summary. How to generate the bases and mantissas of five groups of exponential integers from {h1, h2, h3, h4, h5}. According to the principle, it can be satisfied by taking both the base and the mantissa as ten-digit numbers. The following example is given for illustration:

[0220]

[0221] E(x) = 16 × E(h i ) = 16 × 4.5 = 72,

[0222] E(e) = 72 / 4 = 18,

[0223]

[0224] In this rule, each segment can obtain a field pattern sequence with an average length of 33 characters, which is sufficient to extract 16 characters therefrom as the output summary of this segment. It only needs to perform the expansion work for five relatively small exponential integers in five times, but the advantage is that it can have a lower collision probability. The former has only one set of base and mantissa, but the latter will have a collision only when the same five sets of base and mantissa are found simultaneously.

[0225] Fourth step: Generate the output summary. Expand one larger or five smaller exponential integers generated in the third step in sequence and extract the first 80 characters therefrom as the output summary, or expand the five smaller exponential integers and extract the first sixteen characters from each segment, and then concatenate them into an output of 80 decimal digits as the hash summary.

[0226] V. Example of implementing the hash summary generation technology of the present invention

[0227] The following is to convert the two input messages hello1234 and Hello1234 into decimal characters according to the above base64 encoding table, and then generate two 80-character output summaries respectively according to the above process.

[0228] hello1234 → 333037374053545556 (18 digits)

[0229] Hello1234 → 733037374053545556 (17 digits)

[0230] After expanding them into 80-character strings respectively and completing the first-round homogenization processing, the results are as follows:

[0231]

[0232]

[0233] Next, perform the second-round uniform distribution processing respectively:

[0234] hello1234→0703407106089990610703407106089990610703||4071060899906107034071060899906107034070+3070160999806017043070160999806017043070||0704307016099980601704307016099980601704=3773568105896007653773568105896007653773||4775367916006087635775367916006087635774

[0235] Hello1234→2654211135731111126542111357311111265421||1135731111126542111357311111265421113572+1245621111137531112456211111375311124562||2753111245621111137531112456211111375311=3899832246868642238998322468686422389983||3888842356747653248888423567476532488883

[0236] Next, the third-round uniform distribution process is carried out separately:

[0237]

[0238] The first, second, third, and fourth segments of the above formula are the same, and their final results are detailed in the following two formulas:

[0239]

[0240] And the results of processing its third and fourth segments:

[0241]

[0242] Therefore, the final result obtained is:

[0243] Hello1234 → 73037374053545556730373740535455567303737405354555673037374053545556730373740535 → 37320791155108812372||37320791155108812372||27311991044009021371||27311991044009021371 → 70755695666957647365||70755695666957647365||86965655540492260632||86965655540492260632 → 70755695666957647365707556956669576473658696565554049226063286965655540492260632 Similarly, it can be deduced that:

[0244] hello1234 → 33303737405354555633303737405354555633303737405354555633303737405354555633303737 → 11302688043978730310||11303775965857840310||23121485166258512131||23121594066149612131 → 70744707447075568532695106755477665655432126652544505221115521232333553333311155 In other words, after the first step: information encoding, padding, and homogenization processing, we get:

[0245] hello1234 → 33303737405354555633303737405354555633303737405354555633303737405354555633303737 → 70744707447075568532695106755477665655432126652544505221115521232333553333311155 And:

[0246] Hello1234 → 73037374053545556730373740535455567303737405354555673037374053545556730373740535 → 70755695666957647365707556956669576473658696565554049226063286965655540492260632

[0247] Next, generate the base and the mantissa of the exponent from this result respectively.

[0248] hello1234 → 7074470744707556853269510675547766565543...2126652544505221115521232333553333311155

[0249]

[0250] The base and the mantissa of the exponent are x = 196 and e = 60 respectively. After expansion, it is a field pattern sequence with a length of 138 characters. Take the first 80 characters from it as the output summary:

[0251] 196 60 = 3430554169730956681390165337397838714125... 7330362042754022389906616080973973799161... 2993659096899643116954895432267121541745...

[0254] 424241396155416576(138 digits)

[0255]

[0256] Hello1234 → 7075569566695764736570755695666957647365... 8696565554049226063286965655540492260632

[0259]

[0260] The base and the mantissa of the exponent are 228 and 93 respectively. After expansion, it is a field pattern sequence with a length of 220 characters. The first 80 characters are taken as the output summary:

[0261] 228 93 =1940621198178227352908094027412540778759... 7634646007047559153574431617911592513329... 0368996850920071942579087967724333243141... 6749058595854866710212194436842574820701... 8260817475427654927769310558296648943918...

[0266] 35257951014562562048(220digits)

[0267]

[0268] If it is divided into five sections, five groups of bases and mantissas are generated in sequence, and then 16 length summaries are generated respectively and concatenated into the final output summary.

[0269]

[0270] 74 18 =4427626266116305,911795793508171776

[0271] 82 20 =1889196131813120,32574569023867244773376

[0272] 69 17 =1821522510002046,4195180876374789

[0273] 43 10 =2161148231328424,9

[0274]

[0275] 93 23 =1884116777026212,762307453711666341429524886357

[0276] 92 23 = 1469332310991197,227605027959666916684870975488

[0277] 93 23 = 1884116777026212,762307453711666341429524886357

[0278] 72 18 = 2703864574375502,950215699413336064

[0279] 64 16 = 7922816251426433,7593543950336

[0280] →(1884116777026212||1469332310991197||1884116777026212||2703864574375502||7922816251426433)

[0281] Based on the above results, it can be verified that whether the output digest is generated after expanding through one or five exponents, all the output digests are very different, and there is no correlation between these output digests at all. This proves that the new hash function proposed by this technical solution has sensitivity to minor changes. A minor change in the input data, such as changing a character h→H, leads to a huge change in the output hash value. This property is called the "avalanche effect".

[0282] Hash functions can be used in applications such as digital signatures, message integrity detection, and authentication detection of message origin. At the same time, they also constitute the security guarantee for various cryptographic systems and protocols. The following only provides a flowchart of a hash function as an illustration for the application scenario of digital signatures, and other applications can be inferred accordingly. Refer to Figure 2 and Figure 3As shown below, the following process m represents the information to be transmitted, i.e., the target data, h() represents the hash function of the present invention's technology, and Fs() represents a function that can generate a digital signature D = Fs(M) using the private key of the transmitting end. The transmitting end packs the information m to be transmitted together with the digital signature D and transmits it to the receiver. After the receiver receives X = (m, D) = S, it extracts the m therein to generate a hash digest M' = h(m) at the receiver's local end, and uses the function Fp() that can decrypt the digital signature M' using the public key. Operating this function can obtain D' = Fp(M'), i.e., the target value. Finally, compare whether the D' generated at the receiver is equal to the D in the received X = (m, D). If the two are the same, accept the data m; otherwise, discard it. The entire digital signature process is as shown in the following figure.

[0283] The technology of the present invention proposes to first uniformly scramble the information, and then generate the base and the mantissa of the exponential integer based on this result. Finally, expand it into a field sequence and extract its fragment as the hash output digest. Due to the complex carry relationship of the multiplication operation, expanding the exponent is essentially multiplying the base by itself continuously, resulting in the unpredictability and uniform distribution characteristics of the expanded field sequence, which can meet the characteristics required by an ideal hash function: ① determinacy, ② fast calculation, ③ collision resistance, ④ one-wayness, ⑤ uniformity, ⑥ sensitivity to minor changes. It is organized into Table 4 for easy analysis, comparison, and reference:

[0284]

[0285] Table 4

[0286] The present invention proposes a method that can uniformly scramble the information. Through this method, any character in the information can affect the subsequent work to be performed. This is a brand-new idea. The present invention proposes to use the sequence after the above-mentioned uniform scrambling of the information to generate the base and the mantissa of the exponential integer, which is a new idea. The present invention also proposes to use the field sequence as a hash function to generate a hash digest, which is also a brand-new idea. Considering the computing power, the present invention proposes an elastic mechanism that can segment the calculation of a larger exponential field sequence into the concatenation of multiple smaller exponential field sequences, making it have the potential to be implemented in Internet of Things terminal devices. Overall, the technology of the present invention proposes to first uniformly scramble the information, then generate the base and the mantissa of the exponential integer based on this result, and finally expand it into a field sequence and extract its fragment as the hash output digest. This is a new-generation hash function with a brand-new concept.

[0287] Whether it is the MD sequence or the SHA sequence, four or eight linear functions are first defined, and then each block is processed through multiple rounds of loop operations (based on logic functions and constant tables), gradually updating four or eight variables, and finally generating a hash output. The entire operation process also needs to be paired with multiple constant tables of fixed length to obtain a hash digest with sufficient randomness. However, the technology of the present invention can achieve the same effect only through two key procedures, namely the "information homogenization processing program" and the "expanded field type sequence". Obviously, the technology of the present invention is simple, easy to understand, has a low complexity, and is easier to implement.

[0288] It is difficult to compare the differences between the technical solution of the present invention and the MD sequence or the SHA sequence in terms of collision resistance. Given a Hash function h and y as a message digest, the problem of finding x such that y = h(x) is called the preimage problem. Given a Hash function h and any given message m1, the problem of finding an m2, m1 ≠ m2, such that h(m1) = h(m2) is called the second preimage problem. If it is computationally infeasible to make h(m 1) = h(m2), then h is called a weak collision-free Hash function. Given a Hash function h, the problem of finding any pair of messages m1, m2, m1 ≠ m2, such that h(m1) = h(m2) is called the collision problem. If it is computationally infeasible to make h(m1) = h(m2), then h is called a strong collision-free Hash function. However, the probability of possible collisions can be roughly estimated from the length of the output digest. For a SHA-256 Hash function with a 256-bit output, in the worst case, a collision will occur after 2 256 +1 Hash function calculations, and the average number of times is 2 128 times, where

[0289] 2 256 = 11579208923731619542357098500868790785326998466564056... 4039457584007913129639936

[0291] ≈ 1.158×10 77 <10 80

[0292] 10 80 / 2 256 ≈ 863.6168555094445.

[0293] However, the hash digest of the technical solution of the present invention is 80 decimal characters. In the worst case, a collision will occur after 10 80 +1 Hash function calculations. The probabilities of collision between the two belong to different levels, and the difficulty difference between the two exceeds 860 times.

[0294] In summary, the method provided by the embodiments of the present application includes: obtaining input information, encoding the input information to generate a string; performing information redistribution and homogenization processing on the string to generate a first sequence; determining a base number and a mantissa according to the first sequence, and determining a field type sequence based on the base number and the mantissa; extracting a sequence with a first preset length from the field type sequence to generate an output digest of a hash function. By information redistribution and homogenization, each character change affects the output digest, improving the security of hash processing; on the other hand, generating the output digest through the field type sequence reduces the collision rate; this solution is simple and easy to implement and is convenient for popularization.

[0295] Secondly, refer to the attached Figure 4 Describe a hash digest generation system according to an embodiment of the present invention. The system specifically includes:

[0296] A first module 310, configured to obtain input information and perform encoding processing on the input information to generate a string;

[0297] A second module 320, configured to perform information redistribution and homogenization processing on the string to generate a first sequence;

[0298] A third module 330, configured to determine a base number and a mantissa according to the first sequence, and determine a field type sequence based on the base number and the mantissa;

[0299] A fourth module 340, configured to extract a sequence with a first preset length from the field type sequence to generate an output digest of a hash function.

[0300] It can be seen that the content in the above method embodiments is applicable to the system embodiments of the present application. The functions specifically implemented by the system embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0301] Refer to Figure 5 , an embodiment of the present invention provides a hash digest generation device, including:

[0302] At least one processor 410;

[0303] At least one memory 420, configured to store at least one program;

[0304] When the at least one program is executed by the at least one processor 410, the at least one processor 410 is caused to implement the hash digest generation method described above.

[0305] Similarly, the content in the above method embodiments is applicable to the embodiments of this apparatus. The functions specifically implemented by the embodiments of this apparatus are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0306] An embodiment of the present invention further provides a computer-readable storage medium, which stores a program executable by a processor. The program executable by the processor is used to execute the above hash digest generation method when executed by the processor.

[0307] Similarly, the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0308] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, where the order of various operations is changed and where sub-operations described as part of a larger operation are executed independently.

[0309] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the ordinary skill of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0310] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several programs for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes of various kinds.

[0311] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as an ordered list of executable programs for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by a program execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can retrieve and execute programs from a program execution system, apparatus, or device), or in combination with these program execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with a program execution system, apparatus, or device.

[0312] More specific examples (nonexhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROMs). Additionally, a computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or, if necessary, other appropriate processing, and then stored in a computer memory.

[0313] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0314] In the above description of this specification, the description referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0315] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0316] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.

Claims

1. A method for generating a hash digest, characterized in that, It includes the following steps: Obtain input information, and perform encoding processing on the input information to generate a string; Perform information redistribution and homogenization processing on the string to generate a first sequence; According to the first sequence, determine the base number and the true number, and based on the base number and the true number, determine the field type sequence; Extract a sequence with a first preset length from the field type sequence to generate the output digest of the hash function.

2. The hash digest generation method according to claim 1, characterized in that, The performing information redistribution and homogenization processing on the string to generate a first sequence includes: Add the forward arrangement sequence of the string and the reverse arrangement sequence to obtain a first-round sequence; Divide the first-round sequence into two segments, and add the forward arrangement sequence and the reverse arrangement sequence of the corresponding string for each segment to obtain a second-round sequence; the second-round sequence includes two segments of second sequences; Successively perform splitting and forward-backward addition processing on the strings of each segment of the second sequence, update the quantity and character length of the second sequence until the character length of each segment of the second sequence obtained is less than the second preset length to generate a uniform sequence; Perform expansion processing on the uniform sequence to obtain a first sequence; the length of the first sequence is an integer multiple of the first preset length.

3. The hash digest generation method according to claim 2, wherein The adding the forward arrangement sequence of the string and the reverse arrangement sequence to obtain a first-round sequence includes: If the length of the added sequence is greater than the length of the string, discard the highest-order character to obtain the first-round sequence.

4. The hash digest generation method according to claim 1, wherein The according to the first sequence, determining the base number and the true number, and based on the base number and the true number, determining the field type sequence includes: Split the first sequence according to the first preset length, and add the several split sequences to obtain a third sequence; wherein, the length of the third sequence is the first preset length; Split the third sequence into a fourth sequence and a fifth sequence, and determine the base number according to the fourth sequence and determine the true number according to the fifth sequence; Based on the base number and the true number, determine the field type sequence.

5. The hash digest generation method according to claim 1, wherein The according to the first sequence, determining the base number and the true number, and based on the base number and the true number, determining the field type sequence includes: Split the first sequence according to the first preset length, and add the several split sequences to obtain a third sequence; Split the third sequence into several segments, and perform splitting on each segment. According to the sequences after splitting of each segment, determine the base number and the true number of the corresponding segment to obtain several pairs of base numbers and true numbers; Based on the several pairs of base numbers and true numbers, determine several field type sequences.

6. The hash digest generation method according to claim 2, wherein The method further includes: If the splitting of the second sequence is uniform splitting and the second preset length is 2, determine the number of rounds of splitting and forward-backward addition processing according to the length of the string.

7. A method for encrypted transmission of data, characterized in that, The method includes: Obtain target data and a digital signature; the digital signature is obtained by performing a hash operation and private key encryption on the target data; Perform a hash operation on the target data through the hash digest generation method according to any one of claims 1 to 6 to obtain a target output digest; Decrypt the target output digest through a public key to obtain a target value; If the target value is the same as the digital signature, receive the target data.

8. A hash digest generation system, characterized in that, Comprising: A first module, configured to obtain input information and perform encoding processing on the input information to generate a string; A second module, configured to perform information redistribution and homogenization processing on the string to generate a first sequence; A third module, configured to determine a base number and a true number according to the first sequence, and determine a field type sequence based on the base number and the true number; A fourth module, configured to extract a sequence with a first preset length from the field type sequence to generate an output digest of a hash function.

9. A hash digest generation device, characterized in that, Comprising: At least one processor; At least one memory, configured to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor is used to implement the method according to any one of claims 1 to 7 when executed by the processor.