A data integrity verification system and method based on a hash algorithm
By constructing analytical sets, extracting features, machine learning prediction and introducing random salt values, the problem of insignificant changes in hash values is solved, the risk of hash collision is reduced, and the reliability and security of data integrity verification is enhanced.
Patent Information
- Application Number
- CN202510362529.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-26
AI Technical Summary
When existing hash algorithms process input changes, if the hash value of the output changes small or insignificantly, they cannot intelligently identify this situation, resulting in an increased risk of hash collisions, and forged data may pass verification.
By constructing the analysis set, features related to hash change are extracted, and combined with the prediction of machine learning models, we can intelligently identify potential patterns with insignificant hash change. The random salt value is introduced and the number of hash iterations of the salt value is dynamically adjusted to increase the complexity and randomness of the hash calculation.
Effectively reduce the risk of hash collisions, ensure that the hash value can accurately reflect data changes, enhance the reliability and security of data integrity verification, and improve the system's defense ability against potential attacks.
Smart Images

Figure CN119892376B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data verification, and particularly relates to a data integrity verification system and method based on a hash algorithm. Background Art
[0002] A data integrity verification system based on a hash algorithm is a technology that verifies whether data has changed during transmission or storage by using a hash function. In such a system, the original content of the data is processed by the hash algorithm to generate a hash value of a fixed length (also called a digest). This hash value is the unique "fingerprint" of the data, with strong collision resistance and sensitivity. The transmitted or stored data and its hash value are sent or stored together. When the data recipient or user needs to verify the integrity of the data, the received data is hashed again, and the calculated hash value is compared with the original hash value. If the two are the same, it means that the data has not been tampered with during transmission or storage; otherwise, it indicates that the data may have been tampered with or lost. Common hash algorithms include MD5, SHA-1, SHA-256, etc.
[0003] The prior art has the following deficiencies:
[0004] When the existing hash algorithms process input changes, if the change in the output hash value is small or not significant, the current technology usually cannot intelligently identify this situation. When the input data has strong structurality or regularity, the algorithm may not be able to generate sufficiently diverse hash values, resulting in the same hash value being generated for different inputs. If the change in the input data fails to cause a significant change in the hash value, multiple different data may generate the same hash value, leading to verification failure and even allowing forged data to pass the verification.
[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] The object of the present invention is to provide a data integrity verification system and method based on a hash algorithm. By constructing an analysis set and extracting features related to hash value changes, combined with the prediction of a machine learning model, it can intelligently identify potential patterns with insignificant hash value changes. When the hash value change fails to effectively respond to data changes, by introducing a random salt value and dynamically adjusting the hash iteration times of the salt value, the complexity and randomness of hash calculation are increased, avoiding the situation of insufficient hash value changes caused by low entropy or repeated patterns of data. By enhancing the randomness and adaptability of the hash algorithm, the risk of hash collision is effectively reduced, and it is ensured that in the face of minor data changes, the hash value can accurately reflect these changes, thereby enhancing the reliability and security of data integrity verification and improving the system's defense ability against potential attacks to solve the problems in the above background technology.
[0007] To achieve the above object, the present invention provides the following technical solution: A data integrity verification method based on a hash algorithm, comprising the following steps:
[0008] Process the original input data through a hash algorithm to generate a hash value of a fixed length, collect the generated hash value and construct an analysis set containing the hash values corresponding to all input data;
[0009] After the construction of the analysis set is completed, extract the features related to the hash value changes from it, perform quantization processing on the extracted features, and initially identify potential patterns with insignificant hash value changes;
[0010] Input the quantized features as feature vectors into the trained machine learning model, and predict the change situation of the hash value output result through the model;
[0011] In the case of identifying insignificant hash value changes, combine the generated random salt value with the original data to form new input data, and then process it through the hash algorithm to generate a new hash value, improving the randomness and complexity of the hash result, and dynamically adjusting the hash iteration times of the salt value according to the degree of insignificance of the hash value change to reduce the risk of hash collision.
[0012] Preferably, processing the original input data through a hash algorithm to generate a hash value of a fixed length specifically includes:
[0013] Map the input data to a hash value of a fixed length through the hash algorithm. The hash value is a digital string, and the same input will always produce the same hash value.
[0014] Preferably, features related to the change in the hash value are extracted from the analysis set. Among them, the extracted features include the bit flip rate between adjacent hash values and the similarity degree between adjacent hash values. The bit flip rate between adjacent hash values and the similarity degree between adjacent hash values are quantified to generate a bit flip rate quantization value and an adjacent hash similarity quantization value respectively. Potential patterns with insignificant hash value changes are initially identified through the bit flip rate quantization value and the adjacent hash similarity quantization value.
[0015] Preferably, the quantified bit flip rate quantization value and adjacent hash similarity quantization value are used as feature vectors and input into a trained machine learning model. Based on the machine learning model, a change response index is predicted, and the change situation of the hash value output result is predicted through the change response index.
[0016] Preferably, the change response index generated when predicting the change situation of the hash value output result by the trained machine learning model is compared and analyzed with a preset change response index reference threshold to predict the insignificant change situation of the hash value. The specific steps are as follows: If the change response index is less than the change response index reference threshold, it is determined that the hash value change is insignificant; if the change response index is greater than or equal to the change response index reference threshold, it is determined that the hash value change is significant.
[0017] Preferably, in the case where it is identified that the hash value change is insignificant, the generated random salt value is combined with the original data to form new input data, and then processed through a hash algorithm to generate a new hash value. The specific steps are as follows:
[0018] After it is identified that the hash value change is insignificant, a random salt value is first generated , and it is combined with the original data D to form new input data. The specific expression is:
[0019]
[0020] , where: is an internal parameter used to control the source of randomness, represents a high-entropy random sequence generated under the parameter to ensure the unpredictability of the salt value, represents taking the th power of the result after performing a hash operation on the random sequence to further enhance the complexity of the salt value, represents concatenating and , D is the original data, is the new input to be hashed;
[0021] After obtaining the new input to be hashed After that, multiple iterative calculations are performed through a hashing algorithm, and the number of iterations is determined by the following formula:
[0022]
[0023] , where: is the number of salt value hashing iterations, which determines the number of rounds of repeated hashing processing for the new input to be hashed is the change response exponent, representing a measure of the degree of insignificant change in the hash value is the change response exponent reference threshold is the base multiplication factor, used to control the overall iteration intensity is the exponential amplification parameter, which performs a non-linear transformation on the degree of insignificant change represents floor division, ensuring that the number of iterations is an integer is the logarithmic adjustment factor, which amplifies the impact of insignificant changes on the number of iterations
[0024] Preferably, the specific steps for quantifying the bit flip rate between adjacent hash values to generate a bit flip rate quantization value are as follows:
[0025] For the hash values of a set of input data, a bit flip matrix is constructed to record the bit-by-bit changes between adjacent hash values. The constructed expression is:
[0026]
[0027] , where: represents the bit flip situation of the i th hash value at the j th bit and respectively represent the binary values of the i th and the i +1th hash values at the j th bit represents the exclusive OR operation. If the two bits are the same, the result is 0; if they are different, the result is 1, indicating that the bit has flipped is the weight function, defined as:
[0028]
[0029] , where: M is the length of the hash value mod is the abbreviation of the modulo operation, representing the operation of taking the remainder calculates the result of the bit index j modulo the hash length M , that is, the index j Within the hash value range The loop index within, if the index j exceeds M then it is made to cycle within M by taking the modulus, without exceeding the predetermined bit range, to enhance the influence degree of different bits;
[0030] Calculate the bit flip rate quantization value based on the bit flip matrix BFM, for quantifying the insignificance of hash value changes, and the calculation expression is:
[0031]
[0032] where: is the bit flip rate quantization value, N is the total number of hash samples, that is, the number of hash values in the dataset, Use the sine function to perform non-linear scaling on the bit index, so that bit flips in the middle of the hash value contribute more to the bit flip rate quantization value, while bit flips at the beginning and end have less influence, to avoid patternized hash collisions being ignored, E is the exponential adjustment factor, for amplifying the influence of low flip rates and improving the sensitivity of this bit flip rate quantization value to hash collisions.
[0033] Preferably, the specific steps for quantifying the similarity degree between adjacent hash values to generate the adjacent hash similarity quantization value are as follows:
[0034] Adopt the normalized weighted Hamming similarity to measure the overall similarity of hash values, and the calculation expression of the weighted Hamming similarity is:
[0035]
[0036] where: represents the normalized weighted Hamming similarity, and respectively represent the binary numbers at the k th bit of the hash values calculated from adjacent input data, is the dynamic weight of the k th bit, calculated using the entropy weight method, and is defined as follows: where, represents the probability that 0 or 1 appears at this bit position in multiple hash values;
[0037] Calculate the adjacent hash similarity quantization value according to the normalized weighted Hamming similarity , and the adjacent hash similarity quantization value reflects the similarity of adjacent hash values. The specific calculation expression is:
[0038]
[0039] ,in, Represents the quantized value of adjacent hash similarity, The maximum value of the normalized weighted Hamming similarity among all adjacent hash value calculations.
[0040] A data integrity verification system based on a hash algorithm, including a hash value generation module, a feature extraction and quantification module, a machine learning prediction module, a salt value generation and hash enhancement module, and a dynamic adjustment and collision protection module;
[0041] The hash value generation module processes the original input data through a hash algorithm to generate a hash value of a fixed length, collects the generated hash values and constructs an analysis set, which contains the hash values corresponding to all the input data;
[0042] The feature extraction and quantification module extracts features related to hash value changes from the analysis set after it is built, quantifies the extracted features, and preliminarily identifies potential patterns in which hash value changes are not significant;
[0043] The machine learning prediction module inputs the quantized features as feature vectors into the trained machine learning model, and uses the model to predict the changes in the hash value output results;
[0044] The salt value generation and hash enhancement module, when it is recognized that the hash value change is not significant, combines the generated random salt value with the original data to form new input data, and then processes it through the hash algorithm to generate a new hash value, thereby improving the randomness and complexity of the hash result;
[0045] The dynamic adjustment and collision protection module dynamically adjusts the number of hash iterations of the salt value according to the insignificant degree of hash value change to reduce the risk of hash collision.
[0046] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0047] The present invention constructs an analysis set and extracts features related to hash value changes, combined with the prediction of a machine learning model, to intelligently identify potential patterns where hash value changes are not significant. When hash value changes fail to effectively respond to data changes, the complexity and randomness of the hash calculation are increased by introducing a random salt value and dynamically adjusting the number of hash iterations of the salt value, thereby avoiding situations where insufficient hash value changes are caused by low entropy or repeated patterns of the data. By enhancing the randomness and adaptability of the hash algorithm, the risk of hash collisions is effectively reduced, and it is ensured that the hash value can accurately reflect these changes when facing small data changes, thereby enhancing the reliability and security of data integrity verification and improving the system's defense capabilities against potential attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other accompanying drawings can also be obtained based on these drawings.
[0049] Figure 1 It is a method flowchart of a data integrity verification method based on the hash algorithm of the present invention.
[0050] Figure 2 It is a module schematic diagram of a data integrity verification system based on the hash algorithm of the present invention. Specific embodiments
[0051] Now, the exemplary embodiments will be described more comprehensively with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these exemplary embodiments are provided so that the present disclosure will be more thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.
[0052] The present invention provides a Figure 1 data integrity verification method based on the hash algorithm as shown below, including the following steps:
[0053] Process the original input data through the hash algorithm to generate a hash value of a fixed length, collect the generated hash value and construct an analysis set, which contains the hash values corresponding to all input data;
[0054] The hash algorithm maps the input data (whether it is text, file, or data in other formats) to a hash value of a fixed length, usually a string of numbers, and the same input will always generate the same hash value. The role of this step is to provide a standardized data identifier for subsequent analysis, and can quickly generate the "fingerprint" of the data, laying a foundation for data integrity verification and other subsequent processing. The generated hash value is the unique identifier of the data, so it is the core in the whole solution.
[0055] Collecting the generated hash value and constructing an analysis set that contains the hash values corresponding to all input data is to form a benchmark and comparison set for data changes through the collection and organization of a large number of data hash values. The analysis set can help identify the rules of hash value changes. Especially when dealing with structured data, it can discover potential regularities and patterns, and then provide a rich data basis for subsequent feature extraction and model prediction. Through this step, the distribution characteristics of hash values under different inputs can be understood.
[0056] After the analysis set is constructed, extract the features related to the change of the hash value from it, quantify the extracted features, and preliminarily identify the potential patterns with insignificant hash value changes;
[0057] Extract the features related to the change of the hash value from the analysis set. Among them, the extracted features include the bit flip rate between adjacent hash values (i.e., the proportion of bits that change in the hash value under different inputs) and the similarity degree between adjacent hash values (i.e., the hash values generated for similar input data). Quantify the bit flip rate between adjacent hash values and the similarity degree between adjacent hash values, respectively generate the bit flip rate quantization value and the adjacent hash similarity quantization value, and preliminarily identify the potential patterns with insignificant hash value changes through the bit flip rate quantization value and the adjacent hash similarity quantization value.
[0058] A relatively low bit flip rate between adjacent hash values does indicate insignificant hash value changes. One of the ideal characteristics of a hash algorithm is the avalanche effect, that is, when the input data changes slightly, the output hash value should have a large range of bit flips to ensure the unpredictability and uniform distribution of the hash value. If the bit flip rate is low, it means that even when the input data changes, the different bits of the hash value still remain relatively stable, indicating that the hash algorithm fails to effectively capture the changes in the input data. This phenomenon may be caused by multiple factors, such as insufficient sensitivity of the hash algorithm to specific input patterns, high degree of structuring of the input data, low entropy value, or defects in algorithm design. A low bit flip rate may lead to too high similarity between hash values, thus increasing the risk of hash collisions, making different data may generate the same or highly similar hash values, and then reducing the reliability of the hash algorithm in security applications such as data integrity verification, digital signatures, and encrypted storage.
[0059] The specific steps for quantifying the bit flip rate between adjacent hash values to generate the bit flip rate quantization value are as follows:
[0060] For the hash values of a set of input data, construct a Bit Flip Matrix (BFM) to record the bit-by-bit changes between adjacent hash values. The constructed expression is:
[0061]
[0062] , where: represents the bit flip situation of the i th hash value at the j th bit, and respectively represent the binary values (0 or 1) of the i th and the i +1th hash values at the j th bit, Denotes the exclusive OR operation (XOR). If two bits are the same, the result is 0; if they are different, the result is 1, indicating that this bit is flipped. Is the weight function, defined as:
[0063]
[0064] , where: M Is the length of the hash value (e.g., 256 bits), mod Is the abbreviation of the modulo operation, representing the operation of taking the remainder. Calculates the bit index j Modulo the hash length M The result, that is, the index j Within the range of the hash value The loop index. If the index j Exceeds M , then it is made to loop within M By taking the modulo, without exceeding the predefined bit range. This weight function is used to enhance the influence degree of different bits, making the bits with fewer flips contribute more to the final result, thereby amplifying the influence of the low flip rate and improving the sensitivity.
[0065] This step captures the bit-by-bit changes of adjacent hash values by constructing the BFM matrix and introduces non-linear weights, making the influence of some low-frequency flipped bits more prominent, thus avoiding the smoothing problem that may be caused by simple mean processing. Through the XOR operation, the flipping situation of each bit can be accurately quantified, and the introduction of the weight function can enhance the influence of some key bits and improve the sensitivity of the calculation result to insignificant change patterns.
[0066] Based on the bit flip matrix BFM, calculate the bit flip rate quantization value, which is used to quantify the insignificance of the hash value change. The calculation expression is:
[0067]
[0068] , where: Is the bit flip rate quantization value, N Is the total number of hash samples, that is, the number of hash values in the dataset. Uses the sine function to perform non-linear scaling on the bit index, making the bit flips in the middle of the hash value contribute more to the bit flip rate quantization value, while the bit flips at the beginning and end have less influence, avoiding the omission of some patterned hash collisions. E Is the exponential adjustment factor (e.g., 1.5), which is used to amplify the influence of the lower flip rate and improve the sensitivity of the bit flip rate quantization value to hash collisions.
[0069] From the bit flip rate quantization value, it can be seen that the smaller the performance value of the bit flip rate quantization value generated by quantifying the bit flip rate between adjacent hash values, the lower the bit flip rate of the hash value, that is, the higher the similarity between adjacent hash values and the less significant the change, indicating that the hash algorithm has a weak ability to distinguish input data, which may lead to an increased risk of hash collision. On the contrary, the larger the performance value of the bit flip rate quantization value, the higher the proportion of bit flips of the hash value, indicating that a small change in the input data has caused a large-scale flip of the hash value, indicating that the hash algorithm has a good avalanche effect and can effectively distinguish different input data, improving the resistance to hash collision. Therefore, by monitoring the numerical change of the bit flip rate quantization value, the robustness of the hash algorithm can be evaluated.
[0070] A relatively high degree of similarity between adjacent hash values usually indicates that the change in the hash value is not significant, which means that the hash algorithm lacks sufficient sensitivity to small changes in the input data. Ideally, the hash algorithm should have an avalanche effect, that is, when a small change occurs in the input data, the generated hash value should undergo a large-scale bit flip to ensure the randomness and unpredictability of the output. However, when the hash values generated by adjacent input data (i.e., there are only minor differences between the input data) are highly similar, it indicates that the hash algorithm does not produce a sufficient large change when processing these data, which may lead to an increased risk of hash collision. For example, for two input data that differ by only one character or one byte, if the Hamming distance of their hash values is too small, it indicates that the distribution characteristics of the hash algorithm are insufficient and it may not be sensitive enough to structured data, low-entropy data, or patterned data. This situation may be exploited by attackers to construct forged data to deceive the integrity verification based on hash values.
[0071] The specific steps for quantifying the similarity degree between adjacent hash values to generate the adjacent hash similarity quantization value are as follows:
[0072] When evaluating the similarity degree of adjacent hash values, it is necessary to consider the change situation of each bit, and different weights are assigned to different bits to highlight the stable pattern. For this purpose, the normalized weighted Hamming similarity is used to measure the overall similarity of the hash value, and the calculation expression of the weighted Hamming similarity is:
[0073]
[0074] , where: represents the normalized weighted Hamming similarity, and respectively represent the binary numbers (0 or 1) at the k -th bit of the hash values calculated from adjacent input data, is the dynamic weight of the k -th bit, which is calculated by the entropy weight method and is defined as follows: , where, Indicates the probability that the bit appears as 0 or 1 in multiple hash values. A lower entropy value means less variation in this bit, so a higher weight is assigned;
[0075] The role of this step is to quantify the bit-level similarity between adjacent hash values through the normalized weighted Hamming similarity, and focus on the bits with less variation to detect whether the hash value is sensitive enough to small changes in the input data. If the normalized weighted Hamming similarity is close to 1, it means that the adjacent hash values change little, and the hash algorithm may not fully disperse the input data, resulting in hash collisions or security risks.
[0076] According to the normalized weighted Hamming similarity to calculate the adjacent hash similarity quantization value. The adjacent hash similarity quantization value reflects the similarity of adjacent hash values. The specific calculation expression is:
[0077]
[0078] , where represents the adjacent hash similarity quantization value, the maximum value of the normalized weighted Hamming similarity in all calculations of adjacent hash values;
[0079] The closer the value of the adjacent hash similarity quantization value is to 1, the higher the similarity of adjacent hash values, and the less significant the change in hash values; conversely, the closer the value is to 0, the more significant the change in hash values;
[0080] By calculating the adjacent hash similarity quantization value, the degree of change of hash values can be further quantified, which helps to analyze the stability and adaptability of the hash algorithm, especially when facing similar or highly structured input data. This also provides valuable metrics for further optimizing the hash algorithm to ensure that the hash algorithm can maintain good performance under different data change conditions.
[0081] From the adjacent hash similarity quantization value, it can be seen that the larger the performance value of the adjacent hash similarity quantization value generated by quantifying the similarity degree between adjacent hash values, the higher the similarity between adjacent hash values, which means that the change in hash values is not significant. In other words, when the adjacent hash similarity quantization value is close to 1, it means that even if there are small differences between input data, the generated hash values change little, and the hash algorithm is not sensitive enough to these small input changes, which may lead to an increased risk of collisions. When the adjacent hash similarity quantization value is close to 0, it indicates that the difference between adjacent hash values is large, and the hash values can significantly respond to the input changes, indicating that the hash algorithm has good anti-collision performance and adaptability. Therefore, the adjacent hash similarity quantization value is an important indicator to measure the response degree of hash values to input changes. The smaller the value, the more significant the change in hash values, and vice versa.
[0082] The quantized features are used as feature vectors and input into the trained machine learning model to predict the change in the output result of the hash value through the model.
[0083] The quantized bit flip rate quantization value and the adjacent hash similarity quantization value after quantization processing are used as feature vectors and input into the trained machine learning model. Based on the prediction of the machine learning model, a change response index is generated, and the change in the output result of the hash value is predicted through the change response index.
[0084] The trained machine learning model refers to a model that has undergone a training process, has learned on a large amount of data, and can effectively make accurate predictions based on the input features. In this scenario, the training process refers to using the quantized bit flip rate quantization value and the adjacent hash similarity quantization value after quantization processing as input features and inputting them into the machine learning model for training. During the training process, the model repeatedly learns the relationship between different features in the data and the target output, and continuously optimizes its internal parameters (such as weights and biases) to reduce the error between the prediction result and the actual value. After training, the model can make inferences on new input data and predict the change in the output result of the hash value.
[0085] The main function of the trained machine learning model in this solution is to predict the change in the hash value output by analyzing the hash value change features. These change situations usually refer to whether the hash value changes significantly due to small changes in the input data. The machine learning model uses known sample data (i.e., input features and corresponding output results) during the training process. Based on this data, the model gradually learns how to extract useful information from the bit flip rate quantization value and the adjacent hash similarity quantization value features, and finally forms a rule or model that can predict the hash value change. In this way, the model can not only capture the law of significant changes but also be sensitive to small changes. Through the predicted change response index, the machine learning model can accurately judge whether the output result will change slightly after the input data is processed by the hash algorithm, thereby improving the response ability and reliability of the hash algorithm to changes in the input data.
[0086] In a specific implementation, the trained machine learning model works through the following steps. First, it is trained using sample data that contains known patterns of hash value changes. Each sample data includes input features (such as the quantization value of the bit flip rate and the quantization value of the adjacent hash similarity) and the actual hash value change result (i.e., the degree of output change of the hash value). Through these data, the machine learning model can learn the relationship between the quantization value of the bit flip rate, the quantization value of the adjacent hash similarity, and the hash value change, and adjust its internal parameters to achieve prediction accuracy. Common machine learning algorithms, such as support vector machines (SVM), random forests, neural networks, etc., can be used for this type of prediction problem. The advantage of each algorithm lies in its different way of processing data, which may focus on capturing non-linear relationships, dimensionality reduction, or improving the classification ability of features.
[0087] The trained machine learning model has strong generalization ability. When faced with unseen data, it can still make reasonable predictions based on the patterns learned previously. For example, if the predicted change response index is high, it indicates that the hash value responds weakly to small changes in the input data, and there may be a risk of hash collision or data tampering. When the change response index is low, it means that the hash value can respond well to small changes in the input data, thereby increasing data integrity and security. Therefore, the role of the trained machine learning model is very important. It can not only effectively predict the current data but also provide optimization suggestions based on the actual situation of the hash algorithm, enhancing the performance of the hash algorithm in dealing with small changes. The effectiveness of this prediction process depends on the accurate learning of input features by the model during training and its reasonable inference about future data.
[0088] The machine learning model is not specifically limited here. Any machine learning model that can synthesize and analyze the quantization value of the bit flip rate and the quantization value of the adjacent hash similarity to generate a change response index is acceptable. To implement the technical solution of the present invention, the present invention provides a specific implementation method; the calculation formula for generating the change response index is: , where , are the weight factors of the quantization value of the bit flip rate and the quantization value of the adjacent hash similarity respectively, and , are both greater than 0. The weight factors are used to adjust the influence of the quantization value of the bit flip rate and the quantization value of the adjacent hash similarity on the final result. Specifically, the weight factors and respectively determine the contribution ratios of these two input features in the calculation of the change response index. By setting appropriate weights, the model can adjust their influences according to the different sensitivities of the bit flip rate quantization value and the adjacent hash similarity quantization value to the change of the hash value, so as to more accurately predict the performance of the hash value under minor changes. For example, if the bit flip rate quantization value can better reflect the response of the hash value to minor changes, it can be set to a larger value; if the adjacent hash similarity quantization value is more important for the prediction result, then adjust to enhance its influence. These weight factors are usually learned during the training process of the machine learning model to ensure that the final change response index can effectively reflect the degree of change of the input data.
[0089] It can be seen from the change response index that the smaller the performance value of the bit flip rate quantization value generated by quantifying the bit flip rate between adjacent hash values, and the larger the performance value of the adjacent hash similarity quantization value generated by quantifying the similarity degree between adjacent hash values, that is, the smaller the performance value of the change response index generated when the machine learning model completed through training predicts the change situation of the hash value output result, it indicates that the change of the hash value is less significant, and vice versa indicates that the change of the hash value is more significant.
[0090] Compare and analyze the change response index generated when the machine learning model completed through training predicts the change situation of the hash value output result with the preset change response index reference threshold to predict the insignificant change situation of the hash value. The specific steps are as follows:
[0091] If the change response index is less than the change response index reference threshold, it is determined that the change of the hash value is insignificant; if the change response index is greater than or equal to the change response index reference threshold, it is determined that the change of the hash value is significant.
[0092] In the case of identifying that the change of the hash value is insignificant, combine the generated random salt value with the original data to form new input data, and then process it through the hash algorithm to generate a new hash value, improving the randomness and complexity of the hash result, avoiding insufficient change of the hash value caused by data regularity, and dynamically adjusting the hash iteration times of the salt value according to the insignificant degree of the hash value change to reduce the risk of hash collision;
[0093] In the case of identifying that the change of the hash value is insignificant, combine the generated random salt value with the original data to form new input data, and then process it through the hash algorithm to generate a new hash value. The specific steps are as follows:
[0094] After identifying that the change of the hash value is insignificant, first generate a random salt value , and combine it with the original dataD Combine to form new input data, and the specific expression is:
[0095]
[0096] , where: is an internal parameter used to control the source of randomness (which can be provided by the system clock, environmental noise, or a hardware random source), represents the high-entropy random sequence generated under the parameter , ensuring the unpredictability of the salt value, represents taking the th power of the result after performing a hash operation on the random sequence (which can be achieved by truncation or bit flipping) to further enhance the complexity of the salt value, represents concatenating and , D is the original data, is the new input to be hashed;
[0097] The role of this step is to ensure the randomness and uniqueness of the salt value and combine it with the original data into a new input, thereby breaking the existing regularity of the data and avoiding the problem of insignificant changes in the hash value, providing a higher degree of randomness for subsequent iterations.
[0098] After obtaining the new input to be hashed , perform multiple iterative calculations through a hash algorithm, and the number of iterations is determined by the following formula:
[0099]
[0100] , where: is the number of salt value hash iterations, which determines the number of rounds of repeated hash processing for the new input to be hashed , is the change response exponent, representing a measure of the degree of insignificant change in the hash value. The larger the value, the less obvious the change, is the change response exponent reference threshold, is the base multiplication coefficient, used to control the overall iteration intensity, is the exponential amplification parameter, which performs a non-linear transformation on the degree of insignificant change , represents rounding down to ensure that the number of iterations is an integer, is the logarithmic adjustment factor, which amplifies the impact of insignificant changes on the number of iterations;
[0101] The role of this step is: if the change in the hash value is not significant is relatively large, then It increases correspondingly, thereby increasing the number of iterations, making the hash calculation more complex, and significantly enhancing the difficulty of hash collision;
[0102] If the hash value changes significantly is small, the number of iterations is relatively reduced, saving computing resources and maintaining high security. Through this dynamic adjustment, a balance can be achieved between security and performance, reducing the risk of hash collision and enhancing data integrity.
[0103] Through the above data integrity verification method based on the hash algorithm, the problem that the hash value does not change significantly can be effectively solved, and the anti-collision ability and data integrity of the hash algorithm can be improved. By constructing an analysis set and extracting features related to the change of the hash value, combined with the prediction of the machine learning model, the potential patterns where the hash value does not change significantly can be intelligently identified. When the change of the hash value fails to effectively respond to the data change, by introducing a random salt value and dynamically adjusting the number of hash iterations of the salt value, the complexity and randomness of the hash calculation can be significantly increased, avoiding the situation where the change of the hash value is insufficient due to the low entropy or repeated patterns of the data. This solution effectively reduces the risk of hash collision by enhancing the randomness and adaptability of the hash algorithm, and ensures that when facing minor data changes, the hash value can accurately reflect these changes, thereby enhancing the reliability and security of data integrity verification and improving the system's defense ability against potential attacks.
[0104] The present invention provides a Figure 2 data integrity verification system based on the hash algorithm as shown, including a hash value generation module, a feature extraction and quantization module, a machine learning prediction module, a salt value generation and hash enhancement module, and a dynamic adjustment and collision protection module;
[0105] The hash value generation module processes the original input data through the hash algorithm to generate a hash value with a fixed length, collects the generated hash value and constructs an analysis set, including the hash values corresponding to all input data;
[0106] The feature extraction and quantization module, after the construction of the analysis set is completed, extracts the features related to the change of the hash value from it, performs quantization processing on the extracted features, and initially identifies the potential patterns where the hash value does not change significantly;
[0107] The machine learning prediction module inputs the quantized features as feature vectors into the trained machine learning model, and predicts the change of the hash value output result through the model;
[0108] The salt value generation and hash enhancement module, when it is identified that the hash value does not change significantly, combines the generated random salt value with the original data to form new input data, and then processes it through the hash algorithm to generate a new hash value, enhancing the randomness and complexity of the hash result;
[0109] The dynamic adjustment and collision prevention module dynamically adjusts the hash iteration times of the salt value according to the insignificance degree of the change in the hash value, reducing the hash collision risk.
[0110] The data integrity verification method based on the hash algorithm provided by the embodiments of the present invention is implemented through the above data integrity verification system based on the hash algorithm. For the specific methods and processes of the data integrity verification system based on the hash algorithm, please refer to the embodiments of the data integrity verification method based on the hash algorithm above, which will not be elaborated here.
[0111] The above formulas are all dimensionless and take their numerical calculations. The formula is obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.
[0112] Only some exemplary embodiments of the present invention have been described above by way of illustration. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of the claims of the present invention.
[0113] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0114] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.
[0115] In several embodiments provided by the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0116] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0117] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0118] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0119] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A data integrity verification method based on a hash algorithm, characterized in that: The following steps are involved: The original input data is processed by a hash algorithm to generate a hash value of a fixed length. The generated hash values are collected and an analysis set is constructed, which contains the hash values corresponding to all the input data. After the analysis set is constructed, the features related to the hash value changes are extracted from it, and the extracted features are quantified to preliminarily identify potential patterns in which the hash value changes are not significant; The quantized features are input as feature vectors into the trained machine learning model, and the model is used to predict the changes in the hash value output results. When it is identified that the hash value change is not significant, the generated random salt value is combined with the original data to form new input data, which is then processed through the hash algorithm to generate a new hash value, thereby improving the randomness and complexity of the hash result. The number of hash iterations of the salt value is dynamically adjusted according to the insignificant degree of hash value change to reduce the risk of hash collision. The quantized bit flip rate quantization value and the adjacent hash similarity quantization value after quantization are input into the trained machine learning model as feature vectors, and a change response index is generated based on the prediction of the machine learning model. The change of the hash value output result is predicted by the change response index.
2. A data integrity verification method based on a hash algorithm according to claim 1, characterized in that: The original input data is processed by a hash algorithm to generate a hash value of fixed length, specifically: The hash algorithm maps the input data to a fixed-length hash value. The hash value is a string of numbers, and the same input always produces the same hash value.
3. The data integrity verification method based on the hash algorithm according to claim 1 is characterized in that: Features related to hash value changes are extracted from the analysis set, wherein the extracted features include the bit flip rate between adjacent hash values and the similarity between adjacent hash values. The bit flip rate between adjacent hash values and the similarity between adjacent hash values are quantified to generate a bit flip rate quantization value and an adjacent hash similarity quantization value, respectively. Potential patterns in which hash value changes are not significant are preliminarily identified through the bit flip rate quantization value and the adjacent hash similarity quantization value.
4. The data integrity verification method based on the hash algorithm according to claim 1 is characterized in that: The change response index generated when the trained machine learning model predicts the change of the hash value output result is compared and analyzed with the pre-set change response index reference threshold to predict the insignificant change of the hash value. The specific steps are as follows: if the change response index is less than the change response index reference threshold, the hash value change is determined to be insignificant; if the change response index is greater than or equal to the change response index reference threshold, the hash value change is determined to be significant.
5. A data integrity verification method based on a hash algorithm according to claim 4, characterized in that: When it is recognized that the hash value does not change significantly, the generated random salt value is combined with the original data to form new input data, which is then processed by the hash algorithm to generate a new hash value. The specific steps are as follows: After identifying that the hash value changes are not significant, first generate a random salt value and compare it with the original data D Combine to form new input data. The specific expression is: , in: is an internal parameter used to control the source of randomness, Indicated in the parameter The high entropy random sequence generated under the hood ensures the unpredictability of the salt value. Indicates that after performing a hash operation on a random sequence, the result is taken The power further increases the complexity of the salt value. Indicates that and Cascade, D is the original data, Enter the new hash. After getting the new hash input Afterwards, multiple iterations are performed through the hash algorithm, and the number of iterations is determined by the following formula: , in: The number of salt hash iterations determines the number of new hash inputs. The number of rounds of repeated hashing, is the change response index, which is a measure of the degree to which the hash value changes are not significant. is the reference threshold of the change response index, is the basic multiplication factor, used to control the overall iteration intensity, is the exponential amplification parameter, which has no significant effect on the change Perform nonlinear transformations, Indicates rounding down to ensure that the number of iterations is an integer. is a logarithmic adjustment factor that amplifies the effect of insignificant changes on the number of iterations.
6. A data integrity verification method based on a hash algorithm according to claim 3, characterized in that: The specific steps of quantizing the bit flip rate between adjacent hash values to generate the bit flip rate quantization value are as follows: For the hash value of a set of input data, a bit flip matrix is constructed to record the bit-by-bit changes between adjacent hash values. The constructed expression is: , in: Indicates i The hash value is j The bit flip condition of the bit, and Respectively represent i and i +1 hash value in j The binary value of the bit, It represents XOR operation. If the two bits are the same, the result is 0. If they are different, the result is 1, indicating that the bit is flipped. is the weight function, defined as: , in: M is the length of the hash value, mod It is the abbreviation of modulus operation, which means the operation of taking the remainder. The bit index is calculated j Hash length M The result after modulo, that is, the index j In the hash value range The loop index within, if the index j Exceed M , then by taking the modulus M The range is cycled without exceeding the predetermined bit range, which is used to enhance the influence of different bits; The bit flip rate quantization value is calculated based on the bit flip matrix BFM to quantify the insignificance of the hash value change. The calculation expression is: , in: is the bit flip rate quantized value, N is the total number of hash samples, that is, the number of hash values in the data set, The sine function is used to nonlinearly scale the bit index, so that the bit flip in the middle of the hash value contributes more to the quantized value of the bit flip rate, while the bit flip at the beginning and end has less impact, avoiding the patterned hash collision from being ignored. E It is an exponential adjustment factor used to amplify the impact of low flip rate and increase the sensitivity of the bit flip rate quantization value to hash collision.
7. The data integrity verification method based on the hash algorithm according to claim 3 is characterized in that: The specific steps of quantifying the similarity between adjacent hash values to generate adjacent hash similarity quantization values are as follows: Normalized weighted Hamming similarity is used to measure the overall similarity of hash values. The calculation expression of weighted Hamming similarity is: , in: represents the normalized weighted Hamming similarity, and Respectively represent the hash values calculated from adjacent input data in the k Binary number on bit, For the k The dynamic weight of a bit is calculated using the entropy weight method and is defined as follows: ,in, Indicates the probability of the bit appearing as 0 or 1 in multiple hash values; According to the normalized weighted Hamming similarity To calculate the adjacent hash similarity quantization value, the adjacent hash similarity quantization value reflects the similarity of adjacent hash values. The specific calculation expression is: , in, Represents the quantized value of adjacent hash similarity, The maximum value of the normalized weighted Hamming similarity among all adjacent hash value calculations.
8. A data integrity verification system based on a hash algorithm, used to implement the data integrity verification method based on a hash algorithm as described in any one of claims 1 to 7, characterized in that: It includes hash value generation module, feature extraction and quantification module, machine learning prediction module, salt value generation and hash enhancement module, and dynamic adjustment and collision protection module; The hash value generation module processes the original input data through a hash algorithm to generate a hash value of a fixed length, collects the generated hash values and constructs an analysis set, which contains the hash values corresponding to all the input data; The feature extraction and quantification module extracts features related to hash value changes from the analysis set after it is built, quantifies the extracted features, and preliminarily identifies potential patterns in which hash value changes are not significant; The machine learning prediction module inputs the quantized features as feature vectors into the trained machine learning model, and uses the model to predict the changes in the hash value output results; The salt value generation and hash enhancement module, when it is recognized that the hash value change is not significant, combines the generated random salt value with the original data to form new input data, and then processes it through the hash algorithm to generate a new hash value, thereby improving the randomness and complexity of the hash result; The dynamic adjustment and collision protection module dynamically adjusts the number of hash iterations of the salt value according to the insignificant degree of hash value change to reduce the risk of hash collision.
Citation Information
Patent Citations
Communication data secure transmission system based on block chain
CN119030774A