Decentralized Talent Background Investigation Data Storage and Verification Method and System
The Guomi algorithm SM2 and SM3 hash algorithms generate original data packets with timing weights, combined with the deep learning model and the three-layer Merkel tree structure, the problems of data timing management and verification mechanism in the blockchain talent background investigation system are solved, and the secure storage, credibility and traceability are improved.
Patent Information
- Application Number
- CN202510339190.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing blockchain talent background investigation system lacks effective management of data timing, cannot accurately reflect the time continuity of candidate education and work experience, there are information faults or timing disorders, the verification mechanism lacks dynamic assessment of node trust, cannot effectively prevent the verification behavior of malicious nodes, and the data update lacks flexibility, resulting in confusion in version management.
The national secret algorithm SM2 is used for asymmetric encryption and SM3 hashing algorithm to generate original data packets with timing weights, combined with the deep learning model, and built a hierarchical timing dependency graph, and wrote it to the blockchain network through a three-layer Merkel tree structure. Based on the feature fingerprint identification and real-time node trust scoring mechanism, the consistency verification of zero-knowledge proof is used to generate tamper-proof verification certificates, activate the two-layer consensus smart contract to establish a version control chain, and build a traceability network.
It realizes the secure storage and efficient management of talent background information, improves the credibility and verification efficiency of background investigation data, ensures the accuracy and reliability of verification results, realizes the traceability of the entire data update process, and maintains the transparency and credibility of information updates.
Smart Images

Figure CN119892491B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to background investigation technology, and in particular to a method and system for storing and verifying decentralized talent background investigation data. Background Art
[0002] With the rapid development of the talent market, enterprises need to comprehensively verify the educational background, work experience, and professional skills of candidates during the recruitment process. Traditional talent background investigations mainly rely on manual verification or third-party investigation agencies, which have problems such as low efficiency and high costs. In recent years, with the development of blockchain technology, decentralized data storage and verification solutions have gradually been applied to the field of talent background investigation. Currently, blockchain talent background investigation systems usually use distributed ledger technology to record candidate information, realize data verification through smart contracts, and use encryption algorithms to protect data security.
[0003] Existing blockchain talent background investigation systems lack effective management of data timeliness, cannot accurately reflect the time continuity of candidates' education and work experience, and are prone to information gaps or time sequence chaos, affecting the accuracy of background investigations.
[0004] Traditional blockchain data storage solutions lack a flexible processing mechanism for data updates. When it is necessary to update or correct the stored information, it is often necessary to reconstruct the entire data structure, which is not only inefficient but also may lead to chaos in data version management.
[0005] Existing verification mechanisms generally use simple consensus algorithms, lack dynamic evaluation of the trustworthiness of verification nodes, cannot effectively prevent the verification behavior of malicious nodes, and also lack a complete data traceability mechanism, making it difficult to ensure the credibility and traceability of verification results. Summary of the Invention
[0006] Embodiments of the present invention provide a method and system for storing and verifying decentralized talent background investigation data, which can solve the problems in the prior art.
[0007] In the first aspect of the embodiments of the present invention,
[0008] A method for storing and verifying decentralized talent background investigation data is provided, including:
[0009] Receiving, based on an identity authentication layer, educational experience information, work experience information, and skill certification information submitted by a candidate, asymmetrically encrypting the received information through the national cryptographic algorithm SM2 to obtain a first encrypted data set, and generating an original data packet with time sequence weights for the first encrypted data set in combination with the SM3 hashing algorithm;
[0010] Segment the first encrypted dataset according to the timing weights in the original data packet, construct a hierarchical timing dependency graph. The hierarchical timing dependency graph extracts the timing correlation features between data segments through a deep learning model to generate a second encrypted dataset. The second encrypted dataset contains weighted feature fingerprint identifiers and information segment hash values, and writes the second encrypted dataset into the trusted node group of the blockchain network through a three-layer Merkle tree structure;
[0011] When the trusted node group receives a verification request, locate the target data position based on the feature fingerprint identifier, extract the second encrypted dataset of the target time period according to the hierarchical timing dependency graph, and calculate the real-time trust score of the node at the same time. The real-time trust score of the node is dynamically updated based on the verification historical behavior feature vector of the node. When the overall trust level of the node group reaches the adaptive threshold, start the consistency verification mechanism based on zero-knowledge proof using the feature fingerprint identifier, and generate a tamper-proof verification certificate containing timing integrity proof for the verification result based on the consistency verification mechanism;
[0012] When the hierarchical timing dependency graph detects a data update requirement, calculate the correlation degree between the update request information and the timing weights in the original data packet, and evaluate the update rationality by the deep learning model. After confirmation, locate the data segment to be updated in the second encrypted dataset, activate the double-layer consensus smart contract to establish a version control chain, construct a traceability network in the trusted node group based on the feature fingerprint identifier, and fuse the updated data with the original timing dependency graph through a three-layer Merkle tree structure to generate a new verification link, realizing the traceability of the entire data update process.
[0013] Perform asymmetric encryption on the received information through the national cryptography algorithm SM2 to obtain the first encrypted dataset. Combine the SM3 hash algorithm to generate an original data packet with timing weights for the first encrypted dataset, including:
[0014] Perform asymmetric encryption on the received education experience, work experience, and skill certification information through the national cryptography algorithm SM2. Select the base point G and order n based on the SM2 elliptic curve cryptosystem to generate a key pair (d, Q). Use the public key Q to perform encryption operations on the normalized information, and combine the random number k to generate the first encrypted dataset. The first encrypted dataset contains encrypted structured information;
[0015] Combine the SM3 hash algorithm to perform security processing on the first encrypted dataset. Pad the encrypted data to a multiple of 512 bits, perform iterative operations through the compression function, and introduce timing weight parameters to calculate the timestamp difference and information correlation degree between adjacent data blocks, generating an original data packet with timing weights. The original data packet has the characteristics of timing integrity.
[0016] Asymmetrically encrypt the received education experience, work experience, and skill certification information using the national cryptography algorithm SM2. Based on the SM2 elliptic curve cryptosystem, select the base point G and order n to generate a key pair (d, Q). Use the public key Q to perform an encryption operation on the normalized information, and combine with a random number k to generate the first encrypted dataset, including:
[0017] Receive the education experience information, work experience information, and skill certification information submitted by the user. Classify and encode the education experience information, work experience information, and skill certification information respectively to obtain the structured data of education experience, structured data of work experience, and structured data of skill certification. Integrate the structured data of education experience, structured data of work experience, and structured data of skill certification to construct a standardized data template;
[0018] Under the SM2 elliptic curve cryptosystem, select the base point G and order n. Use a random number generator to obtain the private key d. Perform a point multiplication operation on the private key d and the base point G to generate the public key Q, obtaining the key pair (d, Q). Perform a validity check on the key pair (d, Q) through a key validity verification algorithm;
[0019] Generate a random number k. Use the random number k to perform a point multiplication operation with the base point G to obtain a temporary public key. Perform a point multiplication operation on the temporary public key and the public key Q to obtain a session key. Use a key derivation function to process the session key to generate an encryption key. Use the encryption key to encrypt the standardized data template to generate the first encrypted dataset containing the encrypted data and its corresponding timestamp.
[0020] Extract the second encrypted dataset for the target time period according to the hierarchical time-dependent graph. At the same time, calculate the real-time trust score of the node. The real-time trust score of the node is dynamically updated based on the verification historical behavior feature vector of the node. When the overall trust level of the node group reaches the adaptive threshold, use the feature fingerprint identifier to start the zero-knowledge proof-based consistency verification mechanism. Generate a tamper-proof verification certificate containing the time-sequence integrity proof for the verification result based on the consistency verification mechanism, including:
[0021] The hierarchical time-dependent graph contains a set of time-sequence nodes and a set of time-sequence edges. Extract the second encrypted dataset for the target time period according to the hierarchical time-dependent graph. The second encrypted dataset contains the encrypted data within the target time period and its corresponding time-sequence dependency relationship; Establish a verification historical behavior feature vector for the nodes in the set of time-sequence nodes. Calculate the real-time trust score of the node based on the verification historical behavior feature vector. The real-time trust score of the node is dynamically adjusted as the verification historical behavior feature vector is updated;
[0022] The real-time trust scores of each node in the time series node set are weighted and summed to obtain the overall trust degree of the node group. The overall trust degree of the node group is compared with the adaptive threshold. When the overall trust degree of the node group is greater than the adaptive threshold, a feature fingerprint identifier is generated based on the overall trust degree of the node group and the adaptive threshold;
[0023] Based on the feature fingerprint identifier, a prover commitment is constructed. The prover commitment and the verification rule are combined to generate a zero-knowledge proof. The zero-knowledge proof is used to start the consistency check mechanism to perform consistency check on the second encrypted data set and generate a verification result; According to the verification result, a Merkle tree is constructed to generate a time series integrity proof including the Merkle root node and the verification path. An anti-tampering verification certificate is generated based on the time series integrity proof.
[0024] Constructing a Merkle tree according to the verification result to generate a time series integrity proof including the Merkle root node and the verification path, and generating an anti-tampering verification certificate based on the time series integrity proof includes:
[0025] Hash calculation is performed on each data block in the verification result and the corresponding timestamp sequence to obtain the data block hash value, and the data block hash value and the index information are combined to generate leaf nodes; Pairwise calculation of the leaf nodes is performed to obtain the intermediate node hash value, and the intermediate node hash values are recursively paired and calculated until the root node hash value is generated, and a Merkle tree structure is constructed based on the root node hash value;
[0026] Extract the verification path node set and the path direction sequence in the Merkle tree structure, combine the verification path node set and the path direction sequence to generate a verification path, and combine the verification path and the root node hash value to generate a time series integrity proof;
[0027] An anti-tampering verification certificate is constructed based on the time series integrity proof. The anti-tampering verification certificate includes certificate header information and certificate body content; Hash calculation is performed on the certificate header information and the certificate body content of the anti-tampering verification certificate to obtain a certificate digest, and the certificate digest is signed with a private key to obtain a digital signature, and the digital signature is added to the anti-tampering verification certificate to complete the certificate generation.
[0028] Activate the double-layer consensus smart contract to establish a version control chain. Based on the feature fingerprint identifier, a traceability network is constructed in the trusted node group. The updated data and the original time series dependency graph are fused through a three-layer Merkle tree structure to generate a new verification link, realizing the traceability of the entire process of data update, including:
[0029] Construct the main consensus layer contract and the verification consensus layer contract. The main consensus layer contract records data version control information, and the verification consensus layer contract records data update verification information. Activate the main consensus layer contract and the verification consensus layer contract to form a double-layer consensus smart contract, and establish a version control chain. Perform a hash calculation on the data block to obtain the data block hash value, combine the data block hash value with the timestamp sequence to generate a feature fingerprint identifier, calculate the fingerprint similarity between nodes based on the feature fingerprint identifier, and select trusted nodes according to the fingerprint similarity to construct a traceability network.
[0030] Construct a three-layer Merkle tree structure in the traceability network. Write the updated data information into the data layer of the three-layer Merkle tree structure to generate end nodes, and calculate the temporal dependence relationship based on the end nodes to generate intermediate nodes and write them into the dependence layer of the three-layer Merkle tree structure. Calculate the verification information for the intermediate nodes to obtain the root node and write it into the verification layer of the three-layer Merkle tree structure. Establish a data update path based on the end nodes, intermediate nodes, and root nodes in the three-layer Merkle tree structure.
[0031] Fuse the updated data with the original temporal dependence graph through the data update path to generate a new verification link containing node verification records. The new verification link records the complete process of data update. When receiving a data update request, the main consensus layer contract triggers the verification consensus layer contract to perform verification. After the verification passes, record the update information in the version control chain, and realize the traceability of the entire data update process based on the new verification link.
[0032] Performing a hash calculation on the data block to obtain the data block hash value, combining the data block hash value with the timestamp sequence to generate a feature fingerprint identifier, calculating the fingerprint similarity between nodes based on the feature fingerprint identifier, and selecting trusted nodes according to the fingerprint similarity to construct a traceability network includes:
[0033] Perform a hash calculation on the input data block to obtain the main data block hash value. Perform MD5 and SHA1 hash calculations on the data block to obtain the data block auxiliary hash sequence. Combine the main data block hash value with the data block auxiliary hash sequence and add a random salt value to generate the data block hash value. Collect the timestamp information of the data block at different processing nodes, arrange the timestamp information in chronological order to generate a timestamp sequence, perform encoding processing on the timestamp sequence to obtain a time feature vector, and perform a weighted combination operation on the data block hash value and the time feature vector to generate a feature fingerprint identifier.
[0034] Construct a node feature matrix. The row vectors of the node feature matrix represent the node identifiers of the feature fingerprints, and the column vectors represent different feature dimensions of the feature fingerprints. Calculate the distances between different row vectors in the node feature matrix based on the locality-sensitive hashing algorithm to obtain the initial similarity; perform weighted fusion on the initial similarity and the historical behavior credibility of the nodes to obtain the fingerprint similarity between the nodes. The historical behavior credibility is calculated based on the historical interaction records and verification passing rates of the nodes;
[0035] Set a dynamic node similarity threshold, which is adaptively adjusted according to the network scale and security level. Compare the fingerprint similarity with the dynamic node similarity threshold, and select the nodes with fingerprint similarity higher than the dynamic node similarity threshold as trusted nodes; calculate the trust transfer strength between the trusted nodes. The trust transfer strength is jointly determined by the direct similarity and indirect similarity between the nodes, and construct a weighted connection relationship between the nodes based on the trust transfer strength;
[0036] Construct a hierarchical node traceability network using the weighted connection relationship. The hierarchical node traceability network includes a core trusted layer and an extended verification layer. The core trusted layer consists of the nodes with the highest fingerprint similarity, and the extended verification layer consists of other trusted nodes.
[0037] In the second aspect of the embodiments of the present invention,
[0038] Provide an electronic device, including:
[0039] A processor;
[0040] A memory for storing instructions executable by the processor;
[0041] Wherein, the processor is configured to call the instructions stored in the memory to execute the foregoing method.
[0042] In the third aspect of the embodiments of the present invention,
[0043] Provide a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the foregoing method is implemented.
[0044] The beneficial effects of this application are as follows:
[0045] 1. Perform asymmetric encryption through the national cryptographic algorithm SM2 and generate an original data packet with time-series weights using the SM3 hashing algorithm. Combine a deep learning model to construct a hierarchical time-series dependency graph, and use a three-layer Merkle tree structure to write data into the blockchain network, realizing the secure storage and efficient management of talent background information, and effectively preventing data from being tampered with or forged.
[0046] 2. Based on the feature fingerprint identification and node real-time trust scoring mechanism, combined with the consistency verification of zero-knowledge proof, an anti-tampering verification certificate containing temporal integrity proof is generated, improving the credibility and verification efficiency of background investigation data, and ensuring the accuracy and reliability of verification results.
[0047] 3. By evaluating the rationality of data updates through a deep learning model, establishing a version control chain using a two-layer consensus smart contract, and constructing a traceability network based on feature fingerprint identification, the traceability of the entire process of data updates is achieved, which not only ensures the timeliness of data but also maintains the transparency and credibility of information updates. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a flowchart of the method for decentralized storage and verification of talent background investigation data according to an embodiment of the present invention;
[0049] Figure 2 It is a schematic diagram for evaluating the verification accuracy rate according to an embodiment of the present invention;
[0050] Figure 3 It is a schematic diagram for comparing the hash calculation performance and the feature fingerprint generation efficiency according to an embodiment of the present invention;
[0051] Figure 4 It is a schematic diagram for node trust evaluation according to an embodiment of the present invention;
[0052] Figure 5 It is a schematic diagram of the timestamp sequence information panel according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0055] Figure 1 It is a flowchart of the method for decentralized storage and verification of talent background investigation data according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0056] Based on the identity authentication layer, the received educational experience information, work experience information, and skill certification information submitted by the candidate are asymmetrically encrypted by the national cryptography algorithm SM2 to obtain the first encrypted data set. Combining with the SM3 hashing algorithm, an original data packet with time-sequence weights is generated for the first encrypted data set;
[0057] The first encrypted data set is segmented according to the time-sequence weights in the original data packet to construct a hierarchical time-sequence dependence graph. The hierarchical time-sequence dependence graph extracts the time-sequence correlation features between data segments through a deep learning model to generate the second encrypted data set. The second encrypted data set includes weighted feature fingerprint identifiers and information segment hash values, and the second encrypted data set is written into the trusted node group of the blockchain network through a three-layer Merkle tree structure;
[0058] When the trusted node group receives a verification request, it locates the target data position based on the feature fingerprint identifier, extracts the second encrypted data set of the target time period according to the hierarchical time-sequence dependence graph, and calculates the real-time trust score of the node at the same time. The real-time trust score of the node is dynamically updated based on the verification historical behavior feature vector of the node. When the overall trust level of the node group reaches the adaptive threshold, the zero-knowledge proof-based consistency verification mechanism is started using the feature fingerprint identifier, and a tamper-proof verification certificate including time-sequence integrity proof is generated for the verification result based on the consistency verification mechanism;
[0059] When the hierarchical time-sequence dependence graph detects a data update requirement, the correlation degree between the update request information and the time-sequence weights in the original data packet is calculated, and the update rationality is evaluated by the deep learning model. After confirmation, the data segment to be updated is located in the second encrypted data set, the double-layer consensus intelligent contract is activated to establish a version control chain, a traceability network is constructed in the trusted node group based on the feature fingerprint identifier, and the updated data is fused with the original time-sequence dependence graph through a three-layer Merkle tree structure to generate a new verification link, realizing the traceability of the entire data update process.
[0060] In an optional implementation manner, asymmetrically encrypting the received information by the national cryptography algorithm SM2 to obtain the first encrypted data set, and generating an original data packet with time-sequence weights for the first encrypted data set by combining with the SM3 hashing algorithm includes:
[0061] Asymmetrically encrypt the received educational experience, work experience, and skill certification information by the national cryptography algorithm SM2. Based on the SM2 elliptic curve cryptosystem, select the base point G and the order n to generate the key pair (d, Q), use the public key Q to perform the encryption operation on the normalized information, and combine with the random number k to generate the first encrypted data set. The first encrypted data set includes the encrypted structured information;
[0062] The first encrypted data set is securely processed in combination with the SM3 hash algorithm. The encrypted data is padded with messages in multiples of 512 bits, and iterative operations are performed through the compression function. At the same time, a timing weight parameter is introduced to calculate the timestamp difference and information correlation between adjacent data blocks, and an original data packet with a timing weight is generated. The original data packet has a timing integrity feature.
[0063] First, the received education experience, work experience and skill certification information is standardized and preprocessed. The unstructured text information is converted into structured data in a unified format, including key fields such as time, location, institution, major, position, etc. Data cleaning and format normalization are performed on each field to generate a standardized information data set.
[0064] Then, the key is generated based on the SM2 elliptic curve cryptography. Select the elliptic curve parameters, determine the base point G and the order n. Randomly generate the private key d and calculate the public key Q. Generate a key pair with 256 bits as the security strength. For example, select the elliptic curve parameters as the recommended curve sm2p256v1 to generate a 256-bit private key and a 512-bit public key.
[0065] Then, the standardized information is encrypted using the SM2 public key encryption algorithm. The information is grouped into fixed-length data blocks, each of which is 128 bytes. For each data block, a random number k is generated, and encryption calculation is performed in combination with the public key Q to obtain a ciphertext point. All ciphertext points are combined to form the first encrypted data set. For example, after encrypting an educational experience information "2018-2022 Computer Science Peking University", a 512-byte ciphertext is obtained.
[0066] The first encrypted data set is then processed using the SM3 hash algorithm. First, message padding is performed to pad the data length to an integer multiple of 512 bits. The compression function is then iterated, processing 512-bit data blocks each time. At the same time, a timing weight is introduced to calculate the timestamp difference between adjacent data blocks. For example, if the time difference between two educational experiences is 4 years, the weight is set to 0.8. Finally, a 256-bit hash value is generated.
[0067] Finally, the hash value is combined with the timing weight to generate an original data packet with timing characteristics. The data packet contains information such as encrypted data, hash value, timing weight, etc. The timing integrity verification of data is achieved through timing weight. For example, the generated data packet size is 1024 bytes, containing 512 bytes of ciphertext, 256-bit hash value and 64-bit timing weight.
[0068] Implement secure storage of data through SM2 asymmetric encryption, protect the privacy information of users from being leaked, and improve the security of data transmission and storage. Combine the SM3 hashing algorithm to generate data fingerprints, ensure data integrity, prevent data tampering, and enhance data credibility and anti-counterfeiting capabilities. Introduce a time-series weight mechanism to implement time-series integrity verification of data, effectively prevent data replay attacks, and ensure the time-series correctness and consistency of data.
[0069] In an optional implementation, the received education experience, work experience, and skill certification information are asymmetrically encrypted using the national cryptographic algorithm SM2. Based on the SM2 elliptic curve cryptosystem, the base point G and the order n are selected to generate a key pair (d, Q). The public key Q is used to perform an encryption operation on the standardized information, and a first encrypted data set is generated in combination with the random number k, including:
[0070] Receive the education experience information, work experience information, and skill certification information submitted by the user, classify and encode the education experience information, work experience information, and skill certification information respectively to obtain education experience structured data, work experience structured data, and skill certification structured data, and integrate the education experience structured data, work experience structured data, and skill certification structured data to construct a standardized data template;
[0071] Under the SM2 elliptic curve cryptosystem, select the base point G and the order n, use a random number generator to obtain the private key d, perform a point multiplication operation on the private key d and the base point G to generate the public key Q, obtain the key pair (d, Q), and perform a validity check on the key pair (d, Q) through a key validity verification algorithm;
[0072] Generate a random number k, perform a point multiplication operation on the random number k and the base point G to obtain a temporary public key, perform a point multiplication operation on the temporary public key and the public key Q to obtain a session key, use a key derivation function to process the session key to generate an encryption key, and use the encryption key to encrypt the standardized data template to generate a first encrypted data set including encrypted data and its corresponding timestamp.
[0073] First, receive the education experience information, work experience information, and skill certification information submitted by the user. Classify and encode the education experience information, including fields such as education level, major category, and graduation institution, to generate education experience structured data. Classify and encode the work experience information, including fields such as work unit, position, and work years, to generate work experience structured data. Classify and encode the skill certification information, including fields such as certificate type, issuing agency, and acquisition time, to generate skill certification structured data. Integrate the three types of structured data into a standardized data template.
[0074] Under the SM2 elliptic curve cryptosystem, the coordinates of the base point G are selected as (x, y), and the order n is a 256-bit prime number. A 256-bit random number is generated by a random number generator as the private key d. The private key d is multiplied by the base point G to obtain the coordinates of the public key Q. The generated key pair (d, Q) is verified for validity to check whether the private key d is within the interval [1, n - 1] and whether the public key Q satisfies the elliptic curve equation.
[0075] Generate a random number k, multiply k by the base point G to obtain a temporary public key (x1, y1). Multiply the temporary public key by the public key Q to obtain a session key (x2, y2). Use a key derivation function KDF to process the session key coordinate values to generate an encryption key. Encrypt the standardized data template using the encryption key to obtain ciphertext data. Add a timestamp to the ciphertext data to generate the first encrypted data set.
[0076] For example: The user submits undergraduate education information, including content such as "Computer Science and Technology major, a certain university, graduated in 2020". The work experience includes content such as "Software Engineer, a certain company, 3 years of work experience". The skill certification includes content such as "Senior Programmer Certificate, a certain certification institution, obtained in 2022". After being structured and encrypted, ciphertext data is generated, and a timestamp of "2023-06-01 10:00:00" is added.
[0077] This embodiment provides a data encryption method based on the national cryptographic algorithm SM2. Aiming at the problems of complex key management and low security in the prior art where simple symmetric encryption methods are generally used to process personal information submitted by users, as well as the defect that subsequent processing is difficult due to inconsistent data structures.
[0078] This embodiment first classifies, codes, and structures the education experience, work experience, and skill certification information submitted by the user, and establishes a standardized data template, which solves the problems of inconsistent original data formats and difficulty in unified management and analysis. By establishing a unified data structure template, the standardization and processability of the data are improved.
[0079] Secondly, this embodiment uses the national cryptographic algorithm SM2 for asymmetric encryption, which has higher security compared to the symmetric encryption methods commonly used in the prior art. Generate a key pair through the SM2 elliptic curve cryptosystem, introduce random numbers to increase the encryption strength, generate a session key using point multiplication operations, and combine a key derivation function for key processing to form a multi-level security protection mechanism. This solution avoids the problems of key distribution and management in symmetric encryption and ensures the security of the encryption process at the same time.
[0080] Again, in this embodiment, timestamp information is added to the encrypted data to facilitate the tracking of data timeliness and version management. This design not only ensures the traceability of the data but also facilitates the system to perform data maintenance and update management.
[0081] Through the above technical solution, this embodiment realizes the standardized processing and secure encrypted storage of user personal information, improves the data management efficiency, enhances the information security protection ability, and provides a good foundation for subsequent data analysis and applications. While ensuring data security, this solution also takes into account the practicality and scalability of the system, and has good application value.
[0082] In an alternative embodiment, a second encrypted data set for a target time period is extracted according to the hierarchical time-sequence dependency graph, and at the same time, the real-time trust score of the node is calculated. The real-time trust score of the node is dynamically updated based on the feature vector of the verification historical behavior of the node. When the overall trust degree of the node group reaches the adaptive threshold, the consistency verification mechanism based on zero-knowledge proof is started by using the feature fingerprint identifier. The tamper-proof verification certificate including the time-sequence integrity proof is generated for the verification result based on the consistency verification mechanism, including:
[0083] The hierarchical time-sequence dependency graph includes a set of time-sequence nodes and a set of time-sequence edges. A second encrypted data set for a target time period is extracted according to the hierarchical time-sequence dependency graph. The second encrypted data set includes the encrypted data within the target time period and its corresponding time-sequence dependency relationship; a feature vector of the verification historical behavior is established for the nodes in the set of time-sequence nodes, and the real-time trust score of the node is calculated based on the feature vector of the verification historical behavior. The real-time trust score of the node is dynamically adjusted as the feature vector of the verification historical behavior is updated;
[0084] The real-time trust scores of the nodes in the set of time-sequence nodes are weighted and summed to obtain the overall trust degree of the node group. The overall trust degree of the node group is compared with the adaptive threshold. When the overall trust degree of the node group is greater than the adaptive threshold, a feature fingerprint identifier is generated based on the overall trust degree of the node group and the adaptive threshold;
[0085] A prover commitment is constructed based on the feature fingerprint identifier, and a zero-knowledge proof is generated by combining the prover commitment with the verification rule. The consistency verification mechanism is started by using the zero-knowledge proof to perform consistency verification on the second encrypted data set and generate a verification result; a Merkle tree is constructed according to the verification result to generate a time-sequence integrity proof including the Merkle root node and the verification path, and a tamper-proof verification certificate is generated based on the time-sequence integrity proof.
[0086] The hierarchical time-sequence dependency graph contains a set of time-sequence nodes and a set of time-sequence edges. The time-sequence nodes represent the data generation time points, and the time-sequence edges represent the dependency relationships between the data. First, traverse the hierarchical time-sequence dependency graph according to the target time period, and extract the encrypted data within this time period and its corresponding time-sequence dependency relationships to form a second encrypted data set. Specifically, for each time-sequence node within the time period, obtain its encrypted data and the dependency relationship information represented by the time-sequence edges connected to it.
[0087] Establish a verification historical behavior feature vector for each time-sequence node. The feature vector includes multi-dimensional features such as node verification accuracy, verification timeliness, and verification consistency. Calculate the real-time trust score of the node based on the feature vector, specifically by performing a weighted sum of each dimension of the feature vector. As the node participates in the verification process, dynamically update the feature vector, thereby adjusting the trust score in real time. For example, if a node's verification accuracy is 0.95, verification timeliness is 0.9, and verification consistency is 0.85, and the weights of each dimension are 0.4, 0.3, and 0.3 respectively, then the real-time trust score of this node is 0.904.
[0088] Perform a weighted sum of the real-time trust scores of each node in the set of time-sequence nodes to obtain the overall trust degree of the node group. Set the adaptive threshold to 0.85. When the overall trust degree of the node group is greater than 0.85, generate a feature fingerprint identifier based on the difference between the two. The feature fingerprint identifier uses the SHA256 hash algorithm, and the input is the difference between the overall trust degree of the node group and the adaptive threshold.
[0089] Construct a prover's commitment based on the feature fingerprint identifier, using the Pedersen commitment scheme. Combine the commitment with the preset verification rules to generate a zero-knowledge proof. Start the consistency verification mechanism to verify the data and its dependency relationships in the second encrypted data set, and generate a verification result. Construct the verification result into a Merkle tree to generate a time-sequence integrity proof containing the root node hash value and the verification path. Finally, generate an anti-tampering verification certificate based on the time-sequence integrity proof, and the certificate uses digital signature technology for anti-tampering protection.
[0090] By extracting the encrypted data and dependency relationships of the target time period through the hierarchical time-sequence dependency graph, and dynamically calculating the trust score based on the node verification historical behavior feature vector, the credibility evaluation of the data verification process is realized. By adopting the comparison mechanism of the overall trust degree of the node group and the adaptive threshold, combined with the feature fingerprint identifier and zero-knowledge proof, the security and privacy of the verification process are ensured. By constructing a time-sequence integrity proof through the Merkle tree and generating an anti-tampering verification certificate, the non-tamperability and traceability of the verification result are ensured, improving the reliability of the entire verification mechanism.
[0091] In an alternative embodiment, a Merkle tree is constructed based on the verification result, and a temporal integrity proof including the Merkle root node and the verification path is generated. Generating a tamper-proof verification certificate based on the temporal integrity proof includes:
[0092] Performing a hash calculation on each data block and the corresponding timestamp sequence in the verification result to obtain a data block hash value, and combining the data block hash value with the index information to generate a leaf node; Pairing the leaf nodes in pairs to calculate the intermediate node hash value, and recursively pairing the intermediate node hash values until the root node hash value is generated, and constructing a Merkle tree structure based on the root node hash value;
[0093] Extracting the verification path node set and the path direction sequence in the Merkle tree structure, combining the verification path node set and the path direction sequence to generate a verification path, and combining the verification path with the root node hash value to generate a temporal integrity proof;
[0094] Constructing a tamper-proof verification certificate based on the temporal integrity proof, where the tamper-proof verification certificate includes certificate header information and certificate body content; Performing a hash calculation on the certificate header information and the certificate body content of the tamper-proof verification certificate to obtain a certificate digest, signing the certificate digest with a private key to obtain a digital signature, and adding the digital signature to the tamper-proof verification certificate to complete the certificate generation.
[0095] First, process the data blocks and timestamp sequences in the verification result. Taking a verification result with a set of 10 data blocks as an example, each data block contains transaction information and the corresponding timestamp. Calculate the 32-byte hash value for each data block using the SHA-256 hash algorithm, and combine the hash value with the data block index number to form a leaf node. For example, the hash value of the first data block is "7d8f...3a2b", and the index number is "0", and the combined leaf node is "0-7d8f...3a2b".
[0096] Then pair all the leaf nodes in pairs to calculate the intermediate nodes. Concatenate the hash values of two adjacent leaf nodes and perform the SHA-256 hash operation again to obtain the parent node hash value. For example, the first-level leaf nodes "0-7d8f...3a2b" and "1-9e4c...8f1d" are paired and calculated to obtain the parent node "01-5a3d...2c9e". Calculate recursively in this way until the root node hash value "root-4b2a...9f7d" is finally obtained.
[0097] Extract the verification path information in the Merkle tree structure. On the path from the leaf node to the root node, collect the hash values of sibling nodes used for verification at each layer, and record the path direction (left 0, right 1) at the same time. For example, the verification path of the first leaf node contains three node hash values, "9e4c...8f1d", "2b7a...4e8c", "8c3f...1a9d", and the direction sequence is "011". Combine the node set and the direction sequence to obtain the complete verification path.
[0098] Construct a time-sequential integrity proof based on the root node hash value and the verification path. Package the root node hash value, the verification path node set, and the path direction sequence to generate a proof data structure. The proof structure contains information such as the root node hash value "root-4b2a...9f7d", the verification path node array, and the direction sequence "011".
[0099] Finally, generate a tamper-proof verification certificate. The certificate header contains basic information such as the version number, the signing time, and the validity period. The certificate body contains the content of the time-sequential integrity proof. Calculate the SHA-256 hash of the certificate header and the body content to obtain the certificate digest, and use the RSA private key to sign the digest to obtain the digital signature. Add the signature to the certificate to complete the generation.
[0100] Figure 2 Schematic diagram for evaluating the verification accuracy of the embodiments of the present invention:
[0101] This figure compares the accuracy performance of three verification schemes under different system loads. The "technical solution of the present invention" marked with a triangle represents a hybrid verification mechanism that combines zero-knowledge proof and time-sequential integrity verification, including a node real-time trust scoring system and an adaptive threshold judgment; the "basic zero-knowledge proof scheme" marked with a circle represents a verification scheme that only uses the standard zero-knowledge proof protocol; the "traditional verification scheme" marked with a square represents a conventional blockchain verification scheme that uses digital signatures and hash chains. When the system load rate is 0.2, the technical solution of the present invention maintains a high accuracy of 99.5%, while the basic zero-knowledge proof scheme is 95.0% and the traditional verification scheme is 97.0%. When the load rate increases to 0.6, the technical solution of the present invention still maintains an excellent performance of 98.8%, while the other two schemes drop to 91.0% and 93.0% respectively. In the extreme case of the maximum load rate of 1.0, the technical solution of the present invention still maintains a high accuracy of 98.0%, significantly better than 85.0% of the basic zero-knowledge proof scheme and 87.0% of the traditional verification scheme.
[0102] This embodiment provides a method for generating a time-sequential integrity proof and a tamper-proof verification certificate based on a Merkle tree. Aiming at the problems in the prior art that usually use simple digital signatures or timestamps to verify data integrity, such as low verification efficiency, poor traceability, and difficulty in proving the time-sequential relationship of data.
[0103] In this embodiment, first, hash calculations are performed on the data blocks and timestamp sequences in the verification results, and leaf nodes are constructed in combination with index information, laying a foundation for constructing a Merkle tree subsequently. This method not only ensures the privacy of the original data but also realizes the unique identification of the data, solving the problem that it is difficult to verify data integrity in traditional methods.
[0104] Secondly, in this embodiment, by constructing a Merkle tree structure, the hash values of intermediate nodes and root nodes are generated by means of recursive pairing calculations. Compared with the single hash chain structure in the prior art, the Merkle tree has higher verification efficiency and stronger scalability. Through the design of the tree structure, the computational complexity of the verification process is greatly reduced, and the ability of the system to process large-scale data is improved.
[0105] Thirdly, in this embodiment, the verification path information in the Merkle tree is innovatively extracted, and the verification path node set is combined with the path direction sequence to form a complete temporal integrity proof. This design enables the verification of any node to be carried out independently without obtaining the complete data set, significantly improving the verification efficiency and flexibility.
[0106] Finally, in this embodiment, by constructing a tamper-proof verification certificate, the temporal integrity proof is standardized and encapsulated. By adding certificate header information, calculating the certificate digest and performing digital signatures, a complete trust chain is established. This multi-level security mechanism effectively prevents the content of the certificate from being tampered with and ensures the reliability of the verification results.
[0107] Through the above technical solutions, this embodiment realizes the efficient verification of data integrity and the reliable proof of temporal relationships, significantly improving the efficiency and credibility of data verification. This solution not only solves the problems of low efficiency and poor scalability existing in traditional verification methods but also provides reliable technical support for the long-term preservation and verification of data, having important practical significance.
[0108] In an alternative embodiment, a double-layer consensus smart contract is activated to establish a version control chain, a traceability network is constructed among trusted node groups based on feature fingerprint identification, and the updated data is fused with the original temporal dependency graph through a three-layer Merkle tree structure to generate a new verification link, realizing the traceability of the entire data update process, including:
[0109] Construct the main consensus layer contract and the verification consensus layer contract. The main consensus layer contract records data version control information, and the verification consensus layer contract records data update verification information. Activate the main consensus layer contract and the verification consensus layer contract to form a double-layer consensus smart contract, and establish a version control chain; perform a hash calculation on the data block to obtain the data block hash value, combine the data block hash value with the timestamp sequence to generate a feature fingerprint identifier, calculate the fingerprint similarity between nodes based on the feature fingerprint identifier, and select trusted nodes based on the fingerprint similarity to construct a traceability network;
[0110] Construct a three-layer Merkle tree structure in the traceability network, write the updated data information into the data layer of the three-layer Merkle tree structure to generate end nodes, and calculate the temporal dependency relationship based on the end nodes to generate intermediate nodes and write them into the dependency layer of the three-layer Merkle tree structure; perform verification information calculation on the intermediate nodes to obtain root nodes and write them into the verification layer of the three-layer Merkle tree structure, and establish a data update path based on the end nodes, intermediate nodes, and root nodes in the three-layer Merkle tree structure;
[0111] Fuse the updated data with the original temporal dependency graph through the data update path to generate a new verification link containing node verification records, and the new verification link records the complete process of data update; when receiving a data update request, the main consensus layer contract triggers the verification consensus layer contract to perform verification. After the verification passes, the update information is recorded in the version control chain, and the traceability of the entire process of data update is realized based on the new verification link.
[0112] First, construct a double-layer consensus smart contract. The main consensus layer contract contains fields such as data version number, timestamp, and data block hash, which are used to record version control information; the verification consensus layer contract contains fields such as verification rules and verification results, which are used to record data verification information. The two contracts are associated through a contract call interface to form a double-layer consensus structure. For example, when the data version is updated from V1.0 to V1.1, the main contract records the version number V1.1, the timestamp 20230601, and the data block hash value 0xabc, and the verification contract records the verification rule checksum and the verification result valid.
[0113] Then, construct a traceability network based on feature fingerprints. Calculate the SHA256 hash value for each data block, and splice the hash value with the timestamp to generate a feature fingerprint. Calculate the fingerprint similarity between nodes, and nodes with a similarity exceeding 0.8 are selected as trusted nodes. For example, the fingerprint of node A is "0xabc - 20230601", the fingerprint of node B is "0xabc - 20230602", the similarity is 0.9, and node B is selected as a trusted node. Connections are established between trusted nodes to form a traceability network.
[0114] Then, a three-layer Merkle tree is constructed in the traceability network. The data layer stores the updated data blocks, such as "block_new"; the dependency layer records the temporal dependency relationships between data blocks, such as "block_new depends on block_old"; the verification layer stores verification information. Based on these three-layer structures, a data update path is established: the node block_new in the data layer is associated with the node in the verification layer through the node in the dependency layer.
[0115] Fuse the update path with the original dependency graph. The original dependency graph records the relationships between data blocks, and the new path connects the updated data blocks to the original graph through verification nodes, forming a new verification link. For example, in the original graph, block_old points to block_old2, and after the update, block_new points to block_old through a verification node, achieving traceability.
[0116] When a data update request is received, the main contract triggers the verification contract to execute the rule verification. After verification, the update information is recorded in the version control chain, and the entire update process can be traced based on the verification link.
[0117] Through the double-layer smart contract, the hierarchical management of data version management and verification is realized, improving the maintainability and scalability of the system. Based on the feature fingerprint and the three-layer Merkle tree structure, a trusted data traceability network is constructed, ensuring the traceability and integrity of the data update process. By adopting the verification link fusion mechanism, the association mapping between old and new data is realized, making the data update process transparent and controllable, and enhancing the reliability of the system.
[0118] In an alternative implementation, a hash calculation is performed on the data block to obtain the data block hash value, the data block hash value is combined with the timestamp sequence to generate a feature fingerprint identifier, the fingerprint similarity between nodes is calculated based on the feature fingerprint identifier, and trusted nodes are selected according to the fingerprint similarity to construct the traceability network, including:
[0119] Perform a hash calculation on the input data block to obtain the main hash value of the data block, perform MD5 and SHA1 hash calculations on the data block to obtain the auxiliary hash sequence of the data block, combine the main hash value of the data block with the auxiliary hash sequence of the data block and add a random salt value to generate the data block hash value; collect the timestamp information of the data block at different processing nodes, arrange the timestamp information in chronological order to generate a timestamp sequence, perform encoding processing on the timestamp sequence to obtain a time feature vector, and perform a weighted combination operation on the data block hash value and the time feature vector to generate a feature fingerprint identifier.
[0120] Construct a node feature matrix. The row vectors of the node feature matrix represent the node identifiers of the feature fingerprints, and the column vectors represent different feature dimensions of the feature fingerprints. Calculate the distances between different row vectors in the node feature matrix based on the locality-sensitive hashing algorithm to obtain the initial similarity; perform weighted fusion of the initial similarity and the historical behavior credibility of the nodes to obtain the fingerprint similarity between nodes, where the historical behavior credibility is calculated based on the historical interaction records and verification passing rates of the nodes;
[0121] Set a dynamic node similarity threshold, which is adaptively adjusted according to the network scale and security level. Compare the fingerprint similarity with the dynamic node similarity threshold, and select the nodes with fingerprint similarity higher than the dynamic node similarity threshold as trusted nodes; calculate the trust transfer strength between trusted nodes, which is jointly determined by the direct similarity and indirect similarity between nodes, and construct a weighted connection relationship between nodes based on the trust transfer strength;
[0122] Construct a hierarchical node traceability network using the weighted connection relationship. The hierarchical node traceability network includes a core trusted layer and an extended verification layer. The core trusted layer consists of the nodes with the highest fingerprint similarity, and the extended verification layer consists of other trusted nodes.
[0123] When performing hash calculation on a data block to obtain the data block hash value, first calculate the main hash value of the input data block using the SHA256 algorithm, and at the same time calculate the auxiliary hash value sequences using the MD5 and SHA1 algorithms respectively. For example, for a certain data block, calculate a 64-bit SHA256 main hash value, a 32-bit MD5 hash value, and a 40-bit SHA1 hash value. Concatenate these hash values in the order of the main hash value first and the auxiliary hash values later, and add an 8-bit random salt value to finally generate a 144-bit data block hash value.
[0124] Collect the timestamp information of the data block on different processing nodes, including the data block creation time, modification time, access time, etc. Arrange these timestamps in chronological order. For example, the timestamp sequence of a certain data block on three nodes is "2023-01-01 10:00:00, 2023-01-01 10:05:00, 2023-01-01 10:10:00". Encode the timestamp sequence and convert the time intervals into feature vectors, such as "300, 300" indicating that the adjacent timestamp intervals are both 300 seconds. Perform weighted combination of the data block hash value and the time feature vector to generate a feature fingerprint identifier.
[0125] Construct a node feature matrix, where each row of the matrix corresponds to the feature fingerprint identifier of a node. Use the locality-sensitive hashing algorithm to calculate the Hamming distance between different row vectors in the matrix to obtain the initial similarity between nodes. At the same time, count the historical interaction records and verification pass rates of each node, and calculate the historical behavior reputation of the node. Weight the initial similarity and the historical behavior reputation according to the weights of 0.7 and 0.3 to obtain the final fingerprint similarity between nodes.
[0126] Dynamically adjust the node similarity threshold according to the number of nodes and the security level requirements in the current network. For example, in a network with 100 nodes, set the similarity threshold to 0.8; when the number of nodes increases to 1000, correspondingly increase the threshold to 0.85. Compare the fingerprint similarity between each pair of nodes with the threshold, and select the nodes with similarity higher than the threshold as trusted nodes. Calculate the direct similarity between trusted nodes and the indirect similarity transmitted through other nodes to determine the trust transfer strength between nodes.
[0127] Finally, construct a hierarchical node traceability network. The nodes with the highest fingerprint similarity (such as similarity greater than 0.9) form the core trusted layer, and other trusted nodes (similarity between 0.8 - 0.9) form the extended verification layer. A weighted connection relationship is established between nodes according to the trust transfer strength to form a complete traceability network structure.
[0128] Figure 3 Schematic diagram for comparing the hash calculation performance and feature fingerprint generation efficiency in the embodiments of the present invention:
[0129] This figure comparatively shows three different hash calculation schemes. The "technical solution of the present invention" marked with a triangle represents an innovative algorithm that combines the main hash value, auxiliary hash sequence, and time feature vector; the "single SHA256 hash scheme" marked with a circle represents a traditional single hash algorithm method; the "basic timestamp concatenation scheme" marked with a square represents a simple timestamp connection scheme. Experimental data shows that when processing 2MB of data volume, the technical solution of the present invention only requires 40ms, while the single hash scheme requires 85ms, and the basic timestamp scheme requires 120ms. When the data volume increases to 6MB, the processing times of the three schemes are 100ms, 220ms, and 320ms respectively. At the maximum test data volume of 10MB, the processing time of the technical solution of the present invention is 150ms, significantly better than 320ms of the single hash scheme and 450ms of the basic timestamp scheme. Especially in the large data processing interval from 8MB to 10MB, the processing time of the technical solution of the present invention only increases by 30ms, while the other two schemes increase by 50ms and 70ms respectively, fully demonstrating the performance advantages and stability of the technical solution of the present invention in large-scale data processing scenarios.
[0130] Figure 4Schematic diagram of the data block hash calculation status analysis panel in the embodiment of the present invention:
[0131] This panel shows the complete status of data block hash calculation, using multiple hash algorithms to ensure data security. The main hash value is calculated by the SHA256 algorithm to obtain "a1b2c3d4...e5f6" (64 bits), which serves as the main identity identifier of the data block; the first set of auxiliary hash values is generated by the MD5 algorithm to obtain "g7h8i9j0...k1l2" (32 bits), and the second set of auxiliary hash values is obtained by the SHA1 algorithm to get "m3n4o5p6...q7r8" (40 bits). The system also introduces an 8-bit random salt value "s9t0u1v2" to enhance the randomness and security of hash calculation. The combined application of this multiple hash algorithm not only provides a unique identifier for the data block but also significantly improves the system's anti-attack ability and data integrity protection level through the complementary characteristics of different algorithms. The "Recalculate" button on the panel allows the administrator to update the hash value when needed to ensure the timeliness and accuracy of the data.
[0132] Figure 5 Schematic diagram of the timestamp sequence information panel in the embodiment of the present invention:
[0133] This panel details the complete time track information of the data block in the system. Starting from the creation time "2023-01-01 10:00:00", passing through the modification time "2023-01-01 10:05:00", and finally to the access time "2023-01-01 10:10:00", a complete timestamp sequence is formed. The system automatically calculates the time interval vector [300, 300], indicating that the data block maintains a stable 5-minute (300 seconds) time interval at different processing stages. This regular time interval not only reflects the standardization and normativity of system processing but also provides a reference benchmark for the detection of abnormal behaviors. The "Update Timestamp" function provided by the panel ensures the real-time update of time information, helping system maintenance personnel to timely grasp the processing status and transfer process of the data block.
[0134] This embodiment provides a method for constructing a trusted node traceability network based on feature fingerprint identification. Aiming at the problems in the prior art, such as usually using a single hash algorithm for data integrity verification and only relying on simple threshold judgment to determine the credibility of nodes, there are problems such as insufficient security, one-sided credibility evaluation, and limited traceability ability.
[0135] In this embodiment, the data block features are first calculated through a multiple hashing algorithm, and the data block hash value is constructed by combining the main hash value, the auxiliary hash sequence, and the random salt value. This multi-dimensional feature extraction method greatly improves the uniqueness of the features and the anti-tampering ability compared with the single hashing calculation in the prior art. At the same time, by introducing a timestamp sequence and performing feature encoding, the accurate characterization of the temporal features of the data processing process is realized.
[0136] Secondly, this embodiment innovatively introduces a node feature matrix and a locality-sensitive hashing algorithm, and evaluates the similarity by calculating the feature distance between nodes. Different from the simple binary judgment method in the prior art, this solution constructs a more comprehensive and objective node credibility evaluation mechanism by weighted fusion of the initial similarity and the node historical behavior credibility.
[0137] Thirdly, this embodiment adopts a dynamically adjusted similarity threshold mechanism, and adaptively adjusts the judgment criterion according to the network scale and security requirements. This flexible threshold setting method effectively solves the problem that the traditional fixed threshold is difficult to adapt to the dynamic changes of the network. By calculating the direct and indirect trust transfer strengths between nodes, a more practical trust propagation model is constructed.
[0138] Finally, this embodiment constructs a hierarchical node traceability network, and divides the trusted nodes into a core trusted layer and an extended verification layer. This hierarchical structure design not only improves the reliability of the traceability network, but also enhances the scalability and fault tolerance of the system. Through the weighted connection relationship between nodes, more accurate trust transfer and traceability tracking are realized.
[0139] Through the above technical solutions, this embodiment realizes the reliable traceability of the data processing process and the accurate evaluation of the node credibility, and significantly improves the security and reliability of the system. This solution breaks through the technical bottlenecks such as single feature extraction, rough credibility evaluation, and limited traceability ability in the traditional methods, provides strong technical support for building a highly trusted distributed data processing system, and has important application value.
[0140] In the second aspect of the embodiment of the present invention,
[0141] A kind of electronic device is provided, including:
[0142] A processor;
[0143] A memory for storing instructions executable by the processor;
[0144] Wherein, the processor is configured to call the instructions stored in the memory to execute the foregoing method.
[0145] In the third aspect of the embodiment of the present invention,
[0146] Provided is a computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, implement the foregoing method.
[0147] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0148] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A decentralized method for storing and verifying talent background investigation data, characterized in that, Including: Based on the identity authentication layer, it receives the educational experience information, work experience information, and skill certification information submitted by candidates, and uses the national cryptography algorithm SM2 to perform asymmetric encryption on the received information to obtain the first encrypted data set. In the SM2 elliptic curve cryptosystem, the base point G coordinates are selected as (x, y), and the order n is a 256-bit prime number. A 256-bit random number is generated by a random number generator as the private key d. The private key d is multiplied by the base point G to obtain the coordinates of the public key Q. The generated key pair (d, Q) is verified for validity, verifying whether the private key d is within the interval [1, n - 1] and whether the public key Q satisfies the elliptic curve equation; Generate a random number k, multiply k by the base point G to obtain a temporary public key (x1, y1), multiply the temporary public key by the public key Q to obtain a session key (x2, y2), use a key derivation function KDF to process the session key coordinate values, generate an encryption key, and use the encryption key to encrypt the standardized data template to obtain ciphertext data. Add a time stamp to the ciphertext data to generate the first encrypted data set; Combine the SM3 hash algorithm to generate an original data packet with time series weights for the first encrypted data set; Segment the first encrypted data set according to the time series weights in the original data packet, construct a hierarchical time series dependence graph. The hierarchical time series dependence graph extracts the time series correlation features between data segments through a deep learning model to generate a second encrypted data set. The second encrypted data set contains weighted feature fingerprint identifiers and information segment hash values, and writes the second encrypted data set into the trusted node group of the blockchain network through a three-layer Merkle tree structure; When the trusted node group receives a verification request, it locates the target data location based on the feature fingerprint identifier, extracts the second encrypted data set for the target time period according to the hierarchical time series dependence graph, and calculates the real-time trust score of the node at the same time. The real-time trust score of the node is dynamically updated based on the verification historical behavior feature vector of the node. When the overall trust level of the node group reaches the adaptive threshold, the consistency verification mechanism based on zero-knowledge proof is started using the feature fingerprint identifier, and an anti-tampering verification certificate including time series integrity proof is generated for the verification result based on the consistency verification mechanism; When the hierarchical time series dependence graph detects a data update requirement, calculate the correlation degree between the update request information and the time series weights in the original data packet, and evaluate the update rationality by a deep learning model. After confirmation, locate the data segment to be updated in the second encrypted data set, activate the two-layer consensus smart contract to establish a version control chain, construct a traceability network in the trusted node group based on the feature fingerprint identifier, and fuse the updated data with the original time series dependence graph through a three-layer Merkle tree structure to generate a new verification link to achieve traceability throughout the data update process.
2. The method according to claim 1, wherein Using the national cryptography algorithm SM2 to perform asymmetric encryption on the received information to obtain the first encrypted data set, and combining the SM3 hash algorithm to generate an original data packet with time series weights for the first encrypted data set includes: Asymmetrically encrypt the received education experience, work experience, and skill certification information using the national cryptographic algorithm SM2. Based on the SM2 elliptic curve cryptosystem, select the base point G and the order n to generate a key pair (d, Q). Use the public key Q to perform an encryption operation on the standardized information, and combine it with the random number k to generate the first encrypted data set, which contains the encrypted structured information. Combine the SM3 hashing algorithm to perform security processing on the first encrypted data set. Pad the encrypted data to a multiple of 512 bits, perform iterative operations through a compression function, and introduce a timing weight parameter to calculate the timestamp difference and information correlation degree between adjacent data blocks, generating an original data packet with timing weights, and the original data packet has the characteristic of timing integrity.
3. The method according to claim 2, characterized in that, Asymmetrically encrypt the received education experience, work experience, and skill certification information using the national cryptographic algorithm SM2. Based on the SM2 elliptic curve cryptosystem, select the base point G and the order n to generate a key pair (d, Q). Use the public key Q to perform an encryption operation on the standardized information, and combine it with the random number k to generate the first encrypted data set including: Receive the education experience information, work experience information, and skill certification information submitted by the user, respectively classify and encode the education experience information, work experience information, and skill certification information to obtain education experience structured data, work experience structured data, and skill certification structured data, and integrate the education experience structured data, work experience structured data, and skill certification structured data to construct a standardized data template. Under the SM2 elliptic curve cryptosystem, select the base point G and the order n, use a random number generator to obtain the private key d, perform a point multiplication operation based on the private key d and the base point G to generate the public key Q, obtain the key pair (d, Q), and perform validity verification on the key pair (d, Q) through a key validity verification algorithm. Generate a random number k, perform a point multiplication operation on the random number k and the base point G to obtain a temporary public key, perform a point multiplication operation on the temporary public key and the public key Q to obtain a session key, use a key derivation function to process the session key to generate an encryption key, and use the encryption key to encrypt the standardized data template to generate the first encrypted data set containing the encrypted data and its corresponding timestamp.
4. The method according to claim 1, wherein Extract the second encrypted data set for the target time period according to the hierarchical timing dependency graph, and calculate the real-time trust score of the node at the same time. The real-time trust score of the node is dynamically updated based on the verification historical behavior feature vector of the node. When the overall trust degree of the node group reaches the adaptive threshold, use the feature fingerprint identifier to start the consistency verification mechanism based on zero-knowledge proof, and generate a tamper-proof verification certificate containing the timing integrity proof for the verification result including: The hierarchical time-sequence dependency graph contains a set of time-sequence nodes and a set of time-sequence edges. The second encrypted data set for the target time period is extracted according to the hierarchical time-sequence dependency graph. The second encrypted data set contains the encrypted data within the target time period and its corresponding time-sequence dependency relationships; verification historical behavior feature vectors are established for the nodes in the set of time-sequence nodes, and the real-time trust scores of the nodes are calculated based on the verification historical behavior feature vectors. The real-time trust scores of the nodes are dynamically adjusted as the verification historical behavior feature vectors are updated; The real-time trust scores of each node in the set of time-sequence nodes are weighted and summed to obtain the overall trust degree of the node group. The overall trust degree of the node group is compared with the adaptive threshold. When the overall trust degree of the node group is greater than the adaptive threshold, a feature fingerprint identifier is generated based on the overall trust degree of the node group and the adaptive threshold; A prover commitment is constructed based on the feature fingerprint identifier, and the prover commitment and the verification rule are combined to generate a zero-knowledge proof. The zero-knowledge proof is used to activate the consistency verification mechanism to perform consistency verification on the second encrypted data set and generate a verification result; a Merkle tree is constructed according to the verification result, a time-sequence integrity proof including the Merkle tree root node and the verification path is generated, and a tamper-proof verification certificate is generated based on the time-sequence integrity proof.
5. The method according to claim 4, wherein Constructing a Merkle tree according to the verification result and generating a time-sequence integrity proof including the Merkle tree root node and the verification path, and generating a tamper-proof verification certificate based on the time-sequence integrity proof includes: Performing hash calculation on each data block in the verification result and the corresponding time-stamp sequence to obtain the data block hash value, and combining the data block hash value with the index information to generate leaf nodes; calculating the pairwise pairing of the leaf nodes to obtain the intermediate node hash value, and recursively pairing the intermediate node hash values until the root node hash value is generated, and constructing a Merkle tree structure based on the root node hash value; Extracting the verification path node set and the path direction sequence in the Merkle tree structure, combining the verification path node set and the path direction sequence to generate a verification path, and combining the verification path and the root node hash value to generate a time-sequence integrity proof; Constructing a tamper-proof verification certificate based on the time-sequence integrity proof. The tamper-proof verification certificate contains certificate header information and certificate body content; performing hash calculation on the certificate header information and the certificate body content of the tamper-proof verification certificate to obtain a certificate digest, signing the certificate digest with a private key to obtain a digital signature, and adding the digital signature to the tamper-proof verification certificate to complete the certificate generation.
6. The method according to claim 1, characterized in that Activating a two-layer consensus smart contract to establish a version control chain, constructing a traceability network in the trusted node group based on the feature fingerprint identifier, and fusing the updated data with the original time-sequence dependency graph through a three-layer Merkle tree structure to generate a new verification link, realizing the traceability of the whole process of data update, including: Construct the main consensus layer contract and the verification consensus layer contract. The main consensus layer contract records data version control information, and the verification consensus layer contract records data update verification information. Activate the main consensus layer contract and the verification consensus layer contract to form a double-layer consensus smart contract, and establish a version control chain; perform a hash calculation on the data block to obtain the data block hash value, combine the data block hash value with the timestamp sequence to generate a feature fingerprint identifier, calculate the fingerprint similarity between nodes based on the feature fingerprint identifier, and select trusted nodes according to the fingerprint similarity to construct a traceability network; Construct a three-layer Merkle tree structure in the traceability network, write the updated data information into the data layer of the three-layer Merkle tree structure to generate end nodes, calculate the temporal dependence relationship based on the end nodes to generate intermediate nodes and write them into the dependence layer of the three-layer Merkle tree structure; calculate the verification information for the intermediate nodes to obtain the root node and write it into the verification layer of the three-layer Merkle tree structure, and establish a data update path based on the end nodes, intermediate nodes and root nodes in the three-layer Merkle tree structure; Fuse the updated data with the original temporal dependence graph through the data update path to generate a new verification link containing node verification records, and the new verification link records the complete process of data update; when receiving a data update request, the main consensus layer contract triggers the verification consensus layer contract to perform verification. After the verification is passed, record the update information in the version control chain, and realize the traceability of the entire data update process based on the new verification link.
7. The method according to claim 6, characterized in that Perform a hash calculation on the data block to obtain the data block hash value, combine the data block hash value with the timestamp sequence to generate a feature fingerprint identifier, calculate the fingerprint similarity between nodes based on the feature fingerprint identifier, and select trusted nodes according to the fingerprint similarity to construct a traceability network including: Perform a hash calculation on the input data block to obtain the main data block hash value, perform MD5 and SHA1 hash calculations on the data block to obtain the data block auxiliary hash sequence, combine the main data block hash value with the data block auxiliary hash sequence and add a random salt value to generate the data block hash value; collect the timestamp information of the data block at different processing nodes, arrange the timestamp information in chronological order to generate a timestamp sequence, perform encoding processing on the timestamp sequence to obtain a time feature vector, and perform a weighted combination operation on the data block hash value and the time feature vector to generate a feature fingerprint identifier; Establish a node feature matrix. The row vector of the node feature matrix represents the node identifier of the feature fingerprint identifier, and the column vector represents different feature dimensions of the feature fingerprint identifier. Calculate the distance between different row vectors in the node feature matrix based on the locality-sensitive hashing algorithm to obtain the initial similarity; perform a weighted fusion of the initial similarity and the historical behavior credibility of the node to obtain the fingerprint similarity between nodes, and the historical behavior credibility is calculated according to the historical interaction records and verification passing rate of the node; Set a dynamic node similarity threshold, which is adaptively adjusted according to the network scale and security level. Compare the fingerprint similarity with the dynamic node similarity threshold, and select the nodes with fingerprint similarity higher than the dynamic node similarity threshold as trusted nodes; Calculate the trust transfer strength between trusted nodes, which is jointly determined by the direct similarity and indirect similarity between nodes, and construct a weighted connection relationship between nodes based on the trust transfer strength; Construct a hierarchical node traceability network using the weighted connection relationship. The hierarchical node traceability network includes a core trusted layer and an extended verification layer. The core trusted layer consists of nodes with the highest fingerprint similarity, and the extended verification layer consists of other trusted nodes.
8. A decentralized talent background investigation data storage and verification system for implementing the method described in any one of the foregoing claims 1-7, characterized in that, Including: The first unit is used to receive the education experience information, work experience information, and skill certification information submitted by candidates based on the identity authentication layer, perform asymmetric encryption on the received information through the national cryptographic algorithm SM2 to obtain the first encrypted data set, and generate an original data packet with time-series weights for the first encrypted data set in combination with the SM3 hashing algorithm; The second unit is used to segment the first encrypted data set according to the time-series weights in the original data packet, construct a hierarchical time-series dependence graph, extract the time-series correlation features between data segments through a deep learning model, generate a second encrypted data set, which includes weighted feature fingerprint identifiers and information segment hash values, and write the second encrypted data set into the trusted node group of the blockchain network through a three-layer Merkle tree structure; The third unit is used to, when the trusted node group receives a verification request, locate the target data position based on the feature fingerprint identifier, extract the second encrypted data set for the target time period according to the hierarchical time-series dependence graph, and calculate the real-time trust score of the node at the same time. The real-time trust score of the node is dynamically updated based on the verification historical behavior feature vector of the node. When the overall trust level of the node group reaches the adaptive threshold, use the feature fingerprint identifier to start a consistency verification mechanism based on zero-knowledge proof, and generate a tamper-proof verification certificate including time-series integrity proof for the verification result based on the consistency verification mechanism; The fourth unit is used to, when the hierarchical time-series dependence graph detects a data update requirement, calculate the correlation degree between the update request information and the time-series weights in the original data packet, and evaluate the update rationality by a deep learning model. After confirmation, locate the data segment to be updated in the second encrypted data set, activate a two-layer consensus smart contract to establish a version control chain, construct a traceability network in the trusted node group based on the feature fingerprint identifier, and fuse the updated data with the original time-series dependence graph through a three-layer Merkle tree structure to generate a new verification link, realizing the traceability of the entire data update process.
9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Block chain system, intelligent contract synchronization method, computer equipment and storage medium
CN117333175A
Block chain-based asset processing traceability method and system
CN118505398A