Digital information traceability management system

By using information fingerprint extraction, origin node consensus, flow path mapping, information integrity checking and behavior consensus feedback modules in the digital information traceability management system, the problem of easy tampering of the traceability chain and complex and inefficient traceability process is solved, and efficient, safe and reliable information traceability is achieved.

CN120145334AActive Publication Date: 2025-06-13BEIJING LIUSHEN DATA TECH CO LTD

Patent Information

Application Number
CN202510208644.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The existing digital information traceability technology has the problem of easy tampering with the traceability chain and complex and inefficient traceability process.

Method used

A digital information traceability management system is adopted, which includes an information fingerprint extraction module, an origin node consensus module, a flow path mapping module, an information integrity verification module and a behavior consensus feedback module. Information fingerprints are generated through NLP processing and semantic feature extraction, and source point credential generation and circulation path tracking are used to use consensus node servers to perform information integrity verification, and exception handling decisions are made through consensus node servers and monitoring and handling components.

Benefits of technology

It effectively solves the problem of easy tampering of the traceability chain and complex and inefficient traceability process, and improves the robustness, security, efficiency and applicability of traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145334A_ABST
    Figure CN120145334A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information traceability, in particular to a digital information traceability management system. The system comprises an information fingerprint extraction module, an origin node consensus module, a circulation path mapping module, an information integrity verification module and a behavior consensus feedback module, and can be used for acquiring original information; performing text semantic feature extraction on the original information to obtain semantic features; performing information fingerprint extraction on the semantic features to obtain semantic fingerprints; performing consensus processing on the semantic fingerprint based on an origin node to obtain a consensus result; generating a source point voucher according to the consensus result to obtain the source point voucher; performing circulation behavior triggering according to the source point voucher to obtain a to-be-recorded behavior; and performing content change analysis according to the to-be-recorded behavior to obtain a changed semantic fingerprint. According to the method, through fusion of semantic hash and behavior consensus, the source tracing robustness, safety, efficiency and applicability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information traceability, and particularly to a digital information traceability management system. Background Art

[0002] Digital information traceability management is not a completely new concept. The earliest traceability methods were relatively simple, mainly relying on technologies such as log records and timestamps. Later, in order to improve the anti-tampering ability of traceability information, traditional encryption technologies and hash algorithms were introduced. The emergence of blockchain technology has brought a revolutionary change to digital information traceability.

[0003] However, despite the continuous development of digital information traceability technology, existing methods still have problems such as the traceability chain being easily tampered with and the traceability process being complex and inefficient.

[0004] Problem of the traceability chain being easily tampered with: Although traditional hash algorithms can detect data tampering, they are very sensitive to any minor modification of the data (including format changes, space adjustments, etc.). Even if the core semantics of the information remain unchanged, the hash value will change significantly. This causes the traceability system based on traditional hashes to be prone to misjudgment when facing information format transformation, version iteration, or minor editing, and even leads to the breakage of the traceability chain, reducing the practicality and robustness of traceability.

[0005] Problem of the traceability process being complex and inefficient: Traditional traceability methods mainly rely on literal comparison and operation records of data, lacking the ability to understand the semantic content of information. When information undergoes semantic drift, content evolution, or multiple rounds of dissemination, traditional traceability methods are difficult to accurately identify the true source and evolution trajectory of the information, resulting in a decline in traceability accuracy and efficiency. Summary of the Invention

[0006] Based on this, it is necessary to provide a digital information traceability management system to solve at least one of the above technical problems.

[0007] To achieve the above object, a digital information traceability management system includes the following modules:

[0008] An information fingerprint extraction module, including an NLP processing server and a cache server, is used to obtain the original information; extract text semantic features from the original information to obtain semantic features; extract information fingerprints from the semantic features to obtain semantic fingerprints;

[0009] An origin node consensus module, including a consensus node server, is used to perform consensus processing on the semantic fingerprints based on the origin node to obtain a consensus result; generate an origin certificate according to the consensus result to obtain an origin certificate;

[0010] The transfer path mapping module includes a behavior data collector, which is used to trigger transfer behaviors according to source vouchers to obtain behaviors to be recorded; perform content change analysis on the behaviors to be recorded to obtain the changed semantic fingerprints; perform transfer path mapping based on the behaviors to be recorded and the changed semantic fingerprints to obtain transfer trajectory data;

[0011] The information integrity verification module includes a verification server and an abnormal mode detection engine, which are used to receive verification requests according to source vouchers to obtain information to be verified; perform source fingerprint comparison on the information to be verified to obtain a fingerprint comparison result; perform transfer path backtracking based on the transfer trajectory data to obtain a verified path; perform behavior consistency verification on the verified path and perform abnormal mode recognition to obtain a behavior verification result and an abnormal behavior report; generate a verification report based on the fingerprint comparison result, the behavior verification result, and the abnormal behavior report to obtain a traceability verification report;

[0012] The behavior consensus feedback module includes a consensus node server and a monitoring and handling component, which are used to conduct node consensus voting according to the traceability verification report to obtain a consensus vote repository; form a consensus decision based on a preset consensus algorithm and the consensus vote repository to obtain a draft consensus resolution; execute resolution measures based on the draft consensus resolution to obtain disposal resolution data, so as to implement digital information traceability management operations.

[0013] Preferably, the information fingerprint extraction module includes the following functions:

[0014] Obtain the original information; perform text preprocessing on the original information to obtain preprocessed text;

[0015] Extract semantic features from the preprocessed text to obtain semantic features;

[0016] Construct a semantic vector from the semantic features to obtain a semantic vector;

[0017] Perform semantic hash coding on the semantic vector to obtain a semantic hash code;

[0018] Generate a semantic fingerprint from the semantic hash code to obtain a semantic fingerprint.

[0019] Preferably, the origin node consensus module includes the following functions:

[0020] Encapsulate origin information for the semantic fingerprint to obtain information to be consensus;

[0021] Submit the information to be consensus to the network for broadcasting to obtain a broadcast request;

[0022] Perform preliminary node verification on the broadcast request to obtain information to be verified;

[0023] Conduct consensus proposal and voting on the information to be verified to obtain consensus votes;

[0024] Confirm the consensus result for the consensus vote according to the preset consensus algorithm to obtain the consensus result;

[0025] Generate the source point certificate according to the consensus result to obtain the source point certificate.

[0026] Preferably, the transfer path mapping module includes the following functions:

[0027] Trigger the transfer behavior according to the source point certificate to obtain the behavior to be recorded;

[0028] Analyze the content change according to the behavior to be recorded to obtain the changed semantic fingerprint;

[0029] Construct the transfer event according to the behavior to be recorded and the changed semantic fingerprint to obtain the transfer snapshot;

[0030] Append the transfer record to the transfer snapshot to obtain the record to be uploaded to the chain;

[0031] Write the record to be uploaded to the chain into the distributed ledger according to the preset distributed ledger to obtain the on-chain transfer record;

[0032] Update the transfer track according to the on-chain transfer record and the source point certificate to obtain the transfer track data.

[0033] Preferably, the information integrity verification module includes the following functions:

[0034] Receive the verification request according to the source point certificate to obtain the information to be verified; reconstruct the current semantic fingerprint for the information to be verified to obtain the current semantic fingerprint;

[0035] Compare the current semantic fingerprint with the information to be verified to obtain the fingerprint comparison result;

[0036] Trace back the transfer path according to the information to be verified and the transfer track data to obtain the verified path;

[0037] Verify the behavior consistency of the verified path according to the current semantic fingerprint to obtain the behavior verification result;

[0038] Identify the abnormal mode of the verified path to obtain the abnormal behavior report;

[0039] Summarize the verification results of the fingerprint comparison result, the behavior verification result and the abnormal behavior report to obtain the report to be signed; generate the verification report for the report to be signed to obtain the traceability verification report.

[0040] Preferably, the behavior consensus feedback module includes the following functions:

[0041] Perform verification report distribution on the traceability verification report to obtain the report to be consensus;

[0042] Initiate the consensus process based on the report to be consensus to obtain a consensus proposal;

[0043] Conduct node evaluation and voting based on the report to be consensus and the consensus proposal to obtain disposal ballots;

[0044] Summarize the voting results of the disposal ballots to obtain a consensus ballot warehouse;

[0045] Form a consensus decision based on the preset consensus algorithm and the consensus ballot warehouse to obtain a draft consensus resolution;

[0046] Execute the disposal measures according to the draft consensus resolution to obtain an execution instruction; record the execution result according to the execution instruction to obtain a disposal record;

[0047] Solidify the final resolution of the draft consensus resolution, the consensus ballot warehouse and the disposal record to obtain disposal resolution data.

[0048] Through the NLP processing server and the cache server, the present invention realizes the automated acquisition and efficient processing of original information. By adopting the text semantic feature extraction technology, the system can deeply understand the core semantics of information, transcending the limitations of traditional keyword matching and more accurately capturing the essential content of information. The information fingerprint extraction process compresses complex semantic information into concise semantic fingerprints, realizing efficient indexing and comparison of information content, and laying an efficient and accurate foundation for information recognition and integrity verification in the subsequent traceability process. Using the consensus node server, the semantic fingerprints are processed based on the origin node to ensure the authority and immutability of the information source identity. The introduction of the consensus mechanism makes the generation of the origin certificate no longer rely on the trust endorsement of a single central institution, but is jointly confirmed by the distributed network, enhancing the credibility and public trust of the origin certificate. The generation of the origin certificate provides a reliable anchor point for subsequent transfer path tracking and information integrity verification, ensuring the reliability of the traceability system from the source. With the help of the behavior data collector, the system can real-time sense and record various behaviors during the information transfer process, realizing the automated and refined tracking of the information transfer trajectory. The application of the content change analysis technology enables the system to identify subtle changes in the information content during the transfer process and generate the semantic fingerprint after the change, accurately recording each content evolution. The transfer path mapping process associates the behavior to be recorded with the semantic fingerprint after the change, constructing a complete transfer trajectory data, providing a clear and reliable path record for the full life cycle management and responsibility traceability of information. Through the verification server and the abnormal mode detection engine, the system can receive verification requests according to the origin certificate and perform multi-dimensional and in-depth integrity verification on the information to be verified. The origin fingerprint comparison technology can quickly judge whether the core semantics of the information to be verified are consistent with the original information, initially identifying the risk of content tampering. The transfer path backtracking and behavior consistency verification mechanism can verify whether the transfer trajectory of the information conforms to the recorded path and identify potential abnormal behaviors, such as unauthorized modification or jump-style transfer, effectively improving the accuracy and comprehensiveness of the information integrity verification. The finally generated traceability verification report provides an objective and reliable basis for the credibility assessment and risk warning of information. Using the consensus node server and the monitoring and disposal component, the system can start the node consensus voting based on the traceability verification report, realizing the automated and collaborative disposal decision of abnormal behaviors. The application of the consensus algorithm and the consensus vote warehouse ensures the democracy and fairness of the disposal decision, avoiding the risk of single-point decision-making and enhancing the credibility of the disposal result. The generation of the consensus resolution draft and the execution of the resolution measures realize the closed-loop management from abnormal detection to disposal response, improving the system's rapid response and automated disposal capabilities for information security risks. The finally formed disposal resolution data provides valuable data support for the continuous optimization and security audit of the system, constructing a more robust and reliable digital information traceability management system.Therefore, the present invention provides a digital information traceability management system. Through the innovative integration of semantic hashing and behavior consensus, it not only effectively solves the prominent drawbacks of existing methods in terms of the traceability chain being easily tampered with and the traceability process being complex and inefficient, but also achieves remarkable improvements in aspects such as the robustness, security, efficiency, and applicability of traceability. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic diagram of the step - by - step process of the digital information traceability management system;

[0050] Figure 2 It is a schematic diagram of the detailed implementation steps of the information fingerprint extraction module in the present invention;

[0051] The realization of the object, functional features, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work belong to the scope of protection of the present invention.

[0053] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, so repeated descriptions of them will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0054] It should be understood that although terms such as "first", "second", etc. may be used here to describe each unit, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly, the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.

[0055] To achieve the above - mentioned purpose, please refer to Figures 1 to 2 , a digital information traceability management system, the system includes the following modules:

[0056] The information fingerprint extraction module, including an NLP processing server and a cache server, is used to obtain the original information; extract the text semantic features of the original information to obtain semantic features; extract the information fingerprint from the semantic features to obtain a semantic fingerprint;

[0057] The origin node consensus module, including a consensus node server, is used to perform consensus processing on the semantic fingerprint based on the origin node to obtain a consensus result; generate an origin certificate according to the consensus result to obtain an origin certificate;

[0058] The transfer path mapping module, including a behavior data collector, is used to trigger transfer behaviors according to the origin certificate to obtain behaviors to be recorded; analyze content changes according to the behaviors to be recorded to obtain the changed semantic fingerprint; map the transfer path according to the behaviors to be recorded and the changed semantic fingerprint to obtain transfer trajectory data;

[0059] The information integrity verification module, including a verification server and an abnormal mode detection engine, is used to receive a verification request according to the origin certificate to obtain information to be verified; compare the origin fingerprint of the information to be verified to obtain a fingerprint comparison result; trace back the transfer path according to the transfer trajectory data to obtain a verified path; verify the behavior consistency of the verified path and identify abnormal modes to obtain a behavior verification result and an abnormal behavior report; generate a verification report according to the fingerprint comparison result, the behavior verification result and the abnormal behavior report to obtain a traceability verification report;

[0060] The behavior consensus feedback module, including a consensus node server and a monitoring and handling component, is used to conduct node consensus voting according to the traceability verification report to obtain a consensus vote warehouse; form a consensus decision according to a preset consensus algorithm and the consensus vote warehouse to obtain a draft consensus resolution; execute resolution measures according to the draft consensus resolution to obtain disposal resolution data, so as to realize the digital information traceability management operation.

[0061] In the embodiment of the present invention, refer to Figure 1 As shown, it is a schematic diagram of the step flow of the digital information traceability management system of the present invention. In this example, the digital information traceability management system includes the following modules:

[0062] S1: The information fingerprint extraction module, including an NLP processing server and a cache server, is used to obtain the original information; extract the text semantic features of the original information to obtain semantic features; extract the information fingerprint from the semantic features to obtain a semantic fingerprint;

[0063] In an embodiment of the present invention, by using an NLP processing server and a cache server, first, the original information of various data sources is monitored and received through a preset port. Then, a Jieba or HanLP tokenizer is used for text preprocessing to remove noise and stop words. Next, techniques such as TF-IDF, LDA, NER, and dependency syntactic analysis are adopted to extract semantic features, and these features are transformed into semantic vectors. Finally, the semantic vectors are hashed encoded by the LSH algorithm to generate semantic hash codes, and after attaching metadata, SHA-256 hashing calculation is performed to obtain the final semantic fingerprint as the unique identifier of the information.

[0064] S2: Origin node consensus module, including a consensus node server, which is used to perform consensus processing on the semantic fingerprint based on the origin node to obtain a consensus result; generate an origin certificate according to the consensus result to obtain an origin certificate;

[0065] In an embodiment of the present invention, a consensus node server is used to encapsulate the semantic fingerprint into information to be consensus and submit it to the consensus network through network broadcast. After the consensus node preliminarily verifies the broadcast request, it conducts a consensus proposal and voting. According to the preset consensus algorithm, the consensus result is confirmed for the collected consensus votes, and finally an origin certificate is generated as proof of the trusted identity of the information source.

[0066] S3: Transfer path mapping module, including a behavior data collector, which is used to trigger transfer behaviors according to the origin certificate to obtain behaviors to be recorded; perform content change analysis on the behaviors to be recorded to obtain the changed semantic fingerprint; map the transfer path according to the behaviors to be recorded and the changed semantic fingerprint to obtain transfer trajectory data;

[0067] In an embodiment of the present invention, through a behavior data collector, transfer behavior records are triggered according to the origin certificate. When a transfer behavior is detected, content change analysis is performed. Through fine-grained difference detection techniques (line-level and word-level comparison), content difference segments are obtained, and a changed semantic fingerprint is generated. Then, a transfer event snapshot is constructed, appended to the transfer record, and written into the distributed ledger. Finally, the transfer trajectory data is updated according to the on-chain transfer record and the origin certificate to form a complete transfer path of the information.

[0068] S4: Information integrity verification module, including a verification server and an abnormal mode detection engine, which are used to receive verification requests according to the origin certificate to obtain information to be verified; compare the origin fingerprints of the information to be verified to obtain a fingerprint comparison result; perform transfer path backtracking according to the transfer trajectory data to obtain a verified path; verify the behavior consistency of the verified path and perform abnormal mode recognition to obtain a behavior verification result and an abnormal behavior report; generate a verification report according to the fingerprint comparison result, the behavior verification result, and the abnormal behavior report to obtain a traceability verification report;

[0069] In an embodiment of the present invention, through a verification server and an abnormal mode detection engine, a verification request is received according to the source point certificate, and information to be verified is obtained. The current semantic fingerprint is reconstructed for the information to be verified. Then, the Hamming distance is compared between the current semantic fingerprint and the source point fingerprint in the source point certificate to obtain a fingerprint comparison result. At the same time, the transfer path is traced back and hash verification is performed based on the information to be verified and the transfer trajectory data to obtain a verified path. Next, the behavior consistency verification and abnormal mode recognition are performed on the verified path to obtain a behavior verification result and an abnormal behavior report. Finally, the fingerprint comparison result, the behavior verification result, and the abnormal behavior report are summarized to generate a traceability verification report.

[0070] S5: The behavior consensus feedback module, including a consensus node server and a monitoring and handling component, is used to perform node consensus voting according to the traceability verification report to obtain a consensus ticket warehouse; perform consensus decision formation according to a preset consensus algorithm and the consensus ticket warehouse to obtain a draft consensus resolution; perform resolution measure execution according to the draft consensus resolution to obtain disposal resolution data, so as to implement digital information traceability management operations;

[0071] In an embodiment of the present invention, the consensus node server and the monitoring and handling component are used to distribute the traceability verification report, start the consensus process and generate a consensus proposal. The consensus nodes evaluate and vote based on the report and the proposal to obtain disposal ballots. The disposal ballots are summarized to form a consensus ticket warehouse, and a draft consensus resolution is formed through consensus decision-making according to a preset consensus algorithm (such as the simple majority voting algorithm). Finally, disposal measures are executed according to the draft resolution, the execution result is recorded as a disposal record, and the draft consensus resolution, the consensus ticket warehouse, and the disposal record are finally resolved and solidified to obtain disposal resolution data, completing the digital information traceability management operation.

[0072] As an example of the present invention, refer to Figure 2 shown in Figure 1 the functional flow diagram of the information fingerprint extraction module in

[0073] S11: Obtain the original information; perform text preprocessing on the original information to obtain preprocessed text;

[0074] In the embodiments of the present invention, the operation of obtaining the original information is the primary link in the digital information traceability management process. The system presets an information receiving port, which continuously collects and responds to information requests from the data input channels. The data input channels include but are not limited to: open API interfaces, database connections, file system directories, and web crawler components. When receiving an information input request, the system immediately establishes a data connection and receives the original information data stream from the data source according to a predefined communication protocol, such as the HTTP protocol, TCP / IP protocol, or file transfer protocol. The original information data stream is completely captured by the system in digital form and temporarily stored in the system memory buffer, waiting for subsequent text preprocessing operations. For example, the system can receive text information encapsulated in JSON format by collecting preset API endpoints, or read newly added text files in a specified directory in real time through a file monitoring component, so as to achieve automatic collection of original information from different sources. The text preprocessing stage aims to improve the accuracy and efficiency of subsequent semantic analysis. The system first uses a rule-based word segmentation engine, such as the Jieba word segmenter or HanLP word segmenter, to accurately segment the original information text, splitting the continuous text sequence into independent word units. Subsequently, the system executes a noise elimination program, which removes non-semantic related components in the text according to preset regular expression rules, including but not limited to: various punctuation marks, special characters, HTML / XML tags, and URL links. Further, the system loads a predefined stop word list, which contains high-frequency but low-semantic contribution words such as "de", "shi", "zai", etc. The system precisely matches the word segmentation results with the stop word list and removes all successfully matched stop words, finally outputting the preprocessed text after word segmentation, denoising, and stop word filtering, providing a standardized text data basis for the subsequent semantic feature extraction link.

[0075] S12: Extract semantic features from the preprocessed text to obtain semantic features;

[0076] In the embodiments of the present invention, the semantic feature extraction stage is the core step in constructing the information semantic fingerprint. Its goal is to extract a set of feature vectors from the preprocessed text that can represent the core semantics of the information. The system first uses the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm to calculate the weight of each word in the preprocessed text, thereby quantifying the importance of the word in the information content and forming an initial keyword weight vector. Secondly, the system uses topic model algorithms, such as Latent Dirichlet Allocation (LDA) or Non-Negative Matrix Factorization (NMF), to mine the topic distribution of the preprocessed text, identify the potential topic structure of the text, and output the probability distribution vector of the text on each topic. In addition, the system integrates a Named Entity Recognition (NER) model, such as a NER model based on BiLSTM-CRF or a NER model based on Transformer, to identify named entities such as person names, place names, and organization names in the text and construct a named entity set. Finally, the system uses a dependency parser, such as Stanford Parser or LTP Dependency Parser, to parse the dependency syntactic structure of the sentence and generate a dependency relationship graph. The system integrates the keyword weight vector, the topic probability distribution vector, the named entity set, and the dependency relationship graph to form a multi-dimensional semantic feature data set, comprehensively representing the semantic connotation of the information.

[0077] S13: Construct a semantic vector for the semantic features to obtain a semantic vector;

[0078] In the embodiments of the present invention, the semantic vector construction stage is responsible for converting the extracted multi-dimensional semantic features into a numerical vector representation form that can be processed by a computer. For the keyword weight vector, the system directly uses the weight value output by the TF-IDF algorithm as the vector dimension value to construct the keyword weight vector. For the topic probability distribution, the system directly uses the topic probability value output by the topic model as the vector dimension value to construct the topic distribution vector. For the named entity set, the system pre-constructs a named entity vector space, which can be trained based on word vector models such as Word2Vec, GloVe, or FastText. The system maps each named entity word in the named entity set to this vector space, obtains the corresponding entity vector, and performs average pooling or weighted average pooling operations on all entity vectors to generate a fixed-length entity vector. For the dependency relationship graph, the system uses graph embedding techniques, such as DeepWalk, Node2Vec, or GraphSAGE, to learn the low-dimensional vector representation of the nodes in the graph and embed the entire graph into a vector space of a fixed dimension to generate a relationship vector. The system finally concatenates or weighted fuses the keyword weight vector, the topic distribution vector, the entity vector, and the relationship vector to obtain a comprehensive, high-dimensional semantic vector, which can numerically represent the semantic content of the original information.

[0079] S14: Perform semantic hash encoding on the semantic vector to obtain a semantic hash code;

[0080] In an embodiment of the present invention, the semantic hash coding stage aims to compress the high-dimensional semantic vector into a low-dimensional hash code of a fixed length while maintaining the semantic similarity as much as possible. The system uses the local sensitive hashing (LSH) algorithm for semantic hash coding. The system pre-configures multiple LSH hash function families, each of which contains multiple hash functions. For the input semantic vector, the system uses the hash functions in each hash function family to perform hash calculations respectively to obtain multiple hash bucket numbers. The system combines the hash bucket numbers from different hash function families to form a hash code of fixed length. In order to improve the robustness and precision of the hash, the system adopts a multi-bucket strategy, that is, for each semantic vector, it is not mapped to only one hash bucket, but to multiple related hash buckets. The system can select a specific LSH algorithm such as random hyperplane hashing, MinHash or SimHash for implementation. For example, the system can use 128 random hyperplane hash functions to project the high-dimensional semantic vector into a 128-dimensional binary hash code space, each dimension represents a hash bucket, and finally output a 128-bit binary semantic hash code as a compressed representation of the information semantics.

[0081] S15: Generate a semantic fingerprint for the semantic hash code to obtain a semantic fingerprint;

[0082] In the embodiments of the present invention, the semantic fingerprint generation stage is the final step of the information fingerprint extraction module, and its purpose is to normalize the semantic hash code and generate the final semantic fingerprint identifier. First, the system converts the generated binary semantic hash code into a hexadecimal string representation for easy storage, transmission, and comparison. Subsequently, the system attaches metadata information to the generated hexadecimal semantic hash string to enhance the integrity and traceability of the semantic fingerprint. The attached metadata information includes, but is not limited to: the version identifier of the LSH algorithm used when generating the semantic hash code, the specific parameter configuration of the LSH algorithm (such as the number of hash function families, the size of the hash bucket, the random seed), and the version information of the NLP model used in the semantic feature extraction link. These metadata information are organized in a structured JSON format and concatenated with the hexadecimal semantic hash string to form a composite semantic fingerprint string. To ensure the uniqueness and anti-tampering property of the semantic fingerprint, the system uses a cryptographic hash algorithm, such as the SHA-256 algorithm, to calculate the hash of the composite semantic fingerprint string and generate the final semantic fingerprint hash value. This final hash value is used as the unique semantic identifier of the original information and is recorded by the system for subsequent origin node consensus, transfer path mapping, and information integrity verification. For example, the system can concatenate the JSON string of the version identifier "LSH_V1.2", the parameter configuration "buckets = 256, functions = 128, seed = 12345", and the NLP model version "BERT_NER_V2.1" with the hexadecimal semantic hash code, then calculate its hash value using the SHA-256 algorithm, and finally output a 256-bit hexadecimal string as the semantic fingerprint of the original information. This semantic fingerprint will serve as the unique identity of the information in the entire traceability management system.

[0083] Preferably, the semantic feature extraction of the preprocessed text includes:

[0084] Perform preliminary keyword extraction on the preprocessed text to obtain an initial keyword list;

[0085] Perform topic distribution mining on the preprocessed text to obtain a topic probability distribution;

[0086] Perform named entity recognition on the preprocessed text to obtain a named entity set;

[0087] Perform dependency syntactic analysis on the preprocessed text to obtain a dependency relationship graph;

[0088] Perform semantic feature fusion and enhancement on the initial keyword list, topic probability distribution, named entity set, and dependency relationship graph to obtain semantic features.

[0089] In the embodiments of the present invention, the preliminary keyword extraction aims to quickly screen out the core words in the preprocessed text, laying a foundation for subsequent in-depth semantic analysis. The system adopts a keyword extraction algorithm based on term frequency-inverse document frequency (TF-IDF). The system pre-loads a large-scale general corpus that covers texts in multiple fields. For the input preprocessed text, the system counts the term frequency (TF) of each word and calculates the TF-IDF weight score of each word by combining the pre-computed IDF value. The system sets a keyword quantity threshold, for example, set to the top 20 keywords. The system sorts all words in descending order according to the TF-IDF weight score and selects the words within the threshold ranking to form an initial keyword list.

[0090] Thematic distribution mining aims to reveal the potential thematic structure in the preprocessed text and capture the macro semantic tendency of the text. The system adopts a latent Dirichlet allocation (LDA) thematic model for thematic distribution mining. The system pre-trains the LDA thematic model, and the training corpus is a large-scale domain-related text dataset, such as information security field literature, blockchain technology white papers, etc. The system configures the number of themes of the LDA model, for example, set to 50 themes. For the input preprocessed text, the system inputs it into the pre-trained LDA model. The model analyzes the co-occurrence pattern of words in the text based on the Bayesian inference method and infers the thematic probability distribution of the text on the preset theme set. The system outputs a thematic probability distribution vector, the vector dimension is the number of themes, and each dimension value represents the probability that the text belongs to the corresponding theme.

[0091] Named entity recognition aims to identify entity words with specific meanings in the preprocessed text, such as person names, place names, organization names, etc., and capture the key entity information in the text. The system adopts a deep learning-based named entity recognition (NER) model, such as a NER model based on BiLSTM-CRF or Transformer. The system pre-trains the NER model, and the training data is a professional domain corpus annotated with named entity information, such as information security news, technical blogs, etc. The system configures the entity types recognized by the NER model. For the input preprocessed text, the system inputs it into the pre-trained NER model. The model predicts the entity type label of each word in the text one by one based on the sequence annotation method. The system post-processes the model prediction results, extracts all the words recognized as named entities, and classifies them according to the entity types to construct a named entity set.

[0092] Dependency parsing aims to reveal the grammatical dependencies between words in the preprocessed text and capture the deep structural semantic information of the sentence. The system uses a neural network-based dependency parser, and the system pre-trains the dependency parser with a large-scale Treebank corpus annotated with dependency syntactic relations. For the input preprocessed text, the system performs dependency parsing sentence by sentence. The parser analyzes the dependency relationships between the words in each sentence, such as subject-predicate relationships, verb-object relationships, attributive-middle relationships, etc. The system represents the dependency parsing result of each sentence as a directed graph, where the nodes in the graph represent the words in the sentence, the directed edges represent the dependency relationships between the words, and the labels on the edges indicate the types of dependency relationships. The system integrates the dependency relationship graphs of all sentences to form the overall dependency relationship graph of the preprocessed text. For example, for the sentence "Semantic hashing technology is applied to information traceability", the dependency parser analyzes "Semantic hashing technology" as the subject, "applied" as the predicate, and "to information traceability" as a prepositional object phrase, and annotates the types of dependency relationships between them, such as "subject-predicate relationship", "verb-object relationship", "prepositional object relationship", etc. The system represents and stores this dependency relationship information in the form of a graph structure as an important feature for characterizing the structural semantics of the text.

[0093] The semantic feature fusion and enhancement stage aims to integrate various semantic features extracted in the previous steps to improve the comprehensiveness and robustness of semantic representations. The system adopts a multi-strategy feature fusion method. First, for the initial keyword list, the system uses semantic knowledge bases such as WordNet or HowNet to expand the semantics of keywords, such as synonym expansion, near-synonym expansion, and hypernym expansion, to expand the coverage of the keyword list and enhance the semantic integrity of keyword representations. Second, the system performs dimensionality reduction on the topic probability distribution. For example, dimensionality reduction algorithms such as principal component analysis (PCA) or singular value decomposition (SVD) are used to reduce the dimension of the topic vector, reduce redundant information, and improve computational efficiency. Then, the system performs entity linking on the named entity set. Using entity linking technology, named entities are linked to normalized entities in knowledge graphs or knowledge bases such as Wikidata to obtain richer entity semantic information, such as entity attributes, entity relationships, etc., to enhance the semantic depth of entity representations. Finally, the system fuses the structured information of the dependency relationship graph. The system can extract key subgraph patterns in the dependency relationship graph, such as frequent subgraphs, maximum common subgraphs, etc., as structured semantic features; or, the system can use graph neural networks (GNNs), such as GCN or GAT, to learn the graph embedding representation of the dependency relationship graph and convert the structured information into a low-dimensional vector representation. The system integrates the enhanced keyword list, the dimensionality-reduced topic probability distribution, the named entity set after entity linking, and the dependency relationship graph after graph embedding to form a final set of semantic features. The integration method can be simple concatenation, weighted fusion, or dynamic fusion based on the attention mechanism. The finally obtained set of semantic features will be used as the basis for constructing the information semantic fingerprint to comprehensively and robustly represent the semantic content of the original information. For example, the system can add the expanded vocabulary of the keyword list to the feature vector, concatenate the dimensionality-reduced topic probability vector with the entity vector and the graph embedding vector to form a high-dimensional comprehensive semantic feature vector.

[0094] Preferably, the functions of the origin node consensus module include:

[0095] Encapsulate the origin information of the semantic fingerprint to obtain the information to be consensus;

[0096] Submit the information to be consensus to the network for broadcasting to obtain a broadcast request;

[0097] Perform a preliminary node verification on the broadcast request to obtain the information to be verified;

[0098] Propose and vote on the information to be verified to obtain a consensus vote;

[0099] Confirm the consensus result based on the preset consensus algorithm for the consensus vote to obtain the consensus result;

[0100] Generate the source point certificate according to the consensus result to obtain the source point certificate.

[0101] In the embodiment of the present invention, the origin information encapsulation step aims to integrate the semantic fingerprint with the necessary metadata to form a data packet that can be verified by consensus. The system obtains the semantic fingerprint generated by the information fingerprint extraction module and collects the metadata related to the information origin. The metadata at least includes: the digital identity identifier of the information creator, the timestamp of information creation (accurate to the millisecond level), the information type (such as text, picture, video), and the information access permission control policy. The system encapsulates the semantic fingerprint and the metadata using a predefined structured data format, such as JSON or Protocol Buffer. The encapsulated data packet contains a "semantic_fingerprint" field (storing the semantic fingerprint), a "creator_id" field (storing the creator identity identifier), a "timestamp" field (storing the creation timestamp), an "information_type" field (storing the information type), and an "access_policy" field (storing the access permission control policy). To ensure data integrity and source credibility, the information creator uses its private key to digitally sign the encapsulated data packet based on the elliptic curve digital signature algorithm (ECDSA) or the RSA digital signature algorithm. The digital signature result is added to the data packet as the "signature" field, and finally forms the "information to be consensus", preparing for the subsequent network broadcast submission step.

[0102] The network broadcast submission step is responsible for efficiently and reliably delivering the encapsulated "information to be consensus" to each node in the consensus network. The system adopts network broadcast technologies based on gRPC or ZeroMQ to build an efficient message broadcast channel. The information origin node, as the broadcast initiator, broadcasts the "information to be consensus" and its own node identity through the broadcast channel to the pre-configured list of consensus nodes. The broadcast message adopts a point-to-multipoint communication mode to ensure that the message can be sent to multiple consensus nodes simultaneously. To ensure the reliability of message transmission, the broadcast protocol adopts a message confirmation mechanism and a retransmission mechanism. After receiving the broadcast message, the consensus node sends a reception confirmation message (ACK) to the broadcast initiator. If the broadcast initiator does not receive the ACK within the preset time, the message is retransmitted until all target consensus nodes' ACKs are received or the maximum retransmission times are reached. The content of the broadcast message includes the "information to be consensus" data packet itself and the node identity of the broadcast initiator. The system records the sending logs of all broadcast messages, including message IDs, sending times, target node lists, sending statuses, etc. For example, the origin node sends the encapsulated JSON-format "information to be consensus" data packet and its own node ID to all pre-configured consensus nodes at once through the gRPC broadcast channel and waits to receive the ACK confirmations from each node.

[0103] The node preliminary verification step aims to quickly verify the legality of the request and the correctness of the data format after the consensus node receives the broadcast request, filter out invalid requests, and prepare for the subsequent consensus voting session. After receiving the broadcast request, the consensus node first verifies whether the node identity of the broadcast initiator is in the pre-registered list of trusted nodes. If the initiator's identity fails the verification, the request is directly rejected and an exception log is recorded. After the identity verification passes, the consensus node performs a format verification on the received "information to be consensus" data packet. After the format verification passes, the consensus node uses the public key of the broadcast initiator and, based on the digital signature algorithm, verifies whether the digital signature of the "information to be consensus" data packet is valid. A failed signature verification indicates that the data may have been tampered with, and the consensus node will reject the request and record a tampering warning log. After all verification steps pass, the consensus node marks the "information to be consensus" that has passed the preliminary verification as "information to be verified" and puts it into the consensus queue to wait for the subsequent consensus proposal and voting session. The preliminary verification process should be completed at the millisecond level to ensure the processing efficiency of the consensus network. For example, after receiving the broadcast request, the consensus node first queries the local trusted node list to verify the identity ID of the broadcast initiator, then uses JSON Schema to verify the JSON format of the "information to be consensus", and uses the public key of the initiator to verify the digital signature. After all verifications pass, the information is marked as "information to be verified" and enters the consensus process.

[0104] The consensus proposal and voting steps are the core links of the origin node consensus module, aiming to reach a consensus on the origin of information and the validity of semantic fingerprints through distributed voting. The leader node or rotating node in the consensus network is responsible for initiating the consensus proposal. The proposal content includes the data packet of the "information to be verified" and the identification of the consensus round proposed. After receiving the consensus proposal, the consensus node enters the voting stage. Each consensus node independently evaluates and votes on the "information to be verified" based on its preset consensus strategy and local data. The evaluation content includes: whether the identity of the information creator is trustworthy, whether the information type conforms to the system specification, and whether the semantic fingerprint is consistent with the locally recalculated semantic fingerprint (optional step, which can be selected according to performance requirements). The consensus node generates a voting ballot according to the evaluation result. The ballot content includes at least: the identity identification of the voting node, the voting result ("agree" or "reject"), and the voting timestamp. To ensure the authenticity and non-repudiation of the vote, the consensus node digitally signs the voting ballot with its private key. The signed voting ballot is sent to the consensus coordination node (which can be the leader node or an independent coordination node) through the network. The consensus coordination node is responsible for collecting all the voting ballots of the consensus nodes to form a "consensus ballot" set, providing a data basis for the subsequent consensus result confirmation link.

[0105] The consensus result confirmation step is designed to, after the consensus coordinating node has collected sufficient "consensus votes", statistically analyze the voting results according to a preset consensus algorithm, and finally determine whether consensus is reached and the specific result of the consensus. After receiving the "consensus votes" from all consensus nodes participating in the consensus, the consensus coordinating node first verifies the digital signature of each vote to ensure the authenticity and integrity of the vote and filters out invalid votes. Then, the consensus coordinating node statistically analyzes the valid votes according to a preset consensus algorithm, such as the Practical Byzantine Fault Tolerance (PBFT) algorithm, the Raft algorithm, or the Paxos algorithm. Taking the PBFT algorithm as an example, the consensus coordinating node counts the number of "agree" votes and "reject" votes. If the number of "agree" votes exceeds two-thirds of the total number of votes, it is considered that consensus is reached and the consensus result is "consensus passed". If the number of "reject" votes or "invalid votes" (such as votes not submitted due to timeout) reaches or exceeds one-third of the total number of votes, it is considered that consensus is not reached and the consensus result is "consensus not passed". The consensus result needs to clearly indicate whether consensus is reached and the specific resolution for reaching consensus (such as "consensus passed" or "consensus not passed"). The consensus coordinating node broadcasts the consensus result and all the vote information of the participating votes to notify all consensus nodes, ensuring the transparency and verifiability of the consensus result. For example, after the consensus coordinating node has collected the voting ballots of all consensus nodes, verified the signatures, and counted that the number of "agree" votes is 22, the number of "reject" votes is 2, and the total number of votes is 30. Since the number of "agree" votes exceeds two-thirds (30 * 2 / 3 = 20), consensus is reached, the consensus result is "consensus passed", and this result is broadcast to all consensus nodes.

[0106] The source point credential generation step is the final link of the origin node consensus module, which is designed to generate a trusted digital credential when consensus is reached. When the consensus result is "consensus passed", the consensus coordinating node is responsible for generating the source point credential. To ensure the authority and immutability of the source point credential, the consensus coordinating node uses its private key to digitally sign all the content of the source point credential based on the digital signature algorithm. The digital signature, as an integral part of the credential, ensures the credibility of the credential's source and the integrity of the content. The generated source point credential is stored and published in the form of a structured digital certificate. The digital certificate can adopt the X.509 standard format or a custom JSON format. Once the source point credential is generated, it serves as the authoritative basis for information traceability. The source point credential needs to be persistently stored in a trusted storage medium, such as a distributed ledger or a trusted database, for long-term preservation and query. For example, after consensus is reached, the consensus coordinating node encapsulates the semantic fingerprint, creator ID, creation timestamp, consensus timestamp, list of consensus nodes participating, and the "consensus passed" result in accordance with the X.509 digital certificate standard format, signs it with the private key of the coordinating node, generates the final source point credential, and stores it in the blockchain system.

[0107] Preferably, the transfer path mapping module includes the following functions:

[0108] Trigger transfer actions based on the source point voucher to obtain actions to be recorded;

[0109] Conduct content change analysis based on the actions to be recorded to obtain the changed semantic fingerprint;

[0110] Construct transfer events based on the actions to be recorded and the changed semantic fingerprint to obtain transfer snapshots;

[0111] Append transfer records to the transfer snapshots to obtain records to be uploaded to the chain;

[0112] Write the records to be uploaded to the chain to the distributed ledger according to the preset distributed ledger to obtain on-chain transfer records;

[0113] Update the transfer trajectory based on the on-chain transfer records and the source point voucher to obtain transfer trajectory data.

[0114] In the embodiment of the present invention, the transfer action trigger step aims to monitor and capture various operation actions of users on the traced information in real time, laying a foundation for subsequent transfer path recording. The captured event data is encapsulated into a structured "action to be recorded" data object and sent to the transfer path mapping module for subsequent processing.

[0115] The content change analysis step aims to detect content changes and generate new semantic fingerprints when transfer actions involve information content modification to track the information evolution trajectory. After the system receives the "action to be recorded" data, it first determines the operation type. If the operation type is "MODIFY" (modify) or "QUOTE_MODIFY" (quote and modify), the content change analysis process is triggered. The system obtains the original information content before modification and the information content after modification. The acquisition method can be: retrieve the original version from the information storage system and obtain the modified version from the modification request submitted by the user. The system calls the information fingerprint extraction module to perform a complete semantic fingerprint extraction process on the modified information content. If the operation type is a non-modification operation (such as "FORWARD", "SHARE", "COPY"), the content change analysis process is skipped and the "changed semantic fingerprint" field is empty. The result of the content change analysis, that is, the "changed semantic fingerprint" (if any), will be used together with the "action to be recorded" data to construct transfer event records.

[0116] The transfer event construction step aims to integrate the "behavior to be recorded" data and the "changed semantic fingerprint" (if any) into a complete transfer event record, forming a node on the transfer path. The system receives the "behavior to be recorded" data and the "changed semantic fingerprint" (if any) as inputs. The system creates a structured "transfer snapshot" data object. The "transfer snapshot" data object will serve as a basic unit on the transfer path. For example, for the "behavior to be recorded" data of a user forwarding traced information, the transfer path mapping module creates a "transfer snapshot".

[0117] The transfer record appending step aims to temporarily store the constructed "transfer snapshot" data in the buffer to be written to the distributed ledger, preparing for subsequent batch blockchain operations and improving system efficiency. The system maintains a memory buffer. When the system completes the construction of the "transfer snapshot", it appends the "transfer snapshot" data to the end of the buffer queue. The system can set the size threshold and time threshold of the buffer. When the number of "transfer snapshots" in the buffer reaches the threshold (e.g., 100), or the time interval since the last blockchain operation exceeds the threshold (e.g., 5 minutes), the system triggers a batch blockchain operation. Before the "transfer snapshots" in the buffer are written to the distributed ledger, they are in the "record to be blockchained" state. Through the buffer mechanism, frequent transfer event record operations can be combined into batch write operations, reducing the write pressure on the distributed ledger and enhancing the overall throughput of the system. For example, the system appends the constructed "transfer snapshot" JSON data to the memory buffer queue. When the number of snapshots in the queue reaches 100, or the time since the last blockchain operation exceeds 5 minutes, the system will trigger a batch blockchain operation and write these 100 "transfer snapshot" data to the distributed ledger at once.

[0118] The distributed ledger writing step aims to permanently store the transfer event data in the "records to be written to the chain" in an immutable and traceable manner in a preset distributed ledger system. The system uses a consortium chain or a private chain as the distributed ledger, such as Hyperledger Fabric, Ethereum private chain, etc. The system establishes a connection with the distributed ledger system and calls the transaction writing interface of the ledger system. The system packs all the "transfer snapshots" data in the "records to be written to the chain" buffer into one or more transactions and submits them to the distributed ledger system. Each transaction contains a batch of transfer event records. After receiving the transaction, the distributed ledger system, through the consensus verification of the consensus nodes (such as PBFT consensus), packs the transaction into a new block and appends it to the end of the blockchain. Each transfer event record is assigned a unique transaction hash value or block height on the chain as its on-chain identity identifier. After being written to the distributed ledger, the transfer event record becomes an "on-chain transfer record", which has immutability and traceability. Any modification to the transfer record will not pass the consensus verification, thus ensuring the credibility of the transfer path data. For example, the system packs 100 "transfer snapshots" data in the buffer into one transaction and submits it to the Hyperledger Fabric consortium chain. The Orderer node in the Fabric network sorts and packs the transaction and submits it to the Peer node for endorsement and verification. Finally, the transaction is written into a new block and appended to the Fabric chain, and these 100 transfer event records become "on-chain transfer records" and are permanently stored in the blockchain.

[0119] The transfer track update step aims to retrieve all the "on-chain transfer records" associated with a specific "source point voucher" from the distributed ledger and arrange them in chronological order to form the complete transfer track data of this information. When the transfer path of the information needs to be queried, the system receives a query request, which contains the "source point voucher" identifier of the target information. The system establishes a connection with the distributed ledger system and calls the query interface of the ledger system. The system uses the "source point voucher" identifier as an index to retrieve all the "on-chain transfer records" associated with this "source point voucher" from the distributed ledger. The distributed ledger system returns all the transfer records related to this "source point voucher". The system sorts the retrieved "on-chain transfer records" in ascending order according to the "event_timestamp" field in the records to form a list of transfer events arranged in chronological order. The sorted list of transfer events is the complete "transfer track data" of this information. The transfer track data can be output in a structured format such as JSON or XML.

[0120] Preferably, the content change analysis according to the behavior to be recorded includes:

[0121] Perform an operation type judgment on the record behavior to obtain the operation type result;

[0122] Obtain the content versions according to the operation type result to get the original version data and the modified version data;

[0123] Perform a fine-grained difference detection on the original version data and the modified version data to obtain the content difference segments;

[0124] Perform a semantic impact assessment on the content difference segments to obtain the degree of semantic change;

[0125] Generate a modified semantic fingerprint based on the degree of semantic change to obtain the changed semantic fingerprint.

[0126] In the embodiment of the present invention, the operation type judgment step is the entry of the content change analysis process, aiming to quickly determine whether the "record behavior to be processed" belongs to the content modification operation, so as to decide whether to perform subsequent content change analysis. The system receives the "record behavior to be processed" data as input and parses the "operation_type" field therein. The system predefines a set of operation type enumeration values, including "MODIFY" (modification), "QUOTE_MODIFY" (quote and modify), "FORWARD" (forward), "SHARE" (share), "COPY" (copy), etc. The system precisely matches the value of the "operation_type" field in the "record behavior to be processed" with the predefined modification operation type enumeration values (such as "MODIFY", "QUOTE_MODIFY"). If the match is successful, the operation type result is determined to be a "modification operation" (for example, the boolean value "TRUE"), indicating that the subsequent content change analysis process needs to be performed. If the match fails, the operation type result is determined to be a "non-modification operation" (for example, the boolean value "FALSE"), indicating that there is no need to perform content change analysis and the subsequent steps will be skipped. The operation type judgment result will be used as a process control condition to decide whether to execute subsequent steps such as content version acquisition, difference detection, semantic impact assessment, and semantic fingerprint generation. For example, when the system receives a "record behavior to be processed" with an operation type of "MODIFY", it matches it with the predefined modification operation type "MODIFY", and if the match is successful, the operation type result is determined to be a "modification operation", and the subsequent process will continue with the content change analysis.

[0127] The content version acquisition step is designed to obtain the information content before and after the modification when the operation type is a modification operation, providing a data basis for subsequent difference detection and semantic analysis. The system receives the "operation type result" as input. If the "operation type result" is a "modification operation" (e.g., the boolean value "TRUE"), the system starts the content version acquisition process. The system retrieves the original version data before the modification from the information storage system according to the "source point voucher" identifier of the information being operated on in the "behavior to be recorded". The retrieval method can be: querying the database records according to the source point voucher ID, or loading the original file from the content storage service according to the source point voucher ID. At the same time, the system obtains the modified version data from the modification request submitted by the user. The acquisition method can be: parsing the modified content from the API request parameters, or subscribing to the modified content from the message queue. The system outputs the retrieved original version data and the obtained modified version data as "original version data" and "modified version data" respectively. If the "operation type result" is a "non-modification operation" (e.g., the boolean value "FALSE"), the content version acquisition step is skipped, and both the "original version data" and the "modified version data" are empty. For example, if the operation type judgment result is a "modification operation", the system retrieves the document content before the modification from the database as the "original version data" according to the source point voucher ID "CERT_12345", and obtains the modified document content submitted by the user from the API request as the "modified version data" to prepare for subsequent difference detection.

[0128] The fine-grained difference detection step aims to accurately identify the differences between the original version data and the modified version data, including the specific content segments added, deleted, and modified, providing fine-grained difference information for semantic impact assessment. The system receives "original version data" and "modified version data" as inputs. The system uses line- and word-based text difference comparison algorithms, such as the Myers difference algorithm or the LCS (Longest Common Subsequence) algorithm, for fine-grained difference detection. The system first performs a line-level comparison of the original version data and the modified version data to identify added lines, deleted lines, modified lines, and unmodified lines. For modified lines, the system further performs a word-level comparison to identify added words, deleted words, modified words, and unmodified words. The system outputs the line-level and word-level difference information in a structured "content difference segment" data format. The "content difference segment" data can be represented in JSON or XML format and includes the difference type ("ADD", "DELETE", "MODIFY", "EQUAL"), the difference content (text segment), and the difference location information (line number, word index, etc.). The "content difference segment" data will be used as the input for the semantic impact assessment step. For example, when the system performs fine-grained difference detection on the original document and the modified document and identifies that line 3 is deleted, lines 5-7 are added content, and the word "error" in line 10 is modified to "correct", the system structures this difference information into "content difference segments".

[0129] The semantic impact assessment step aims to analyze the impact degree of the "content difference segment" on the overall semantics of the information, quantify the magnitude of semantic changes, and provide a reference for subsequent disposal decisions. The system receives the "content difference segment" data as input. The system uses a semantic similarity-based method for semantic impact assessment. First, the system separately calls the information fingerprint extraction module for the original version data and the modified version data, recalculates their semantic fingerprints, and obtains the "original version semantic fingerprint" and the "modified version semantic fingerprint". Then, the system uses a semantic similarity calculation method, such as cosine similarity, Hamming distance, or Jaccard similarity coefficient, to calculate the similarity score between the "original version semantic fingerprint" and the "modified version semantic fingerprint". The higher the similarity score, the smaller the degree of semantic change; the lower the similarity score, the greater the degree of semantic change. The system can output the semantic similarity score, or map the similarity score to a predefined semantic change degree level (such as "no change", "slight change", "medium change", "significant change") as the output result of the "degree of semantic change". And the generation of the final traceability verification report. For example, the system calculates the semantic fingerprint of the original document as "SF_ORIGIN", the semantic fingerprint of the modified document as "SF_MODIFIED", and then calculates the cosine similarity score of the two semantic fingerprints as 0.95. The system presets a similarity threshold, such as 0.9. If the similarity score is higher than 0.9, the degree of semantic change is considered to be "slight change". Therefore, the system outputs the "degree of semantic change" as "slight change".

[0130] The modified semantic fingerprint generation step is the final step of the content change analysis process, aiming to generate a new semantic fingerprint for the modified information content, which is used to update the transfer path record and support subsequent traceability and verification based on the modified content. The system receives the "degree of semantic change" as input. The system judges the "degree of semantic change". If the "degree of semantic change" reaches the preset "significant change" level threshold (for example, the semantic similarity score is lower than 0.7, or the degree of semantic change level is "significant change"), the system considers that the content modification has caused a substantial change in the core semantics of the information, and a completely new semantic fingerprint needs to be generated for the modified information. The system calls the information fingerprint extraction module to re-execute the complete semantic fingerprint extraction process on the "modified version data" to obtain the "changed semantic fingerprint". If the "degree of semantic change" does not reach the "significant change" level threshold (for example, the semantic similarity score is higher than 0.7, or the degree of semantic change level is "no change" or "minor change"), the system considers that the content modification has not caused a substantial change in the core semantics of the information, and the semantic fingerprint of the original information can be reused without generating a new "changed semantic fingerprint". At this time, the "changed semantic fingerprint" field is empty. The generated "changed semantic fingerprint" (if any) will be used as part of the transfer event record to update the transfer path of the information and be used in subsequent traceability and verification processes. For example, the "degree of semantic change" output by the semantic impact assessment step is "significant change", and the system determines that a new semantic fingerprint needs to be generated. The system calls the information fingerprint extraction module to recalculate the semantic fingerprint for the modified document content to obtain a new semantic fingerprint "SF_NEW", and outputs "SF_NEW" as the "changed semantic fingerprint".

[0131] Preferably, the fine-grained difference detection of the original version data and the modified version data includes:

[0132] Perform line-level preprocessing on the original version data and the modified version data to obtain the line-segmented original version and the line-segmented modified version;

[0133] Perform line-level difference comparison on the line-segmented original version and the line-segmented modified version to obtain the line-level difference result;

[0134] Perform word-level preprocessing on the line-level difference result to obtain the original line to be compared at the word level and the modified line to be compared at the word level;

[0135] Perform word-level difference comparison on the original line to be compared at the word level and the modified line to be compared at the word level to obtain the word-level difference result;

[0136] Perform difference segment structuring on the line-level difference result and the word-level difference result to obtain the content difference segment.

[0137] In the embodiments of the present invention, the line-level preprocessing operation aims to split the original version data and the modified version data into independent sequences of text lines, laying a foundation for subsequent line-level difference comparison. The text line splitting program receives the original version data and the modified version data as inputs. The program uses the line feed character as the line delimiter, traverses the entire original version data and the modified version data, and when a line feed character is encountered, extracts the text content before the line feed character as an independent text line. The program arranges all the extracted text lines in the order in which they appear in the original text, forming a line-split original version list and a line-split modified version list respectively. For example, if the original version data is a string containing multiple lines of text, the line-level preprocessing program splits the string into multiple string elements based on the line feed character and stores these string elements in a list to form the line-split original version; the same processing process is performed on the modified version data to obtain the line-split modified version. The preprocessed line-split version data provides structured input data for subsequent line-level difference comparison operations.

[0138] The line-level difference comparison operation aims to identify the text lines that have changed between the original version and the modified version and record the change types. The line-level difference comparison program receives the line-split original version list and the line-split modified version list as inputs. The program uses the Longest Common Subsequence (LCS) algorithm to compare the two line lists and identify the added lines, deleted lines, and modified lines. The LCS algorithm uses dynamic programming to find the longest common subsequence in the two sequences, and the lines not included in the longest common subsequence are determined to be different lines. The program records the difference type of each line, such as "added" (lines added in the modified version), "deleted" (lines deleted in the original version), "modified" (lines with changed content), and "unmodified" (lines with unchanged content). The line-level difference results are output in the form of structured data, such as JSON format, containing the text content of each line and its corresponding difference type label. For example, if a certain line in the original version is deleted in the modified version, in the line-level difference results, the text content of this line will be marked as the "deleted" type; if the content of a certain line changes in the modified version, this line will be marked as the "modified" type, and the content of the original version line and the content of the modified version line will be recorded at the same time.

[0139] The word-level preprocessing operation focuses on the text lines marked as "modified" in the line-level difference comparison results, preparing for the subsequent word-level difference comparison. The word-level preprocessing program receives the line-level difference results as input. The program traverses the line-level difference results and filters out the line records with the difference type of "modified". For each line record of the "modified" type, the program extracts the corresponding original version line text and the modified version line text respectively. The word-level preprocessing program performs word segmentation operations on the extracted original version line text and modified version line text. The word segmentation operation adopts a statistical-based word segmentation model, such as the HMM model or the CRF model, to split the continuous text line into independent word units. The preprocessing program removes the stop words in the word segmentation results. The stop word list is predefined and contains high-frequency but semantically meaningless words such as "of", "is", "in", etc. The program outputs the processed word lists, which are used as the original line to be compared at the word level and the modified line to be compared at the word level respectively, providing word-level input data for the subsequent word-level difference comparison. For example, if a line in the line-level difference results is marked as "modified", and the original version line is "Information security is an important issue", and the modified version line is "Network security is the core issue", then the word-level preprocessing program will perform word segmentation and stop word removal operations on these two lines of text respectively, and output the word list ["information", "security", "important", "issue"] as the original line to be compared at the word level, and the word list ["network", "security", "core", "issue"] as the modified line to be compared at the word level.

[0140] The word-level difference comparison operation targets the original line to be compared at the word level and the modified line to be compared at the word level after word-level preprocessing, identifies the changes at the word level, and records the change types. The word-level difference comparison program receives the list of original lines to be compared at the word level and the list of modified lines to be compared at the word level as input. For each pair of the original line list and the modified line list, the program uses an edit distance algorithm, such as the Levenshtein distance algorithm or the Wagner-Fischer algorithm, to calculate the edit distance between the two word sequences. The edit distance algorithm calculates the minimum number of single-character edit operations (including insertion, deletion, or replacement) required to convert one string to another through dynamic programming. Based on the results of the edit distance algorithm, the program identifies the added words, deleted words, and replaced words. The program records the difference type of each word, such as "added" (words added in the modified version line), "deleted" (words deleted in the original version line), "replaced" (replaced words), and "unchanged" (words remaining the same). The word-level difference results are output in a structured data form, such as a nested JSON format. Based on the line-level difference results, for the lines of the "modified" type, the difference information at the word level is detailedly recorded, including the original words, the modified words, and the difference type markers.

[0141] The differential segment structuring operation is responsible for integrating the results of line-level and word-level differential comparisons into structured content differential segments, facilitating subsequent semantic impact assessment and traceability verification report generation by the system. The differential segment structuring program takes the line-level and word-level differential results as input. The program traverses each row record in the line-level differential result and performs different processing based on the type of difference in the row.

[0142] When the difference type of the row record is "unmodified", the program determines that no change has occurred to the content of this row, so there is no need to generate a differential segment, and the program continues to process the next row record.

[0143] When the difference type of the row record is "added", the program determines that this row is a newly added row in the modified version. Then the program creates an "added row segment". This segment structurally represents the information of the added row and contains the following fields: the "segment type" field, marked as "added row"; the "line number" field, recording the line number of the added row in the modified version; the "content" field, storing the complete text content of the added row. For example, if the line-level differential result indicates that the 5th row is a newly added row in the modified version and the content is "This system uses blockchain technology", then the program generates an added row segment with the "segment type" being "added row", the "line number" being 5, and the "content" being "This system uses blockchain technology".

[0144] When the difference type of the row record is "deleted", the program determines that this row is a deleted row in the original version. Then the program creates a "deleted row segment". This segment structurally represents the information of the deleted row and contains the following fields: the "segment type" field, marked as "deleted row"; the "line number" field, recording the line number of the deleted row in the original version; the "content" field, storing the complete text content of the deleted row. For example, if the line-level differential result indicates that the 10th row is deleted in the original version and the content is "Traditional centralized management mode", then the program generates a deleted row segment with the "segment type" being "deleted row", the "line number" being 10, and the "content" being "Traditional centralized management mode".

[0145] When the difference type of the current line record is "modification", the program determines that the content of this line has been modified. The program further parses the word-level difference results to obtain the difference information at the word level corresponding to this line. The program creates a "modified line segment". This segment structurally represents the information of the modified line and contains the following fields: a "segment type" field marked as "modified line"; a "line number" field that records the line numbers of the modified line in the original version and the modified version (the line numbers are usually the same); a "line-level difference type" field marked as "modification"; and a "word difference list" field that stores the list of difference information at the word level. For the "word difference list" field of the "modified line segment", the program further traverses each word difference record corresponding to this line in the word-level difference results. If the word difference type is "not modified", the program ignores this word and continues to process the next word. If the word difference type is "added", the program creates an "added word segment" and adds it to the "word difference list". The "added word segment" contains the following fields: a "word type" field marked as "added word"; a "word content" field that stores the added word text; and a "word position" field that records the word index position of the added word in the line of the modified version. If the word difference type is "deleted", the program creates a "deleted word segment" and adds it to the "word difference list". The "deleted word segment" contains the following fields: a "word type" field marked as "deleted word"; a "word content" field that stores the deleted word text; and a "word position" field that records the word index position of the deleted word in the line of the original version. If the word difference type is "replaced", the program creates a "replaced word segment" and adds it to the "word difference list". The "replaced word segment" contains the following fields: a "word type" field marked as "replaced word"; an "original word" field that stores the original word text to be replaced; a "modified word" field that stores the modified word text after replacement; and a "word position" field that records the word index position of the replaced word in the line of the original version and the word index position of the replaced word in the line of the modified version.

[0146] The program arranges all the generated "added line segments", "deleted line segments", and "modified line segments" in the order of their line numbers in the text to form a final list of content difference segments. The list of content difference segments is output in a structured data format, such as a JSON array, and each array element is a difference segment object. The structured content difference segments provide detailed and accurate change information for the subsequent semantic impact assessment module, facilitating the system to analyze the potential impact of content changes on the semantic integrity of information and the traceability chain.

[0147] Preferably, the information integrity verification module includes the following functions:

[0148] Receive a verification request based on the source point certificate to obtain the information to be verified; reconstruct the current semantic fingerprint of the information to be verified to obtain the current semantic fingerprint;

[0149] Compare the current semantic fingerprint with the information to be verified to obtain a fingerprint comparison result;

[0150] Trace back the transfer path based on the information to be verified and the transfer track data to obtain the verified path;

[0151] Verify the behavior consistency of the verified path according to the current semantic fingerprint to obtain a behavior verification result;

[0152] Identify the abnormal mode of the verified path to obtain an abnormal behavior report;

[0153] Summarize the verification results of the fingerprint comparison result, the behavior verification result, and the abnormal behavior report to obtain a report to be signed; generate a verification report for the report to be signed to obtain a traceability verification report.

[0154] In the embodiment of the present invention, the verification request receiving program continuously collects information of a preset verification request receiving port. This port follows a predefined communication protocol, such as the HTTPS protocol, and provides a verification service interface externally. When a verification request is received, the system immediately parses the request data packet. The request data packet must contain a source point certificate, which is the identity identifier for information traceability and is used to associate the information to be verified with the original information. The system first verifies the validity of the source point certificate, such as checking the certificate signature, validity period, and whether it is a legal certificate issued by the system. After the certificate verification passes, the system extracts the information to be verified from the request data packet. The information to be verified is the digital information content that needs to be verified for integrity, and its data format is consistent with the data format of the original information. The system transfers the information to be verified that has been successfully received and initially verified, and the corresponding source point certificate, to the subsequent semantic fingerprint reconstruction link for processing.

[0155] The current semantic fingerprint reconstruction program receives the information to be verified as input. This program reuses the processing logic and algorithms of the information fingerprint extraction module to recalculate and generate the semantic fingerprint of the information to be verified. The program first performs text preprocessing operations on the information to be verified, including word segmentation, denoising, and stop word removal, to obtain the preprocessed text. Subsequently, the program extracts semantic features from the preprocessed text. The extraction process uses exactly the same semantic feature extraction methods as in the original information fingerprint extraction stage, such as the TF-IDF algorithm, topic models, named entity recognition, and dependency syntactic analysis, to obtain a multi-dimensional semantic feature set. The program then converts the semantic feature set into a semantic vector representation. The vector construction method is consistent with that in the original information fingerprint extraction stage, such as using word vector models, graph embedding techniques, etc. Finally, the program performs semantic hash encoding on the semantic vector. The encoding algorithm and parameter configuration are the same as those in the original information fingerprint extraction stage, such as using the locality-sensitive hashing algorithm, to obtain a fixed-length semantic hash code. The reconstruction program encapsulates the generated semantic hash code and metadata information such as the algorithm version and parameter configuration used during the reconstruction process to generate the current semantic fingerprint. The current semantic fingerprint serves as the semantic identity identifier of the information to be verified.

[0156] The source point fingerprint comparison program receives the current semantic fingerprint, the information to be verified, and the source point voucher as input. The program extracts the source point semantic fingerprint of the original information from the source point voucher. The comparison program uses the Hamming distance calculation method to calculate the Hamming distance between the current semantic fingerprint and the source point semantic fingerprint. The Hamming distance measures the number of different characters at corresponding positions between two binary strings of equal length. The system presets a fingerprint comparison threshold, which is set in advance according to the system's tolerance for semantic similarity. The program compares the calculated Hamming distance with the preset threshold. If the Hamming distance is less than or equal to the preset threshold, it is determined that the current semantic fingerprint matches the source point semantic fingerprint successfully, and the fingerprint comparison result is "consistent". If the Hamming distance is greater than the preset threshold, it is determined that the current semantic fingerprint does not match the source point semantic fingerprint successfully, and the fingerprint comparison result is "inconsistent". The fingerprint comparison result is output in the form of a boolean value, and the Hamming distance value can be attached at the same time.

[0157] The transfer path backtracking program receives the information to be verified and the transfer trajectory data as inputs. First, the program queries all transfer event records associated with the source point voucher in the transfer trajectory data storage system according to the source point voucher in the information to be verified. The transfer trajectory data storage system, such as a distributed ledger or a relational database, stores all transfer path information of the information starting from the origin node. The query operation is indexed based on the source point voucher to retrieve all on-chain transfer records or database records related to the source point voucher. The program sorts the retrieved transfer event records in chronological order to restore the complete transfer path of the information from the source point to the current state. The program conducts a preliminary verification on the backtracked transfer path, such as checking the chronological order of the transfer event timestamps, the integrity of the event association relationship, and whether there are broken chains. Further, the program conducts an integrity verification on each transfer event record in the transfer path. Assuming that each transfer event record contains the hash value of the event data and the hash value of the previous event (forming a chain structure), the program will recalculate the hash value of the current event data and compare it with the hash value stored in the record to verify whether the event data has been tampered with. At the same time, the program verifies whether the hash value of the previous event stored in the current event record is consistent with the actual hash value of the previous event to verify the continuity and integrity of the transfer chain. If the hash verification of any transfer event fails, the transfer path is determined to be an invalid path and the transfer path backtracking operation fails. If the hash verification of all transfer events is successful and the chronological order and association relationship of the transfer events are correct, the program marks the transfer path as a "verified path". The verified path is output in a structured data form, such as an ordered list of transfer event records, where each element in the list represents a transfer event that has passed the integrity verification, including detailed information such as the event occurrence time, event type, event operator, semantic fingerprint before change, semantic fingerprint after change, and event data hash value. The verified path, as a reliable record of the information transfer history, is passed to the subsequent behavior consistency verification link.

[0158] The behavior consistency verification program receives the current semantic fingerprint and the verified path as inputs. The program traverses each transfer event record in the verified path. For each event, the program analyzes the event type and the semantic fingerprint change information recorded in the event. The program compares the "semantic fingerprint before change" recorded in the current event record with the "semantic fingerprint after change" in the previous event record to verify whether the evolution track of the semantic fingerprint is logical. For transfer events of the "modification" type, the program further compares the current semantic fingerprint with the "semantic fingerprint after change" recorded in the event record. Ideally, if the transfer path record is complete and the information has not been tampered with, the information state pointed to by the final "verified path" should be consistent with the content of the "information to be verified". Therefore, the "current semantic fingerprint" obtained by reconstructing the "information to be verified" should match the "semantic fingerprint after change" recorded in the last "modification" event (or the starting event, if there is no modification) in the "verified path". The program compares the "current semantic fingerprint" with the "semantic fingerprint after change" of the last valid event in the "verified path". The comparison method can use semantic fingerprint similarity calculation, such as Hamming distance calculation, and a similarity threshold is set. If the similarity is higher than the threshold, it is determined that the behavior consistency verification passes, and the behavior verification result is "consistent". If the similarity is lower than the threshold, it is determined that the behavior consistency verification fails, and the behavior verification result is "inconsistent", indicating that there is a deviation between the actual information state and the transfer path recorded in the traceability system, and there may be abnormal situations such as unexpected modification of information content or incomplete transfer path record. The behavior verification result is output in the form of a boolean value, and at the same time, the semantic fingerprint similarity value and the detailed information of the inconsistency can be attached.

[0159] The anomaly pattern recognition engine receives the verified path and the behavior verification result as input. The engine predefines multiple anomaly behavior patterns, such as: "Unauthorized modification" pattern: There is a modification event by an unauthorized user in the verified path; "Jump - style transfer" pattern: Necessary intermediate transfer links are missing in the verified path; "Content tampering suspicion" pattern: The behavior consistency verification result is "inconsistent", and the fingerprint similarity is lower than the significant threshold; "Malicious sharing and diffusion" pattern: Information is shared to high - risk or unauthorized target objects; "Abnormal operation time" pattern: The occurrence time of the transfer event does not match the normal working hours or operation habits. The engine traverses the verified path and, in combination with the behavior verification result, detects whether there are predefined anomaly behavior patterns one by one. The detection methods include rule matching, statistical analysis, machine - learning anomaly detection algorithms, etc. For example, for the "Unauthorized modification" pattern, the engine checks the identity of the operator of each modification event, compares it with the permission list, and determines whether it is an authorized user. For the "Content tampering suspicion" pattern, the engine directly refers to the result of the behavior consistency verification. For the "Jump - style transfer" pattern, the engine analyzes the transfer event sequence and checks whether there is a missing expected intermediate state or link. If any anomaly behavior pattern is detected, the engine generates a corresponding anomaly behavior report. The anomaly behavior report details the anomaly type, the occurrence time of the anomaly, the transfer events involved, the potential risk level, and the recommended handling measures. If no anomaly behavior pattern is detected, an empty anomaly behavior report is generated, indicating that no anomaly is found. The anomaly behavior report is output in the form of structured data. For example, if the anomaly pattern recognition engine analyzes the verified path and finds that the identity of the operator of a "modification" event is an "unauthorized user", the engine determines that the "Unauthorized modification" anomaly pattern is triggered and generates an anomaly behavior report. The report type is "Unauthorized modification", the risk level is "high", and the recommended measure is "Immediately roll back the modification operation and lock the unauthorized user account".

[0160] The verification result summary and report generation program receives the fingerprint comparison result, the behavior verification result, and the anomaly behavior report as input. The program integrates and summarizes these three parts of verification results to form a complete report to be signed. The report to be signed is organized in the form of structured data, such as JSON or XML format,

[0161] After the report to be signed is generated, the verification report generation process is entered. The verification report generation program receives the report to be signed as input. To ensure the immutability and authority of the traceability verification report, the system will digitally sign the report to be signed. The digital signature process uses asymmetric encryption technology, such as the RSA algorithm or the ECDSA algorithm. The system uses the private key of the verification server to encrypt the hash value of the report to be signed to generate a digital signature. The digital signature and the original text of the report to be signed together constitute the final traceability verification report. The data format of the traceability verification report can be a PDF document, a JSON file, or an XML file. The report contains the complete content of the report to be signed and additional digital signature information. To facilitate the verification of the authenticity of the traceability verification report, the report usually contains the public key information of the verification server or provides a way to obtain the public key. The verification party that receives the traceability verification report can use the public key of the verification server to decrypt the digital signature in the report to obtain the decrypted hash value, and recalculate the hash value of the report text to compare whether the two hash values are the same. If they are the same, it indicates that the traceability verification report has not been tampered with and is indeed issued by the verification server, ensuring the authenticity and credibility of the report. The finally generated traceability verification report is stored by the system and can be pushed to the requester according to the needs of the verification requester, or report query and download services are provided. For example, the traceability verification report can be presented in PDF format. The title page of the report shows the title of "Traceability Verification Report", including the summary of verification information, fingerprint comparison conclusion, behavior verification conclusion, content of abnormal behavior report, summary of verified transfer path, overall evaluation of verification conclusion, etc. The last page of the report is attached with the digital signature of the verification server and a QR code for obtaining the public key to ensure the authority and verifiability of the report.

[0162] Preferably, the behavior consensus feedback module includes the following functions:

[0163] Distribute the verification report for the traceability verification report to obtain the report to be consensus;

[0164] Start the consensus process according to the report to be consensus to obtain a consensus proposal;

[0165] Conduct node evaluation and voting according to the report to be consensus and the consensus proposal to obtain disposal ballots;

[0166] Summarize the voting results of the disposal ballots to obtain a consensus ballot warehouse;

[0167] Form a consensus decision according to the preset consensus algorithm and the consensus ballot warehouse to obtain a draft consensus resolution;

[0168] Execute the disposal measures according to the draft consensus resolution to obtain an execution instruction; record the execution result according to the execution instruction to obtain a disposal record;

[0169] Finalize the disposal resolution by solidifying the consensus resolution draft, the consensus vote repository, and the disposal records to obtain the disposal resolution data.

[0170] In an embodiment of the present invention, the verification report distribution program is responsible for delivering the traceability verification report generated by the information integrity verification module to each node in the consensus network to initiate the subsequent consensus decision-making process. The distribution program first receives the traceability verification report from the information integrity verification module. The program then constructs a message containing the report, where the message header includes the unique identifier of the report, the report type (traceability verification report), and the target consensus network identifier. The message body encapsulates the complete content of the traceability verification report, and the report content is in a predefined structured data format, such as JSON or ProtocolBuffers. The distribution program broadcasts the constructed message to all participating nodes in the consensus network according to the preset consensus network communication protocol, such as a point-to-point communication protocol based on gRPC or ZeroMQ. After each node in the consensus network receives the broadcast message, it parses the message header, identifies the message type as the traceability verification report, and stores the report content in the message body in the local message queue as the report to be consensus, waiting for subsequent consensus process handling. For example, the verification report distribution program encapsulates the generated traceability verification report into a gRPC message, and the target address list of the message is the pre-configured list of IP addresses of the consensus network nodes. The program loops through the target address list and sends gRPC messages to the target nodes one by one to ensure that each consensus node receives the report to be consensus.

[0171] The consensus process startup program monitors the local message queue of the consensus network nodes in real time to detect whether a new report to be consensus arrives. When a new report to be consensus is detected, the program retrieves the report from the message queue. The program parses the content of the report to be consensus and extracts key information such as the overall evaluation of the verification conclusion, the fingerprint comparison result, the behavior verification result, and the abnormal behavior report in the report. The program then generates a consensus proposal, which is an instruction to initiate the consensus voting process. The consensus proposal includes summary information of the report to be consensus, such as the report ID, the information of the verification object, the overall evaluation of the verification conclusion, etc., as well as the goal and scope of the consensus. The goal of the consensus is usually to form a consensus resolution based on the conclusion of the traceability verification report. For example, if the verification report indicates that the information integrity verification fails, the goal of the consensus may be to determine the disposal measures to be taken. The scope of the consensus defines the scope of nodes participating in the consensus voting, such as all consensus nodes or nodes with specific roles. The consensus process startup program broadcasts the generated consensus proposal to the consensus network to notify all relevant nodes to initiate the consensus voting process for the traceability verification report.

[0172] The node evaluation and voting process runs on each node in the consensus network. The process receives the consensus proposal broadcast by the consensus process initiator and obtains the reports to be consensus stored locally. First, according to the objectives and scope of the consensus proposal, the process determines whether it is a node participating in this consensus vote. If it is determined to be a participating node, the process begins to evaluate the traceability verification report. The evaluation process is mainly based on information such as the overall evaluation of the verification conclusion, fingerprint comparison results, behavior verification results, and abnormal behavior reports in the report to be consensus. The node makes an independent judgment on the conclusion of the report according to the preset evaluation strategy and the information it holds, and forms a disposal opinion based on the judgment result. The disposal opinion usually manifests as a vote selection for different disposal measures, such as "maintain the status quo", "give an alarm", "manual intervention", "data rollback", etc. The process encapsulates the node's disposal opinion and node identity information to generate a disposal ballot. The disposal ballot is signed using digital signature technology to ensure the immutability and traceability of the ballot. The node evaluation and voting process sends the generated disposal ballot to the voting result summary process to participate in the subsequent voting result statistics and consensus decision-making. For example, after a consensus node receives a consensus proposal for a certain traceability verification report, the node process automatically reads the report and analyzes that the report conclusion is "the information integrity verification fails, and there is an unauthorized modification exception". According to the preset risk assessment strategy, the node process determines that the risk level of this abnormal behavior is "high", votes to select the disposal measures of "data rollback" and "give an alarm", generates a disposal ballot containing the vote selection and node signature, and sends it to the voting result summary node.

[0173] The voting result aggregation program is responsible for collecting the disposal votes of all participating nodes in the consensus network, aggregating and counting them to form a consensus vote repository. The voting result aggregation program continuously monitors the preset voting reception port to receive the disposal votes sent by each consensus node. After the program receives a disposal vote, it first verifies the digital signature of the vote to ensure the validity of the vote and the credibility of its source. The votes that pass the verification are stored in the local consensus vote repository data structure. The consensus vote repository usually adopts data structures such as hash tables or linked lists, and stores all valid votes for the report indexed by the report ID. The program continuously collects votes until the preset voting deadline is reached or the preset minimum number of voting nodes threshold is met. When the voting end condition is satisfied, the voting result aggregation program stops receiving new votes and starts statistical analysis of the votes in the consensus vote repository to provide a data basis for the subsequent consensus decision-making process. For example, the voting result aggregation program is deployed on a designated node in the consensus network and is responsible for receiving the voting information of all consensus nodes. The program receives a "data rollback" vote for report XXX sent by node A, and after verifying the signature, stores the vote in the consensus vote repository with report XXX as the key value. The program continuously receives votes from other nodes such as node B and node C until the voting time ends. The program stops receiving votes and starts counting the voting results.

[0174] The consensus decision-making program receives the consensus vote repository as input and calculates and makes decisions on the voting results in the consensus vote repository according to the preset consensus algorithm to form a draft consensus resolution. The program first extracts all valid votes from the consensus vote repository. The program then conducts statistical analysis on the voting results according to the preset consensus algorithm. The program calculates the voting data in the consensus vote repository according to the selected consensus algorithm to obtain the consensus determination result. The consensus determination result clearly indicates the finally adopted disposal measures, such as "perform data rollback operation", "send alarm notification", "initiate manual review process", or "maintain the current status". The program integrates the consensus determination result and the voting statistical data (such as the vote distribution of various disposal measures, the list of participating voting nodes, etc.) to generate a draft consensus resolution. The draft consensus resolution is the preliminary result of the consensus decision-making process and needs to go through subsequent execution and solidification links before it can finally come into effect. The draft consensus resolution is output in a structured data form, such as JSON format, containing information such as the type of consensus algorithm, the consensus determination result, the voting statistical data, the list of voting nodes, and the generation time of the resolution draft. For example, if the consensus algorithm is simple majority voting and the consensus vote repository shows that the "alarm" measure has received the majority of votes, the consensus decision-making program generates a draft consensus resolution with the content "Consensus Resolution: Adopt the disposal measure 'alarm', Voting Results: 'alarm' 5 votes, 'Manual Intervention' 3 votes, 'Maintain Status Quo' 2 votes, Participating Voting Nodes: [Node A, Node B, Node C, Node D, Node E, Node F, Node G, Node H, Node I, Node J]".

[0175] The disposal measure execution program receives the consensus resolution draft as input. The program parses the consensus resolution draft and identifies the final disposal measures formed by the consensus decision. The program calls the corresponding execution module or component according to different types of disposal measures to perform specific disposal operations. For example, if the consensus resolution is "data rollback", the program calls the data rollback module and passes relevant parameters (such as the ID of the information to be rolled back, the rollback target version, etc.) to perform the data rollback operation and restore the information data to a previous version state. If the consensus resolution is "alarm", the program calls the alarm notification component and sends an alarm notification to the relevant responsible person or monitoring system according to the preset alarm policy. The notification content includes information such as the alarm type, alarm object, alarm level, and the link to the traceability verification report. The execution program records the results of the disposal measure execution and generates a disposal record. The disposal record details the type of disposal measure executed, the execution time, the execution result (success or failure), the log information during the execution process, and the operator information, etc. The disposal record provides a basis for subsequent auditing and traceability. The execution instruction is the specific command that drives the execution of the disposal measure. According to different disposal measures, the content and form of the execution instruction are also different. For example, if the disposal measure is "data rollback", the execution instruction may include an SQL rollback statement or an API call instruction. If the disposal measure is "alarm", the execution instruction may include an email sending instruction or a message queue push instruction. The disposal record is output in a structured data form, such as JSON format, including information such as the type of disposal measure, the execution time, the execution status, the execution log, and the operator. For example, if the consensus resolution draft instructs to perform a "data rollback" operation, the disposal measure execution program calls the data rollback module and records the execution process log "Starting data rollback... Data version verification passed... Data rollback successful...", and finally generates a disposal record, recording the disposal measure as "data rollback", the execution status as "success", the execution time as the current time, and including the detailed execution log.

[0176] The final resolution solidification program receives the draft consensus resolution, consensus vote warehouse and disposal records as input. The program stores the key data in the consensus decision-making process, such as the draft consensus resolution, consensus vote warehouse and disposal records, persistently to form the final disposal resolution data. The final resolution data is usually stored in a highly reliable and tamper-proof data storage medium, such as a distributed ledger system or a dedicated audit database. The data solidification process uses hash anchoring technology to anchor the hash value of the disposal resolution data to the blockchain or other trusted timestamp service to ensure the immutability of the data and the authority of the timestamp. The final solidified disposal resolution data serves as the final output result of the behavioral consensus feedback link in the information traceability management process, providing irrefutable evidence for subsequent audits, supervision and dispute resolution. The disposal resolution data contains the complete process information of the consensus decision, including consensus proposals, voting ballots, voting result statistics, consensus resolution drafts, final disposal measures and disposal execution records, etc., forming a complete decision-making chain to achieve transparency and traceability of the behavioral consensus feedback process. The disposal resolution data is stored in the form of structured data.

[0177] Preferably, the consensus decision-making process based on a preset consensus algorithm and consensus ticket warehouse includes:

[0178] Count the voting results of the consensus vote warehouse to obtain voting statistics;

[0179] Apply the consensus algorithm according to the preset consensus algorithm and voting statistics to obtain the consensus judgment result;

[0180] Determine the disposal measures based on the consensus determination results and obtain the final disposal measures;

[0181] A consensus resolution draft is generated based on the consensus determination results, final disposal measures, and voting statistics to obtain a consensus resolution draft.

[0182] In an embodiment of the present invention, a voting result statistics program receives consensus ticket warehouse data as input, which contains the disposal votes of all participating consensus nodes. The program first traverses the consensus ticket warehouse, and for each ballot, parses the identity of the voting node recorded therein and the disposal measures selected by it. The program uses a hash mapping table or dictionary data structure, with the disposal measure type as the key and the vote counter as the value to initialize the data structure. The program traverses all ballots, and whenever it encounters a ballot, it extracts the disposal measure type selected by it, and increases the vote counter value of the corresponding disposal measure type in the hash mapping table by one. After the program completes the traversal of all ballots, the final vote count results of various disposal measures are stored in the hash mapping table. The voting statistics data is output in the form of structured data, such as JSON format, including each disposal measure type and its corresponding number of votes.

[0183] The consensus algorithm application module receives voting statistics data and a preset consensus algorithm rule as inputs. In this embodiment, the preset consensus algorithm is the simple majority voting algorithm. The rule of the simple majority voting algorithm is that among all alternative disposal measures, the disposal measure with the highest number of votes will be determined as the consensus decision result. The consensus algorithm application module first parses the voting statistics data to obtain the number of votes for each disposal measure. The program compares the number of votes for various disposal measures to find the disposal measure with the highest number of votes. If there are multiple disposal measures with the same and highest number of votes, the program makes a selection according to the preset priority rule. The priority rule can be set in advance. For example, it can be sorted according to the risk level or execution cost of the disposal measure, and the measure with the highest priority is selected. If there is no priority rule, the program can randomly select one of the disposal measures with the highest number of votes. The program outputs the finally selected disposal measure as the consensus determination result.

[0184] The disposal measure determination module receives the consensus determination result as an input. The consensus determination result is a text string output by the consensus algorithm application module, indicating the type of disposal measure selected by the consensus network voting, such as "alarm", "manual intervention", or "data rollback", etc. The disposal measure determination module directly converts the consensus determination result into a final disposal measure instruction executable by the system. The mapping relationship between the disposal measure type and the specific execution instruction is pre-configured inside the module. For example, if the consensus determination result is "alarm", the disposal measure determination module maps it to the instruction of "send alarm notification"; if the consensus determination result is "manual intervention", it is mapped to the instruction of "start manual review process"; if the consensus determination result is "data rollback", it is mapped to the instruction of "execute data rollback operation". The final disposal measure is output in the form of an instruction recognizable and executable within the system, such as an enumeration type, an instruction object, or an API call parameter, etc. For example, the consensus determination result is "alarm". The disposal measure determination module queries the preset disposal measure mapping table and finds that the instruction corresponding to "alarm" is "SEND_ALARM_NOTIFICATION". The module outputs the final disposal measure instruction as the enumeration value SEND_ALARM_NOTIFICATION, and this instruction will be passed to the disposal measure execution module to drive the execution of the alarm notification function.

[0185] The consensus resolution draft generation module receives the consensus determination result, the final disposal measure, and the vote statistics data as inputs. The module integrates these input data to generate a structured consensus resolution draft document. The consensus resolution draft document adopts a predefined format, such as JSON or XML, and contains the following key information: the "consensus algorithm type" field, which records the consensus algorithm adopted for this consensus decision, such as the "simple majority voting algorithm"; the "consensus determination result" field, which records the determination result of the consensus algorithm, that is, the type of disposal measure selected by voting, such as "alarm"; the "final disposal measure" field, which records the disposal measure instruction finally determined by the system, such as the enumerated value SEND_ALARM_NOTIFICATION; the "vote statistics data" field, which embeds the complete vote statistics data.

[0186] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed within the present invention.

[0187] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features invented herein.

Claims

1. A digital information traceability management system, characterized in that: Includes the following modules: The information fingerprint extraction module includes an NLP processing server and a cache server, which are used to obtain original information; extract text semantic features from the original information to obtain semantic features; Extract information fingerprints from semantic features to obtain semantic fingerprints; The origin node consensus module includes a consensus node server, which is used to perform consensus processing on the semantic fingerprint based on the origin node to obtain a consensus result; generate a source point certificate according to the consensus result to obtain a source point certificate; The flow path mapping module includes a behavior data collector, which is used to trigger the flow behavior according to the source point credential to obtain the behavior to be recorded; perform content change analysis according to the behavior to be recorded to obtain the semantic fingerprint after the change; Map the flow path based on the behavior to be recorded and the semantic fingerprint after the change to obtain the flow trajectory data; The information integrity verification module includes a verification server and an abnormal pattern detection engine, which is used to receive a verification request based on the source point credential to obtain the information to be verified; Perform source point fingerprint comparison on the information to be verified to obtain fingerprint comparison results; trace back the flow path according to the flow trajectory data to obtain the verified path; Verify the behavior consistency of the verified path and identify abnormal patterns to obtain behavior verification results and abnormal behavior reports; Generate a verification report based on the fingerprint comparison results, behavior verification results, and abnormal behavior reports to obtain a traceability verification report; The behavioral consensus feedback module includes a consensus node server and a monitoring and disposal component, which is used to conduct node consensus voting based on the traceability verification report to obtain a consensus vote warehouse; to form a consensus decision based on a preset consensus algorithm and a consensus vote warehouse to obtain a consensus resolution draft; to execute resolution measures based on the consensus resolution draft to obtain disposal resolution data, so as to realize digital information traceability management operations.

2. The digital information traceability management system according to claim 1 is characterized in that: The information fingerprint extraction module includes the following functions: Acquire original information; perform text preprocessing on the original information to obtain preprocessed text; Extract semantic features from the preprocessed text to obtain semantic features; Constructing semantic vectors for semantic features to obtain semantic vectors; Perform semantic hash encoding on the semantic vector to obtain a semantic hash code; The semantic hash code is subjected to semantic fingerprint generation to obtain a semantic fingerprint.

3. The digital information traceability management system according to claim 2 is characterized in that: The said extraction of semantic features from the preprocessed text includes: Perform preliminary keyword extraction on the preprocessed text to obtain an initial keyword list; Perform topic distribution mining on the preprocessed text to obtain topic probability distribution; Perform named entity recognition on the preprocessed text to obtain a named entity set; Perform dependency syntactic analysis on the preprocessed text to obtain a dependency graph; The semantic features are fused and enhanced on the initial keyword list, topic probability distribution, named entity set and dependency graph to obtain semantic features.

4. The digital information traceability management system according to claim 1 is characterized in that: The origin node consensus module includes the following functions: Encapsulate the origin information of the semantic fingerprint to obtain the information to be agreed upon; Submit the consensus information for network broadcast and obtain a broadcast request; Perform preliminary node verification on the broadcast request to obtain the information to be verified; Propose and vote on the information to be verified to obtain consensus votes; The consensus result of the consensus ballot is confirmed according to the preset consensus algorithm to obtain the consensus result; The source point certificate is generated according to the consensus result to obtain the source point certificate.

5. The digital information traceability management system according to claim 1 is characterized in that: The flow path mapping module includes the following functions: Trigger the flow behavior based on the source voucher to obtain the behavior to be recorded; Perform content change analysis based on the behavior to be recorded to obtain the semantic fingerprint after the change; Construct flow events based on the behavior to be recorded and the semantic fingerprint after the change to obtain a flow snapshot; Append the circulation record to the circulation snapshot to obtain the record to be uploaded to the chain; Write the records to be uploaded to the distributed ledger according to the preset distributed ledger to obtain the on-chain circulation records; The flow track is updated according to the on-chain flow records and source point credentials to obtain the flow track data.

6. The digital information traceability management system according to claim 5 is characterized in that: The content change analysis according to the behavior to be recorded includes: Determine the operation type of the behavior to be recorded and obtain the operation type result; Obtain the content version according to the operation type result to obtain the original version data and the modified version data; Perform fine-grained difference detection on the original version data and the modified version data to obtain content difference fragments; Conduct semantic impact assessment on content difference segments to obtain the degree of semantic change; The semantic fingerprint is generated after the semantic change degree is modified to obtain the changed semantic fingerprint.

7. The digital information traceability management system according to claim 6 is characterized in that: The fine-grained difference detection of the original version data and the modified version data includes: Perform row-level preprocessing on the original version data and the modified version data to obtain the row-segmented original version and the row-segmented modified version; Perform row-level difference comparison on the original version of the row split and the modified version of the row split to obtain the row-level difference result; Perform word-level preprocessing on the row-level difference results to obtain the original row to be compared at the word level and the modified row to be compared at the word level; Perform word-level difference comparison on the original line to be compared at the word level and the line after being modified at the word level to compare at the word level, and obtain the word-level difference result; The line-level difference results and the word-level difference results are structured into difference segments to obtain content difference segments.

8. The digital information traceability management system according to claim 1 is characterized in that: The information integrity verification module includes the following functions: Receive the verification request based on the source point credential and obtain the information to be verified; Reconstruct the current semantic fingerprint of the information to be verified to obtain the current semantic fingerprint; Perform source point fingerprint comparison on the current semantic fingerprint and the information to be verified to obtain the fingerprint comparison result; According to the information to be verified and the flow trajectory data, the flow path is traced back to obtain the verified path; Perform behavioral consistency verification on the verified path according to the current semantic fingerprint to obtain the behavioral verification result; Perform abnormal pattern recognition on the verified path and obtain abnormal behavior report; The verification results of fingerprint comparison results, behavior verification results and abnormal behavior reports are summarized to obtain a report to be signed; a verification report is generated for the report to be signed to obtain a traceability verification report.

9. The digital information traceability management system according to claim 1 is characterized in that: The behavior consensus feedback module includes the following functions: Distribute the traceability verification report and obtain the consensus report; The consensus process is initiated based on the consensus report to obtain a consensus proposal; Conduct node evaluation and voting based on the consensus report and consensus proposal to obtain disposal votes; Summarize the voting results of the disposal ballots to obtain the consensus vote bank; A consensus decision is formed based on the preset consensus algorithm and consensus votes to obtain a consensus resolution draft; Execute disposal measures according to the draft consensus resolution and obtain execution instructions; record execution results according to the execution instructions and obtain disposal records; The consensus resolution draft, consensus vote bank and disposal records are solidified as final decisions to obtain disposal resolution data.

10. The digital information traceability management system according to claim 9, characterized in that: The consensus decision-making process based on the preset consensus algorithm and consensus ticket warehouse includes: Count the voting results of the consensus vote warehouse to obtain voting statistics; Apply the consensus algorithm according to the preset consensus algorithm and voting statistics to obtain the consensus judgment result; Determine the disposal measures based on the consensus determination results and obtain the final disposal measures; A consensus resolution draft is generated based on the consensus determination results, final disposal measures, and voting statistics to obtain a consensus resolution draft.

Citation Information

Patent Citations

  • Data traceability analysis system and method

    CN118350055A

  • Block chain-based asset processing traceability method and system

    CN118505398A

  • Video conferencing system for performing a video conference and for passing resolutions

    EP4102832A1

Cited By

  • Medium and small financial institution customer service full-flow digital system and method

    CN120430797A

  • A full-process digital customer service system and method for small and medium-sized financial institutions

    CN120430797B

  • Packaging defect eliminating method based on visual inspection

    CN120571781A

  • Packaging defect elimination method based on visual inspection

    CN120571781B

  • Logistics package traceability data transmission system and method based on block chain

    CN120782350A