A judicial system secret data safe flow method based on blockchain technology
By building a trusted transmission channel and intelligent analysis framework using blockchain technology, the security and privacy protection issues in cross-level and cross-departmental judicial data exchange have been resolved. This has enabled efficient, traceable, and intelligent collaboration of classified data, thereby improving the security and efficiency of the judicial data system.
Patent Information
- Application Number
- CN202510630917.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Existing judicial data exchange suffers from problems such as interface fragmentation, high transmission security risks, privacy leaks, and lack of full lifecycle management in cross-level, cross-departmental, and cross-network scenarios, making it difficult to achieve high-security, high-efficiency, and traceable transfer of classified data.
A trusted transmission channel is built using blockchain technology. End-to-end encryption is achieved through the SM4 encryption algorithm and the TLS 1.3 protocol. Differential privacy mechanism is used for de-identification, and PBFT consensus mechanism is used for cross-node evidence storage. Access control is implemented using the XACML policy framework and RBAC/ABAC model. A judicial knowledge graph is built for intelligent analysis, and a federated learning framework is used for cross-domain model training, thereby achieving secure data sharing and compliant management.
It enables the end-to-end tamper-proof, traceable, and shared exchange of classified data across agencies, ensuring the security and privacy of data transmission, supporting intelligent analysis and collaborative training across departments, and improving the security and efficiency of the judicial data system.
Smart Images

Figure CN120567451B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of judicial data security sharing and trusted transfer technology, and in particular to a method for secure transfer of classified data in the judicial system based on blockchain technology. Background Technology
[0002] The vigorous advancement of court informatization in recent years has accumulated massive amounts of case data, creating an environment for judicial big data research and providing a new perspective and height for the ideal of consistent judgments in similar cases. The construction of smart courts under the big data perspective starts from the perspective of big data technology, combines basic legal knowledge with judges' practical experience in adjudicating cases, and uses mathematical, statistical, and cloud computing methods to establish intelligent analysis and application models, exploring and researching quantitative standards for judicial data. It represents the cross-application of big data and cloud computing technologies in the judicial field.
[0003] While there is no clear definition of methods for quantifying judicial data, a certain consensus has been reached in practice. Judges, through guidance from classic cases at higher institutions and peer training, gradually develop consensus-based, experiential judgment methods for similar cases. Once these consensus-based, somewhat vague indicators are quantified, they can become data that computers can understand, thus enabling a shift from experience-based judgment to scientific decision-making. This also provides regulatory authorities with an effective oversight tool, playing a crucial role in preventing inconsistent judgments in similar cases and resolving social conflicts. On the other hand, refined case information management helps courts comprehensively understand their current situation, allows the public to directly experience judicial fairness, and enables legal professionals to analyze case issues more deeply.
[0004] The construction of smart courts under the big data perspective begins with the intelligent identification and judgment of similar cases. Judgment methods have been solicited and validated from a large pool of judges. Based on this, the common characteristics and key factors of similar cases are further explored, along with the calculation of similarity, and a similarity judgment model is established. Then, through comprehensive computer exploration of judicial big data, differences in the judgment results of similar cases are identified. An attempt is made to use expert discussions and fuzzy logic calculations to find the subjective and objective reasons for these differences. Ultimately, the influencing factors of similar cases and the changing patterns of judgments are explored. This provides guidance for judges' future judgments, offers more comprehensive analytical tools for trial management, and provides richer dimensions for thinking about similar cases.
[0005] With the improvement of citizens' legal awareness and the lowering of the threshold for filing cases, the number of cases in people's courts has increased significantly in recent years. The types and circumstances of cases have become more complex. At the same time, the continuously growing judicial expectations and social responsibilities have also placed higher demands on the trial work of the courts.
[0006] Current judicial data exchange primarily relies on manual collection or dedicated network connections, resulting in issues such as system fragmentation, inconsistent interfaces, and inconsistent data formats, making it difficult to meet the high-frequency collaborative needs of multiple institutions. Furthermore, because the data contains a large amount of confidential information, such as party identities, case details, and evidence materials, privacy is easily exposed and there is a risk of tampering and leakage during data transfer. Traditional encryption methods and access control mechanisms are insufficient to meet the high requirements of "full lifecycle security" in modern judicial collaboration scenarios. In addition, the lack of a unified and reliable operational auditing and abnormal behavior identification mechanism makes it difficult to effectively trace data flow and quickly respond to risk events in complex distributed environments.
[0007] On the other hand, although artificial intelligence and big data analytics have made progress in areas such as case identification, structuring of judgments, and knowledge graph construction, the lack of deep integration with data security mechanisms has led to the "data silo" problem in intelligent models, limiting the ability to implement cross-departmental model co-construction and knowledge collaboration. Judicial data, as a highly sensitive resource, urgently requires a new security mechanism that enables cross-domain intelligent analysis and collaborative training without sacrificing privacy.
[0008] Therefore, there is an urgent need for a judicial data transfer method with the capabilities of "trusted exchange, privacy protection, intelligent identification, and compliance management" to achieve highly secure, efficient, and traceable transfer of classified data across institutions and systems. This would provide a secure and controllable digital foundation for the judicial system, supporting the in-depth development of similar case judgments, trial supervision, and smart justice. Summary of the Invention
[0009] The technical problem this invention aims to solve is: how to overcome the problems of interface fragmentation, high transmission security risks, privacy leakage risks, and lack of full life-cycle management in the context of cross-level, cross-department, and cross-network judicial data sharing, and how to build a trusted transmission channel for classified data through blockchain technology to achieve secure desensitization, end-to-end encrypted transmission, blockchain evidence storage, and fine-grained access control for multi-source heterogeneous data in the judicial system, ensuring that classified data among the judicial, procuratorial, and judicial departments is tamper-proof, traceable, and shared throughout the entire process.
[0010] To achieve the above objectives, the present invention employs the following technical solution:
[0011] This invention provides a method for the secure transfer of classified data in the judicial system based on blockchain technology, comprising the following steps:
[0012] Step 1: Collect multi-source judicial data, including structured, semi-structured, and unstructured data;
[0013] Step 2: Use a large language model to identify sensitive information in the data and perform desensitization processing based on a differential privacy mechanism;
[0014] Step 3: Construct an end-to-end encrypted channel using the SM4 encryption algorithm and TLS 1.3 protocol, and complete data transmission by combining X.509 certificates, biometric authentication, and ECC key negotiation;
[0015] Step 4: Generate a SHA-256 hash fingerprint for the encrypted data, and write the hash fingerprint and transmission metadata into the blockchain system to achieve cross-node notarization through the PBFT consensus mechanism;
[0016] Step 5: Classify the data using a multimodal classification engine, including text semantic analysis, video keyframe feature extraction, and speech-to-text processing;
[0017] Step 6: Adopt a hybrid access control model, integrating the XACML policy framework and Kafka message queue to achieve real-time response and authentication of access requests. Support dynamic access management of time, space, and device attributes through the joint mechanism of RBAC and ABAC, and bind user identity and device by combining geographical location and timestamp two-factor authentication.
[0018] Step 7: Adopt a hierarchical encryption storage strategy to encrypt sensitive data and store it in a private cloud, supporting homomorphic encryption retrieval. General data is stored in a distributed redundant manner using erasure coding, and metadata is built in a graph database.
[0019] Step 8: Construct a judicial knowledge graph, extract legal entities and semantic relationships from judgment documents based on the BERT model, and form a graph network containing four types of entities and three types of relationships to support semantic association query and decision assistance;
[0020] Step 9: During the system operation phase, the operation log is recorded based on the blockchain audit system, and abnormal behavior is identified by combining LSTM neural network and dynamic threshold detection. Risk assessment and graded emergency response are carried out through the DREAD model.
[0021] Between steps 8 and 9, there is also a step: cross-regional model training based on the federated learning framework, gradient aggregation through secure multi-party computation (MPC) and Paillier encryption, and secure training that protects the original data from leaving the region by combining differential privacy.
[0022] In the above scheme, step 1 includes the following sub-steps:
[0023] S1.1: Connects to the court trial system, the procuratorate business system, and the judicial administration management system, supporting three data connection methods: standardized API calls, message middleware push, and scheduled data retrieval;
[0024] S1.2: Synchronously collect structured data, semi-structured data, and unstructured data, wherein the structured data includes case number, party information, and judgment result fields; the semi-structured data includes court hearing audio and video files and electronic transcripts; and the unstructured data includes scanned evidence materials, PDF format judgment documents, and audio evidence files.
[0025] S1.3: The collected data is formatted and pre-parsed. Structured data is directly parsed into a set of standard fields. Audio and video data are converted into text data through speech-to-text service. Image and video files are visually analyzed through keyframe extraction tools.
[0026] In the above scheme, step 2 includes the following sub-steps:
[0027] S2.1: Use a pre-trained language model to perform Named Entity Recognition (NER) on the collected data to identify sensitive fields including the person's name, ID number, phone number, address, and bank account.
[0028] S2.2: Implement differential privacy desensitization processing based on preset sensitive field levels. For strongly identifying fields, irreversible desensitization is performed using random perturbation and hash replacement. For fields that need to retain semantics, recoverable desensitization is performed using character replacement and token masking.
[0029] S2.3: Attach a differential privacy budget value label to the de-identified fields and record the de-identification policy in the metadata for subsequent access decisions by the access control module.
[0030] In the above scheme, step 3 includes the following sub-steps:
[0031] S3.1: Establish an end-to-end encrypted channel based on the TLS 1.3 protocol, complete the identity authentication of both communicating parties through X.509 two-way digital certificates, introduce a biometric identification mechanism in the high-privilege node, and combine elliptic curve cryptography (ECC) to complete key negotiation and generate a symmetric session key;
[0032] S3.2: The data body is encrypted in blocks using the SM4-CBC mode, and the HMAC-SM3 algorithm is used to calculate the integrity check value of the encrypted data. The receiving end verifies the consistency of the check value before decryption.
[0033] In the above scheme, step 4 includes the following sub-steps:
[0034] S4.1: Calculate a unique SHA-256 hash digest for the encrypted data content as a digital fingerprint, which is used for end-to-end consistency verification and blockchain registration;
[0035] S4.2: Package the hash digest, sender node ID, receiver node ID, and timestamp into a transaction record and associate it with the data transmission task metadata;
[0036] S4.3: Write the transaction record to the Hyperledger Fabric consortium chain composed of nodes from the court, procuratorate, and judicial bureau, and confirm it through multi-judicial node voting via the PBFT consensus mechanism to complete cross-chain evidence storage;
[0037] S4.4: Based on smart contracts, the on-chain status of transaction records, the consistency of node signatures, and the integrity of original files are verified on-chain, generating a unique transaction ID and binding it to the transmission task log.
[0038] In the above scheme, step 5 includes the following sub-steps:
[0039] S5.1: Construct a multi-channel asynchronous processing architecture, and establish preprocessing pipelines for structured data, text data, audio data, and video data respectively;
[0040] S5.2: Normalize structured data through JSON Schema mapping and regular expression parsing to generate a standard field set;
[0041] S5.3: Input text data using the semantic parsing module of DeepSeek-v3, and generate structured semantic vectors through entity recognition and keyword tag extraction;
[0042] S5.4: After the audio data is transcribed into text using the iFlytek ASR model, the semantic parsing process in step S5.3 is reused to generate semantic vectors;
[0043] S5.5: The video data is processed by the inter-frame difference algorithm to extract key frames, and the visual features are extracted by the VGG16 network to generate a 4096-dimensional visual vector.
[0044] S5.6: Input the multimodal vector into the classifier based on the Gated Decision Tree, and output the case type, classification level, and archive path label.
[0045] In the above scheme, step 6 includes the following sub-steps:
[0046] S6.1: Embed the Policy Decision Point (PDP) module in the server kernel. Through the front-end proxy component, user access requests are converted into standardized XACML Request XML format and delivered as event messages to the Kafka message queue. The KafkaProducer then pushes the messages to the PDP instance clusters divided according to business dimensions, realizing asynchronous and concurrent distribution of policy judgment tasks.
[0047] S6.2: Construct a multi-channel Kafka Topic, and automatically classify and route the access request to the corresponding PDP instance for processing according to the business type. The business types include case file viewing, audio and video download, and case file export.
[0048] S6.3: Adopts a hybrid model of role-based access control (RBAC) and attribute-based access control (ABAC). Roles are defined from the unified authentication platform of the judicial authorities, and attribute dimensions include time window, access geolocation, device fingerprint, network type, and target data sensitivity level.
[0049] S6.4: Before the client initiates access, the local security module embeds the current UTC timestamp. with geographic coordinates Generate signature digest using the HMAC-SM3 algorithm Attached to the access request, the server verifies whether the timestamp and location parameter match the allowed time period in the XACML policy template. The IP / GPS whitelist of judicial office areas, among which SM3 cryptographic hash function, This represents a pre-shared symmetric key. The symbol indicates data concatenation;
[0050] S6.5: Complete TLS 1.3 channel identity authentication through the national cryptographic standard X.509 digital certificate, and dynamically generate an access token by combining the organization role information. The token carries the policy path hash and the request context hash fingerprint.
[0051] S6.6: The access determination result and its corresponding policy path, attribute matching conditions, and user context hash digest are structured and packaged, written into the consortium blockchain through the Hyperledger Fabric smart contract interface, and jointly confirmed and stored by multiple judicial nodes using the PBFT consensus mechanism.
[0052] In the above scheme, step 7 includes the following sub-steps:
[0053] S7.1: Storage types are divided according to data sensitivity levels. High-sensitivity data is encrypted using the AES-256 algorithm and stored in a private cloud. A ciphertext indexing mechanism is built based on the Paillier homomorphic encryption algorithm, and ciphertext keyword retrieval is achieved by combining Bloom Filter and inverted index. Medium and low-sensitivity data are generated into k data fragments and r redundant fragments through erasure coding and Reed-Solomon encoding, and distributed and stored on different physical nodes.
[0054] S7.2: Store metadata in the Neo4j graph database to construct a judicial semantic graph containing case nodes, party nodes, document nodes and related relationships, supporting graph traversal and pattern matching queries;
[0055] S7.3: Design multi-layered smart contracts in the Hyperledger Fabric consortium blockchain, including access control subcontracts, audit triggering subcontracts and version traceability subcontracts. Encapsulate data operation rules through standard ABI interfaces. The version traceability subcontract records the data version hash chain based on the Merkle tree and reconstructs the Merkle root to verify the change path during each update.
[0056] S7.4: The access control subcontract verifies the permission attributes of the calling entity, dynamically matches the strategy expression with off-chain hybrid authorization data, and returns a rejection record when authorization fails and writes it to the asynchronous audit log;
[0057] S7.5: Data storage behavior, version change events and audit trigger records are written into the blockchain through the PBFT consensus mechanism, and the evidence is jointly confirmed by judicial nodes.
[0058] S7.1: Storage types are divided according to data sensitivity levels. High-sensitivity data is encrypted using the AES-256 algorithm and stored in a private cloud. A ciphertext indexing mechanism is built based on the Paillier homomorphic encryption algorithm, and ciphertext keyword retrieval is achieved by combining Bloom Filter and inverted index. Medium and low-sensitivity data are generated into k data fragments and r redundant fragments through erasure coding and Reed-Solomon encoding, and distributed and stored on different physical nodes.
[0059] S7.2: Store metadata in the Neo4j graph database to construct a judicial semantic graph containing case nodes, party nodes, document nodes and related relationships, supporting graph traversal and pattern matching queries;
[0060] S7.3: Design multi-layered smart contracts in the Hyperledger Fabric consortium blockchain, including access control subcontracts, audit triggering subcontracts and version traceability subcontracts. Encapsulate data operation rules through standard ABI interfaces. The version traceability subcontract records the data version hash chain based on the Merkle tree and reconstructs the Merkle root to verify the change path during each update.
[0061] S7.4: The access control subcontract verifies the permission attributes of the calling entity, dynamically matches the strategy expression with off-chain hybrid authorization data, and returns a rejection record when authorization fails and writes it to the asynchronous audit log;
[0062] S7.5: Data storage behavior, version change events and audit trigger records are written into the blockchain through the PBFT consensus mechanism, and the evidence is jointly confirmed by judicial nodes.
[0063] In the above scheme, step 8 includes the following sub-steps:
[0064] S8.1: Perform natural language cleaning on judgment documents, including noise reduction, sentence segmentation, word segmentation, and part-of-speech tagging. Input the pre-trained BERT model to complete the extraction of legal entities. The entity types include four categories: court, plaintiff, defendant, and case.
[0065] S8.2: Construct a semantic edge structure based on legal logical relationships, define three core relationships: "litigation-acceptance-being sued", and form a multi-level knowledge network centered on case nodes;
[0066] S8.3: Map the extracted entities and relationships to the Neo4j graph database, bind a unique ID and attribute set to each node, and define the relationship type and occurrence timestamp for each edge;
[0067] S8.4: Based on the graph structure, semantic association query is realized, which supports traversing the associated plaintiff, defendant, court entity and reference legal provisions information with the case as the entry point, and the query results are attached with access control tags.
[0068] In the above scheme, step 9 includes the following sub-steps:
[0069] S9.1: Deploy a lightweight log agent on each business node of the system to capture user operation logs. The log fields should include at least the operation type, resource identifier, timestamp, user digital certificate fingerprint and device hash fingerprint, and generate log digests using the SM3 algorithm.
[0070] S9.2: Construct an anomaly detection model based on LSTM neural network, independently train behavioral baselines according to user role classification (judge, prosecutor, clerk), input is the operation sequence vector within the time window, output is the operation category prediction distribution, and determine the deviation of behavior by calculating the KL divergence value between the prediction result and the actual operation.
[0071] S9.3: A dynamic 3σ threshold detection mechanism is adopted. A normal distribution model is established based on the user's historical operation frequency and resource access mode. When the deviation of a single operation exceeds μ±3σ or the KL divergence value of a continuous operation sequence exceeds the preset threshold, an abnormal alarm event is triggered.
[0072] S9.4: Integrates the CVE vulnerability database and the legal industry threat signature database, uses the DREAD model to assess the risk of alert events, and outputs risk scores and classifications.
[0073] High risk (≥7 points): Immediately isolate the user session, freeze permissions, and initiate a judicial notification mechanism;
[0074] Medium risk (4-6.9 points): Restrict access to sensitive resources and mark audit trails;
[0075] Low risk (<4 points): Increase the frequency of log recording and dynamically adjust the weights of the behavior model;
[0076] S9.5: Encapsulate the metadata of abnormal events, risk assessment results, and response operation records into blockchain transactions, and write them into the Hyperledger Fabric consortium blockchain through the PBFT consensus mechanism. The evidence storage nodes include court audit departments, procuratorate supervision nodes, and judicial technology verification nodes.
[0077] The blockchain-based method for secure transfer of classified data in the judicial system, through the integration of multiple technologies and innovative mechanism design, offers the following advantages in ensuring data security, privacy protection, transfer efficiency, and compliance management:
[0078] 1. End-to-end trust and tamper-proof protection
[0079] The cross-node evidence storage mechanism based on blockchain technology (step 4) ensures the integrity and immutability of data in all stages of transmission, storage, and processing through hash fingerprinting on-chain and the PBFT consensus algorithm. Combined with on-chain verification of operational behavior by smart contracts (step 4.4), a trusted traceability chain for the entire lifecycle of judicial data is formed, effectively solving the problems of tampering risk and liability determination in traditional judicial data exchange.
[0080] 2. Dynamic, fine-grained privacy protection
[0081] A hierarchical desensitization strategy is implemented for sensitive information using a large language model combined with differential privacy (step 2). This eliminates the risk of leakage of personally identifiable data while retaining necessary semantic features to support subsequent business processing. By attaching differential privacy budget labels (step 2.3) and associating them with metadata, a dynamic balance between desensitization strength and data availability is achieved, providing a compliance foundation for cross-organizational data sharing.
[0082] 3. Multi-dimensional secure transmission mechanism
[0083] An end-to-end encrypted channel based on the Chinese national cryptographic algorithm (SM4) and the TLS 1.3 protocol is constructed (step 3). Combined with ECC key negotiation and biometric authentication, the security of the transmission link is enhanced. A block encryption and integrity verification mechanism is adopted (step 3.2) to prevent data from being stolen or tampered with during transmission, meeting the high security requirements of judicial data in cross-network transmission scenarios.
[0084] 4. Intelligent dynamic permission management
[0085] By using the XACML policy framework and a hybrid permission model (step 6), role-based authentication (RBAC) and context-based authentication (ABAC) are integrated to achieve real-time dynamic authentication based on time, space, and device fingerprints. Combined with two-factor authentication binding (step 6.4), the consistency between user identity and operating environment is ensured, preventing permission abuse or unauthorized access, and improving the precision of access control for confidential data.
[0086] 5. Multimodal data processing and knowledge fusion
[0087] The text and audio / video data are structured using a multimodal classification engine (step 5), and semantic modeling of legal entities and relationships is achieved by combining it with a judicial knowledge graph (step 8). Key case elements are extracted using the BERT model (step 8.1) and a relational network is constructed to provide intelligent support for case analysis and judgment reference, thereby enhancing the value mining capability of judicial data in cross-departmental collaboration.
[0088] 6. Privacy-preserving collaborative intelligence training
[0089] Cross-domain model training is achieved based on the federated learning framework (step 9), and secure multi-party computation (MPC) and homomorphic encryption ensure that the original data does not leave the domain. Differential privacy noise is injected during gradient aggregation (step 9.3) to protect data privacy while promoting collaborative optimization of cross-institutional models and solving the problem of insufficient model generalization ability caused by judicial data silos.
[0090] 7. Closed-loop risk prevention and control system
[0091] By combining blockchain audit logs with the LSTM anomaly detection model (step 9), operational sequences deviating from normal behavior are identified in real time. This, combined with the DREAD risk assessment model (step 9.4), triggers a tiered response mechanism, achieving closed-loop management from anomaly detection and risk assessment to emergency response, thereby enhancing the proactive defense and compliance governance capabilities of the judicial data system.
[0092] 8. Efficient storage and intelligent retrieval
[0093] A hierarchical encryption storage strategy is adopted (step 7) to implement homomorphic encryption retrieval for highly sensitive data, balancing security and availability. A judicial semantic graph is constructed through a graph database (step 7.2) to support complex relational queries and graph traversal analysis, optimize the retrieval efficiency of unstructured data, and provide a multi-dimensional analytical perspective for judicial decision-making.
[0094] 9. Through technological coupling and closed-loop process design, the various steps of this invention form a security-enhanced collaborative system covering the entire data lifecycle. Steps 1 to 3 construct a front-end security barrier of "collection-de-identification-encryption": After standardized preprocessing of multi-source heterogeneous data (step 1), sensitive fields are identified and differential privacy de-identification is implemented through a large language model (step 2), which eliminates the risk of privacy leakage while preserving the semantic usability of the data; combined with SM4 encryption and TLS channel (step 3), an end-to-end secure transmission link is constructed to ensure the confidentiality of data flow across networks. Steps 4 to 6 form the central control layer of "evidence storage-classification-permission": Hash fingerprint on-chain evidence storage (step 4) provides an immutable trust anchor for the whole process of data traceability; the multimodal classification engine (step 5) intelligently parses and generates tags for text, audio and video data, providing fine-grained attribute tags for subsequent permission control (step 6); and the dynamic hybrid permission model (RBAC+ABAC) implements precise access control based on classification tags, spatiotemporal attributes and device fingerprints, forming a closed-loop feedback mechanism of "data tag-driven permission strategy". Steps 7 through 10 construct a deep application system of "storage-knowledge-risk control": the joint design of hierarchical encrypted storage (step 7) and judicial knowledge graph (step 8) enables efficient retrieval of encrypted data through semantic association in the encrypted state; the federated learning framework (step 9), based on the differential privacy mechanism of step 2 and the encrypted channel of step 3, realizes cross-domain model collaborative training and knowledge sharing, solving the data silo problem; and blockchain auditing (step 10), through LSTM anomaly detection and DREAD risk assessment, reverse-optimizes the permission policy of step 6 and the encryption strength threshold of step 3, forming a dynamic security reinforcement cycle. Through the deep interweaving of technical elements (such as blockchain notarization providing a trusted source for classification labels, knowledge graphs providing semantic features for federated learning, and dynamic permission labels providing input parameters for risk models), the four core capabilities of data security, privacy protection, intelligent analysis, and compliance management are synergistically amplified, constructing an integrated solution of "trusted transmission-intelligent processing-proactive defense" for cross-domain judicial data transfer.
[0095] In summary, this invention, through the deep integration of blockchain and various security technologies, enables trusted exchange, dynamic protection, and intelligent collaboration of judicial classified data during cross-level and cross-departmental transfers. It provides a secure and controllable technical foundation for the construction of smart courts, while also meeting the compliance, efficiency, and traceability requirements of judicial data sharing. Attached Figure Description
[0096] Figure 1 This is the overall framework flowchart of the present invention.
[0097] Figure 2 This is the knowledge graph design diagram of the present invention.
[0098] Figure 3 This is a diagram of the unified data platform structure of the present invention.
[0099] Figure 4 This is a data transmission flow diagram of the present invention.
[0100] Figure 5 This is a schematic diagram of the access control system of the present invention. Detailed Implementation
[0101] This invention relates to the field of judicial data secure sharing and trusted transfer technology, and in particular to a method for secure transfer of classified data across levels, departments, and networks within the judicial system based on blockchain technology. This method supports highly secure data sharing among multiple business systems such as adjudication, prosecution, and judicial administration. It features non-intrusive data acquisition and processing capabilities, content-based sensitive information desensitization, end-to-end encrypted transmission, on-chain evidence storage, traceable auditing, and risk perception control. It is applicable to the full lifecycle security management and transfer of multimodal classified data, including electronic documents, court hearing audio and video, electronic evidence, and case files, and is widely used in judicial business collaboration scenarios such as electronic document exchange, court hearing information sharing, and sentence reduction and parole.
[0102] The purpose of this invention is to solve the problems of insufficient security and traceability of existing judicial classified data during the transfer of data across levels, departments, and networks.
[0103] To achieve the above objectives, the present invention employs the following technical solution:
[0104] This invention provides a method for the secure transfer of classified data in the judicial system based on blockchain technology, comprising the following steps:
[0105] Step 1: Collect multi-source judicial data, including structured, semi-structured, and unstructured data;
[0106] Step 2: Use a large language model to identify sensitive information in the data and perform desensitization processing based on a differential privacy mechanism;
[0107] Step 3: Construct an end-to-end encrypted channel using the SM4 encryption algorithm and TLS 1.3 protocol, and complete data transmission by combining X.509 certificates, biometric authentication, and ECC key negotiation;
[0108] Step 4: Generate a SHA-256 hash fingerprint for the encrypted data, and write the hash fingerprint and transmission metadata into the blockchain system to achieve cross-node notarization through the PBFT consensus mechanism;
[0109] Step 5: Classify the data using a multimodal classification engine, including text semantic analysis, video keyframe feature extraction, and speech-to-text processing;
[0110] Step 6: Adopt a hybrid access control model, integrating the XACML policy framework and Kafka message queue to achieve real-time response and authentication of access requests. Support dynamic access management of time, space, and device attributes through the joint mechanism of RBAC and ABAC, and bind user identity and device by combining geographical location and timestamp two-factor authentication.
[0111] Step 7: Adopt a hierarchical encryption storage strategy to encrypt sensitive data and store it in a private cloud, supporting homomorphic encryption retrieval. General data is stored in a distributed redundant manner using erasure coding, and metadata is built in a graph database.
[0112] Step 8: Construct a judicial knowledge graph, extract legal entities and semantic relationships from judgment documents based on the BERT model, and form a graph network containing four types of entities and three types of relationships to support semantic association query and decision assistance;
[0113] Step 9: During the system operation phase, the operation log is recorded based on the blockchain audit system, and abnormal behavior is identified by combining LSTM neural network and dynamic threshold detection. Risk assessment and graded emergency response are carried out through the DREAD model.
[0114] In the above scheme, step 1 includes the following sub-steps:
[0115] S1.1: Connects to the court trial system, the procuratorate business system, and the judicial administration management system, supporting three data connection methods: standardized API calls, message middleware push, and scheduled data retrieval;
[0116] S1.2: Synchronously collect structured data, semi-structured data, and unstructured data, wherein the structured data includes case number, party information, and judgment result fields; the semi-structured data includes court hearing audio and video files and electronic transcripts; and the unstructured data includes scanned evidence materials, PDF format judgment documents, and audio evidence files.
[0117] S1.3: The collected data is formatted and pre-parsed. Structured data is directly parsed into a set of standard fields. Audio and video data are converted into text data through speech-to-text service. Image and video files are visually analyzed through keyframe extraction tools.
[0118] In the above scheme, step 2 includes the following sub-steps:
[0119] S2.1: Use a pre-trained language model to perform Named Entity Recognition (NER) on the collected data to identify sensitive fields including the person's name, ID number, phone number, address, and bank account.
[0120] S2.2: Implement differential privacy desensitization processing based on preset sensitive field levels. For strongly identifying fields, irreversible desensitization is performed using random perturbation and hash replacement. For fields that need to retain semantics, recoverable desensitization is performed using character replacement and token masking.
[0121] S2.3: Attach a differential privacy budget value label to the de-identified fields and record the de-identification policy in the metadata for subsequent access decisions by the access control module.
[0122] In the above scheme, step 3 includes the following sub-steps:
[0123] S3.1: Establish an end-to-end encrypted channel based on the TLS 1.3 protocol, complete the identity authentication of both communicating parties through X.509 two-way digital certificates, introduce a biometric identification mechanism in the high-privilege node, and combine elliptic curve cryptography (ECC) to complete key negotiation and generate a symmetric session key;
[0124] S3.2: The data body is encrypted in blocks using the SM4-CBC mode, and the HMAC-SM3 algorithm is used to calculate the integrity check value of the encrypted data. The receiving end verifies the consistency of the check value before decryption.
[0125] In the above scheme, step 4 includes the following sub-steps:
[0126] S4.1: Calculate a unique SHA-256 hash digest for the encrypted data content as a digital fingerprint, which is used for end-to-end consistency verification and blockchain registration;
[0127] S4.2: Package the hash digest, sender node ID, receiver node ID, and timestamp into a transaction record and associate it with the data transmission task metadata;
[0128] S4.3: Write the transaction record to the Hyperledger Fabric consortium chain composed of nodes from the court, procuratorate, and judicial bureau, and confirm it through multi-judicial node voting via the PBFT consensus mechanism to complete cross-chain evidence storage;
[0129] S4.4: Based on smart contracts, the on-chain status of transaction records, the consistency of node signatures, and the integrity of original files are verified on-chain, generating a unique transaction ID and binding it to the transmission task log.
[0130] In the above scheme, step 5 includes the following sub-steps:
[0131] S5.1: Construct a multi-channel asynchronous processing architecture, and establish preprocessing pipelines for structured data, text data, audio data, and video data respectively;
[0132] S5.2: Normalize structured data through JSON Schema mapping and regular expression parsing to generate a standard field set;
[0133] S5.3: Input text data using the semantic parsing module of DeepSeek-v3, and generate structured semantic vectors through entity recognition and keyword tag extraction;
[0134] S5.4: After the audio data is transcribed into text using the iFlytek ASR model, the semantic parsing process in step S5.3 is reused to generate semantic vectors;
[0135] S5.5: The video data is processed by the inter-frame difference algorithm to extract key frames, and the visual features are extracted by the VGG16 network to generate a 4096-dimensional visual vector.
[0136] S5.6: Input the multimodal vector into the classifier based on the Gated Decision Tree, and output the case type, classification level, and archive path label.
[0137] In the above scheme, step 6 includes the following sub-steps:
[0138] S6.1: Embed the Policy Decision Point (PDP) module in the server kernel. Through the front-end proxy component, user access requests are converted into standardized XACML Request XML format and delivered as event messages to the Kafka message queue. The KafkaProducer then pushes the messages to the PDP instance clusters divided according to business dimensions, realizing asynchronous and concurrent distribution of policy judgment tasks.
[0139] S6.2: Construct a multi-channel Kafka Topic, and automatically classify and route the access request to the corresponding PDP instance for processing according to the business type. The business types include case file viewing, audio and video download, and case file export.
[0140] S6.3: Adopts a hybrid model of role-based access control (RBAC) and attribute-based access control (ABAC). Roles are defined from the unified authentication platform of the judicial authorities, and attribute dimensions include time window, access geolocation, device fingerprint, network type, and target data sensitivity level.
[0141] S6.4: Before the client initiates access, the local security module embeds the current UTC timestamp. with geographic coordinates Generate signature digest using the HMAC-SM3 algorithm Attached to the access request, the server verifies whether the timestamp and location parameter match the allowed time period in the XACML policy template. The IP / GPS whitelist of judicial office areas, among which SM3 cryptographic hash function, This represents a pre-shared symmetric key. The symbol indicates data concatenation;
[0142] S6.5: Complete TLS 1.3 channel identity authentication through the national cryptographic standard X.509 digital certificate, and dynamically generate an access token by combining the organization role information. The token carries the policy path hash and the request context hash fingerprint.
[0143] S6.6: The access determination result and its corresponding policy path, attribute matching conditions, and user context hash digest are structured and packaged, written into the consortium blockchain through the Hyperledger Fabric smart contract interface, and jointly confirmed and stored by multiple judicial nodes using the PBFT consensus mechanism.
[0144] In the above scheme, step 7 includes the following sub-steps:
[0145] S7.1: Storage types are divided according to data sensitivity levels. High-sensitivity data is encrypted using the AES-256 algorithm and stored in a private cloud. A ciphertext indexing mechanism is built based on the Paillier homomorphic encryption algorithm, and ciphertext keyword retrieval is achieved by combining Bloom Filter and inverted index. Medium and low-sensitivity data are generated into k data fragments and r redundant fragments through erasure coding and Reed-Solomon encoding, and distributed and stored on different physical nodes.
[0146] S7.2: Store metadata in the Neo4j graph database to construct a judicial semantic graph containing case nodes, party nodes, document nodes and related relationships, supporting graph traversal and pattern matching queries;
[0147] S7.3: Design multi-layered smart contracts in the Hyperledger Fabric consortium blockchain, including access control subcontracts, audit triggering subcontracts and version traceability subcontracts. Encapsulate data operation rules through standard ABI interfaces. The version traceability subcontract records the data version hash chain based on the Merkle tree and reconstructs the Merkle root to verify the change path during each update.
[0148] S7.4: The access control subcontract verifies the permission attributes of the calling entity, dynamically matches the strategy expression with off-chain hybrid authorization data, and returns a rejection record when authorization fails and writes it to the asynchronous audit log;
[0149] S7.5: Data storage behavior, version change events and audit trigger records are written into the blockchain through the PBFT consensus mechanism, and the evidence is jointly confirmed by judicial nodes.
[0150] S7.1: Storage types are divided according to data sensitivity levels. High-sensitivity data is encrypted using the AES-256 algorithm and stored in a private cloud. A ciphertext indexing mechanism is built based on the Paillier homomorphic encryption algorithm, and ciphertext keyword retrieval is achieved by combining Bloom Filter and inverted index. Medium and low-sensitivity data are generated into k data fragments and r redundant fragments through erasure coding and Reed-Solomon encoding, and distributed and stored on different physical nodes.
[0151] S7.2: Store metadata in the Neo4j graph database to construct a judicial semantic graph containing case nodes, party nodes, document nodes and related relationships, supporting graph traversal and pattern matching queries;
[0152] S7.3: Design multi-layered smart contracts in the Hyperledger Fabric consortium blockchain, including access control subcontracts, audit triggering subcontracts and version traceability subcontracts. Encapsulate data operation rules through standard ABI interfaces. The version traceability subcontract records the data version hash chain based on the Merkle tree and reconstructs the Merkle root to verify the change path during each update.
[0153] S7.4: The access control subcontract verifies the permission attributes of the calling entity, dynamically matches the strategy expression with off-chain hybrid authorization data, and returns a rejection record when authorization fails and writes it to the asynchronous audit log;
[0154] S7.5: Data storage behavior, version change events and audit trigger records are written into the blockchain through the PBFT consensus mechanism, and the evidence is jointly confirmed by judicial nodes.
[0155] In the above scheme, step 8 includes the following sub-steps:
[0156] S8.1: Perform natural language cleaning on judgment documents, including noise reduction, sentence segmentation, word segmentation, and part-of-speech tagging. Input the pre-trained BERT model to complete the extraction of legal entities. The entity types include four categories: court, plaintiff, defendant, and case.
[0157] S8.2: Construct a semantic edge structure based on legal logical relationships, define three core relationships: "litigation-acceptance-being sued", and form a multi-level knowledge network centered on case nodes;
[0158] S8.3: Map the extracted entities and relationships to the Neo4j graph database, bind a unique ID and attribute set to each node, and define the relationship type and occurrence timestamp for each edge;
[0159] S8.4: Based on the graph structure, semantic association query is realized, which supports traversing the associated plaintiff, defendant, court entity and reference legal provisions information with the case as the entry point, and the query results are attached with access control tags.
[0160] In the above scheme, step S9 specifically includes:
[0161] S9.1: Deploy a lightweight log agent on each business node of the system to capture user operation logs. The log fields should include at least the operation type, resource identifier, timestamp, user digital certificate fingerprint and device hash fingerprint, and generate log digests using the SM3 algorithm.
[0162] S9.2: Construct an anomaly detection model based on LSTM neural network, independently train behavioral baselines according to user role classification (judge, prosecutor, clerk), input is the operation sequence vector within the time window, output is the operation category prediction distribution, and determine the deviation of behavior by calculating the KL divergence value between the prediction result and the actual operation.
[0163] S9.3: A dynamic 3σ threshold detection mechanism is adopted. A normal distribution model is established based on the user's historical operation frequency and resource access mode. When the deviation of a single operation exceeds μ±3σ or the KL divergence value of a continuous operation sequence exceeds the preset threshold, an abnormal alarm event is triggered.
[0164] S9.4: Integrates the CVE vulnerability database and the legal industry threat signature database, uses the DREAD model to assess the risk of alert events, and outputs risk scores and classifications.
[0165] High risk (≥7 points): Immediately isolate the user session, freeze permissions, and initiate a judicial notification mechanism;
[0166] Medium risk (4-6.9 points): Restrict access to sensitive resources and mark audit trails;
[0167] Low risk (<4 points): Increase the frequency of log recording and dynamically adjust the weights of the behavior model;
[0168] S9.5: Encapsulate the metadata of abnormal events, risk assessment results, and response operation records into blockchain transactions, and write them into the Hyperledger Fabric consortium blockchain through the PBFT consensus mechanism. The evidence storage nodes include court audit departments, procuratorate supervision nodes, and judicial technology verification nodes.
[0169] Example 1
[0170] To achieve the above objectives, the present invention provides the following technical solution:
[0171] This invention provides a method for the secure transfer of classified judicial data based on blockchain technology. By constructing an integrated mechanism for collection, desensitization, transmission, classification, authorization, storage and auditing, it enables the reliable, controllable and traceable transfer of classified judicial data among multiple departments.
[0172] This invention connects to the business systems of courts, procuratorates, and judicial administration through a unified data exchange platform, aggregating multimodal judicial data, including structured (such as case documents), semi-structured (such as audio and video), and unstructured (such as electronic evidence). To ensure data privacy and security upon entering the system, sensitive fields (such as ID numbers and names) are automatically identified using a large language model (Deepseek-v3) after collection, and data anonymization is completed based on a differential privacy mechanism.
[0173] During data transfer, the system employs the national standard SM4 algorithm in conjunction with the TLS 1.3 protocol to construct an end-to-end encrypted channel. User identity is jointly verified using X.509 digital certificates and biometric identification, and key negotiation is completed based on elliptic curve cryptography. Data in transmission is encrypted using SM4-CBC, combined with the HMAC-SM3 algorithm for integrity verification. Simultaneously, the system generates an SHA-256 hash fingerprint for each data object and writes it to the Hyperledger Fabric consortium blockchain. Cross-chain notarization is completed by multiple judicial nodes using the PBFT consensus algorithm, ensuring the trustworthiness and immutability of data transfer at each stage.
[0174] In one embodiment of the present invention, the system constructs an end-to-end encrypted transmission path based on the national cryptographic algorithm SM4 and the TLS 1.3 protocol during the cross-departmental and cross-network transmission of classified judicial data. Considering the heterogeneous structures of judicial internal and external networks and the poor compatibility between the national cryptographic algorithm and the TLS standard protocol, the system designs a dual-path switchable transmission mechanism. This mechanism supports dynamic switching between SM4-CBC and SM4-GCM modes during the TLS handshake phase. A nonce random generator (based on timestamps and session IDs) is used to prevent replay attacks and ensure the stability of the encrypted path. Regarding key management, the client and server complete key negotiation based on an ECC algorithm (such as the SM2 curve) and employ a session state monitoring mechanism. When the session lifespan exceeds a threshold (e.g., 5 minutes), a key update process is automatically triggered to ensure the long-term confidentiality of transmitted data. Identity authentication adopts an improved two-factor strategy. During the TLS handshake phase, basic identity authentication is completed using an X.509 national cryptographic certificate, while SM3 biometric fingerprint verification is superimposed to ensure the binding of the person, certificate, and device, preventing certificate sharing and forgery risks.
[0175] Furthermore, regarding data integrity and traceability, this invention employs the Hyperledger Fabric blockchain system to perform structured notarization on each transmitted data item. Each data object (including audio / video, case files, documents, etc.) generates a unique SHA-256 hash fingerprint after de-identification and encapsulation, and constructs a transaction structure, embedding metadata such as operator ID, timestamp, access IP, and operation type. This data is then written to the blockchain via a dedicated Chaincode call interface. The blockchain nodes consist of consensus mechanisms comprised of courts, procuratorates, and judicial bureaus, employing the PBFT consensus mechanism to achieve endorsement by 2 / 3 of the nodes and complete transaction confirmation. After successful on-chain processing, the system writes the corresponding block number and transaction hash back to the original data system as a "trusted label" embedded in the database metadata field, enabling on-chain and off-chain linkage and traceability.
[0176] To improve processing efficiency after data transfer, the system constructs a multimodal content recognition and classification engine. Structured and text data undergo semantic understanding and label extraction using a pre-trained language model; video data acquires visual features through keyframe extraction and a VGG16 network; audio is transcribed and incorporated into a unified text processing flow, and finally, classification results are output through a decision tree algorithm, enabling the archiving, labeling, and rapid retrieval of classified data.
[0177] In one embodiment of the present invention, to improve the archiving efficiency and intelligent retrieval capabilities of judicial classified data after its transfer process, a content recognition and classification engine for multimodal data is constructed. This engine designs a multi-channel asynchronous processing architecture, establishing preprocessing and feature extraction pipelines for different modal data types (structured, text, audio, video, etc.), and fusing them in a unified intermediate representation layer. Structured data (such as case metadata and litigation stage information) is automatically normalized through JSON Schema mapping and regular expression parsing; text data is input to a semantic parsing module based on the DeepSeek-v3 large language model to complete entity recognition, case type judgment, and keyword tag extraction, outputting a structured semantic vector. Audio data is transcribed by the iFlytek ASR model and undergoes the same text processing flow to ensure semantic consistency.
[0178] The video data processing workflow employs an inter-frame difference algorithm for keyframe extraction, avoiding redundant frames from interfering with subsequent analysis. The extracted keyframes, after OpenCV image enhancement and normalization, are input into a VGG16 network for convolutional feature extraction, generating 4096-dimensional visual vector representations. All modality vectors are uniformly fed into a classifier based on a Gated Decision Tree architecture. This classifier automatically adjusts the weights of different modalities during the training phase to address incomplete data or missing modalities. The classification output includes multi-dimensional labels such as case type, confidentiality level, archiving path, and tag set, which are mapped into the storage system via a chained indexer, enabling automatic archiving and rapid semantic retrieval of confidential data. This multimodal recognition scheme effectively solves key challenges in judicial practice, such as low efficiency in unstructured data processing, difficulty in modality fusion, and unstable tag extraction, providing a semantic foundation for subsequent access control and risk monitoring.
[0179] In terms of access control, the system integrates the XACML policy framework and Kafka message channel to achieve real-time response and authentication of access requests. Based on a hybrid RBAC and ABAC model, the system supports a dynamic authorization mechanism based on contextual information such as time, location, and device. It further combines timestamp + geolocation two-factor authentication and user digital certificates to achieve fine-grained control over data access behavior.
[0180] In a specific embodiment of this invention, to address core technical issues such as permission drift, policy execution delay, user identity spoofing, and ambiguous permission boundaries in cross-network and cross-level system access to judicial data, a distributed dynamic permission control mechanism integrating an XACML policy engine and a Kafka real-time message channel is proposed. In the access control architecture, the Policy Decision Point (PDP) module is embedded in the server-side kernel. Before accessing the network, all user resource access requests are converted into standardized XACML Request XML format by a front-end proxy component and delivered as event messages to the Kafka message queue. The Kafka Producer is responsible for pushing these messages to PDP instance clusters divided according to business dimensions, achieving asynchronous and concurrent policy judgment task distribution. To address the decision performance bottleneck under large-scale access, the system constructs a multi-channel Kafka Topic to automatically classify and route access requests according to business type (such as case file viewing, audio / video download, case file export, etc.). Requests of different categories are processed concurrently by the corresponding PDP instances, supporting policy caching and asynchronous pre-parsing of policy templates, thereby improving policy judgment speed. In terms of permission model design, the system integrates role-based access control (RBAC) and attribute-based access control (ABAC). Roles are defined by the unified authentication platform of the judicial authorities, while attribute dimensions include precise time window (UTC time synchronized by NTP), access geolocation (based on GPS / Wi-Fi or public IP lookup), device fingerprint (generated through browser identifier and system registry snapshot), network type (internal / external network identifier), and target data sensitivity level (marked by data tags).
[0181] This invention introduces a two-factor context verification process into the authentication mechanism: Before the client initiates access, the local security module automatically embeds the current UTC timestamp and geographic coordinates, and calculates and generates a signature digest using HMAC-SM3, which is then appended to the access request. Upon receiving the request, the server matches and verifies the timestamp and location parameters against the allowed time period defined in the XACML policy template and the IP / GPS whitelist for judicial office areas, ensuring the spatial and temporal legitimacy of the request. If the time or location verification fails, access is denied, and an abnormal behavior event is automatically generated in the Kafka security audit channel for subsequent risk analysis. During the session, the user must authenticate their identity using a TLS 1.3 channel with a national cryptographic standard X.509 digital certificate. Combined with the bound institutional role information, the access control module dynamically generates an access token. This token only carries the policy path hash (Policy ID + matching chain digest) and the request context hash fingerprint, avoiding direct exposure of permission details and effectively preventing session hijacking and man-in-the-middle attacks. Ultimately, all access decision results (Allow / Deny) and their corresponding policy paths, attribute matching conditions, and user context hash digests will be structured and packaged, written to the consortium blockchain through the Hyperledger Fabric smart contract interface, and jointly confirmed and stored by multiple nodes such as courts and procuratorates using the PBFT consensus mechanism, thus constructing a complete access control audit chain to ensure that decision-making behavior is verifiable, traceable, and tamper-proof.
[0182] In the data storage stage, the system employs a hierarchical encryption strategy. Highly sensitive data (such as dossiers) is encrypted using AES-256 and stored in the private cloud, supporting ciphertext retrieval based on homomorphic encryption. General data is stored in a distributed, redundant manner using erasure coding and Reed-Solomon encoding to enhance disaster recovery capabilities. Metadata is built into the Neo4j graph database to support complex queries and analyses involving multiple entity relationships. On-chain business rules are encapsulated in a standard ABI interface via smart contracts and combined with a Merkle tree structure to verify and trace data change paths.
[0183] In one embodiment of this invention, a hierarchical encryption storage strategy based on data sensitivity is constructed to address challenges such as varying data sensitivity levels, diverse access needs, and stringent compliance requirements in judicial scenarios. Specifically, highly sensitive data (such as electronic case files and expert reports) is symmetrically encrypted using the AES-256 algorithm before being stored, ensuring the confidentiality and compliance of static data storage. To meet the queryability requirements under encrypted data conditions, a ciphertext indexing mechanism based on the Paillier homomorphic encryption algorithm is introduced. A ciphertext keyword mapping is designed using a Bloom filter and an inverted index structure, supporting fast keyword-based matching and retrieval without decrypting the original data, thus solving the technical bottleneck of "inability to quickly locate highly sensitive encrypted judicial data." For unstructured data with medium to low sensitivity levels, the system introduces a redundant distributed storage mechanism combining erasure coding and Reed-Solomon encoding. The original data is divided into k data fragments and r redundant fragments, distributed and stored on different physical nodes, improving disaster recovery capabilities in the event of node failure or attack. System metadata (such as case number, association, document status, etc.) is modeled in the Neo4j graph database. Complex judicial semantic graphs are constructed through nodes (cases, parties, documents) and edges (belonging, reference, trial), supporting graph traversal and pattern matching, enabling efficient querying of case-related elements and semantic-level upstream and downstream reasoning.
[0184] To ensure the verifiability of on-chain judicial data status and the traceability of version change paths, this invention designs a multi-layered smart contract structure based on Hyperledger Fabric. This contract structure includes three core modules: an access control subcontract (AccessPolicyContract), an audit trigger subcontract (AuditTriggerContract), and a version traceability subcontract (VersionTraceContract). It encapsulates business rules and data operation interfaces through a standard ABI interface and supports asynchronous calls from off-chain systems via RESTful APIs. In specific implementation, before writing data to the chain, the system first verifies the current caller's permission attributes (such as department level, authentication status, access time range, etc.) through the AccessPolicyContract subcontract. Internally, the contract combines off-chain ABAC and RBAC hybrid authorization data to dynamically match policy expressions. If authorization is successful, the data writing process begins; otherwise, the contract aborts, returns a rejection record, and writes the unauthorized attempt event to the asynchronous audit log. After successful data writing, the VersionTraceContract subcontract constructs Merkle tree leaf nodes for the current data version, recording metadata such as the original hash, timestamp, and operator identity in key-value format. Simultaneously, with each data update, the contract automatically reconstructs the Merkle root and links the change path to the previous and subsequent versions in a hash chain, writing this to the on-chain Key-Value State Database (StateDB). This ensures that any version backtracking can verify its authenticity and consistency based on the hash path. When specific security event trigger conditions are met (such as frequent access to certain types of data, cross-domain login, abnormal data content offset, etc.), the AuditTriggerContract subcontract automatically generates an audit task and calls the Fabric event mechanism (Chaincode Event) to push it to the judicial supervision node, notifying the audit module to conduct in-depth verification and off-chain emergency response. The entire contract system operates under the PBFT consensus mechanism. Every data authorization, update, and audit trigger execution is confirmed by at least 2 / 3 of the member nodes' signatures, effectively preventing single points of malicious activity and improving the credibility and technical compliance of on-chain governance of confidential data.
[0185] At the data semantic mining layer, the system uses the BERT model to extract legal entities and relationships from judgment documents, constructs a judicial knowledge graph centered on cases, realizes the structured expression of 4 types of core entities and 3 types of semantic relationships, stores them in a graph database, and supports subsequent semantic-level queries and data sharing access control.
[0186] To ensure secure system operation and compliance, a blockchain auditing system is integrated during the operation and maintenance phase to record all operation logs, including operation type, execution time, and digital signature subject. LSTM neural networks are used to perform sequence modeling of behavioral patterns, combined with a 3σ standard deviation dynamic threshold detection mechanism, automatically triggering risk response procedures when abnormal behavior is detected. The risk assessment system integrates a CVE vulnerability database and an industry threat signature database, and uses a DREAD model to assess risk levels, automatically performing tiered emergency response operations such as data isolation and permission freezing, forming a closed-loop risk control system.
[0187] In the specific implementation of this invention, the system addresses the complex semantic structure and non-standard expression of legal entities and their relationships in judicial judgments by proposing a semantic extraction and knowledge graph construction method based on the BERT model. This method first uses a pre-trained Chinese BERT model to perform contextual encoding on the document text and then extracts four core legal entities—"plaintiff," "defendant," "case," and "court"—using Named Entity Recognition (NER) technology. Based on this, the system models the context between entity pairs using a multi-classifier, extracting three types of event relationships: "litigation," "acceptance," and "being sued," thereby constructing a case-centric judicial knowledge graph structure. The graph not only includes entities and their semantic relationships but also further associates additional attribute information, such as the plaintiff's and defendant's legal status and litigation capacity, and semantic tags like the case's occurrence time and trial outcome, giving the knowledge graph stronger semantic expression and reasoning capabilities. All structured triples and attribute information are ultimately organized using a node-edge model and stored in the Neo4j graph database, supporting complex semantic queries and graph traversal operations. This solution ensures the accuracy of entity extraction while systematically addressing the problem of the difficulty in structuring the semantics of judicial documents. Furthermore, it achieves efficient modeling, storage, and retrieval of judicial semantic information through a customized legal graph structure design.
[0188] In one embodiment of the present invention, to address the technical challenges of log falsification, difficulty in tracing behavior, and delayed risk response in traditional judicial data operation and maintenance, the system integrates a behavior auditing and risk control mechanism driven by blockchain and deep learning. First, the system deploys a lightweight log capture agent on each business node to record user behavior operation logs, including fields such as request source IP, user ID, and operation type, and uses SM2 signatures to bind the identity of the operation subject. After unified structuring, the collected logs are submitted to the Hyperledger Fabric blockchain for evidence storage.
[0189] To identify potential anomalous behaviors, the system built a behavior modeling engine based on LSTM neural networks.
[0190] In one embodiment of the present invention, to identify potential abnormal behavior, the system constructs a behavior modeling engine based on an LSTM neural network. This engine takes on-chain log sequences as input, with log fields including user ID, timestamp, operation type, resource type, access IP, and other information. The system uses a preprocessing procedure to convert the log sequences into time-series vectors of user operation categories (e.g., operation categories such as login, querying case files, exporting videos, and downloading evidence are expressed using one-hot encoding, with time information embedded through time difference or periodic position information).
[0191] During the training phase, a training sample set is constructed according to the user dimension. An LSTM model is independently trained for each account type (e.g., judges, court clerks, prosecutors) to capture the temporal dependencies and operational patterns of their actions. The model takes a fixed-window-length sequence of operations as input and uses the difference between the predicted distribution of the next possible operation and the actual operation as the discrimination criterion. The training objective employs a combination of supervised and semi-supervised strategies: supervised training is performed using cross-entropy loss on data labeled as normal behavior, while prediction bias (e.g., reconstruction error) is used to control the model update rate for unlabeled data to avoid introducing abnormal patterns. Ultimately, the system learns a set of behavioral pattern vectors for each account type and calculates the prediction bias using a sliding window. When abnormal behavior deviating from the threshold (e.g., high-frequency calls, unauthorized access) is detected, an early warning signal is immediately sent to the risk control system. Simultaneously, a composite risk assessment engine runs in the background, connecting in real-time to the CVE vulnerability database and the judicial industry threat signature database. After an early warning is triggered, this engine automatically calculates a risk score based on parameters such as the exposed asset surface, attack difficulty, and impact scope, and then executes an emergency response strategy based on the score classification.
[0192] High risk (score ≥ 7): Immediately isolate the user session involved, freeze their relevant access permissions and resources, issue a patch repair command to the affected nodes, and initiate a notification mechanism with judicial authorities for manual intervention and source tracing;
[0193] Medium risk (score 4-6.9): Temporarily freeze access permissions to some high-sensitivity resources, mark users as audit targets, record behavior logs and notify the system administrator for verification, and perform pre-patch push if necessary;
[0194] Low risk (score < 4): Record event logs, increase the weight of the behavior model and strengthen monitoring. The risk level will be automatically increased if the subsequent behavior deviation continues to worsen.
[0195] This mechanism realizes a closed-loop response path in the judicial data system, from anomaly detection and level assessment to handling feedback, which significantly improves the security self-healing capability and attack response efficiency of classified data systems.
[0196] Example 2
[0197] The various business platforms within the judicial system (including the court trial system, the procuratorate business system, and the judicial administration management system) are interconnected through a unified data exchange platform to achieve centralized data collection. This platform supports multiple data connection methods, including standardized API calls, message middleware push (such as Kafka), and scheduled data retrieval, to adapt to the different system connection capabilities and data update frequencies. The system can simultaneously collect the following three types of multimodal data:
[0198] 1. Structured data: such as case number, party information, judgment result, timestamp, etc., which come from the court's trial database and the procuratorial organ's business system;
[0199] 2. Semi-structured data: such as audio and video files, spreadsheets, and electronic transcripts from court proceedings;
[0200] 3. Unstructured data: including scanned evidence materials, photographs, PDF-format court documents, and audio recordings of evidence in audio files.
[0201] After data collection, the system automatically performs format unification and pre-parsing processing on the data. Structured data is directly parsed into a standard field set; audio and video data are transcribed using the ASR service integrated with the iFlytek Open Platform, converting speech content into text; keyframes of image and video files are extracted using OpenCV tools for subsequent visual analysis. In the data preprocessing stage, the Deepseek-v3 large language model deployed on a local server is used to perform entity recognition and context analysis on the data content. The model encodes the input text data and uses the Named Entity Recognition (NER) algorithm to identify the following sensitive fields: name of the person involved, ID number, phone number, address, company name, bank account, etc. The identified sensitive fields then enter the differential privacy processing module. Based on the preset sensitivity field levels and de-identification strategies, the system adopts different processing methods for different types of data:
[0202] For strongly identifying fields such as ID card numbers and phone numbers, irreversible desensitization methods such as random perturbation and hash replacement are used. For fields that partially retain semantic meaning (such as court names and work units), recoverable desensitization methods such as character replacement and token masking are used. Differential privacy budget values are automatically appended to the processed fields and recorded in metadata tags for future reference by the access control module.
[0203] The data that has been anonymized needs to be transmitted across departments between various judicial systems. To ensure that the data is not stolen, tampered with or counterfeited during transmission, the system is designed with multiple security mechanisms, builds an end-to-end encrypted communication channel and achieves verifiable transmission protection throughout the entire process.
[0204] First, on the data transmission link, the system employs the SM4 symmetric encryption algorithm based on Chinese national cryptographic algorithms, combined with the TLS 1.3 transport layer security protocol to construct an encrypted channel. Before communication, both parties establish a secure connection through a TLS handshake process, introducing the following key security mechanisms during the handshake phase:
[0205] 1. Two-way authentication: Using two-way digital certificates based on the X.509 standard, it enables legitimate identity verification for both communicating parties (such as court data gateways and procuratorate service nodes);
[0206] 2. Biometric identification-assisted authentication: Facial recognition or fingerprint authentication mechanisms are introduced at specific high-privilege nodes (judicial data control center) and bound to digital certificates to prevent key theft;
[0207] 3. Key exchange mechanism: Elliptic curve cryptography (ECC) is used to complete key exchange, ensuring the establishment of symmetric session keys under asymmetric conditions.
[0208] After the data is encrypted at the session layer, it enters the data content layer encryption stage. The data body is encrypted in blocks using the SM4-CBC mode, and the key is the symmetric key generated during the TLS handshake stage. After encryption, the data is appended with an integrity check value calculated by the HMAC-SM3 algorithm, which is used to perform consistency verification at the receiving end before decryption, preventing man-in-the-middle attacks and data tampering.
[0209] After each data file is encrypted, the system calculates a unique SHA-256 hash digest value for the data content, which serves as the "digital fingerprint" of the data file for end-to-end consistency verification and blockchain registration.
[0210] To ensure traceability and non-repudiation of transmission activities, the system packages the hash fingerprint, sender and receiver node IDs, timestamps, and other metadata of each data transmission into a transaction record and uploads it to the Hyperledger Fabric consortium blockchain system. The Fabric network includes multiple blockchain nodes representing judicial departments (courts, procuratorates, and judicial bureaus), and uses the PBFT fault-tolerant consensus mechanism for multi-node voting confirmation. This ensures that each upload is certified by the consensus of these three authorities, preventing single-point tampering or forgery. Each data transaction during transmission generates a unique transaction ID, which is bound to the transmission task record. The system can use the Fabric blockchain's smart contract to query whether the transaction was successfully uploaded to the blockchain, whether the node signatures are consistent, and whether the original file has been modified.
[0211] This invention integrates an intelligent classification engine based on multimodal deep learning for content recognition and automatic classification of structured, semi-structured, and unstructured data. The classification engine performs preprocessing based on the data source:
[0212] For text-based data (such as case documents, court transcripts, etc.), the system uses Deepseek-v3 to perform semantic analysis and feature extraction on the text, and completes document topic classification and key tag annotation.
[0213] For video data (such as court recordings), OpenCV is used to sample video frames and extract keyframe images; then, VGG16 convolutional neural network is used to extract image features, and finally, the visual features are input into a decision tree classifier for semantic label assignment.
[0214] For audio data (such as witness recordings), the system first converts speech to text through the iFlytek Open Platform, and then performs subsequent semantic analysis and classification according to the text processing path.
[0215] All category tags are individually bound to the original data. This tag information also serves as the foundation for subsequent modules such as access control, search optimization, and data visualization analysis, achieving semantic consistency in data processing and control. The classification results, along with the tag information, generate SHA-256 hash digests and are stored on the blockchain as evidence. This process ensures that the intelligent classification behavior has auditable credentials, facilitating later retrospective processing and meeting the interpretability and compliance requirements of the judicial system.
[0216] To achieve fine-grained access control and dynamic response control during access to classified data, the system integrates a policy engine based on XACML and combines a hybrid access control model of RBAC (role-based) and ABAC (attribute-based) to adapt to complex judicial business needs.
[0217] When the system receives an access request, it first forwards the request information (including visitor identity, access time, requested resource, device information, etc.) to the permission policy engine in real time via a Kafka message queue. The permission judgment process includes the following steps:
[0218] User identity and role binding: Visitors complete identity authentication through X.509 digital certificates, and the system automatically resolves their roles (such as judge, prosecutor, court clerk, lawyer, etc.).
[0219] Attribute-aware analysis: The system collects and analyzes contextual attributes in access requests, such as access time period, geographical location (via GPS module), access device type, and current network environment;
[0220] Policy evaluation and decision-making: The XACML policy engine matches the preset access control policies, combines the role and permission boundaries of RBAC with the attribute matching logic of ABAC for joint evaluation, and finally obtains control results such as "allow", "deny" or "further authentication required".
[0221] Dynamic two-factor authentication: For situations where there is uncertainty in accessing sensitive data or policy judgments, the system forcibly introduces a two-factor authentication mechanism of timestamp + geolocation. This requires users to complete identity verification through a bound legitimate terminal within a specified time and authorized location, thereby improving the credibility of access behavior.
[0222] The authorization result is recorded in the form of a digital signature and simultaneously written into the blockchain audit module, serving as a valid credential for subsequent behavior tracing and compliance auditing.
[0223] To meet the requirements of secure storage and efficient retrieval of classified data within the judicial system, this system employs a tiered encryption strategy and a smart contract-driven data lifecycle management mechanism. Specifically, this includes the following mechanisms:
[0224] 1. Hierarchical encryption storage mechanism
[0225] Data is categorized based on its sensitivity into sensitive data (such as case files and investigation materials) and general data (such as publicly available court judgments and meeting minutes):
[0226] Sensitive data is encrypted using AES-256 symmetric encryption and then uniformly uploaded to a private cloud platform deployed in a dedicated network environment. Simultaneously, homomorphic encryption technology (such as BFV or CKKS) is introduced, enabling some keyword searches and comparisons to be performed in encrypted form, avoiding plaintext exposure.
[0227] General data uses erasure coding and Reed-Solomon coding for distributed redundant slice storage, which is distributed across different physical nodes to improve disaster recovery and resilience.
[0228] Metadata and permission indexes are built on the Neo4j graph database, where nodes and edges represent data entities and their logical relationships, supporting complex semantic retrieval and access control assisted judgment.
[0229] 2. Blockchain On-Chain Mechanism
[0230] Each instance of storing classified data generates a unique SHA-256 hash digest, which is then synchronously uploaded to the Hyperledger Fabric consortium blockchain.
[0231] The blockchain nodes consist of courts, procuratorates, and judicial administration departments, and employ the PBFT consensus algorithm to ensure consistency of evidence storage. Each data write, read, modification, and deletion is accompanied by a metadata change event, the summary of which is recorded as an on-chain transaction. The block structure uses a Merkle tree mechanism to calculate the root hash, ensuring that any tampering with any data record in the block can be quickly detected.
[0232] 3. Smart Contract Management Mechanism
[0233] The system pre-deploys a series of smart contracts developed based on Fabric Chaincode, which execute according to standardized ABI interface definitions. Contract functionalities include:
[0234] Data access timeliness rules definition: can be set such as "read-only for the first 3 years after release from prison, destroyed after 3 years";
[0235] Access behavior triggering rules: It is possible to set "access to specific levels of data requires joint authorization from two levels of institutions";
[0236] Automatic destruction mechanism: The contract calls an external time oracle to determine the validity period of the data. After the data expires, the data erasure process is executed, and the destruction behavior is recorded on the blockchain for future reference.
[0237] All contract operations are protected by digital signatures and on-chain verification mechanisms, recording the operating entity and timestamp, forming a legally compliant chain of responsibility.
[0238] To enhance the semantic organization and interpretability of judicial data, this system constructs a case-centric legal knowledge graph to achieve a structured expression of judgments and their key information. This technical solution mainly includes four stages: entity extraction, relation identification, graph construction, and graph database storage.
[0239] The system collects text data such as judgments, indictments, and court transcripts from the court's adjudication system. The data is first processed by natural language cleaning (noise removal, sentence segmentation, word segmentation, and part-of-speech tagging) and then input into a pre-trained BERT model for contextual semantic modeling.
[0240] Entity types include, but are not limited to:
[0241] 1. Court-related entities: name, region, jurisdiction, judges;
[0242] 2. Subject entities: Plaintiff, Defendant, and supplementary attributes such as their identity information, education level, litigation capacity, and legal standing;
[0243] 3. Case-related elements: time of the incident, location of the incident, applicable legal provisions, trial outcome, and case closure time.
[0244] Based on the semantic and legal logical relationships between entities, the system constructs a graph edge structure and defines semantic paths such as "plaintiff → lawsuit → court", "court → acceptance → case", "case → reference → legal provisions", and "case → participation → defendant", which fully reflect the subject-object relationship in the judicial process.
[0245] The system maps the aforementioned entities and relationships to nodes and edges in a knowledge graph, forming a multi-level knowledge network with "cases" as the core node. The overall knowledge graph structure consists of the following four types of nodes and three types of edges:
[0246] Entity types (nodes): Case, Plaintiff, Defendant, Court
[0247] Relationship types (edges): Litigation, Acceptance, Being Litigated
[0248] This graph uses the Neo4j graph database as its storage medium, leveraging its native graph structure to support efficient entity querying and relational reasoning. Each node is bound to a unique ID and its set of attributes, and each edge defines a specific relationship type and occurrence time. The graph structure is sustainably expandable, supporting the future addition of more dimensions of legal entities and relationships, such as chains of evidence, litigation stages, and appeal relationships.
[0249] To ensure the system's operation is visible, controllable, and auditable, this system has built an operation and maintenance monitoring mechanism based on blockchain auditing, anomaly detection, risk assessment, and automatic response, covering operation log uploading to the blockchain, security audit analysis, and multi-dimensional risk management.
[0250] 1. On-chain auditing of operational behavior
[0251] All operational and data access behaviors, such as uploading, downloading, querying, permission changes, and authentication, are logged in a unified audit agent to generate complete logs.
[0252] The log content includes operation type, timestamp, user identity (digital signature), resource identifier, access IP, and device information. Each log entry calculates a SHA-256 hash digest, which is then encapsulated into blocks in chronological order and uploaded to Hyperledger Fabric. The PBFT consensus mechanism ensures the log records are non-repudiable and tamper-proof, facilitating subsequent compliance audits and traceability.
[0253] 2. Abnormal Behavior Detection and Intelligent Alarm
[0254] By introducing an LSTM (Long Short-Term Memory) neural network to model continuous behavioral log data, potential high-risk behavioral patterns can be identified.
[0255] Construct multi-dimensional time-series features (such as access frequency, data call types, operation paths, etc.) to train a baseline model to identify "normal behavior flows," and perform real-time comparisons during the deployment phase. Once a user's behavior deviates (e.g., daily access frequency exceeds the long-term average + 3σ threshold), an anomaly alarm is triggered. Alarm events are logged on the chain for traceability.
[0256] 3. Risk assessment and automatic response mechanism
[0257] The system integrates the CVE vulnerability database with a custom risk signature database for the judicial industry, performing dynamic threat modeling at the network, identity, and application layers. Risk modeling employs the DREAD model, assigning a score to each risk and updating it dynamically. Once a high-risk vulnerability or abnormal behavior closely matches attack characteristics is detected, the system automatically calculates its current risk value. Based on the risk level (high, medium, low), a tiered emergency response mechanism is activated, specifically including: isolating the affected user sessions, freezing access to specific privileged resources, issuing patch commands, and initiating a notification mechanism with judicial authorities.
[0258] All risk contingency operations are also confirmed using digital signatures and written into the blockchain audit system.
[0259] This invention provides a solution to the problem of secure transmission and sharing of business data based on the interconnection and interoperability of data among the judicial, procuratorial, and legal departments. It enables secure sharing and exchange of judicial data across levels, departments, and networks. During file transmission and sharing, the system strengthens trust verification between transmission objects and authorization authentication mechanisms for requesting objects. Ensuring strict authentication and authorization are required for establishing data transmission and file sharing channels, different security levels are established as needed to improve the reliability of system terminals and solve the security problems of cross-level, cross-departmental, and cross-network judicial data sharing and exchange within the judicial, procuratorial, and legal systems. This will enable the secure and reliable flow and exchange of legal data among the judicial, procuratorial, and legal departments according to case-handling needs. Based on blockchain technology, it supports the daily sharing of necessary case-related legal data within the legal system, such as the exchange of case trial information, electronic documents, and information on criminal sentence reduction and parole.
Claims
1. A method for secure transfer of classified data in the judicial system based on blockchain technology, characterized in that, Includes the following steps: Step 1: Collect multi-source judicial data, including structured, semi-structured, and unstructured data; Step 2: Use a large language model to identify sensitive information in the data and perform desensitization processing based on a differential privacy mechanism; Step 3: Construct an end-to-end encrypted channel using the SM4 encryption algorithm and TLS 1.3 protocol, and complete the data transmission obtained in Step 2 by combining X.509 certificate, biometric authentication and ECC key negotiation; Step 4: Generate a SHA-256 hash fingerprint for the data encrypted by the SM4 encryption algorithm in Step 3, and write the hash fingerprint and transmission metadata into the blockchain system to achieve cross-node evidence storage through the PBFT consensus mechanism; Step 5: Classify the data using a multimodal classification engine, including text semantic analysis, video keyframe feature extraction, and speech-to-text processing, and output case type, confidentiality level, and archiving path labels. These labels provide fine-grained attribute labels for subsequent step 6, access control. Step 6: Upon receiving a user access request, a hybrid access control model is adopted, integrating the XACML policy framework and Kafka message queue to achieve real-time response and authentication of access requests. The dynamic access management of time, space, and device attributes is supported through the joint mechanism of RBAC and ABAC, and the user identity and device are bound by two-factor authentication of geographic location and timestamp. Step 7: Adopt a hierarchical encryption storage strategy to encrypt sensitive data and store it in a private cloud, supporting homomorphic encryption retrieval. General data is stored in a distributed redundant manner using erasure coding, and metadata is built in a graph database. Step 8: Construct a judicial knowledge graph, extract legal entities and semantic relationships from judgment documents based on the BERT model, and form a graph network containing four types of entities and three types of relationships to support semantic association query and decision assistance; Step 9: Record operation logs based on the blockchain audit system, combine LSTM neural network and dynamic threshold detection to identify abnormal behavior, and conduct risk assessment and graded emergency response through the DREAD model.
2. The method according to claim 1, characterized in that, Step 1 includes the following sub-steps: S1.1: Connects to the court trial system, the procuratorate business system, and the judicial administration management system, supporting three data connection methods: standardized API calls, message middleware push, and scheduled data retrieval; S1.2: Synchronously collect structured data, semi-structured data, and unstructured data, wherein the structured data includes case number, party information, and judgment result fields, and the unstructured data includes scanned evidence materials, court hearing audio and video files, PDF format judgment documents, and audio evidence files; S1.3: The collected data is formatted and pre-parsed. Structured data is directly parsed into a set of standard fields. Audio and video data are converted into text data through speech-to-text service. Image and video files are visually analyzed through keyframe extraction tools.
3. The method according to claim 1, characterized in that, Step 2 includes the following sub-steps: S2.1: Use a pre-trained language model to perform Named Entity Recognition (NER) on the collected data to identify sensitive fields including the person's name, ID number, phone number, address, and bank account. S2.2: Implement differential privacy desensitization processing based on preset sensitive field levels. For strongly identifying fields, irreversible desensitization is performed using random perturbation and hash replacement. For fields that need to retain semantics, recoverable desensitization is performed using character replacement and token masking. S2.3: Attach a differential privacy budget value label to the de-identified fields and record the de-identification policy in the metadata for subsequent access decisions by the access control module.
4. The end-to-end encrypted transmission step in the method for secure transfer of classified data according to claim 1, characterized in that, Step 3 includes the following sub-steps: S3.1: Establish an end-to-end encrypted channel based on the TLS 1.3 protocol, complete the identity authentication of both communicating parties through X.509 two-way digital certificates, introduce a biometric identification mechanism in high-privilege nodes, and combine elliptic curve cryptography (ECC) to complete key negotiation and generate symmetric session keys; S3.2: The data body is encrypted in blocks using the SM4-CBC mode, and the HMAC-SM3 algorithm is used to calculate the integrity check value of the encrypted data. The receiving end verifies the consistency of the check value before decryption.
5. The blockchain evidence storage step in the method for secure transfer of classified data according to claim 1, characterized in that, Step 4 includes the following sub-steps: S4.1: Calculate a unique SHA-256 hash digest for the encrypted data content as a digital fingerprint, which is used for end-to-end consistency verification and blockchain registration; S4.2: Package the hash digest, sender node ID, receiver node ID, and timestamp into a transaction record and associate it with the data transmission task metadata; S4.3: Write the transaction record to the Hyperledger Fabric consortium chain composed of nodes from the court, procuratorate, and judicial bureau, and confirm it through multi-judicial node voting via the PBFT consensus mechanism to complete cross-chain evidence storage; S4.4: Based on smart contracts, the on-chain status of transaction records, the consistency of node signatures, and the integrity of original files are verified on-chain, generating a unique transaction ID and binding it to the transmission task log.
6. The blockchain evidence storage step in the method for secure transfer of classified data according to claim 1, characterized in that, Step 5 includes the following sub-steps: S5.1: Construct a multi-channel asynchronous processing architecture, and establish preprocessing pipelines for structured data, text data, audio data, and video data respectively; S5.2: Normalize structured data through JSON Schema mapping and regular expression parsing to generate a standard field set; S5.3: Input text data using the semantic parsing module of DeepSeek-v3, and generate structured semantic vectors through entity recognition and keyword tag extraction; S5.4: After the audio data is transcribed into text using the iFlytek ASR model, the semantic parsing process in step S5.3 is reused to generate semantic vectors; S5.5: The video data is processed by the inter-frame difference algorithm to extract key frames, and the visual features are extracted by the VGG16 network to generate a 4096-dimensional visual vector. S5.6: Input multimodal vectors into a classifier based on a decision tree with a gating mechanism, and output case type, confidentiality level, and archiving path label.
7. The method according to claim 1, characterized in that, Step 6 includes the following sub-steps: S6.1: Embed the Policy Decision Point (PDP) module in the server kernel. Through the front-end proxy component, user access requests are converted into standardized XACML Request XML format and delivered as event messages to the Kafka message queue. The KafkaProducer then pushes the messages to the PDP instance clusters divided according to business dimensions, realizing asynchronous and concurrent distribution of policy judgment tasks. S6.2: Construct a multi-channel Kafka Topic, and automatically classify and route the access request to the corresponding PDP instance for processing according to the business type. The business types include case file viewing, audio and video download, and case file export. S6.3: Adopts a hybrid model of role-based access control (RBAC) and attribute-based access control (ABAC). Roles are defined from the unified authentication platform of the judicial authorities, and attribute dimensions include time window, access geolocation, device fingerprint, network type, and target data sensitivity level. S6.4: Before the client initiates access, the local security module embeds the current UTC timestamp. with geographic coordinates Generate signature digest using the HMAC-SM3 algorithm Attached to the access request, the server verifies whether the timestamp and location parameter match the allowed time period in the XACML policy template. The IP / GPS whitelist of judicial office areas, among which Represents the cryptographic hash function. This represents a pre-shared symmetric key. The symbol indicates data concatenation; S6.5: Complete TLS 1.3 channel identity authentication through the national cryptographic standard X.509 digital certificate, and dynamically generate an access token by combining the organization role information. The token carries the policy path hash and the request context hash fingerprint. S6.6: The access determination result and its corresponding policy path, attribute matching conditions, and user context hash digest are structured and packaged, written into the consortium blockchain through the Hyperledger Fabric smart contract interface, and jointly confirmed and stored by multiple judicial nodes using the PBFT consensus mechanism.
8. The method according to claim 1, characterized in that, Step 7 includes the following sub-steps: S7.1: Storage types are divided according to data sensitivity levels. High-sensitivity data is encrypted using the AES-256 algorithm and stored in a private cloud. A ciphertext indexing mechanism is built based on the Paillier homomorphic encryption algorithm, and ciphertext keyword retrieval is achieved by combining Bloom Filter and inverted index. Medium and low-sensitivity data are generated into k data fragments and r redundant fragments through erasure coding based on Reed-Solomon encoding, and distributed and stored on different physical nodes. S7.2: Store metadata in the Neo4j graph database to construct a judicial semantic graph containing case nodes, party nodes, document nodes and related relationships, supporting graph traversal and pattern matching queries; S7.3: Design multi-layered smart contracts in the Hyperledger Fabric consortium blockchain, including access control subcontracts, audit triggering subcontracts and version traceability subcontracts. Encapsulate data operation rules through standard ABI interfaces. The version traceability subcontract records the data version hash chain based on the Merkle tree and reconstructs the Merkle root to verify the change path during each update. S7.4: The access control subcontract verifies the permission attributes of the calling entity, dynamically matches the strategy expression with off-chain hybrid authorization data, and returns a rejection record when authorization fails and writes it to the asynchronous audit log; S7.5: Data storage behavior, version change events and audit trigger records are written into the blockchain through the PBFT consensus mechanism, and the evidence is jointly confirmed by judicial nodes.
9. The method according to claim 1, characterized in that, Step 8 includes the following sub-steps: S8.1: Perform natural language cleaning on judgment documents, including noise reduction, sentence segmentation, word segmentation, and part-of-speech tagging. Input the pre-trained BERT model to complete the extraction of legal entities. The entity types include four categories: court, plaintiff, defendant, and case. S8.2: Construct a semantic edge structure based on legal logical relationships, define three core relationships: "litigation-acceptance-being sued", and form a multi-level knowledge network centered on case nodes; S8.3: Map the extracted entities and relationships to the Neo4j graph database, bind a unique ID and attribute set to each node, and define the relationship type and occurrence timestamp for each edge; S8.4: Based on the graph structure, semantic association query is realized, which supports traversing the associated plaintiff, defendant, court entity and reference legal provisions information with the case as the entry point, and the query results are attached with access control tags.
10. The method according to claim 1, characterized in that, Step 9 includes the following sub-steps: S9.1: Deploy a lightweight log agent on each business node of the system to capture user operation logs. The log fields should include at least the operation type, resource identifier, timestamp, user digital certificate fingerprint and device hash fingerprint, and generate log digests using the SM3 algorithm. S9.2: Construct an anomaly detection model based on LSTM neural network, independently train behavioral baselines according to user role classification {judge, prosecutor, clerk}, input is the operation sequence vector within the time window, output is the operation category prediction distribution, and determine the deviation of behavior by calculating the KL divergence value between the prediction result and the actual operation. S9.3: A dynamic 3σ threshold detection mechanism is adopted. A normal distribution model is established based on the user's historical operation frequency and resource access mode. When the deviation of a single operation exceeds μ±3σ or the KL divergence value of a continuous operation sequence exceeds the preset threshold, an abnormal alarm event is triggered. S9.4: Integrates the CVE vulnerability database and the threat signature database of the judicial industry, and uses the DREAD model to assess the risk of alert events, outputting risk scores and classifications: High risk: If the score is ≥7, immediately isolate the user session, freeze permissions, and initiate a judicial notification mechanism; Medium risk: Score 4-6.9, restrict access to sensitive resources and mark audit trails; Low risk: Score < 4 points, increase log recording frequency and dynamically adjust behavioral model weights; S9.5: Encapsulate the metadata of abnormal events, risk assessment results, and response operation records into blockchain transactions, and write them into the Hyperledger Fabric consortium blockchain through the PBFT consensus mechanism. The evidence storage nodes include court audit departments, procuratorate supervision nodes, and judicial technology verification nodes.
Citation Information
Patent Citations
Protection system and method for judicial protection of juveniles
CN118013551A
Education digital identity authentication method of distributed heterogeneous system based on block chain
CN119538223A