Judicial system confidential data security circulation method based on block chain technology

Through blockchain technology, the trusted transmission channel of judicial data is built, and data security and privacy protection problems in cross-level, cross-departmental and cross-network scenarios are solved, efficient, traceable flow and intelligent analysis of confidential data are realized, and secure and controllable data sharing in the judicial system is supported.

CN120567451AActive Publication Date: 2025-08-29UESTC (SHENZHEN) ADVANCED RES INST +1

Patent Information

Application Number
CN202510630917.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-29
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The existing judicial data exchange has problems such as interface fragmentation, high transmission security risks, hidden dangers of privacy leakage and lack of full life cycle control in cross-level, cross-departmental and cross-network scenarios, making it difficult to achieve high security, high efficiency, and traceable flow of confidential data.

Method used

Blockchain technology is used to build a trusted transmission channel, end-to-end encryption is realized through SM4 encryption algorithm and TLS1.3 protocol, desensitization is carried out in combination with differential privacy mechanism, identity verification is used using X.509 certificates and biometric authentication, and cross-node evidence storage is realized through the PBFT consensus mechanism, combining hybrid permission control and hierarchical encrypted storage, a judicial knowledge graph is built for intelligent analysis.

Benefits of technology

It realizes the full process of tamper-proof, traceable sharing and exchange of confidential data across institutions, ensures data security and privacy protection, supports cross-departmental intelligent collaboration and compliance control, and improves the efficiency and reliability of judicial data flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005405267610000011
    Figure HDA0005405267610000011
  • Figure HDA0005405267610000012
    Figure HDA0005405267610000012
  • Figure HDA0005405267610000021
    Figure HDA0005405267610000021
Patent Text Reader

Abstract

The invention relates to the technical field of judicial data security, and discloses a judicial system confidential data security circulation method based on a block chain technology. Collecting multi-source judicial data, identifying sensitive information through a large language model, and performing differential privacy desensitization processing; constructing an SM4 encryption and TLS 1.3 end-to-end secure channel, generating a data hash fingerprint, and writing the data hash fingerprint into an alliance chain evidence based on a PBFT consensus mechanism; designing a multi-modal classification engine to perform feature extraction and intelligent classification; establishing a hybrid authority model to integrate an XACML strategy and a Kafka queue, and implementing dynamic authority control in combination with an RBAC / ABAC mechanism; a hierarchical encryption storage architecture is constructed, and homomorphic encryption retrieval and erasure code distributed storage are adopted; constructing a judicial knowledge graph based on a BERT model; and deploying a block chain auditing system, and combining LSTM anomaly detection and a DREAD risk assessment model to form a closed-loop risk control system. According to the method, the problems of security risk and privacy disclosure in judicial data cross-department circulation are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of secure sharing and trusted transfer of judicial data, and in particular to a method for secure transfer of confidential data in a judicial system based on blockchain technology. Background Art

[0002] The vigorous advancement of court informatization in recent years has accumulated vast amounts of case data, creating an environment for judicial big data research and offering a new perspective and perspective on the ideal of similar cases being judged alike. The construction of smart courts within the context of big data is based on big data technology. Combining basic legal knowledge with judges' actual case experience, it utilizes mathematics, statistics, and cloud computing to establish intelligent analysis and application models, exploring and researching quantitative standards for judicial data. This represents the cross-application of big data and cloud computing technologies in the judicial field.

[0003] While methods for quantifying judicial data haven't yet been clearly defined, there's a certain degree of consensus in practice. Each judge's judgments on similar cases gradually develop a consensus-based approach through classic case guidance from senior institutions, peer exchange and training, and other channels. Once these consensus-based, fuzzy indicators are quantified, they become computer-understandable data, enabling the transition from empirical judgment to scientific decision-making. This also provides regulatory authorities with an effective oversight tool, playing an important role in avoiding different judgments in similar cases and resolving social conflicts. Furthermore, refined case information management helps courts gain a comprehensive understanding of their current status, allows the public to intuitively experience judicial fairness, and enables legal professionals to more deeply analyze case issues.

[0004] The development of smart courts under the big data perspective begins with the intelligent identification and judgment of similar cases. We have solicited and validated judgment methods from a large number of judges. Based on this foundation, we further explore the common characteristics and key factors of similar cases, calculate their similarity, and establish a similarity judgment model. Through comprehensive computer exploration of judicial big data, we identify discrepancies in the judgment results of similar cases. We also explore the subjective and objective reasons for these discrepancies, using expert discussions and computer fuzzy calculations. Ultimately, we explore the influencing factors of similar cases and the patterns of judgment change. This provides guidance for judges' future judgments, a more comprehensive analytical tool for trial management, and a richer perspective for the study of similar cases.

[0005] With the improvement of citizens' legal awareness and the lowering of the threshold for filing cases, the number of cases in the people's courts has shown a significant upward trend in recent years, and the types and circumstances of cases have become more complicated. At the same time, the continued growth of judicial expectations and social responsibilities has also placed higher demands on the courts' trial work.

[0006] Existing judicial data exchange relies primarily on manual collection or dedicated network connection, resulting in system fragmentation, inconsistent interfaces, and inconsistent data formats, making it difficult to meet the business needs of high-frequency collaboration among multiple institutions. Furthermore, because the data contains a large amount of confidential information, such as the identities of the parties, case details, and evidentiary materials, it is extremely vulnerable to privacy exposure and the risk of tampering and leakage during data circulation. Traditional encryption methods and permission control mechanisms are unable to meet the high requirements for "full lifecycle security" in modern judicial collaboration scenarios. Furthermore, the lack of a unified and trusted operation audit and abnormal behavior identification mechanism makes it difficult to effectively trace data flows and quickly respond to risk events in complex distributed environments.

[0007] On the other hand, while artificial intelligence and big data analytics have made progress in identifying similar cases, structuring judicial documents, and building knowledge graphs, they lack deep integration with data security mechanisms. This has led to intelligent models facing the "data silo" problem, limiting the ability to build cross-departmental models and collaborate on knowledge. As a highly sensitive resource, judicial data urgently requires a new security mechanism that enables cross-domain intelligent analysis and collaborative training without sacrificing privacy.

[0008] Therefore, there is an urgent need for a judicial data circulation method with the capabilities of "trusted exchange, privacy protection, intelligent identification, and compliance control", which can realize the high-security, high-efficiency, and traceable circulation of confidential data across institutions and systems, thereby providing a safe and controllable digital foundation for the judicial system, supporting the in-depth development of similar cases with the same judgment, trial supervision, and smart judicial construction. Summary of the Invention

[0009] The technical problem to be solved by the present invention is: how to overcome the problems of interface fragmentation, high transmission security risks, privacy leakage risks and lack of full life cycle management in the existing technology in the scenario of judicial data sharing across levels, departments and networks, and build a trusted transmission channel for confidential data through blockchain technology to achieve secure desensitization, end-to-end encrypted transmission, blockchain evidence storage and fine-grained authority control of multi-source heterogeneous data in the judicial system, so as to ensure the full process of tamper-proof and traceable sharing and exchange of confidential data among the three departments of law, prosecution and justice.

[0010] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0011] The present invention provides a method for securely transferring confidential data in a judicial system based on blockchain technology, comprising the following steps:

[0012] Step 1: Collect multi-source judicial data, including structured, semi-structured, and unstructured data;

[0013] Step 2: Use a large language model to identify sensitive information in the data and perform desensitization based on a differential privacy mechanism;

[0014] Step 3: Use the SM4 encryption algorithm and TLS 1.3 protocol to build an end-to-end encrypted channel, and combine X.509 certificates, biometric authentication, and ECC key negotiation to complete data transmission;

[0015] Step 4: Generate a SHA-256 hash fingerprint for the encrypted data, write the hash fingerprint and transmission metadata into the blockchain system, and implement cross-node evidence storage through the PBFT consensus mechanism;

[0016] Step 5: Classify the data using a multimodal classification engine, including text semantic analysis, video keyframe feature extraction, and speech-to-text processing;

[0017] Step 6: Adopt a hybrid permission control model, integrating the XACML policy framework with the Kafka message queue to achieve real-time response and authentication to access requests. This model supports dynamic permission management based on time, space, and device attributes through a combined RBAC and ABAC mechanism. Furthermore, it uses geolocation and timestamp-based two-factor authentication to bind user identities to devices.

[0018] Step 7: Adopt a hierarchical encryption storage strategy. Sensitive data is encrypted and stored in a private cloud with support for homomorphic encryption retrieval. General data is distributed and redundantly stored using erasure codes, and metadata is built in a graph database.

[0019] Step 8: Construct a judicial knowledge graph. Based on the BERT model, extract the legal entities and semantic relationships in the judgment documents, forming a graph network containing four types of entities and three types of relationships to support semantic associative query and decision-making assistance.

[0020] Step 9: Conduct cross-region model training based on the federated learning framework, implement gradient aggregation through secure multi-party computation (MPC) and Paillier encryption, and combine differential privacy to ensure secure training without leaving the original data domain.

[0021] Step 10: During the system operation phase, the blockchain audit system records operation logs, combines LSTM neural networks with dynamic threshold detection to identify abnormal behaviors, and uses the DREAD model to conduct risk assessment and graded emergency response.

[0022] In the above scheme, step 1 includes the following sub-steps:

[0023] S1.1: Connects to the court trial system, procuratorate business system, and judicial administration management system, supporting three data connection methods: standardized API call, message middleware push, and scheduled data pull;

[0024] S1.2: Synchronously collect structured data, semi-structured data, and unstructured data. Structured data includes fields such as case number, party information, and judgment outcome. Semi-structured data includes courtroom audio and video files and electronic transcripts. Unstructured data includes scanned evidence materials, PDF-formatted judgment documents, and audio evidence files.

[0025] S1.3: Collected data is formatted and pre-parsed. Structured data is directly parsed into standard field sets. Audio and video data is converted into text data through a speech transcription service. Images and video files are visually analyzed using keyframe extraction tools.

[0026] In the above scheme, step 2 includes the following sub-steps:

[0027] S2.1: Use a pre-trained language model to perform named entity recognition (NER) on the collected data to identify sensitive fields such as the party’s name, ID number, phone number, address, and bank account number.

[0028] S2.2: Implement differential privacy desensitization based on preset sensitive field levels. For strongly identifying fields, use irreversible desensitization using random perturbation and hash replacement. For fields that need to retain semantic meaning, use recoverable desensitization using character replacement and token masking.

[0029] S2.3: Attach a differential privacy budget value label to the desensitized field and record the desensitization policy in the metadata for access decisions in subsequent permission control modules.

[0030] In the above scheme, step 3 includes the following sub-steps:

[0031] S3.1: Establish an end-to-end encrypted channel based on the TLS 1.3 protocol, authenticate both communicating parties using X.509 two-way digital certificates, introduce biometric recognition mechanisms in high-authority nodes, and use elliptic curve cryptography (ECC) to negotiate keys and generate symmetric session keys.

[0032] S3.2: Use SM4-CBC mode to encrypt the data body in blocks, and use the HMAC-SM3 algorithm to calculate the integrity check value of the encrypted data. The receiving end verifies the consistency of the check value before decryption.

[0033] In the above scheme, step 4 includes the following sub-steps:

[0034] S4.1: Calculate and generate a unique SHA-256 hash digest of the encrypted data content as a digital fingerprint for full-process consistency verification and blockchain registration;

[0035] S4.2: Packing the hash summary, sender node ID, receiver node ID, and timestamp into a transaction record and associating it with the data transmission task metadata;

[0036] S4.3: Write the transaction record to the Hyperledger Fabric consortium chain composed of nodes from the court, procuratorate, and judicial bureau. Use the PBFT consensus mechanism for multi-judicial node voting confirmation to complete cross-chain evidence storage.

[0037] S4.4: Based on the smart contract, the transaction record’s on-chain status, node signature consistency, and original file integrity are verified on-chain, and a unique transaction ID is generated and bound to the transfer task log.

[0038] In the above scheme, step 5 includes the following sub-steps:

[0039] S5.1: Build a multi-channel asynchronous processing architecture and establish preprocessing pipelines for structured data, text data, audio data, and video data respectively;

[0040] S5.2: Normalize structured data through JSON Schema mapping and regular expression parsing to generate a standard field set;

[0041] S5.3: Input the text data into the semantic parsing module based on DeepSeek-v3, and generate structured semantic vectors through entity recognition and keyword label extraction;

[0042] S5.4: After the audio data is transcribed into text using the iFlytek ASR model, the semantic parsing process of step S5.3 is reused to generate a semantic vector.

[0043] S5.5: Use the inter-frame difference algorithm to extract key frames from the video data, extract visual features through the VGG16 network, and generate a 4096-dimensional visual vector;

[0044] S5.6: Input the multimodal vector into the classifier based on the gated decision tree, and output the case type, confidentiality level, and filing path label.

[0045] In the above scheme, step 6 includes the following sub-steps:

[0046] S6.1: A policy decision point (PDP) module is embedded in the server kernel. The front-end proxy component converts user access requests into the standardized XACML Request XML format and delivers them as event messages to the Kafka message queue. The KafkaProducer pushes these messages to the PDP instance clusters divided by business dimensions, enabling asynchronous and concurrent policy decision task distribution.

[0047] S6.2: Build a multi-channel Kafka topic to automatically classify and route access requests to the corresponding PDP instance based on their service type, including case file viewing, audio and video downloading, and file exporting.

[0048] S6.3: A hybrid model of role-based access control (RBAC) and attribute-based access control (ABAC) is adopted, where role definitions are derived from the unified authentication platform of judicial authorities, and attribute dimensions include time windows, access geographic location, device fingerprint, network type, and target data sensitivity level;

[0049] S6.4: Before the client initiates access, the local security module embeds the current UTC timestamp t current Generate signature summary D=H through HMAC-SM3 algorithm with geographic coordinates (lat, lng) SM3 (K secret ||t current ||lat||lng) is appended to the access request, and the server verifies whether the timestamp and location parameters match the allowed time period [t start , t end ] and the judicial office area IP / GPS whitelist, including H SM3 SM3 cryptographic hash function, K secret Indicates a pre-shared symmetric key, and the || sign indicates data concatenation;

[0050] S6.5: Complete TLS 1.3 channel identity authentication using the national cryptographic standard X.509 digital certificate, and dynamically generate an access token based on the organization role information. The token carries the policy path hash and request context hash fingerprint.

[0051] S6.6: The access decision results and their corresponding policy paths, attribute matching conditions, and user context hash summaries are packaged and structured, written into the consortium chain through the Hyperledger Fabric smart contract interface, and jointly confirmed and stored by multiple judicial nodes using the PBFT consensus mechanism.

[0052] In the above scheme, step 8 includes the following sub-steps:

[0053] S7.1: Data storage types are divided according to their sensitivity levels. Highly sensitive data is encrypted using the AES-256 algorithm and stored in a private cloud. A ciphertext indexing mechanism is built based on the Paillier homomorphic encryption algorithm, and ciphertext keyword retrieval is achieved by combining BloomFilter and inverted index lists. For medium and low-sensitivity data, k data fragments and r redundant fragments are generated using erasure coding and Reed-Solomon encoding, and distributed storage is performed on different physical nodes.

[0054] S7.2: Store metadata in a Neo4j graph database to construct a judicial semantic graph containing case nodes, party nodes, document nodes, and their relationships, supporting graph traversal and pattern matching queries;

[0055] S7.3: Design a multi-layered smart contract in the Hyperledger Fabric consortium chain, including an access control subcontract, an audit trigger subcontract, and a version traceability subcontract. Encapsulate data operation rules through a standard ABI interface. The version traceability subcontract records the data version hash chain based on a Merkle tree and reconstructs the Merkle root during each update to verify the change path.

[0056] S7.4: The access control sub-contract verifies the permission attributes of the calling subject, dynamically matches the policy expression based on the off-chain hybrid authorization data, and returns a rejection record and writes it to the asynchronous audit log when authorization fails;

[0057] S7.5: Data storage behavior, version change events, and audit trigger records are written into the blockchain through the PBFT consensus mechanism, and the judicial nodes jointly confirm the evidence.

[0058] S7.1: Data storage types are divided according to their sensitivity levels. Highly sensitive data is encrypted using the AES-256 algorithm and stored in a private cloud. A ciphertext indexing mechanism is built based on the Paillier homomorphic encryption algorithm, and ciphertext keyword retrieval is achieved by combining BloomFilter and inverted index lists. For medium and low-sensitivity data, k data fragments and r redundant fragments are generated using erasure coding and Reed-Solomon encoding, and distributed storage is performed on different physical nodes.

[0059] S7.2: Store metadata in the Neo4.j graph database to construct a judicial semantic graph containing case nodes, party nodes, document nodes, and their relationships, supporting graph traversal and pattern matching queries;

[0060] S7.3: Design a multi-layered smart contract in the Hyperledger Fabric consortium chain, including an access control subcontract, an audit trigger subcontract, and a version traceability subcontract. Encapsulate data operation rules through a standard ABI interface. The version traceability subcontract records the data version hash chain based on a Merkle tree and reconstructs the Merkle root during each update to verify the change path.

[0061] S7.4: The access control sub-contract verifies the permission attributes of the calling subject, dynamically matches the policy expression based on the off-chain hybrid authorization data, and returns a rejection record and writes it to the asynchronous audit log when authorization fails;

[0062] S7.5: Data storage behavior, version change events, and audit trigger records are written into the blockchain through the PBFT consensus mechanism, and the judicial nodes jointly confirm the evidence.

[0063] In the above scheme, step 8 includes the following sub-steps:

[0064] S8.1: Perform natural language cleaning on judgment documents, including denoising, sentence segmentation, and part-of-speech tagging. This is then fed into a pre-trained BERT model to extract legal entities. Entity types include court, plaintiff, defendant, and case.

[0065] S8.2: Construct a semantic edge structure based on legal logical relationships, define the three core relationships of "litigation-acceptance-being sued", and form a multi-level knowledge network centered on the case node;

[0066] S8.3: Map the extracted entities and relationships to the Neo4j graph database, binding each node to a unique ID and a set of attributes, and defining each edge with a relationship type and occurrence timestamp;

[0067] S8.4: Implement semantic associative query based on graph structure, support traversal of related plaintiffs, defendants, court entities and reference legal information using cases as the entry point, and attach access control tags to query results.

[0068] In the above scheme, step 9 includes the following sub-steps:

[0069] S9.1: Deploy a lightweight log agent on each business node in the system to capture user operation logs. The log fields should include at least the operation type, resource identifier, timestamp, user digital certificate fingerprint, and device hash fingerprint. Generate a log summary using the SM3 algorithm.

[0070] S9.2: Build an anomaly detection model based on an LSTM neural network. Independently train behavioral baselines by user role (judge, prosecutor, clerk). The input is the operation sequence vector within the time window, and the output is the predicted distribution of operation categories. The degree of behavioral deviation is determined by calculating the KL divergence between the predicted results and the actual operation.

[0071] S9.3: A dynamic 3σ threshold detection mechanism is used to establish a normal distribution model based on the user's historical operation frequency and resource access pattern. When the real-time monitoring detects that the deviation of a single operation exceeds μ±3σ or the KL divergence value of a continuous operation sequence exceeds the preset threshold, an abnormal alarm event is triggered;

[0072] S9.4: Integrate the CVE vulnerability database and the judicial industry threat signature library, conduct risk assessment on alarm events using the DREAD model, and output risk scores and classifications:

[0073] High risk (≥7 points): Immediately isolate the user session, freeze permissions, and initiate legal notification.

[0074] Medium risk (4-6.9 points): Restrict access to sensitive resources and mark audit trails;

[0075] Low risk (<4 points): Increase logging frequency and dynamically adjust behavior model weights;

[0076] S10.5: Encapsulate the abnormal event metadata, risk assessment results, and response operation records into blockchain transactions, and write them into the Hyperledger Fabric consortium chain through the PBFT consensus mechanism. The evidence storage nodes include the court audit department, the procuratorate supervision node, and the judicial technology verification node.

[0077] The blockchain-based method for securely transferring confidential data within the judicial system provided by this invention, through multi-technical integration and innovative mechanism design, has the following beneficial effects in ensuring data security, privacy protection, transfer efficiency, and compliance control:

[0078] 1. Full-process trust and tamper-proof protection

[0079] A cross-node evidence storage mechanism based on blockchain technology (step 4) ensures the integrity and immutability of data at all stages of transmission, storage, and processing through the on-chain hash fingerprint and PBFT consensus algorithm. Combined with the on-chain verification of operational actions by smart contracts (step 4.4), a trusted traceability chain for the entire lifecycle of judicial data is formed, effectively addressing the tampering risks and liability issues that exist in traditional judicial data exchange.

[0080] 2. Dynamic fine-grained privacy protection

[0081] Using a large language model combined with differential privacy desensitization (step 2), a hierarchical desensitization strategy is implemented for sensitive information, eliminating the risk of leaking personally identifiable data while retaining essential semantic features to support subsequent business processing. By attaching differential privacy budget tags (step 2.3) and associating them with metadata, a dynamic balance between desensitization intensity and data availability is achieved, providing a compliant foundation for cross-institutional data sharing.

[0082] 3. Multi-dimensional secure transmission mechanism

[0083] Establish an end-to-end encrypted channel based on the National Security Algorithm (SM4) and the TLS 1.3 protocol (Step 3), combining ECC key negotiation and biometric authentication to enhance the security of the transmission link. Implement block encryption and integrity verification mechanisms (Step 3.2) to prevent data theft or tampering during transmission, meeting the high security requirements of judicial data in cross-network transmission scenarios.

[0084] 4. Intelligent permission dynamic management and control

[0085] Through the XACML policy framework and hybrid permission model (step 6), role attributes (RBAC) and context attributes (ABAC) are integrated to achieve real-time dynamic authentication based on time, space, and device fingerprints. Combined with two-factor authentication binding (step 6.4), this ensures the consistency of user identity and operating environment, prevents abuse of permissions or unauthorized access, and improves the accuracy of access control for confidential data.

[0086] 5. Multimodal Data Processing and Knowledge Fusion

[0087] Utilize a multimodal classification engine (step 5) to structure text, audio, and video data, and integrate it with the judicial knowledge graph (step 8) to achieve semantic modeling of legal entities and relationships. Use the BERT model to extract key case elements (step 8.1) and construct an association network, providing intelligent support for similar case analysis and adjudication references, enhancing the ability to mine the value of judicial data in cross-departmental collaboration.

[0088] 6. Privacy-preserving collaborative intelligent training

[0089] Cross-domain model training is implemented based on a federated learning framework (Step 9), ensuring that the original data remains within the domain through secure multi-party computation (MPC) and homomorphic encryption. Differential privacy noise is injected during gradient aggregation (Step 9.3). This not only protects data privacy but also promotes cross-institutional model collaborative optimization, addressing the problem of insufficient model generalization caused by judicial data silos.

[0090] 7. Closed-loop risk prevention and control system

[0091] Through blockchain audit logs and the LSTM anomaly detection model (step 10), operational sequences that deviate from normal behavior are identified in real time. Combined with the DREAD risk assessment model (step 10.4), a hierarchical response mechanism is triggered, achieving a closed-loop control system from anomaly discovery, risk assessment, and emergency response, enhancing the proactive defense and compliance governance capabilities of the judicial data system.

[0092] 8. Efficient storage and intelligent retrieval

[0093] Adopt a hierarchical encryption storage strategy (step 7) and implement homomorphic encryption retrieval for highly sensitive data, balancing security and usability. Construct a judicial semantic graph through a graph database (step 7.2), supporting complex association queries and graph traversal analysis, optimizing the retrieval efficiency of unstructured data, and providing a multi-dimensional analytical perspective for judicial decision-making.

[0094] 9. Each step of the present invention forms a set of security-enhanced collaborative systems covering the entire life cycle of data through technical coupling and closed-loop process design. Steps 1 to 3 build a front-end security barrier of "collection-desensitization-encryption": After the multi-source heterogeneous data is standardized and pre-processed (step 1), the sensitive fields are identified through a large language model (step 2) and differential privacy desensitization is implemented, which not only eliminates the risk of privacy leakage but also retains the semantic availability of the data; combined with SM4 encryption and TLS channel (step 3), an end-to-end secure transmission link is built to ensure the confidentiality of data flowing across the network. Steps 4 to 6 form the central control layer of "storage-classification-authority": hash fingerprint chain storage (step 4) provides an unalterable trust anchor for full-process data tracing; the multimodal classification engine (step 5) intelligently analyzes and generates labels for text, audio and video data, and provides fine-grained attribute labels for subsequent authority control (step 6), while the dynamic hybrid authority model (RBAC+ABAC) implements precise access control based on classification labels, spatiotemporal attributes and device fingerprints, forming a closed-loop feedback mechanism of "data label-driven authority strategy". Steps 7 through 10 build a deep application system encompassing "storage, knowledge, and risk control." The combined design of hierarchical encrypted storage (Step 7) and the judicial knowledge graph (Step 8) enables efficient retrieval of encrypted data, even in ciphertext form, through graph semantic associations. The federated learning framework (Step 9), leveraging the differential privacy mechanism of Step 2 and the encrypted channel of Step 3, enables cross-domain model collaborative training and knowledge sharing, addressing the data silo challenge. Blockchain auditing (Step 10) utilizes LSTM anomaly detection and DREAD risk assessment to reversely optimize the permission policy of Step 6 and the encryption strength threshold of Step 3, forming a dynamic security reinforcement loop. Through the deep interweaving of technical elements (e.g., blockchain evidence provides a trusted source for classification labels, the knowledge graph provides semantic features for federated learning, and dynamic permission labels provide input parameters for risk models), these steps ultimately achieve the synergistic amplification of four core capabilities: data security, privacy protection, intelligent analysis, and compliance management. This results in an integrated solution for the cross-domain flow of judicial data: trusted transmission, intelligent processing, and proactive defense.

[0095] In summary, the present invention realizes the trusted exchange, dynamic protection and intelligent collaboration of judicial confidential data in the process of cross-level and cross-departmental circulation through the deep integration of blockchain and various security technologies, providing a safe and controllable technical foundation for the construction of smart courts, while meeting the compliance, efficiency and traceability requirements of judicial data sharing. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] Figure 1 It is the overall framework process design diagram of the present invention.

[0097] Figure 2 It is the knowledge graph design diagram of the present invention.

[0098] Figure 3 This is a structural diagram of the unified data platform of the present invention.

[0099] Figure 4 It is the transmission data flow diagram of the present invention.

[0100] Figure 5 It is a schematic diagram of authority control of the present invention. DETAILED DESCRIPTION

[0101] The present invention relates to the technical field of secure sharing and trusted transfer of judicial data, and in particular to a method for secure transfer of confidential data in a cross-level, cross-departmental, and cross-network judicial system based on blockchain technology. The method supports high-security data sharing among multiple business systems such as trial, prosecution, and judicial administration, and has features such as non-invasive data collection and processing capabilities, content-based sensitive information desensitization, end-to-end encrypted transmission, on-chain evidence storage, traceable auditing, and risk-aware control. It is suitable for the full life cycle security management and transfer of multimodal confidential data such as electronic documents, court trial audio and video, electronic evidence, and case files, and is widely used in judicial business collaboration scenarios such as electronic document exchange, court trial process information sharing, and sentence reduction and parole.

[0102] The purpose of the present invention is to solve the problem of insufficient security and traceability of existing judicial confidential data during the cross-level, cross-departmental and cross-network circulation process.

[0103] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0104] The present invention provides a method for securely transferring confidential data in a judicial system based on blockchain technology, comprising the following steps:

[0105] Step 1: Collect multi-source judicial data, including structured, semi-structured, and unstructured data;

[0106] Step 2: Use a large language model to identify sensitive information in the data and perform desensitization based on a differential privacy mechanism;

[0107] Step 3: Use the SM4 encryption algorithm and TLS 1.3 protocol to build an end-to-end encrypted channel, and combine X.509 certificates, biometric authentication, and ECC key negotiation to complete data transmission;

[0108] Step 4: Generate a SHA-256 hash fingerprint for the encrypted data, write the hash fingerprint and transmission metadata into the blockchain system, and implement cross-node evidence storage through the PBFT consensus mechanism;

[0109] Step 5: Classify the data using a multimodal classification engine, including text semantic analysis, video keyframe feature extraction, and speech-to-text processing;

[0110] Step 6: Adopt a hybrid permission control model, integrating the XACML policy framework with the Kafka message queue to achieve real-time response and authentication to access requests. This model supports dynamic permission management based on time, space, and device attributes through a combined RBAC and ABAC mechanism. Furthermore, it uses geolocation and timestamp-based two-factor authentication to bind user identities to devices.

[0111] Step 7: Adopt a hierarchical encryption storage strategy. Sensitive data is encrypted and stored in a private cloud with support for homomorphic encryption retrieval. General data is distributed and redundantly stored using erasure codes, and metadata is built in a graph database.

[0112] Step 8: Construct a judicial knowledge graph. Based on the BERT model, extract the legal entities and semantic relationships in the judgment documents, forming a graph network containing four types of entities and three types of relationships to support semantic associative query and decision-making assistance.

[0113] Step 9: During the system operation phase, the blockchain audit system records operation logs, combines LSTM neural networks with dynamic threshold detection to identify abnormal behaviors, and uses the DREAD model to conduct risk assessment and graded emergency response.

[0114] In the above scheme, step 1 includes the following sub-steps:

[0115] S1.1: Connects to the court trial system, procuratorate business system, and judicial administration management system, supporting three data connection methods: standardized API call, message middleware push, and scheduled data pull;

[0116] S1.2: Synchronously collect structured data, semi-structured data, and unstructured data. Structured data includes fields such as case number, party information, and judgment outcome. Semi-structured data includes courtroom audio and video files and electronic transcripts. Unstructured data includes scanned evidence materials, PDF-formatted judgment documents, and audio evidence files.

[0117] S1.3: Collected data is formatted and pre-parsed. Structured data is directly parsed into standard field sets. Audio and video data is converted into text data through a speech transcription service. Images and video files are visually analyzed using keyframe extraction tools.

[0118] In the above scheme, step 2 includes the following sub-steps:

[0119] S2.1: Use a pre-trained language model to perform named entity recognition (NER) on the collected data to identify sensitive fields such as the party’s name, ID number, phone number, address, and bank account number.

[0120] S2.2: Implement differential privacy desensitization based on preset sensitive field levels. For strongly identifying fields, use irreversible desensitization using random perturbation and hash replacement. For fields that need to retain semantic meaning, use recoverable desensitization using character replacement and token masking.

[0121] S2.3: Attach a differential privacy budget value label to the desensitized field and record the desensitization policy in the metadata for access decisions in subsequent permission control modules.

[0122] In the above scheme, step 3 includes the following sub-steps:

[0123] S3.1: Establish an end-to-end encrypted channel based on the TLS 1.3 protocol, authenticate both communicating parties using X.509 two-way digital certificates, introduce biometric recognition mechanisms in high-authority nodes, and use elliptic curve cryptography (ECC) to negotiate keys and generate symmetric session keys.

[0124] S3.2: Use SM4-CBC mode to encrypt the data body in blocks, and use the HMAC-SM3 algorithm to calculate the integrity check value of the encrypted data. The receiving end verifies the consistency of the check value before decryption.

[0125] In the above scheme, step 4 includes the following sub-steps:

[0126] S4.1: Calculate and generate a unique SHA-256 hash digest of the encrypted data content as a digital fingerprint for full-process consistency verification and blockchain registration;

[0127] S4.2: Packing the hash summary, sender node ID, receiver node ID, and timestamp into a transaction record and associating it with the data transmission task metadata;

[0128] S4.3: Write the transaction record to the Hyperledger Fabric consortium chain composed of nodes from the court, procuratorate, and judicial bureau. Use the PBFT consensus mechanism for multi-judicial node voting confirmation to complete cross-chain evidence storage.

[0129] S4.4: Based on the smart contract, the transaction record’s on-chain status, node signature consistency, and original file integrity are verified on-chain, and a unique transaction ID is generated and bound to the transfer task log.

[0130] In the above scheme, step 5 includes the following sub-steps:

[0131] S5.1: Build a multi-channel asynchronous processing architecture and establish preprocessing pipelines for structured data, text data, audio data, and video data respectively;

[0132] S5.2: Normalize structured data through JSON Schema mapping and regular expression parsing to generate a standard field set;

[0133] S5.3: Input the text data into the semantic parsing module based on DeepSeek-v3, and generate structured semantic vectors through entity recognition and keyword label extraction;

[0134] S5.4: After the audio data is transcribed into text using the iFlytek ASR model, the semantic parsing process of step S5.3 is reused to generate a semantic vector.

[0135] S5.5: Use the inter-frame difference algorithm to extract key frames from the video data, extract visual features through the VGG16 network, and generate a 4096-dimensional visual vector;

[0136] S5.6: Input the multimodal vector into the classifier based on the gated decision tree, and output the case type, confidentiality level, and filing path label.

[0137] In the above scheme, step 6 includes the following sub-steps:

[0138] S6.1: A policy decision point (PDP) module is embedded in the server kernel. The front-end proxy component converts user access requests into the standardized XACML Request XML format and delivers them as event messages to the Kafka message queue. The KafkaProducer pushes these messages to the PDP instance clusters divided by business dimensions, enabling asynchronous and concurrent policy decision task distribution.

[0139] S6.2: Build a multi-channel Kafka topic to automatically classify and route access requests to the corresponding PDP instance based on their service type, including case file viewing, audio and video downloading, and file exporting.

[0140] S6.3: A hybrid model of role-based access control (RBAC) and attribute-based access control (ABAC) is adopted, where role definitions are derived from the unified authentication platform of judicial authorities, and attribute dimensions include time windows, access geographic location, device fingerprint, network type, and target data sensitivity level;

[0141] S6.4: Before the client initiates access, the local security module embeds the current UTC timestamp t current Generate signature summary D=H through HMAC-SM3 algorithm with geographic coordinates (lat, lng) SM3 (K secret ||t current ||lat||lng) is appended to the access request, and the server verifies whether the timestamp and location parameters match the allowed time period [tstart , t end ] and the judicial office area IP / GPS whitelist, including H SM3 SM3 cryptographic hash function, K secret Indicates a pre-shared symmetric key, and the || sign indicates data concatenation;

[0142] S6.5: Complete TLS 1.3 channel identity authentication using the national cryptographic standard X.509 digital certificate, and dynamically generate an access token based on the organization role information. The token carries the policy path hash and request context hash fingerprint.

[0143] S6.6: The access decision results and their corresponding policy paths, attribute matching conditions, and user context hash summaries are packaged and structured, written into the consortium chain through the Hyperledger Fabric smart contract interface, and jointly confirmed and stored by multiple judicial nodes using the PBFT consensus mechanism.

[0144] In the above scheme, step 8 includes the following sub-steps:

[0145] S7.1: Data storage types are divided according to their sensitivity levels. Highly sensitive data is encrypted using the AES-256 algorithm and stored in a private cloud. A ciphertext indexing mechanism is constructed based on the Paillier homomorphic encryption algorithm, and ciphertext keyword retrieval is achieved by combining Bloom filters and inverted indexes. Medium and low-sensitivity data is generated using erasure coding and Reed-Solomon encoding to generate k data fragments and r redundant fragments, which are then distributed and stored on different physical nodes.

[0146] S7.2: Store metadata in a Neo4j graph database to construct a judicial semantic graph containing case nodes, party nodes, document nodes, and their relationships, supporting graph traversal and pattern matching queries;

[0147] S7.3: Design a multi-layered smart contract in the Hyperledger Fabric consortium chain, including an access control subcontract, an audit trigger subcontract, and a version traceability subcontract. Encapsulate data operation rules through a standard ABI interface. The version traceability subcontract records the data version hash chain based on a Merkle tree and reconstructs the Merkle root during each update to verify the change path.

[0148] S7.4: The access control sub-contract verifies the permission attributes of the calling subject, dynamically matches the policy expression based on the off-chain hybrid authorization data, and returns a rejection record and writes it to the asynchronous audit log when authorization fails;

[0149] S7.5: Data storage behavior, version change events, and audit trigger records are written into the blockchain through the PBFT consensus mechanism, and the judicial nodes jointly confirm the evidence.

[0150] S7.1: Data storage types are divided according to their sensitivity levels. Highly sensitive data is encrypted using the AES-256 algorithm and stored in a private cloud. A ciphertext indexing mechanism is constructed based on the Paillier homomorphic encryption algorithm, and ciphertext keyword retrieval is achieved by combining Bloom filters and inverted indexes. Medium and low-sensitivity data is generated using erasure coding and Reed-Solomon encoding to generate k data fragments and r redundant fragments, which are then distributed and stored on different physical nodes.

[0151] S7.2: Store metadata in a Neo4j graph database to construct a judicial semantic graph containing case nodes, party nodes, document nodes, and their relationships, supporting graph traversal and pattern matching queries;

[0152] S7.3: Design a multi-layered smart contract in the Hyperledger Fabric consortium chain, including an access control subcontract, an audit trigger subcontract, and a version traceability subcontract. Encapsulate data operation rules through a standard ABI interface. The version traceability subcontract records the data version hash chain based on a Merkle tree and reconstructs the Merkle root during each update to verify the change path.

[0153] S7.4: The access control sub-contract verifies the permission attributes of the calling subject, dynamically matches the policy expression based on the off-chain hybrid authorization data, and returns a rejection record and writes it to the asynchronous audit log when authorization fails;

[0154] S7.5: Data storage behavior, version change events, and audit trigger records are written into the blockchain through the PBFT consensus mechanism, and the judicial nodes jointly confirm the evidence.

[0155] In the above scheme, step 8 includes the following sub-steps:

[0156] S8.1: Perform natural language cleaning on judgment documents, including denoising, sentence segmentation, and part-of-speech tagging. This is then fed into a pre-trained BERT model to extract legal entities. Entity types include court, plaintiff, defendant, and case.

[0157] S8.2: Construct a semantic edge structure based on legal logical relationships, define the three core relationships of "litigation-acceptance-being sued", and form a multi-level knowledge network centered on the case node;

[0158] S8.3: Map the extracted entities and relationships to the Neo4j graph database, binding each node to a unique ID and a set of attributes, and defining each edge with a relationship type and occurrence timestamp;

[0159] S8.4: Implement semantic associative query based on graph structure, support traversal of related plaintiffs, defendants, court entities and reference legal information using cases as the entry point, and attach access control tags to query results.

[0160] In the above solution, step S9 specifically includes:

[0161] S9.1: Deploy a lightweight log agent on each business node in the system to capture user operation logs. The log fields should include at least the operation type, resource identifier, timestamp, user digital certificate fingerprint, and device hash fingerprint. Generate a log summary using the SM3 algorithm.

[0162] S9.2: Build an anomaly detection model based on an LSTM neural network. Independently train behavioral baselines by user role (judge, prosecutor, clerk). The input is the operation sequence vector within the time window, and the output is the predicted distribution of operation categories. The degree of behavioral deviation is determined by calculating the KL divergence between the predicted results and the actual operation.

[0163] S9.3: A dynamic 3σ threshold detection mechanism is used to establish a normal distribution model based on the user's historical operation frequency and resource access pattern. When the real-time monitoring detects that the deviation of a single operation exceeds μ±3σ or the KL divergence value of a continuous operation sequence exceeds the preset threshold, an abnormal alarm event is triggered;

[0164] S9.4: Integrate the CVE vulnerability database and the judicial industry threat signature library, conduct risk assessment on alarm events using the DREAD model, and output risk scores and classifications:

[0165] High risk (≥7 points): Immediately isolate the user session, freeze permissions, and initiate legal notification.

[0166] Medium risk (4-6.9 points): Restrict access to sensitive resources and mark audit trails;

[0167] Low risk (<4 points): Increase logging frequency and dynamically adjust behavior model weights;

[0168] S9.5: Encapsulate abnormal event metadata, risk assessment results, and response operation records into blockchain transactions, and write them into the Hyperledger Fabric consortium chain through the PBFT consensus mechanism. The evidence storage nodes include the court audit department, the procuratorate supervision node, and the judicial technology verification node.

[0169] Example 1

[0170] To achieve the above object, the present invention is implemented through the following technical solutions:

[0171] The present invention provides a method for securely transferring confidential data in the judicial system based on blockchain technology. By constructing an integrated collection, desensitization, transmission, classification, authorization, storage and auditing mechanism, it achieves a credible, controllable and traceable transfer of confidential judicial data among multiple departments.

[0172] This invention connects to business systems such as courts, procuratorates, and judicial administration through a unified data exchange platform, aggregating multimodal judicial data, including structured (such as case documents), semi-structured (such as audio and video), and unstructured (such as electronic evidence). To ensure the privacy and security of data upon entry into the system, sensitive fields (such as ID numbers and names) are automatically identified using a large language model (Deepseek-v3) upon collection, and data desensitization is performed using a differential privacy mechanism.

[0173] During data transmission, the system uses the national secret SM4 algorithm and the TLS 1.3 protocol to establish an end-to-end encrypted channel. User identity is verified through a combination of X.509 digital certificates and biometric identification, and key negotiation is performed based on elliptic curve cryptography. Data in transit is encrypted using the SM4-CBC method, combined with the HMAC-SM3 algorithm for integrity verification. The system also generates a SHA-256 hash fingerprint for each data object and writes it to the Hyperledger Fabric consortium chain. Cross-chain evidence is stored by multiple judicial nodes using the PBFT consensus algorithm, ensuring the trustworthiness and tamper-proof nature of data flow across all links.

[0174] In one embodiment of the present invention, the system constructs an end-to-end encrypted transmission path based on the national secret SM4 algorithm and the TLS1.3 protocol during the transmission of confidential judicial data across departments and networks. Taking into account the heterogeneous structures of the judicial internal and external networks, the poor compatibility of the national secret algorithm with the TLS standard protocol, and other issues, the system designs a dual-path switchable transmission mechanism that supports dynamic switching of SM4-CBC and SM4-GCM modes during the TLS handshake phase, and avoids replay attacks through the Nonce random generator (based on timestamp and session ID) to ensure the stability of the encryption path. In terms of key management, the client and the server complete key negotiation based on the ECC algorithm (such as the SM2 curve), and adopt a session status monitoring mechanism to automatically trigger the key update process when the session survival exceeds a threshold (such as 5 minutes) to ensure the long-term confidentiality of the transmitted data. Identity authentication adopts an improved two-factor strategy. While using the X.509 national secret certificate to complete basic identity authentication during the TLS handshake phase, the biometric fingerprint SM3 hash check is superimposed to ensure that the three are bound together, preventing certificate sharing and counterfeiting risks.

[0175] In addition, in terms of data integrity and traceability, the present invention uses the Hyperledger Fabric blockchain system to perform structured evidence storage for each transmitted data. Each data object (including audio and video, files, documents, etc.) generates a unique SHA-256 hash fingerprint after desensitization and packaging, and constructs a transaction structure, embeds metadata such as operator ID, timestamp, access IP, operation type, etc., and writes to the blockchain through a dedicated Chaincode call interface. The blockchain nodes are composed of consensus organizations such as courts, procuratorates, and judicial bureaus, and use the PBFT consensus mechanism to achieve 2 / 3 node endorsement and complete transaction confirmation. After successful chaining, the system writes the corresponding block number and transaction hash back to the original data system as a "trusted label" embedded in the database meta field, realizing on-chain and off-chain linkage and traceability.

[0176] To improve post-data processing efficiency, the system has built a multimodal content recognition and classification engine. Pre-trained language models are used to perform semantic understanding and label extraction for structured and textual data. Keyframe extraction and the VGG16 network are used to extract visual features from video data. Audio is transcribed and incorporated into a unified text processing pipeline, ultimately achieving classification output using a decision tree algorithm. This enables the archiving, labeling, and rapid retrieval of confidential data.

[0177] In one embodiment of the present invention, in order to improve the archiving efficiency and intelligent retrieval capabilities of judicial confidential data after the circulation process, the system constructs a content recognition and classification engine for multimodal data. The engine designs a multi-channel asynchronous processing architecture, establishes pre-processing and feature extraction pipelines for different modal types of data (structured, text, audio, video, etc.), and integrates them in a unified intermediate representation layer. Structured data (such as case meta-information and litigation stage information) is automatically normalized through JSON Schema mapping and regular parsing; text data is input into the semantic parsing module based on the DeepSeek-v3 large language model to complete entity recognition, case type judgment and keyword label extraction, and output structured semantic vectors. The audio data is transcribed by the iFlytek ASR model and goes through the same text processing process to ensure semantic consistency.

[0178] The video data processing flow uses an inter-frame difference algorithm to extract key frames to avoid redundant frames interfering with subsequent analysis. After the extracted key frames are enhanced and normalized by OpenCV, they are input into the VGG16 network for convolutional feature extraction to generate a 4096-dimensional visual vector representation. All modal vectors are uniformly fed into a classifier based on the Gated Decision Tree architecture, which automatically adjusts the weights of different modalities during the training phase to cope with incomplete data or missing modalities. The classification output includes multi-dimensional labels such as case type, confidentiality level, archiving path, and label set, which are mapped into the storage system through a chain indexer to achieve automatic archiving and rapid semantic retrieval of confidential data. This multimodal recognition solution effectively solves key problems in judicial practice such as low efficiency in unstructured data processing, difficulty in modal fusion, and unstable label extraction, providing a semantic basis for subsequent authority control and risk monitoring.

[0179] In terms of permission control, the system integrates the XACML policy framework with Kafka message channels to achieve real-time response and authentication to access requests. Based on a hybrid RBAC and ABAC model, the system supports dynamic authorization mechanisms based on contextual information such as time, location, and device. Furthermore, it combines timestamp + geolocation two-factor authentication and user digital certificates to achieve fine-grained control over data access behavior.

[0180] In a specific embodiment of the present invention, in order to solve the core technical problems of permission drift, policy execution delay, user identity forgery and blurred permission boundaries in cross-network and cross-level system access to judicial data, a distributed dynamic permission control mechanism that integrates the XACML policy engine and the Kafka real-time message channel is proposed. In the access control architecture, the system embeds the policy decision point (PDP) module into the server kernel. All user resource access requests are converted into the standardized XACML Request XML format by the front-end proxy component before entering the network and delivered as event messages to the Kafka message queue. The Kafka Producer is responsible for pushing them to the PDP instance cluster divided by business dimensions, realizing asynchronous and concurrent policy judgment task distribution. To solve the decision-making performance bottleneck under large-scale access, the system constructs a multi-channel Kafka Topic to automatically classify and route access requests according to business type (such as case file viewing, audio and video downloading, file exporting, etc.). Different types of requests are processed concurrently by the corresponding PDP instances, supporting policy caching and asynchronous pre-parsing of policy templates, thereby improving the policy judgment speed. In terms of permission model design, the system integrates role-based access control (RBAC) and attribute-based access control (ABAC). The role definition comes from the unified authentication platform of the judicial authorities, and the attribute dimensions include precise time window (UTC time synchronized by NTP), access geographic location (based on GPS / Wi-Fi or public IP reverse lookup), device fingerprint (generated through browser identification and system registry snapshot), network type (internal / external network identification) and target data sensitivity level (marked by data tags).

[0181] The present invention introduces a two-factor context verification process in the authentication mechanism: before the client initiates access, the local security module automatically embeds the current UTC timestamp and geographic coordinates, and generates a signature digest through HMAC-SM3 calculation and attaches it to the access request. After receiving the request, the server matches and verifies the timestamp and location parameters with the allowed time period defined in the XACML policy template and the judicial office area IP / GPS whitelist to ensure the spatial and temporal legitimacy of the request behavior. If the time or location factor verification fails, access is denied, and an abnormal behavior event is automatically generated in the Kafka security audit channel for subsequent risk analysis. During the session stage, users need to perform TLS1.3 channel identity authentication through the national secret standard X.509 digital certificate. Combined with the bound institutional role information, the permission control module dynamically generates an access token (Access Token). The token only carries the policy path hash (Policy ID + matching chain summary) and the request context hash fingerprint to avoid directly exposing permission details and effectively prevent session hijacking and man-in-the-middle attacks. Ultimately, all access decision results (Allow / Deny) and their corresponding policy paths, attribute matching conditions, and user context hash summaries will be structured and packaged, written into the consortium chain through the smart contract interface of Hyperledger Fabric, and jointly confirmed and stored by multiple nodes such as courts and procuratorates using the PBFT consensus mechanism to build a complete access control audit chain, ensuring that decision-making behaviors are verifiable, traceable, and cannot be tampered with.

[0182] For data storage, the system employs a hierarchical encryption strategy. Highly sensitive data (such as files) is encrypted with AES-256 before being stored in a private cloud, supporting ciphertext retrieval based on homomorphic encryption. General data is stored in a distributed, redundant manner using erasure coding and Reed-Solomon encoding to enhance disaster recovery capabilities. Metadata is built into the Neo4j graph database to support complex queries and analysis across multi-entity relationships. On-chain business rules are encapsulated via smart contracts using a standard ABI interface, and combined with a Merkle tree structure, data change paths can be verified and traced.

[0183] In one embodiment of the present invention, in response to the challenges of different data sensitivity levels, differentiated access requirements, and strict compliance and evidence requirements in judicial scenarios, a set of hierarchical encryption storage strategies based on data sensitivity is constructed. Specifically, highly sensitive data (such as electronic files and appraisal reports) are symmetrically encrypted using the AES-256 algorithm before being stored in the warehouse to ensure the confidentiality and compliance of static data storage. In order to meet the queryability requirements under the data encryption state, a ciphertext indexing mechanism based on the Paillier homomorphic encryption algorithm is introduced, and the ciphertext keyword mapping is designed in combination with the Bloom Filter and the inverted table structure. Without decrypting the original data, it supports fast matching retrieval based on keywords, solving the technical bottleneck of "judicial highly sensitive encrypted data cannot be quickly located". For unstructured data of medium and low sensitivity levels, the system introduces a redundant distributed storage mechanism that combines erasure codes and Reed-Solomon encoding. The original data is divided into k data fragments and r redundant fragments, which are distributed and stored in different physical nodes, improving the disaster recovery capability in the event of node failure or attack. System metadata (such as case numbers, associations, document status, etc.) are modeled in the NeO4j graph database. A complex judicial semantic graph is constructed through nodes (cases, parties, documents) and edges (ownership, references, trials), supporting graph traversal and pattern matching, and achieving efficient query of case-related elements and semantic-level upstream and downstream reasoning.

[0184] In order to ensure the verifiability of the judicial data status on the chain and the traceability of the version change path, the present invention designs a multi-level smart contract structure based on Hyperledger Fabric. The contract structure includes three core modules: access control sub-contract (AccessPolicyContract), audit trigger sub-contract (AuditTriggerContract) and version tracing sub-contract (VersionTraceContract), which encapsulates business rules and data operation interfaces through standard ABI interfaces, and supports asynchronous calls in the off-chain system through RESTful API. In specific implementation, before writing data to the chain, the system first verifies the permission attributes (such as department level, identity authentication status, access time interval, etc.) of the current calling subject through the AccessPolicyContract sub-contract. The contract combines the ABAC and RBAC hybrid authorization data provided by the off-chain, dynamically matches the policy expression, and if the authorization is successful, enters the data writing process; otherwise, the contract terminates, returns a rejection record, and writes the unauthorized attempt event into the asynchronous audit log. After data is successfully written, the VersionTraceContract subcontract constructs a Merkle tree node for the current data version, recording metadata such as the original hash, timestamp, and operator identity in key-value format. Simultaneously, with each data update, the contract automatically reconstructs the Merkle root and links the change path to the previous and next versions in a hash chain format, writing it to the on-chain Key-Value State Database (StateDB). This ensures that any version backtracking can be verified for authenticity and consistency based on the hash path. When specific security event trigger conditions are met (such as frequent access to certain data types, cross-domain logins, or abnormal data content deviations), the AuditTriggerContract subcontract automatically generates an audit task and invokes the Fabric event mechanism (Chaincode Event) to push it to the judicial supervision node, notifying the audit module for in-depth verification and off-chain emergency response. The entire contract system operates under the PBFT consensus mechanism. Every data authorization, update, and audit trigger execution is signed and confirmed by at least two-thirds of the member nodes, effectively preventing single points of malicious activity and improving the credibility and technical compliance of on-chain governance of confidential data.

[0185] At the data semantic mining layer, the system uses the BERT model to extract legal entities and relationships from judicial documents, construct a judicial knowledge graph centered on cases, and achieve structured expression of four types of core entities and three types of semantic relationships, which are stored in a graph database to support subsequent semantic-level queries and data sharing permission control.

[0186] To ensure system security and compliance, a blockchain audit system is integrated into the operations and maintenance phase to log all operations, including operation type, execution time, and digital signature subject information. An LSTM neural network is used to sequentially model behavioral patterns, combined with a 3σ standard deviation dynamic threshold detection mechanism to automatically trigger risk response processes when abnormal behavior is detected. The risk assessment system integrates the Common Vulnerability Database (CVE) and industry threat signature databases, and uses the DREAD model to assess risk levels. It automatically implements tiered emergency response actions such as data isolation and permission freezing, forming a closed-loop risk control system.

[0187] In the specific implementation of the present invention, the system proposes a semantic extraction and knowledge graph construction method based on the BERT model in view of the characteristics of judicial judgment documents such as complex semantic structure and non-standard expression of legal entities and their relationships. The method first uses the Chinese pre-trained BERT model to contextually encode the document text, and combines the named entity recognition (NER) technology to extract the four core legal entities of "plaintiff", "defendant", "case" and "court". On this basis, the system uses multiple classifiers to model the context between entity pairs, extracts the three types of event relationships of "litigation", "acceptance" and "being sued", and thus constructs a judicial knowledge graph structure centered on the case. The graph not only includes entities and their semantic relationships, but also further associates additional attribute information, such as the plaintiff and defendant's subject qualifications, litigation capacity, case occurrence time and trial results and other semantic labels, so that the knowledge graph has stronger semantic expression and reasoning capabilities. All structured triples and attribute information are finally organized in a node-side model and stored in the Neo4j graph database, supporting complex semantic queries and graph traversal operations. While ensuring the accuracy of entity extraction, this solution systematically solves the problem of difficult structural expression of judicial document semantics, and realizes efficient modeling, storage and query of judicial semantic information through customized legal graph structure design.

[0188] In one embodiment of the present invention, to address technical challenges such as log falsification, difficulty in tracing behavior, and delayed risk response in traditional judicial data operations and maintenance, the system integrates a behavioral audit and risk control mechanism based on a joint drive of blockchain and deep learning. First, the system deploys a lightweight log capture agent at each business node to record user behavior operation logs, including fields such as the request source IP, user ID, and operation type, and uses SM2 signatures to bind the identity of the operating subject. After the collected logs are uniformly structured, they are submitted to the Hyperledger Fabric blockchain for storage.

[0189] To identify potential abnormal behaviors, the system builds a behavior modeling engine based on LSTM neural network.

[0190] In one embodiment of the present invention, to identify potential anomalous behavior, the system builds a behavioral modeling engine based on an LSTM neural network. This engine takes on-chain log sequences as input. Log fields include user ID, timestamp, operation type, resource type, access IP address, and other information. Through a preprocessing process, the system converts the log sequences into time series vectors representing user operation categories (e.g., operations such as login, file query, video export, and evidence download are represented using a one-hot encoding method, with time information embedded using time difference or period position information).

[0191] During the training phase, a training set of samples is constructed based on the user dimension. An independent LSTM model is trained for each account type (e.g., judge, clerk, prosecutor, etc.) to capture the temporal dependencies and operational patterns of their operations. The model takes a fixed-window-length sequence of operations as input and uses the difference between the predicted distribution of the next possible operation and the actual operation as the discrimination criterion. The training objective utilizes a combination of supervised and semi-supervised strategies: supervised training is performed using a cross-entropy loss on data labeled as normal behavior. For unlabeled data, prediction bias (e.g., reconstruction error) is used to control the model update rate to avoid introducing anomalous patterns. Ultimately, the system learns a set of behavioral pattern vectors for each account type and calculates the prediction bias using a sliding window. When abnormal behavior (e.g., high-frequency calls, unauthorized access) deviates from a threshold, an alert is immediately issued to the risk control system. Simultaneously, the system runs a hybrid risk assessment engine in the background, integrating real-time data from the CVE vulnerability database and the legal industry threat signature database. Once an alert is triggered, the engine automatically calculates a risk score using the DREAD model based on parameters such as the exposure of the operational assets, the difficulty of the attack, and the scope of impact. Emergency response strategies are then implemented based on the score classification:

[0192] High risk (score ≥ 7): Immediately isolate the user session involved, freeze their access to relevant resources, issue a patch to the affected nodes, and initiate a judicial notification mechanism for manual intervention and tracing.

[0193] Medium risk (score 4-6.9): Temporarily freeze access to some highly sensitive resources, mark the user as an audit target, record behavior logs, notify system administrators for verification, and perform patch pre-push if necessary;

[0194] Low risk (score <4): Record the event log, increase the weight of the behavior model and strengthen monitoring. If the subsequent behavior deviation continues to intensify, the risk level will be automatically increased;

[0195] This mechanism realizes a closed-loop response path from anomaly detection, level assessment to disposal feedback in the judicial data system, significantly improving the security self-healing capability and attack response efficiency of confidential data systems.

[0196] Example 2

[0197] All business platforms within the judicial system (including the court trial system, the procuratorate business system, and the judicial administration management system) are connected through a unified data exchange platform, enabling centralized data collection. This platform supports multiple data connection methods, including standardized API calls, message middleware push (such as Kafka), and scheduled data pull, to accommodate the connection capabilities and data update frequencies of different systems. The system can simultaneously collect the following three types of multimodal data:

[0198] 1. Structured data: such as case number, party information, judgment results, timestamps, etc., which comes from the court trial database and the procuratorate's business system;

[0199] 2. Semi-structured data: such as audio and video files, spreadsheets, and electronic transcripts of court proceedings;

[0200] 3. Unstructured data: including scanned evidence materials, photos, judgment documents in PDF format, and evidence recordings in audio files, etc.

[0201] After the collection is completed, the system automatically unifies the format and pre-parses the above data. Structured data is directly parsed into a set of standard fields; audio and video data are transcribed through the ASR service integrated with the iFlytek open platform to convert voice content into text; image and video files use OpenCV tools to extract key frames for subsequent visual analysis. In the data preprocessing stage, the large language model Deepseek-v3 deployed on the local server is used to perform entity recognition and context analysis on the data content. The model encodes the input text data and uses the named entity recognition (NER) algorithm to identify the following sensitive fields: the name of the party, ID number, telephone number, address, unit name, bank account, etc. The identified sensitive fields then enter the differential privacy processing module. The system adopts different processing methods for different types of data based on the preset sensitive field level and desensitization strategy:

[0202] For strongly identifying fields like ID numbers and phone numbers, we use irreversible desensitization methods like random perturbation and hash replacement. For fields that partially retain semantic meaning (such as court names and work units), we use recoverable desensitization methods like character replacement and token masking. The processed fields are automatically assigned a differential privacy budget value and recorded in the metadata tag for subsequent reference by the permission control module.

[0203] The data that has been desensitized needs to be transmitted across departments between various judicial systems. To ensure that the data is not stolen, tampered with or forged during transmission, the system has designed multiple security mechanisms to build end-to-end encrypted communication channels and achieve verifiable transmission guarantees throughout the entire process.

[0204] First, on the data transmission link, the system uses the SM4 symmetric encryption algorithm based on the national secret algorithm, combined with the TLS1.3 transport layer security protocol to build an encrypted channel. Before the two parties communicate, a secure connection is established through the TLS handshake process. The following key security mechanisms are introduced during the handshake phase:

[0205] 1. Two-way identity authentication: Using X.509 standard two-way digital certificates to achieve legal identity verification of both communicating parties (such as the court data gateway and the procuratorate service node);

[0206] 2. Biometric-assisted authentication: Introducing facial recognition or fingerprint authentication mechanisms at specific high-authority nodes (judicial data control centers) and binding them to digital certificates to prevent key theft;

[0207] 3. Key negotiation mechanism: Use the elliptic curve cryptography (ECC) algorithm to complete key exchange and ensure the establishment of symmetric session keys under asymmetric conditions.

[0208] After the data is encrypted at the session layer, it enters the data content layer encryption stage. The data body is encrypted in blocks using the SM4-CBC mode. The key is the symmetric key generated in the TLS handshake phase. The encrypted data is appended with an integrity check value calculated by the HMAC-SM3 algorithm for consistency verification before decryption at the receiving end to prevent man-in-the-middle attacks and data tampering.

[0209] After each data file is encrypted, the system will calculate a unique SHA-256 hash summary value for the data content as the "digital fingerprint" of the data file, which will be used for full-process consistency verification and blockchain registration.

[0210] To ensure traceability and non-repudiation of data transmission, the system packages metadata, including hash fingerprints, sender and receiver node IDs, and timestamps, into transaction records for each data transfer and uploads them to the Hyperledger Fabric consortium blockchain system. The Fabric network incorporates blockchain nodes from multiple judicial departments (courts, procuratorates, and judicial bureaus), and utilizes a PBFT fault-tolerant consensus mechanism for multi-node voting confirmation. This ensures that each data upload is authenticated by consensus among the three governing bodies, preventing single-point tampering or record falsification. Each data transaction during the transmission process generates a unique transaction ID and is associated with the transmission task record. The system can use Fabric's smart contracts to verify whether the transaction was successfully uploaded, whether the node signatures are consistent, and whether the original file has been modified.

[0211] The present invention integrates an intelligent classification engine based on multimodal deep learning to perform content recognition and automatic classification of structured, semi-structured and unstructured data. The classification engine performs pre-processing based on the data source:

[0212] For text data (such as case documents, court transcripts, etc.), the system performs semantic analysis and feature extraction on the text based on Deepseek-v3, and completes document topic classification and key tag annotation.

[0213] For video data (such as court trial videos), OpenCV is used to sample video frames and extract key frame images; then the VGG16 convolutional neural network is used to extract image features, and finally the visual features are input into the decision tree classifier for semantic label assignment.

[0214] For audio data (such as witness recordings), the system first completes speech-to-text conversion through the iFlytek open platform, and then completes subsequent semantic analysis and classification according to the text processing path.

[0215] All classification labels are individually bound to the original data. This information serves as the foundation for subsequent modules like permission management, search optimization, and data visualization and analysis, achieving semantic unification of data processing and management. SHA-256 hash summaries are generated for both the classification results and the label information and stored on-chain. This process ensures auditable proof of intelligent classification, facilitates subsequent retrospective processing, and meets the interpretability and compliance requirements of the judicial system.

[0216] In order to achieve fine-grained permission management and dynamic response control in the process of accessing confidential data, the system integrates an XACML-based policy engine and combines the hybrid permission model of RBAC (role-based) and ABAC (attribute-based) to adapt to complex judicial business needs.

[0217] When the system receives an access request, it first forwards the request information (including visitor identity, access time, requested resources, device information, etc.) to the permission policy engine in real time through the Kafka message queue. The permission judgment process includes the following steps:

[0218] User identity and role binding: Visitors complete identity authentication through X.509 digital certificates, and the system automatically resolves their roles (such as judge, prosecutor, clerk, lawyer, etc.);

[0219] Attribute-aware analysis: The system collects and analyzes contextual attributes in access requests, such as access time period, geographic location (via GPS module), access device type, current network environment, etc.

[0220] Policy evaluation and decision-making: The XACML policy engine matches the preset access control policy, combines RBAC role permission boundaries with ABAC attribute matching logic for joint evaluation, and ultimately arrives at a control result such as "allow," "deny," or "further authentication required."

[0221] Dynamic two-factor authentication: For situations where there is uncertainty in sensitive data access or policy judgment, the system forcibly introduces a timestamp + geolocation two-factor authentication mechanism, which requires users to complete identity confirmation through a bound legal terminal within the specified time and authorized location, thereby improving the credibility of access behavior.

[0222] The authorization results are recorded with digital signatures and synchronously written into the blockchain audit module as valid evidence for subsequent behavior tracing and compliance audits.

[0223] To meet the needs of secure storage and efficient access to confidential data in the judicial system, this system adopts a hierarchical encryption strategy and a smart contract-driven data lifecycle management mechanism. Specifically, it includes the following mechanisms:

[0224] 1. Hierarchical encryption storage mechanism

[0225] Data is divided into sensitive data (such as case files and investigative materials) and general data (such as public judgment documents and meeting minutes) according to its sensitivity:

[0226] Sensitive data is symmetric-encrypted with AES-256 before being uploaded to a private cloud platform deployed in a dedicated network environment. At the same time, homomorphic encryption (such as BFV or CKKS) is introduced to enable some keyword searches and comparisons to be performed in ciphertext, preventing the exposure of plaintext.

[0227] General data is stored in distributed redundant slices using erasure codes and Reed-Solomon codes, distributed across different physical nodes to improve disaster recovery and anti-destruction capabilities.

[0228] The metadata and permission indexes are built in the Neo4j graph database. Nodes and edges represent data entities and their logical associations, respectively, supporting complex semantic retrieval and access control auxiliary judgment.

[0229] 2. Blockchain on-chain mechanism

[0230] Each storage behavior of confidential data generates a unique SHA-256 hash summary and is synchronized to the Hyperledger Fabric consortium chain:

[0231] Blockchain nodes are composed of courts, procuratorates, and judicial administrative departments, using the PBFT consensus algorithm to ensure evidence consistency. Every data write, read, modify, and delete is accompanied by a metadata change event, the summary of which is recorded as an on-chain transaction. The Merkle tree mechanism is used within the block structure to calculate the root hash, ensuring that any tampering with any data record in the block can be quickly detected.

[0232] 3. Smart Contract Control Mechanism

[0233] The system pre-deploys a series of smart contracts developed based on Fabric Chaincode, and defines execution rules according to the standardized ABI interface. The contract functions include:

[0234] Definition of data access timeliness rules: You can set settings such as "read-only within 3 years after the prison term, and destroyed after 3 years";

[0235] Access behavior triggering rules: You can set "access to specific level of data requires joint authorization from two levels of institutions";

[0236] Automatic destruction mechanism: The contract calls an external time oracle to determine the validity period of the data. After the expiration, the data erasure process is executed and the destruction behavior is uploaded to the chain for future reference.

[0237] All contract operations are legally guaranteed through digital signatures and on-chain verification mechanisms, which record the operating entity and timestamp to form a chain of responsibility for judicial compliance.

[0238] To improve the semantic organization and interpretability of judicial data, this system constructs a case-centered legal knowledge graph to achieve a structured representation of judicial documents and their key information. This technical solution mainly includes four steps: entity extraction, relationship identification, graph construction, and graph database storage:

[0239] The system collects text data such as judgment documents, prosecution opinions, and trial transcripts from the court's adjudication system, which are first processed through natural language cleaning (noise removal, sentence segmentation, and part-of-speech tagging), and then input into the pre-trained BERT model for contextual semantic modeling.

[0240] Entity types include but are not limited to:

[0241] 1. Court entities: name, region, jurisdiction, judges;

[0242] 2. Subject entities: plaintiff, defendant, and their identity information, educational level, litigation capacity, subject qualifications, and other supplementary attributes;

[0243] 3. Case entities: time of incident, location of incident, reference laws, trial results, case closing time and other elements.

[0244] Based on the semantic and legal logical relationship between entities, the system constructs a graph edge structure and defines semantic paths such as "plaintiff → court", "court → acceptance → case", "case → reference → law", "case → participation → defendant", etc., which fully reflect the subject-object relationship in the judicial process.

[0245] The system maps the above entities and relationships into nodes and edges in the graph, forming a multi-level knowledge network with "case" as the core node. The overall knowledge graph structure consists of the following four types of nodes and three types of edges:

[0246] Entity type (node): case, plaintiff, defendant, court

[0247] Relationship type (edge): litigation, acceptance, and being sued

[0248] The graph uses the Neo4j graph database as its storage medium, leveraging its native graph structure to support efficient entity queries and association reasoning. Each node is associated with a unique ID and its attribute set, while each edge defines the specific relationship type and occurrence time. The graph structure is continuously extensible, supporting the future addition of legal entities and associations in more dimensions, such as evidence chains, litigation stages, and appeals.

[0249] To ensure the visualization, controllability and auditability of system operation, this system builds an operation and maintenance monitoring mechanism based on blockchain audit + anomaly detection + risk assessment + automatic response, covering operation log chain, security audit analysis and multi-dimensional risk management.

[0250] 1. On-chain audit of operational behavior

[0251] All operation and maintenance operations and data access behaviors, such as uploads, downloads, queries, permission changes, and authentication, are recorded and fully logged through the unified audit agent:

[0252] Log content includes the operation type, timestamp, user identity (digital signature), resource identifier, access IP address, and device information. A SHA-256 hash is calculated for each log entry, chronologically encapsulated into blocks, and uploaded to Hyperledger Fabric. The PBFT consensus mechanism ensures that log records are non-repudiable and tamper-proof, facilitating subsequent compliance audits and accountability.

[0253] 2. Abnormal behavior identification and intelligent alerting

[0254] LSTM (Long Short-Term Memory) neural networks are introduced to model continuous behavior log data and identify potential high-risk behavior patterns:

[0255] We construct multi-dimensional time series features (such as access frequency, data call type, and operation path), train a baseline model to identify "normal behavior flows," and conduct real-time comparisons during the deployment phase. If a user's behavior deviates (for example, daily access frequency exceeds the long-term average + 3σ threshold), an anomaly alarm is triggered. Alarm events are recorded on the chain for traceability.

[0256] 3. Risk assessment and automatic response mechanism

[0257] The system integrates the CVE vulnerability library with a custom risk signature library for the judicial industry to perform dynamic threat modeling at the network, identity, and application layers. Risk modeling utilizes the DREAD model, assigning a score to each risk and dynamically updating it. Once a high-risk vulnerability or abnormal behavior is discovered that closely matches an attack signature, the system automatically calculates the current system risk value. Based on the risk level (high, medium, or low), a graded emergency response mechanism is initiated, including isolating the user session involved, freezing access to specific privileged resources, issuing patch repair commands, and initiating a judicial notification mechanism.

[0258] All risk emergency operations are also confirmed with digital signatures and written into the blockchain audit system.

[0259] The embodiment of the present invention is a solution to the problem of secure transmission and sharing of business data based on the interconnection and intercommunication of data among the three departments of law, procuratorate and justice, which enables the secure sharing and exchange of judicial data across levels, departments and networks. In the process of file transmission and sharing in the system, the trust verification between the transmission objects and the authorization authentication mechanism for the objects sending the request are strengthened. It is necessary to ensure that the establishment of channels for data transmission and file sharing requires strict authentication and authorization, establish different levels of security levels as needed, improve the reliability of system terminals, and solve the security problems of judicial data sharing and exchange across levels, departments and networks in the law, procuratorate and justice systems. It will enable legal data to be safely and reliably circulated and exchanged among law, procuratorate, justice and other departments according to the needs of case handling, and support the daily sharing of necessary legal data involved in the case, such as case trial information exchange, electronic official document exchange, and prisoner commutation and parole information, in the legal system based on blockchain technology.

Claims

1. A method for securely transferring confidential data in the judicial system based on blockchain technology, characterized in that: The following steps are involved: Step 1: Collect multi-source judicial data, including structured, semi-structured, and unstructured data; Step 2: Use a large language model to identify sensitive information in the data and perform desensitization based on a differential privacy mechanism; Step 3: Use the SM4 encryption algorithm and TLS 1.3 protocol to build an end-to-end encrypted channel, and combine X.509 certificates, biometric authentication, and ECC key negotiation to complete data transmission; Step 4: Generate a SHA-256 hash fingerprint for the encrypted data, write the hash fingerprint and transmission metadata into the blockchain system, and implement cross-node evidence storage through the PBFT consensus mechanism; Step 5: Classify the data using a multimodal classification engine, including text semantic analysis, video keyframe feature extraction, and speech-to-text processing; Step 6: Adopt a hybrid permission control model, integrating the XACML policy framework with the Kafka message queue to achieve real-time response and authentication to access requests. This model supports dynamic permission management based on time, space, and device attributes through a combined RBAC and ABAC mechanism. Furthermore, it uses geolocation and timestamp-based two-factor authentication to bind user identities to devices. Step 7: Adopt a hierarchical encryption storage strategy. Sensitive data is encrypted and stored in a private cloud with support for homomorphic encryption retrieval. General data is distributed and redundantly stored using erasure codes, and metadata is built in a graph database. Step 8: Construct a judicial knowledge graph. Based on the BERT model, extract the legal entities and semantic relationships in the judgment documents, forming a graph network containing four types of entities and three types of relationships to support semantic associative query and decision-making assistance. Step 9: During the system operation phase, the blockchain audit system records operation logs, combines LSTM neural networks with dynamic threshold detection to identify abnormal behaviors, and uses the DREAD model to conduct risk assessment and graded emergency response.

2. The method according to claim 1, characterized in that Step 1 includes the following sub-steps: S1.1: Connects to the court trial system, procuratorate business system, and judicial administration management system, supporting three data connection methods: standardized API call, message middleware push, and scheduled data pull; S1.2: Synchronously collect structured data, semi-structured data, and unstructured data. Structured data includes fields such as case number, party information, and judgment outcome. Semi-structured data includes courtroom audio and video files and electronic transcripts. Unstructured data includes scanned evidence materials, PDF-formatted judgment documents, and audio evidence files. S1.3: Collected data is formatted and pre-parsed. Structured data is directly parsed into standard field sets. Audio and video data is converted into text data through a speech transcription service. Images and video files are visually analyzed using keyframe extraction tools.

3. The method according to claim 1, characterized in that Step 2 includes the following sub-steps: S2.1: Use a pre-trained language model to perform named entity recognition (NER) on the collected data to identify sensitive fields such as the party’s name, ID number, phone number, address, and bank account number. S2.2: Implement differential privacy desensitization based on preset sensitive field levels. For strongly identifying fields, use irreversible desensitization using random perturbation and hash replacement. For fields that need to retain semantic meaning, use recoverable desensitization using character replacement and token masking. S2.3: Attach a differential privacy budget value label to the desensitized field and record the desensitization policy in the metadata for access decisions in subsequent permission control modules.

4. The end-to-end encrypted transmission step in the method for secure transfer of confidential data according to claim 1 is characterized in that: Step 3 includes the following sub-steps: S3.1: Establish an end-to-end encrypted channel based on the TLS 1.3 protocol, authenticate both communicating parties using X.509 two-way digital certificates, introduce biometric recognition mechanisms in high-authority nodes, and use elliptic curve cryptography (ECC) to negotiate keys and generate symmetric session keys. S3.2: Use SM4-CBC mode to encrypt the data body in blocks, and use the HMAC-SM3 algorithm to calculate the integrity check value of the encrypted data. The receiving end verifies the consistency of the check value before decryption.

5. The blockchain evidence storage step in the method for secure transfer of confidential data according to claim 1 is characterized in that: Step 4 includes the following sub-steps: S4.1: Calculate and generate a unique SHA-256 hash digest of the encrypted data content as a digital fingerprint for full-process consistency verification and blockchain registration; S4.2: Packing the hash summary, sender node ID, receiver node ID, and timestamp into a transaction record and associating it with the data transmission task metadata; S4.3: Write the transaction record to the Hyperledger Fabric consortium chain composed of nodes from the court, procuratorate, and judicial bureau. Use the PBFT consensus mechanism for multi-judicial node voting confirmation to complete cross-chain evidence storage. S4.4: Based on the smart contract, the transaction record’s on-chain status, node signature consistency, and original file integrity are verified on-chain, and a unique transaction ID is generated and bound to the transfer task log.

6. The blockchain evidence storage step in the method for secure transfer of confidential data according to claim 1 is characterized in that: Step 5 includes the following sub-steps: S5.1: Build a multi-channel asynchronous processing architecture and establish preprocessing pipelines for structured data, text data, audio data, and video data respectively; S5.2: Normalize structured data through JSON Schema mapping and regular expression parsing to generate a standard field set; S5.3: Input the text data into the semantic parsing module based on DeepSeek-v3, and generate structured semantic vectors through entity recognition and keyword label extraction; S5.4: After the audio data is transcribed into text using the iFlytek ASR model, the semantic parsing process of step S5.3 is reused to generate a semantic vector. S5.5: Use the inter-frame difference algorithm to extract key frames from the video data, extract visual features through the VGG16 network, and generate a 4096-dimensional visual vector; S5.6: Input the multimodal vector into the classifier based on the gated decision tree, and output the case type, confidentiality level, and filing path label.

7. The method according to claim 1, characterized in that Step 6 includes the following sub-steps: S6.1: A policy decision point (PDP) module is embedded in the server kernel. The front-end proxy component converts user access requests into the standardized XACML Request XML format and delivers them as event messages to the Kafka message queue. The KafkaProducer pushes these messages to the PDP instance clusters divided by business dimensions, enabling asynchronous and concurrent policy decision task distribution. S6.2: Build a multi-channel Kafka topic to automatically classify and route access requests to the corresponding PDP instance based on their service type, including case file viewing, audio and video downloading, and file exporting. S6.3: A hybrid model of role-based access control (RBAC) and attribute-based access control (ABAC) is adopted, where role definitions are derived from the unified authentication platform of judicial authorities, and attribute dimensions include time windows, access geographic location, device fingerprint, network type, and target data sensitivity level; S6.4: Before the client initiates access, the local security module embeds the current UTC timestamp t current Generate signature summary D=H through HMAC-SM3 algorithm with geographic coordinates (lat, lng) SM3 (K secret ||t current ||lat||lng) is appended to the access request, and the server verifies whether the timestamp and location parameters match the allowed time period [t start , t end ] and the judicial office area IP / GPS whitelist, including H SM3 SM3 cryptographic hash function, K secret Indicates a pre-shared symmetric key, and the || sign indicates data concatenation; S6.5: Complete TLS 1.3 channel identity authentication using the national cryptographic standard X.509 digital certificate, and dynamically generate an access token based on the organization role information. The token carries the policy path hash and request context hash fingerprint. S6.6: The access decision results and their corresponding policy paths, attribute matching conditions, and user context hash summaries are packaged and structured, written into the consortium chain through the Hyperledger Fabric smart contract interface, and jointly confirmed and stored by multiple judicial nodes using the PBFT consensus mechanism.

8. The method according to claim 1, characterized in that Step 7 includes the following sub-steps: The following sub-steps are included: S7.1: Data storage types are divided according to their sensitivity levels. Highly sensitive data is encrypted using the AES-256 algorithm and stored in a private cloud. A ciphertext indexing mechanism is constructed based on the Paillier homomorphic encryption algorithm, and ciphertext keyword retrieval is achieved by combining Bloom filters and inverted indexes. Medium and low-sensitivity data is generated using erasure coding and Reed-Solomon encoding to generate k data fragments and r redundant fragments, which are then distributed and stored on different physical nodes. S7.2: Store metadata in a Neo4j graph database to construct a judicial semantic graph containing case nodes, party nodes, document nodes, and their relationships, supporting graph traversal and pattern matching queries; S7.3: Design a multi-layered smart contract in the Hyperledger Fabric consortium chain, including an access control subcontract, an audit trigger subcontract, and a version traceability subcontract. Encapsulate data operation rules through a standard ABI interface. The version traceability subcontract records the data version hash chain based on a Merkle tree and reconstructs the Merkle root during each update to verify the change path. S7.4: The access control sub-contract verifies the permission attributes of the calling subject, dynamically matches the policy expression based on the off-chain hybrid authorization data, and returns a rejection record and writes it to the asynchronous audit log when authorization fails; S7.5: Data storage behavior, version change events, and audit trigger records are written into the blockchain through the PBFT consensus mechanism, and the judicial nodes jointly confirm the evidence.

9. The method according to claim 1, characterized in that Step 8 includes the following sub-steps: S8.1: Perform natural language cleaning on judgment documents, including denoising, sentence segmentation, and part-of-speech tagging. This is then fed into a pre-trained BERT model to extract legal entities. Entity types include court, plaintiff, defendant, and case. S8.2: Construct a semantic edge structure based on legal logical relationships, define the three core relationships of "litigation-acceptance-being sued", and form a multi-level knowledge network centered on the case node; S8.3: Map the extracted entities and relationships to the Neo4j graph database, binding each node to a unique ID and a set of attributes, and defining each edge with a relationship type and occurrence timestamp; S8.4: Implement semantic associative query based on graph structure, support traversal of related plaintiffs, defendants, court entities and reference legal information using cases as the entry point, and attach access control tags to query results.

10. The method according to claim 1, characterized in that Step 10 includes the following sub-steps: S9.1: Deploy a lightweight log agent on each business node in the system to capture user operation logs. The log fields should include at least the operation type, resource identifier, timestamp, user digital certificate fingerprint, and device hash fingerprint. Generate a log summary using the SM3 algorithm. S9.2: Build an anomaly detection model based on an LSTM neural network. Independently train behavioral baselines by user role (judge, prosecutor, clerk). The input is the operation sequence vector within the time window, and the output is the predicted distribution of operation categories. The degree of behavioral deviation is determined by calculating the KL divergence between the predicted results and the actual operation. S9.3: A dynamic 3σ threshold detection mechanism is used to establish a normal distribution model based on the user's historical operation frequency and resource access pattern. When the real-time monitoring detects that the deviation of a single operation exceeds μ±3σ or the KL divergence value of a continuous operation sequence exceeds the preset threshold, an abnormal alarm event is triggered; S9.4: Integrate the CVE vulnerability database and the judicial industry threat signature library, conduct risk assessment on alarm events using the DREAD model, and output risk scores and classifications: High risk (≥7 points): Immediately isolate the user session, freeze permissions, and initiate legal notification. Medium risk (4-6.9 points): Restrict access to sensitive resources and mark audit trails; Low risk (<4 points): Increase logging frequency and dynamically adjust behavior model weights; S9.5: Encapsulate abnormal event metadata, risk assessment results, and response operation records into blockchain transactions, and write them into the Hyperledger Fabric consortium chain through the PBFT consensus mechanism. The evidence storage nodes include the court audit department, the procuratorate supervision node, and the judicial technology verification node.

Citation Information

Patent Citations

  • Protection system and method for judicial protection of juveniles

    CN118013551A

  • Security management and retrieval method based on Internet of Vehicles data

    CN119150349A

  • Education digital identity authentication method of distributed heterogeneous system based on block chain

    CN119538223A

  • Distributed multi-mode data cross-trust-domain data sharing method and system

    CN119696850A

  • System and computer method including a blockchain-mediated agreement engine

    US20210336796A1

Cited By

  • Data storage method for artificial intelligence learning mode

    CN120723945A

  • Data storage method for artificial intelligence learning mode

    CN120723945B

  • Unified authentication and data authority management and control method and system for big data component

    CN120811764A

  • A unified authentication and data authority management method and system for a big data component

    CN120811764B

  • System security and data protection mechanism method

    CN120930171A