AI-driven zero trust data exchange and approval system for laboratory instruments

CN121690729BActive Publication Date: 2026-09-29BEIJING NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511858314.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-09-29
Estimated Expiration
2045-12-10

AI Technical Summary

Technical Problem

[0006]针对现有技术的不足,本发明提供了面向实验室仪器的AI驱动零信任数据交换与审批系统,解决了实验室科研数据在跨网交换过程中,传统安全手段干扰精密仪器运行、静态审批效率低下以及数据真伪难以精准鉴别的难题

Benefits of technology

1.本发明通过构建包含被动式边缘镜像采集装置的感知接入层与零信任执行控制层,实现了对仪器原始数据的非侵入式源头管控。该设置将数据采集与仪器业务运行环境物理隔离,避免了数据在上链前的篡改风险,配合基于AI决策的动态通道管控,在完全不干扰精密科研仪器正常运行的前提下,确保了跨网数据交换的全生命周期安全性与真实性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121690729B_ABST
    Figure CN121690729B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of network security and discloses an AI-driven zero-trust data exchange and approval system for laboratory instruments, which comprises a sensing access layer, an intelligent computing layer, an execution control layer and an operation and maintenance prediction layer. The sensing access layer passively collects instrument original data through edge access devices; the intelligent computing layer generates digital fingerprints by using an AIGC fingerprint identification engine combined with an attention mechanism and dynamically calculates an access strategy by using a reinforcement learning module; the execution control layer implements accurate management and control by using a zero-trust gateway and a self-adaptive sandbox; and the operation and maintenance prediction layer realizes threat collaborative prediction under privacy protection by using federated learning. The application solves the problems of difficult scientific research data source authentication and low static approval efficiency, guarantees the lossless operation of instruments and the safety of the whole life cycle of data, and realizes the adaptive evolution of a safety strategy and the active defense of unknown threats.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cybersecurity technology, specifically to an AI-driven zero-trust data exchange and approval system for laboratory instruments. Background Technology

[0002] Currently, various research laboratories are accelerating their digital transformation, with sophisticated instruments generating massive amounts of high-value experimental data daily. This data, as a core asset for scientific research and innovation, needs to be frequently transferred between relatively isolated instrument networks and external data processing centers. Ensuring the integrity, authenticity, and compliance of data during cross-network exchange has become a crucial step in guaranteeing the credibility of research results.

[0003] To address the aforementioned data exchange needs, existing solutions typically employ a combination of physical isolation and perimeter protection. The usual practice is to deploy a traditional firewall at the lab network egress, allowing researchers to copy files via USB storage or upload them via FTP servers. System administrators pre-configure access control lists based on IP addresses and ports to isolate specific network segments. File content review often relies on local antivirus software installed on the instrument's host computer for scanning and removal, followed by manual application for data export permissions through a form, which is then manually approved by approvers at each level.

[0004] However, this traditional model has revealed many shortcomings in practical applications. Many sophisticated instruments run on outdated operating systems, and installing third-party antivirus or management agent software can easily consume system resources, leading to data acquisition timing jitter or even system crashes. Relying solely on traditional hash algorithms for integrity verification is too mechanical and cannot identify benign operations such as format conversions. Once data has undergone normal processing, the hash value changes, making it difficult to identify genuine data tampering. Static access control rules lack flexibility, often causing approval bottlenecks and slowing down research progress when faced with sudden, high-frequency data exchange demands. Furthermore, the security defenses of each laboratory are isolated, making it impossible to share threat signatures without leaking sensitive data, and making it difficult to cope with customized, unknown attacks targeting the research field.

[0005] Therefore, this invention provides an AI-driven zero-trust data exchange and approval system for laboratory instruments to address the shortcomings of existing technologies. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides an AI-driven zero-trust data exchange and approval system for laboratory instruments. It solves the problems of traditional security measures interfering with the operation of precision instruments, low efficiency of static approval, and difficulty in accurately identifying the authenticity of data during cross-network exchange of laboratory research data.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solution: an AI-driven zero-trust data exchange and approval system for laboratory instruments, comprising a perception access layer, an intelligent computing layer, an execution control layer, and an operation and maintenance prediction layer; The perception access layer is used to connect user terminals and scientific instruments, including edge access devices and interactive software running on user mobile terminals. The edge access devices are used to collect raw output data and operating status data of scientific instruments, and the interactive software is used to provide identity authentication interfaces and initiate data transmission requests. The intelligent computing layer is used to process data and output decision instructions, including an AI data classification engine, an AIGC fingerprint recognition engine, and a reinforcement learning dynamic authorization module. The AI ​​data classification engine is used to parse and classify file content, the AIGC fingerprint recognition engine is used to generate digital fingerprints, and the reinforcement learning dynamic authorization module is used to calculate access control policies. The execution control layer is used to implement security policies, including a zero-trust gateway, an adaptive sandbox, and a blockchain evidence storage node. The zero-trust gateway is used to manage the data transmission channel according to the decision instructions. The adaptive sandbox is used to isolate and analyze suspicious files. The blockchain evidence storage node is used to record the digital fingerprint and access logs. The operation and maintenance prediction layer is used for status monitoring and threat analysis, including a federated learning prediction center and a situational awareness platform. The federated learning prediction center is used to generate a global threat prediction model, and the situational awareness platform is used to identify abnormal behavior patterns.

[0008] By adopting the above technical solutions, the system's layered architecture design evolves traditional perimeter defense into a data-centric zero-trust defense system. The perception and access layer manages data sources through physical access at the edge, preventing data tampering risks before being uploaded to the blockchain. The intelligent computing layer utilizes artificial intelligence algorithms to transform static rules into dynamic policies, enabling real-time adjustments to authorization based on environmental context. The execution control layer, combined with the operation and maintenance prediction layer, achieves proactive defense against unknown threats and comprehensive network situational awareness. This architecture solves the challenges of complex laboratory instruments, diverse data formats, and difficulties in unified supervision, ensuring the security of research data throughout its entire lifecycle while also facilitating efficient research workflows.

[0009] Preferably, the edge access device includes a data mirroring interface module, a local AI computing power module, and an encrypted communication module. The data mirroring interface module is equipped with a physical interface for parallel connection to the data output line of the scientific instrument via a physical cable, and passively mirrors the raw data stream generated by the scientific instrument. The data mirroring interface module integrates a signal conditioning circuit for filtering and shaping electrical signals. The local AI computing power module is connected to the data mirroring interface module and is used to receive the preprocessed data stream and calculate the feature vector of the data using a stored feature extraction model. The encrypted communication module integrates a hardware encryption chip for hardware-level encrypted encapsulation of the processed feature data before sending it to the zero-trust gateway.

[0010] By adopting the above technical solution, and utilizing passive physical mirroring acquisition technology, raw data streams can be acquired without interfering with the normal operation of precision scientific instruments or consuming resources of the instrument's host computer system. Combined with local AI computing power modules at the edge, feature extraction can be completed at the source of data generation, avoiding bandwidth pressure caused by direct transmission of massive amounts of raw data. At the same time, hardware-level encryption ensures the confidentiality and integrity of data during transmission.

[0011] Preferably, the AIGC fingerprint recognition engine is configured to perform fingerprint generation and authentication processes. In the generation process, the system decomposes the read data stream into metadata and content data. For the metadata, a hash algorithm is used to generate feature vectors; for the content data, it is converted into a standardized multidimensional tensor and input into a convolutional neural network incorporating an attention mechanism. This network calculates normalized attention weights for feature regions through attention layers, retaining only the feature regions with the highest weights for weighted aggregation, thereby generating a content feature vector. Finally, the content feature vector and the metadata feature vector are concatenated serially to generate a globally unique digital fingerprint vector and write it to the blockchain's evidence storage node. In the authentication process, the Hamming distance between the current fingerprint vector of the file under test and the original evidence storage fingerprint vector on the chain is calculated, which is the sum of the number of bits that differ in value at the same position in the two vectors. Based on the comparison result of the Hamming distance and a preset threshold, original genuine data, suspected edited data, and forged data are distinguished, and release, sandbox cleaning, or blocking operations are performed.

[0012] By adopting the above technical solution, the system, through the introduction of an attention-based deep learning model, can automatically focus on key regions (such as spectral peaks and lattice edges) in the experimental data that have the highest information entropy or the most significant structure, ignoring background noise. This makes the generated fingerprint highly sensitive to malicious data tampering, while maintaining a certain degree of robustness to benign operations such as format conversion. Combining metadata hashing and blockchain notarization, a dual-modal anti-tampering mechanism is constructed, effectively solving the problems of difficulty in establishing ownership and tracing the source of scientific research data.

[0013] Preferably, the reinforcement learning dynamic authorization module constructs an environment interaction model based on a deep Q-network. This model defines a multi-dimensional state space including user characteristics, device health, data attributes, and historical behavior, as well as a discrete action space including permission granting, sandbox detection, and blocking. The deep Q-network receives the current state vector, outputs the value assessment value of each action, and selects the action with the highest value as the decision. The system is configured with a composite reward function, which is a weighted sum of a security penalty term and an efficiency penalty term: if a security event is detected within the time window after the action is executed, a high negative penalty is applied; if it causes the user to wait, a negative penalty proportional to time is applied. The network iteratively updates its parameters to minimize the error between the predicted value and the target value, thereby maximizing the long-term cumulative reward.

[0014] By adopting the above technical solution and utilizing reinforcement learning algorithms, the system can automatically learn the optimal security strategy through interaction with the environment without predefined fixed rules. The composite reward function guides the model to find a dynamic balance between strict security defenses and research efficiency, avoiding both business stagnation caused by over-defense and security vulnerabilities resulting from lax policies, thus achieving adaptive evolution of the security strategy.

[0015] Preferably, the federated learning prediction center executes a threat prediction process based on graph neural networks. The system models the laboratory network environment as a graph structure, where nodes represent physical entities and edges represent communication relationships. Using graph neural network algorithms, it updates the hidden state representation of the current node by aggregating information from neighboring nodes, thereby capturing local network structure patterns and anomaly associations. Employing a federated learning framework, each edge access device performs local training using local data, only uploading encrypted model parameter gradients to the center for aggregation.

[0016] By adopting the above technical solutions, and transforming network logs into a graph structure and utilizing graph neural networks for mining, the system can identify abnormal subgraph structures lurking around normal nodes, effectively predicting the lateral movement path of viruses. The federated learning mechanism enables various laboratory nodes to share attack characteristics and model parameters without sharing sensitive raw data, achieving collaborative defense across the entire network and early detection of unknown threats.

[0017] Preferably, the adaptive sandbox possesses dynamic resource allocation and abnormal behavior detection capabilities. Dynamic resource allocation refers to adjusting the computing resources and simulation time of the sandbox virtual machine based on the file's reputation score. Abnormal behavior detection utilizes a variational autoencoder-based model, which includes an encoder and a decoder, to map the input system call sequence to latent spatial distribution parameters and attempt to reconstruct it. When the loss value, consisting of the reconstruction error of the input sequence and the divergence of the latent distribution, exceeds an anomaly threshold, the file is determined to be malware.

[0018] By adopting the above technical solution, and utilizing an unsupervised learning method based on variational autoencoders, the sandbox no longer relies on traditional signature matching. Instead, it identifies anomalies by learning the behavioral distribution of normal scientific research software. This enables the system to effectively detect unknown virus variants or zero-day attack payloads. Simultaneously, the dynamic resource allocation mechanism improves the sandbox's detection efficiency, ensuring that high-risk files are fully analyzed while low-risk files can pass through quickly.

[0019] Preferably, the intelligent computing layer further includes an intent recognition and request structuring module and a policy optimization engine. The intent recognition module uses a natural language processing model to parse user voice or text commands, extract key entities, and automatically generate structured request forms. The policy optimization engine converts access control rules into Boolean expressions, identifies and masks conflicts and redundant conflicts through set operations, and automatically generates new rules based on security priority weights to clarify actions in logically overlapping areas.

[0020] By adopting the above technical solutions, the complex approval process is significantly simplified and the user experience is improved due to the introduction of natural language processing technology. The policy optimization engine, based on Boolean logic mathematical verification, ensures that the access control policy library remains logically consistent and efficient as the number of rules increases, thus resolving security risks caused by human configuration errors.

[0021] This invention provides an AI-driven, zero-trust data exchange and approval system for laboratory instruments. It offers the following advantages: 1. This invention achieves non-intrusive source control of raw instrument data by constructing a perception access layer containing a passive edge mirroring acquisition device and a zero-trust execution control layer. This setup physically isolates data acquisition from the instrument's operational environment, avoiding the risk of data tampering before being uploaded to the blockchain. Combined with AI-based dynamic channel control, it ensures the security and authenticity of cross-network data exchange throughout its entire lifecycle without interfering with the normal operation of precision scientific instruments.

[0022] 2. This invention combines an AIGC fingerprint recognition engine based on an attention-based convolutional neural network with a blockchain-based evidence storage node to establish a tamper-proof identification mechanism sensitive to data content. By extracting key feature regions to generate a weighted content vector and combining it with metadata hashing on the blockchain, the system can accurately identify minor malicious tampering behaviors while maintaining robustness to benign format conversions. This effectively solves the problem of verifying the ownership and traceability of scientific research experimental data and the inability of traditional hash verification to distinguish between malicious editing and normal processing.

[0023] 3. This invention utilizes a reinforcement learning dynamic authorization module and a federated learning prediction center to achieve adaptive evolution of security policies and collaborative defense under privacy protection. The reinforcement learning model automatically balances security detection intensity and business flow efficiency through a composite reward function, avoiding approval blockages or security vulnerabilities caused by fixed rules. Simultaneously, the federated learning mechanism supports the aggregation of threat features across the entire network without sharing original sensitive data, enhancing the system's proactive defense capabilities against unknown virus variants and advanced persistent threats. Attached Figure Description

[0024] Figure 1 This is an architecture diagram of the AI-driven zero-trust data exchange and approval system for laboratory instruments according to the present invention. Figure 2 This is a flowchart of the attention-based dual-modal data fingerprint generation and identification process of the present invention. Figure 3 This is a logical diagram of the reinforcement learning-based dynamic authorization decision-making model of the present invention; Figure 4 This is a flowchart of the threat prediction process based on federated graph neural networks of the present invention; Figure 5 This is a flowchart of the natural language approval, strategy optimization, and sandbox testing process of this invention. Detailed Implementation

[0025] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] See attached document Figure 1 This invention provides an AI-driven zero-trust data exchange and approval system for laboratory instruments, which includes a perception access layer, an intelligent computing layer, an execution control layer, and an operation and maintenance prediction layer.

[0027] The sensing and access layer is located at the forefront of the system, directly facing user terminals and scientific instruments. It comprises multiple edge access devices distributed across various laboratory nodes and interactive software running on user mobile devices. The edge access devices connect directly to the data output ports of the scientific instruments via physical interfaces, enabling real-time acquisition of raw output data and operational status data. The interactive software runs on the user's smartphone or tablet, providing an authentication interface and a platform for initiating data transmission requests. The primary functions of the sensing and access layer are to complete physical access to the data source, preliminary feature extraction, and initial verification of user identity. The edge access devices and user terminals are connected to the upper-layer network infrastructure via a local area network (LAN) or wireless network.

[0028] The intelligent computing layer is the core data processing center of the system, deployed on a local server cluster or private cloud platform. It includes an AI data classification engine, an AIGC fingerprint recognition engine, and a reinforcement learning dynamic authorization module. The AI ​​data classification engine receives data from the perception access layer and uses a pre-trained natural language processing model to parse and classify file content. The AIGC fingerprint recognition engine is responsible for extracting deep features from the collected data and generating digital fingerprints for authenticity verification. The reinforcement learning dynamic authorization module receives user requests and environmental status information and uses a deep neural network model to calculate access control policies. The intelligent computing layer communicates bidirectionally with the perception access layer and the execution control layer via a high-speed internal network, receiving input data and outputting decision instructions.

[0029] The execution control layer is responsible for implementing security policies and directly managing data flow and access permissions. It comprises a zero-trust gateway, an adaptive sandbox, and blockchain evidence nodes. The zero-trust gateway is deployed on critical paths of the lab network, opening or closing data transmission channels based on decisions from the intelligent computing layer. The adaptive sandbox is an isolated, virtualized operating environment used for dynamic behavioral analysis and virus detection of files deemed suspicious. The blockchain evidence nodes maintain a distributed, immutable ledger, recording digital fingerprints and access logs for all data. Based on instructions from the intelligent computing layer, the execution control layer allows, blocks, or redirects data packets passing through the gateway to the sandbox.

[0030] The Operations and Maintenance Prediction Layer is responsible for long-term system status monitoring and potential threat analysis. This layer includes a Federated Learning Prediction Center and a Situational Awareness Platform. The Federated Learning Prediction Center coordinates local models distributed across various edge access devices to aggregate parameters and generate a global threat prediction model. The Situational Awareness Platform collects security logs and alarm information from the entire network and uses time-series analysis algorithms to identify abnormal behavior patterns. The Operations and Maintenance Prediction Layer connects to other layers of the system through an out-of-band management network, continuously updating the security rule base and prediction model parameters.

[0031] An edge access device is an embedded hardware device deployed near the physical end of scientific instruments. This device includes a data mirroring interface module, a local AI computing power module, and an encrypted communication module.

[0032] The data mirroring interface module is equipped with multiple physical interfaces, including a universal serial bus interface, an Ethernet interface, and an industrial serial communication interface. This module connects in parallel to the data output lines of scientific instruments via physical cables, passively mirroring the raw data stream generated by the instruments without interfering with their normal operation. The module integrates signal conditioning circuitry to filter and shape the acquired electrical signals, ensuring the integrity of the acquired data.

[0033] The local AI computing module connects to the data mirroring interface module to receive the preprocessed data stream. The local AI computing module includes an embedded tensor processing unit or graphics processing unit, possessing the computational capability to perform deep neural network inference. The local AI computing module stores a lightweight feature extraction model for real-time computation of feature vectors at the edge. The local AI computing module is also equipped with high-speed random access memory for temporarily storing acquired data fragments for computation.

[0034] The encrypted communication module connects to the local AI computing module and is responsible for sending the processed feature data to the intelligent computing layer. The encrypted communication module integrates a hardware encryption chip compliant with national commercial cryptography standards, supporting SM2, SM3, and SM4 algorithms. All outgoing data packets are encrypted and encapsulated at the hardware layer and sent to the zero-trust gateway via Ethernet or fiber optic interface. The encrypted communication module also includes an authentication unit that stores the device's unique digital certificate for two-way authentication with other components of the system.

[0035] When a terminal device attempts to access the laboratory network, the system executes a network access process that includes digital certificate verification, two-factor authentication, and environmental baseline scanning.

[0036] When a terminal device connects to a network access point via wired or wireless means, a digital certificate verification step is triggered first. The zero-trust gateway sends an identity request message to the terminal device. The terminal device responds to the request, submitting a digital certificate stored in a local security chip or USB key. The zero-trust gateway verifies the certificate's validity period, signature integrity, and revocation status. If the certificate is invalid or no certificate is provided, the gateway directly discards all data packets from the terminal at the link layer, preventing it from establishing a network connection.

[0037] After the digital certificate verification is successful, the system initiates the two-factor authentication process. The system sends a one-time dynamic password or biometric verification request to the user's linked mobile terminal. The user completes the verification operation through the mobile terminal and sends the verification result back to the system. The system compares the verification result to confirm the true identity of the current operator. This step ensures that even if the terminal device is lost or the certificate is stolen, unauthorized holders cannot pass the identity verification.

[0038] After successful two-factor authentication, the system performs an environmental baseline scan. The system distributes a lightweight scanning agent to the terminal device. This agent scans the terminal device's operating system version, patch installation status, antivirus software running status, and specific process lists. The agent generates a baseline report from the scan results and uploads it to the system. The system compares the baseline report with the preset security baseline policy. If the terminal device's baseline score is below a preset threshold, the system restricts its network access to the remediation area, allowing only access to the patch server and antivirus software update server. Only after the terminal device remediates the violations and passes the review is its access to scientific instruments and core data granted.

[0039] See attached document Figure 2 First, the heterogeneous raw data collected from scientific instruments undergoes preprocessing and feature separation. Since scientific instruments output various file formats, including but not limited to CSV text, TIFF images, and RAW mass spectrometry data, the system first reads the binary header information of the data stream using a file parser to identify the file type and locate the starting offset of the data segment. Based on the file structure definition, the system then inputs the data object... The process is divided into two parts: metadata and content data. The metadata includes the file generation timestamp, instrument serial number, environmental parameters (such as temperature, pressure, and flow rate), and file header description information. The content data contains the actual experimental observation numerical matrix or image pixel matrix. The system extracts the metadata into a structured text stream and converts the content data into a standardized multidimensional tensor, which is then input into the subsequent feature extraction module.

[0040] For the separated content data portion, the system employs a convolutional neural network incorporating an attention mechanism for deep feature extraction. This network structure aims to select the most representative key feature regions from massive experimental data, rather than averaging the entire dataset. Specifically, the system inputs a standardized content data tensor into a multi-layer convolutional network to generate a high-dimensional feature map. To focus on regions with the highest information entropy or most significant structure in the data (e.g., lattice edges in electron microscopy images or strong ion peaks in mass spectra), the system introduces an attention layer after the convolutional layer. The attention layer calculates the energy score for each spatial location or channel and obtains the attention weights by normalization using the Softmax function. For the th feature map... Each feature region has a normalized attention weight. The calculation formula is as follows: ; in, Indicates the total number of feature regions; Indicates the first The original energy score of each feature region is determined by the feature response values ​​output by the convolutional network. The system then calculates the energy score based on the weights. The feature regions are sorted, and only the top 196 feature regions by weight are retained. The feature vectors of these high-response regions are then weighted and aggregated to generate content feature vectors. This process effectively filters out background noise and redundant information in the data, ensuring that the generated feature vectors are highly sensitive to minor data tampering, while maintaining a certain degree of robustness to format conversion or resampling operations.

[0041] After feature extraction, the system performs fingerprint vector synthesis. First, the system uses the SHA-256 algorithm to hash the extracted metadata portion, extracting the first 128 bits to generate the metadata feature vector. Subsequently, the system generates a 256-bit content feature vector. Metadata feature vector Serial concatenation is performed to generate a globally unique digital fingerprint vector with a total length of 384 bits. This fingerprint vector serves as a digital identity for the data and does not contain any original experimental data, thus meeting data privacy protection and cross-border transmission compliance requirements. The system will generate the fingerprint vector... The generation of a timestamp, user digital signature, and instrument ID is encapsulated into a data packet for evidence storage. This data packet is written into the block body of the consortium blockchain through a consensus mechanism, generating an immutable on-chain evidence storage record. This is used for subsequent data lifecycle traceability and authenticity verification.

[0042] When the system receives a data exchange request or file transfer instruction, it initiates the authenticity verification process in real time. The system first recalculates the fingerprint vector of the file to be tested according to the steps described above. And retrieve the corresponding original evidence fingerprint vector from the blockchain ledger based on the file ID. The system calculates the Hamming distance between two binary vectors. Hamming distance is the sum of the number of digits that differ in value at the same position between two vectors. The calculation formula is as follows: ; in, The vector represents the first Bit binary value; This represents the XOR operation. The system calculates the result... The value, based on a preset three-level threshold logic, executes differentiated control actions: when When the file to be tested is determined to be original, genuine data, the system performs a release operation, allowing data transmission; when When the system determines that the file to be tested is suspected of being edited data or has suffered minor formatting loss, it performs a sandbox cleaning operation, importing the file into an isolated environment for further testing; If the system determines that the file under test is forged data, AI-generated data, or maliciously tampered with, it will perform a blocking operation, cut off the transmission connection, trigger a security alarm, and record the violation in the audit log.

[0043] See attached document Figure 3 This invention employs reinforcement learning technology to construct a dynamic access control system, optimizing the authorization strategy in real time through the interaction between the agent and the environment. First, reinforcement learning environment modeling is performed, and the system defines its state space. Let be a continuous vector space containing multidimensional features. At any time... Environmental conditions It is composed of feature vectors from four core dimensions: user feature vector Device health vector Data attribute vector and historical behavior feature vectors Among them, user feature vector The code encodes the user's role permissions, project group affiliation, and current authentication level; device health vector. It includes the terminal's security baseline score, patch status, and real-time running process information; data attribute vectors. It covers the data's security classification label, file type, and fingerprint authentication results; historical behavior feature vectors. This records the user's or device's access frequency, number of failures, and abnormal operation sequences over a past period. The system defines the action space. For discrete decision sets Among them, actions In response to the direct release policy, the system immediately establishes an encrypted transmission channel; action In accordance with the sandbox detection strategy, the system redirects the data stream to an isolated environment for dynamic analysis; Action In accordance with the corresponding blocking policy, the system immediately disconnects the connection and logs the violation. The agent then determines the appropriate action based on the observed state. Output action To maximize long-term cumulative rewards.

[0044] To ensure system security while maintaining operational efficiency for scientific research, this invention designs a composite reward function. This function consists of a weighted sum of a security penalty term and an efficiency penalty term, guiding the agent to learn an optimal strategy that neither misses threats nor causes excessive delays. At time... Immediate reward from environmental feedback after an action is performed. The calculation formula is as follows: ; in, This is a security event indication function. If a virus infection, data breach, or unauthorized access is detected within a preset time window after the action is executed, then... The value is 1 otherwise. A constant coefficient of -100 represents a higher penalty weight for safety incidents, forcing the agent to prioritize avoiding high-risk decisions. This indicates the user's waiting time as a result of this decision, in seconds. If the release action is executed, Approaching 0; if sandbox detection is performed, This equals the actual detection time; if a blocking action is executed and proven to be a false alarm, This includes the time spent on user appeals and manual review. The coefficient -1 represents a linear penalty for time cost. This is achieved by minimizing the negative value of the reward function (i.e., maximizing...). The model can automatically converge to a balance point that minimizes user waiting time while ensuring no safety incidents occur.

[0045] The system employs a Deep Q-Network (DQN) algorithm to approximate the optimal action-value function. The DQN network consists of an input layer, multiple hidden layers, and an output layer. The input layer receives the environmental state. The output layer corresponds to the Q-value of each action in the action space. The system maintains an experience replay buffer to store historical transition quadruples. During training, the system randomly samples a mini-batch of samples from the buffer and iteratively updates the Q-value network parameters using the Bellman equation. The update objective for the Q-value follows the formula: ; in, Indicates the environmental state Take action below The current estimated value; The learning rate controls the step size of model parameter updates and determines the extent to which newly acquired information overwrites old information. This is a discount factor, ranging from 0 to 1, used to weigh the importance of immediate rewards against future long-term rewards. Indicates the next state Under these conditions, the optimal action is chosen to obtain the maximum expected value. By iteratively minimizing the mean square error between the predicted Q-value and the target Q-value, the deep neural network gradually masters the mapping relationship from complex state characteristics to the optimal security strategy, achieving proactive defense against unknown threats and adaptive access to normal business operations.

[0046] See attached document Figure 4 First, the complex network environment within the laboratory is modeled as graph-structured data. The system defines an undirected graph. This represents the entire network topology. The set of nodes... Each node in This represents a physical entity within the network, specifically including scientific instruments, experimental control PCs, data servers, printers, and handheld terminal devices. Each node All are associated with an initial feature vector This vector contains the device's static attributes (such as operating system type, IP address range, and open port list) and dynamic states (such as CPU utilization, memory usage, and current connection count). Edge set Each edge in Representing two nodes and The communication relationships or data transmission behaviors that exist between the edges. Edge weights. Quantitative values ​​are assigned based on communication frequency, data transmission volume, and protocol type. Through this graphical modeling, the system can correlate discrete security events, intuitively reflecting potential attack paths and lateral movement trajectories of viruses within the network.

[0047] After constructing the network graph, the system utilizes a Graph Neural Network (GNN) algorithm for deep feature mining and aggregation. The core idea of ​​GNN is to use a message-passing mechanism to allow each node to update its hidden state representation using information from its neighbors, thereby capturing local network structure patterns and anomalies. For any node in the graph... The system in the Feature aggregation and updating in layered neural networks follow the formula below: ; in, Represents a node In the Hidden layer feature vectors of layer, initial state Equal to the original feature vector of the node ; Represents a node The set of first-order neighbor nodes, i.e., all nodes connected to the node There are directly connected nodes; AGG represents the set of feature vectors of all neighboring nodes in the previous layer. For aggregation functions, in this embodiment, mean aggregation or max pooling aggregation is used to summarize the feature information of neighboring nodes into a vector of fixed length. For the first The learnable weight matrix of the layer is used to linearly transform the aggregated information; This is a non-linear activation function, such as the ReLU function. Through multiple stacking layers, the nodes... The final embedding vector will contain the structural and attribute information of its multi-hop neighbors, enabling the model to identify abnormal subgraph structures lurking around normal nodes and predict the probability of a node being infected by its neighbors.

[0048] Considering the high sensitivity of laboratory data and privacy protection requirements, this system employs a federated learning framework for collaborative model training. During training, each laboratory subnet acts as an independent local client, utilizing locally collected network traffic logs and graph data for local training of the GNN model. After local training is complete, each client does not upload any raw log data or graph structure; instead, it only encrypts and sends the calculated model parameter gradients or updated weight matrices to the central server. Upon receiving parameter updates from each client, the central server uses a federated averaging algorithm (FedAvg) to weight and aggregate the parameters, generating updated global model parameters. Subsequently, the central server distributes the global parameters back to each local client, overwriting the old local model. This process iterates until the model converges. In this way, the system achieves the sharing of attack characteristics across the entire network while ensuring that the original data remains within the domain. Each laboratory node can not only defend against locally known threats but also instantly acquire new attack patterns discovered by other nodes (such as the lateral movement characteristics of new ransomware) through the global model, thereby generating a dynamic attack probability graph covering the entire network and enabling joint defense and early blocking of unknown threats.

[0049] See attached document Figure 5 To address the problem of low research efficiency caused by cumbersome forms in traditional approval systems, this invention designs an intelligent intent recognition and request structuring module. When a user needs to initiate a data exchange request, they input natural language commands through the voice acquisition interface of a mobile application, or directly enter descriptive text in a text box. The system first calls an automatic speech recognition engine to convert analog audio signals into digital text streams. Subsequently, the system uses a pre-trained natural language processing model to perform semantic parsing on the text. This model is based on a bidirectional encoder representation transformer architecture and has been fine-tuned on a laboratory-specific corpus.

[0050] The system automatically extracts key elements from unstructured natural language text using named entity recognition technology. The model labels each word in the text sequence, identifying entity fragments representing the source file name, target recipient, data purpose, and urgency. For example, for an instruction to send yesterday's chromatography data to Professor Li for project completion, the system automatically identifies the chromatography data as the operation object, Professor Li as the target entity, and project completion as the business context. After identification, the system fills these discrete entities into a predefined JSON-formatted approval template, automatically generating a structured request form containing the source IP address, target user ID, data fingerprint hash value, and application reason. This request form is echoed back to the user for final confirmation before being submitted to the policy engine, ensuring the accuracy of machine understanding.

[0051] As laboratory data exchange rules accumulate, static access control lists often become redundant or exhibit logical conflicts. This system incorporates a Boolean logic-based policy optimization engine. This engine first transforms each access control rule into a standard Boolean expression. The source IP address range, destination port number, and protocol type in the rule are mapped to a set of Boolean variables. The system constructs a multi-dimensional Boolean space, with each rule corresponding to a hyperrectangular region within that space.

[0052] The system uses set operations to detect spatial relationships between rules, identifying two main types of conflicts: blocking and redundancy. When the IP range covered by a high-priority blocking rule completely encompasses the IP range of another low-priority allowing rule, it is determined to be a blocking conflict. The system automatically marks the blocked rule as invalid and recommends that the administrator delete it. When two rules have completely identical conditions and the same actions, they are determined to be a redundant conflict. The system automatically merges these two rules to reduce the retrieval overhead of the policy database. In addition, for related conflicts that have partial logical overlap but opposite actions, the system automatically generates a new high-priority rule based on preset security priority weights to clarify the action in the overlapping area, thereby eliminating the uncertainty of policy execution and ensuring the logical consistency of the firewall rule set.

[0053] For files deemed suspicious or slightly edited by the authenticity verification module, the system imports them into an adaptive sandbox environment for isolated execution. Unlike traditional static resource sandboxes, the sandbox of this invention has dynamic resource allocation capabilities. The system calculates a suspicion score based on the file's type, size, and source credibility. Based on this score, the system dynamically adjusts the number of CPU cores, memory size, and simulation time allocated to the sandbox virtual machine. For complex files with high suspicion, the system allocates more computing resources and extends the observation time to lure in malicious code with anti-detection capabilities.

[0054] In terms of detection logic, the sandbox not only performs traditional signature matching but also introduces an anomaly detection model based on variational autoencoders specifically for identifying unknown virus variants. The system collects a large number of system call sequences from normal scientific research software runtime as training data. The variational autoencoder consists of an encoder and a decoder. The encoder processes the input system call sequences... The mapping is to Gaussian distribution parameters in the latent space, and the decoder attempts to extract these parameters from the latent vectors. Reconstruct the original input The training objective of the model is to minimize the KL divergence between the reconstruction error and the latent distribution. Loss function. The definition is as follows: ; in, This represents a sequence vector of input system call behaviors. Represents the sequence vector output by the network reconstruction; This is the reconstruction error term, used to measure the network's ability to reconstruct the input; KL divergence is used to measure the posterior distribution of the encoder output. Compared with the standard normal prior distribution The differences between them; This is the regularization coefficient, used to adjust the weights of the two terms. During the detection phase, if the behavior sequence generated by the test file running in the sandbox is input into the model, the total loss value is calculated. If the behavior exceeds a preset anomaly threshold, it indicates that the behavior pattern cannot be effectively reconstructed by the model, meaning it deviates from the normal behavior distribution of software. Based on this, the system determines that the file is an unknown malware variant or a zero-day attack payload and immediately triggers blocking and cleaning processes to prevent it from penetrating into the core network.

Claims

1. An AI-driven zero-trust data exchange and approval system for laboratory instruments, characterized in that: include: The system comprises a perception and access layer, an intelligent computing layer, an execution and control layer, and an operation and maintenance prediction layer. The perception access layer is used to connect user terminals and scientific instruments, including edge access devices and interactive software running on user mobile terminals. The edge access devices are used to collect raw output data and operating status data of scientific instruments, and the interactive software is used to provide identity authentication interfaces and initiate data transmission requests. The intelligent computing layer is used to process data and output decision instructions, including an AI data classification engine, an AIGC fingerprint recognition engine, and a reinforcement learning dynamic authorization module. The AI ​​data classification engine is used to parse and classify file content, the AIGC fingerprint recognition engine is used to generate digital fingerprints, and the reinforcement learning dynamic authorization module is used to calculate access control policies. The execution control layer is used to implement security policies, including a zero-trust gateway, an adaptive sandbox, and a blockchain evidence storage node. The zero-trust gateway is used to manage the data transmission channel according to the decision instructions. The adaptive sandbox is used to isolate and analyze suspicious files. The blockchain evidence storage node is used to record the digital fingerprint and access logs. The operation and maintenance prediction layer is used for status monitoring and threat analysis, including a federated learning prediction center and a situational awareness platform. The federated learning prediction center is used to generate a global threat prediction model, and the situational awareness platform is used to identify abnormal behavior patterns. The AIGC fingerprint recognition engine is configured to execute the fingerprint generation process: First, the binary header information of the data stream is read using a file parser, and the read data stream is decomposed into metadata and content data. For the metadata portion, a hash algorithm is used to perform calculations and the first few bits are truncated to generate a metadata feature vector; For the content data portion, the content data portion is converted into a standardized multidimensional tensor and input into a convolutional neural network with an attention mechanism. The attention layer of the convolutional neural network calculates the normalized attention weights of the feature regions, and the feature regions with the highest weights are retained for weighted aggregation to generate a content feature vector. The content feature vector and the metadata feature vector are concatenated in series to generate a globally unique digital fingerprint vector, and the globally unique digital fingerprint vector is encapsulated and written into the blockchain evidence storage node. The AIGC fingerprint recognition engine is also configured to perform a genuine / counterfeit authentication process: Upon receiving a data exchange request, the current fingerprint vector of the file under test is calculated by concatenating the content feature vector extracted by a convolutional neural network with an attention mechanism and the metadata feature vector generated by a hash algorithm. The corresponding original evidence fingerprint vector is then retrieved from the blockchain evidence storage node based on the file ID. The Hamming distance between the current fingerprint vector and the original stored fingerprint vector is calculated by XORing and summing the binary bits of the two vectors. Based on the comparison result between the Hamming distance and the preset threshold, the following control actions are executed: when the Hamming distance is less than or equal to the first preset threshold, it is determined to be original true data and a release operation is triggered; when the Hamming distance is greater than the first preset threshold and less than or equal to the second preset threshold, it is determined to be suspected edited data and an operation to import the file into the adaptive sandbox is triggered; when the Hamming distance is greater than the second preset threshold, it is determined to be forged data and a blocking operation is triggered. The reinforcement learning dynamic authorization module constructs an environment interaction model based on a deep Q-network. The environmental interaction model defines the state space as a vector space containing multidimensional features, including user feature vectors, device health vectors, data attribute vectors, and historical behavior feature vectors. The environmental interaction model defines the action space as a discrete set of decisions, which includes direct release strategy, sandbox detection strategy, and blocking strategy. The deep Q-network is used to receive the current state vector, output the value assessment value of each action in the action space, and select the action with the largest value assessment value as the decision instruction to be output to the execution control layer.

2. The AI-driven zero-trust data exchange and approval system for laboratory instruments according to claim 1, characterized in that, The edge access device includes a data mirroring interface module, a local AI computing power module, and an encrypted communication module. The data mirroring interface module is equipped with a physical interface for parallel connection to the data output line of the scientific instrument via a physical cable, and passively mirrors the raw data stream generated by the scientific instrument. The data mirroring interface module integrates a signal conditioning circuit for filtering and shaping electrical signals. The local AI computing power module is connected to the data mirroring interface module. The local AI computing power module is used to receive the preprocessed data stream and use the stored feature extraction model to calculate the feature vector of the data. The encrypted communication module is connected to the local AI computing power module. The encrypted communication module integrates a hardware encryption chip, which is used to encapsulate the processed feature data in hardware layer encryption and send it to the zero-trust gateway.

3. The AI-driven zero-trust data exchange and approval system for laboratory instruments according to claim 1, characterized in that, The reinforcement learning dynamic authorization module is configured with a composite reward function to optimize the deep Q-network; The composite reward function is composed of a weighted sum of a safety penalty term and an efficiency penalty term. The security penalty item is configured as follows: when a security event is detected within a preset time window after the action is performed, a first negative value is taken; otherwise, a zero value is taken. The efficiency penalty item is configured to: calculate a second negative value based on the user operation waiting time caused by the decision instruction; The deep Q-network iteratively updates its network parameters to maximize the cumulative value of the composite reward function.

4. The AI-driven zero-trust data exchange and approval system for laboratory instruments according to claim 1, characterized in that, The federated learning prediction center is configured to perform a threat prediction process based on graph neural networks: The laboratory network environment is modeled as a graph structure, where nodes represent physical entities and edges represent communication relationships. The features of the node are mined using a graph neural network algorithm, and the hidden state representation of the current node is updated by aggregating information from neighboring nodes. A federated learning framework is used for collaborative training. The edge access device acts as a local client and uses local data for local training. It then sends the encrypted model parameter gradients to the federated learning prediction center. The federated learning prediction center uses a federated averaging algorithm to weight and aggregate the received parameters, generate global model parameters, and send them back to the edge access device.

5. The AI-driven zero-trust data exchange and approval system for laboratory instruments according to claim 1, characterized in that, The adaptive sandbox is configured to have dynamic resource allocation capabilities and abnormal behavior detection capabilities. The dynamic resource allocation capability refers to calculating a suspicion score based on the file type, size, and source credibility, and adjusting the number of processor cores, memory size, and simulation time length allocated to the sandbox virtual machine based on the suspicion score; The abnormal behavior detection capability refers to the ability to identify unknown virus variants using a variational autoencoder-based model. The variational autoencoder includes an encoder and a decoder, which are used to map the input system call sequence to potential spatial distribution parameters and reconstruct them. When the loss value consisting of the reconstruction error of the input sequence and the divergence of the potential distribution exceeds a preset abnormal threshold, the file is determined to be malware.

6. The AI-driven zero-trust data exchange and approval system for laboratory instruments according to claim 1, characterized in that, The intelligent computing layer also includes an intent recognition and request structuring module; The intent recognition and request structuring module is configured to receive natural language instructions, convert audio into a text stream using an automatic speech recognition engine, and perform semantic parsing on the text stream using a natural language processing model. Key elements are extracted from the text stream using named entity recognition technology. These key elements include the source file name, the target recipient, the purpose of the data, and the urgency level. Fill the key elements into the predefined template to automatically generate a structured request form containing the source address, target user identifier, and application reason.

7. The AI-driven zero-trust data exchange and approval system for laboratory instruments according to claim 1, characterized in that, The intelligent computing layer also includes a strategy optimization engine; The policy optimization engine is configured to convert access control rules into Boolean expressions and map each rule to a region in a multidimensional Boolean space. Spatial relationships between rules are detected through set operations to identify masked conflicts and redundant conflicts. When the scope of a high-priority blocking rule includes the scope of a low-priority allowing rule, it is determined to be a blocking conflict and the blocked rule is marked as invalid. When two rules define the same conditions and have the same actions, they are considered redundant and conflicting, and the rules are merged. When rules have logical overlap but opposite actions, a new high-priority rule is generated based on a preset security priority weight. The new high-priority rule is used to clarify the actions in the overlapping area.

Citation Information

Patent Citations

  • Zero-trust access method and device for Internet of Things terminal

    CN120263515A

  • Systems and methods for trustworthy electronic authentication using a computing device

    US11405189B1