Intelligent Traceability System for the Entire Laboratory Testing Process Based on Reinforcement Learning

The intelligent traceability system for the entire laboratory testing process, which utilizes reinforcement learning and employs a multi-objective deep deterministic strategy gradient algorithm and dynamic weight adjustment, solves the problems of data loss and delayed evidence storage caused by resource competition in existing technologies, and achieves efficient and reliable traceability of the laboratory testing process.

CN121073409BActive Publication Date: 2026-03-06连云港海关综合技术中心
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511630656.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-03-06
Estimated Expiration
2045-11-10

AI Technical Summary

Technical Problem

Existing laboratory testing traceability systems are prone to data loss or delayed evidence storage when resource competition is fierce. They lack intelligent decision-making mechanisms, cannot adapt to complex scenarios, resulting in low traceability efficiency and serious waste of resources, and cannot meet the needs of high-precision testing.

Method used

A reinforcement learning-based intelligent traceability system for the entire laboratory testing process is adopted. It dynamically balances resource load and traceability integrity through a multi-objective deep deterministic strategy gradient algorithm, and combines dynamic weight adjustment and model iterative optimization to achieve adaptive switching of resource scheduling decisions and data storage modes.

Benefits of technology

It significantly improves the efficiency and reliability of laboratory testing processes, reduces resource consumption under high load, ensures the integrity of traceability data under low load, adapts to changes in the laboratory environment, and provides a reliable and efficient traceability solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121073409B_ABST
    Figure CN121073409B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent traceability system for the entire laboratory testing process based on reinforcement learning, belonging to the field of intelligent management technology for laboratory testing processes. The system includes: a status data acquisition module; a reinforcement learning intelligent decision-making module; a traceability scheduling and execution module; a hybrid architecture storage and verification module; and a model dynamic optimization module. By introducing a multi-objective deep deterministic strategy gradient algorithm, combined with the system's comprehensive resource load rate formula, dynamic multi-objective reward function formula, and multi-objective action value function formula, this invention achieves dual optimization of resource utilization and traceability accuracy. The system can dynamically adjust the traceability strategy according to real-time resource load conditions. Under high load, it automatically switches to fingerprint storage mode to reduce resource consumption, while under low load, it adopts full storage mode to ensure the integrity of traceability data. This intelligent scheduling mechanism significantly improves the efficiency and reliability of the entire laboratory testing process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent management technology for laboratory testing processes, specifically to an intelligent traceability system for the entire laboratory testing process based on reinforcement learning. Background Technology

[0002] With the digitalization and automation of laboratory testing processes, the traceability management of testing data has become a crucial link in ensuring the reliability of results. Traditional traceability systems often rely on static rules or single indicators to trigger data storage, making it difficult to dynamically balance resource consumption and traceability integrity. Especially when sample volumes surge or instruments are operating under high load, the system is prone to data loss or delayed storage due to resource competition, affecting the reliability and compliance of the testing process. Furthermore, the variability of the laboratory environment requires traceability strategies to be adaptable in real time, but existing technologies lack intelligent decision-making mechanisms to cope with complex scenarios.

[0003] Existing traceability technologies suffer from three major drawbacks: First, they rely on fixed thresholds to trigger evidence storage mode switching, failing to consider the dynamic coupling relationship between computing and storage resources, which can easily lead to resource mismatch. Second, traceability priority scheduling depends on manually preset rules and cannot be dynamically adjusted based on the real-time status of nodes, potentially missing crucial traceability steps. Third, they lack algorithm-driven self-optimization capabilities, requiring manual adjustment of system parameters and making it difficult to adapt to environmental changes during long-term laboratory operation. These shortcomings result in low traceability efficiency, significant resource waste, and an inability to meet the demands of high-precision testing scenarios. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an intelligent traceability system for the entire laboratory testing process based on reinforcement learning. This invention dynamically balances resource load and traceability integrity through a multi-objective deep deterministic strategy gradient algorithm. The system collects resource and traceability data in real time during the sample receiving, testing, and output stages, automatically generates evidence storage mode switching and node scheduling decisions, and combines dynamic weight adjustment and model iterative optimization mechanisms. The system can adapt to changes in the laboratory environment, significantly improving traceability efficiency and reliability, and is suitable for high-precision scenarios such as medical testing and environmental monitoring.

[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a reinforcement learning-based intelligent traceability system for the entire laboratory testing process, comprising:

[0006] Status data acquisition module: used to collect resource load data and traceability integrity correlation data during the sample receiving stage, instrument detection stage and data output stage, and transmit the collected status data to the reinforcement learning intelligent decision-making module;

[0007] The reinforcement learning intelligent decision-making module is based on a multi-objective deep deterministic policy gradient algorithm. It processes the received state data by combining the system's comprehensive resource load rate formula, dynamic multi-objective reward function formula, and multi-objective action value function formula to generate source tracing scheduling decisions. At the same time, the source tracing scheduling decisions are transmitted to the source tracing scheduling execution module, and the historical calculation data generated during the data processing is transmitted to the model dynamic optimization module.

[0008] Source tracing scheduling and execution module: Based on the received source tracing scheduling decision, it performs the full-data storage or fingerprint storage mode switching operation for source tracing data, as well as the source tracing node priority scheduling operation; after the operation is completed, it feeds back the execution result to the reinforcement learning intelligent decision module, and transmits the storage index of the source tracing data to the hybrid architecture storage and verification module;

[0009] Hybrid architecture storage and verification module: Based on the blockchain-cloud hybrid architecture, the module stores the traceability data corresponding to the evidence index and performs hash verification on the traceability data to complete the accuracy verification. The hybrid architecture storage and verification module feeds back the verification results to the reinforcement learning intelligent decision-making module.

[0010] Model Dynamic Optimization Module: This module receives historical computational data from the reinforcement learning intelligent decision-making module, iteratively optimizes the policy network weights and mathematical calculation formulas of the multi-objective deep deterministic policy gradient algorithm, and pushes the optimized parameters to the reinforcement learning intelligent decision-making module to update its data processing logic.

[0011] Furthermore, the resource load data collected by the status data acquisition module includes storage resource load rate and computing resource load rate, and the collected traceability integrity association data includes traceability node coverage, original hash value of traceability data, and data generation timestamp; the status data acquisition module collects traceability node coverage through a sample progress sensor, collects storage resource load rate through a cloud storage server resource monitoring submodule, and collects computing resource load rate through an edge computing node resource monitoring submodule, and the status data acquisition module uses the MQTT protocol to transmit status data with a transmission latency of ≤100ms.

[0012] Furthermore, the formula for the overall system resource load rate in the reinforcement learning intelligent decision-making module is: ,in, for The system's overall resource load rate at any given time is used to quantify the overall utilization of the system's storage and computing resources. for Real-time storage resource load rate, i.e. The ratio of the used storage capacity to the total storage capacity at any given time is obtained by the status data acquisition module through the cloud storage server resource monitoring submodule. for Calculate resource load rate in real time, i.e. The real-time utilization rate of CPU or GPU is collected by the status data acquisition module through the edge computing node resource monitoring submodule; 0.6 is the weighting coefficient of storage resource load rate, and 0.4 is the weighting coefficient of computing resource load rate.

[0013] Furthermore, the formula for the timeliness decay factor of the source data in the reinforcement learning intelligent decision-making module is as follows: ,in, is the timeliness decay factor of traceability data at time t, with a value range of 0 to 1, used to quantify the impact of the time difference between the generation and storage of traceability data on its integrity. The difference between the data generation time and the evidence storage time at time t is expressed in seconds. The generation time is automatically marked by the detection instrument when the data is generated, and the evidence storage time is automatically recorded by the hybrid architecture storage and verification module when the data is written. This is a time-sensitive threshold, ranging from 2 to 5 seconds. The attenuation coefficient, ranging from 1.0 to 1.5, was determined by the model dynamic optimization module based on the correlation analysis of timeliness and traceability value of 50,000 sets of data with different delays.

[0014] Furthermore, the formula for the dynamic weight coefficients in the reinforcement learning intelligent decision-making module is as follows: ,in, The target weight for traceability integrity at time t, with a value ranging from 0 to 1, is used to indicate the degree to which traceability integrity is prioritized in the current scenario; The weight of the resource consumption control target at time t, with a value ranging from 0 to 1, is related to... Complementarity is used to indicate the degree to which resource consumption is prioritized in the current scenario; This is the weighting adjustment coefficient, with a value ranging from 2.0 to 3.0; This represents the system's maximum tolerable load rate, set at 90%. Let t be the overall system resource load rate at time t.

[0015] Furthermore, the multi-dimensional source tracing integrity formula and the source tracing integrity increment formula in the reinforcement learning intelligent decision-making module are as follows: ; ,in, The score for the completeness of the multi-dimensional traceability at time t ranges from 0 to 1 and is used to comprehensively evaluate the coverage, accuracy, and timeliness of the traceability data. The weight for the coverage of traceability nodes ranges from 0.4 to 0.6. The weight for the product of accuracy and timeliness of traceability data ranges from 0.4 to 0.6. + =1; The traceability node coverage at time t is the ratio of the number of traceability nodes with existing evidence to the total number of traceability nodes that need to be stored at time t. The accuracy of the traceability data at time t is the ratio of the number of datasets whose hash verification matches the traceability data at time t to the total number of datasets. This is obtained by the hybrid architecture storage and verification module by comparing the original hash value of the traceability data with the hash value after storage. The timeliness decay factor of the traceability data at time t; It is the comprehensive score of source traceability completeness at time t, with a value range of 0 to 1; for The real-time multi-dimensional traceability integrity score is stored in the historical cache unit of the reinforcement learning intelligent decision-making module, and the storage period is consistent with the system data processing period.

[0016] Furthermore, the resource consumption elasticity penalty formula in the reinforcement learning intelligent decision-making module is as follows: ,in, The resource consumption elasticity penalty value at time t, with a value ranging from 0 to 1, is used to quantify the impact of the current resource consumption level on the system. This is the storage resource elasticity coefficient, with a value of 0.4, used to represent the weight of the impact of storage resource consumption on the system. The resource elasticity coefficient is set to 0.6, which represents the weight of the impact of resource consumption on the system. This represents the maximum load rate of storage resources, with a value of 100%, which is the load rate when storage resources are fully utilized. This is the maximum load rate of computing resources, set to 100%, which is the load rate when the CPU or GPU is fully utilized. Let t be the storage resource load rate; Calculate the resource load rate at time t.

[0017] Furthermore, the formula for the dynamic multi-objective reward function in the reinforcement learning intelligent decision-making module is as follows: ,in, The dynamic multi-objective reward value at time t, ranging from -1 to 1, is used to provide feedback signals for the multi-objective depth deterministic policy gradient algorithm. Let t be the weight of the traceability integrity target. The weight of the resource consumption control target at time t; The score for the completeness of multi-dimensional source tracing at time t; Let t be the resource consumption elasticity penalty value; The value is the action impact coefficient, which is 0.1 when performing full evidence storage, 0.3 when performing fingerprint evidence storage, and 0.2 when performing high-priority node scheduling. The source traceability integrity increment at time t.

[0018] Furthermore, the multi-objective action value function formula in the reinforcement learning intelligent decision-making module is as follows: ,in, for The optimal action value of performing an action at a given time state; for System status at all times; for Actions that are performed at all times; for Real-time dynamic multi-objective reward value; This is a discount factor with a value of 0.9, used to weigh the importance of immediate rewards against future rewards; The candidate actions for time t+1 include four types of actions: full evidence storage, fingerprint evidence storage, high-priority node scheduling, and low-priority node delayed scheduling. State at time t+1 Next, execute candidate actions The optimal action value; for The stability coefficient of the state at any given time. The value is 1 when the fluctuation range between the system state at time t and the system state at time t is ≤5%, and the value is 0.8 when the fluctuation range is >5%. The fluctuation range is calculated by comparing the changes in parameters such as resource load rate and traceability node coverage in the system state at two time points.

[0019] Furthermore, when the traceability scheduling execution module performs the full-data storage / fingerprint storage mode switching operation, it switches to fingerprint storage mode when the storage resource load rate is >80% or the computing resource load rate is >80%; it switches to full-data storage mode when the storage resource load rate is <40% and the computing resource load rate is <40%. The hybrid architecture storage and verification module uses Hyperledger Fabric blockchain to store the full data of high-priority traceability data, and uses Alibaba Cloud OSS cloud storage to store the SHA-256 fingerprint and key index of low-priority traceability data. The iterative optimization cycle of the model dynamic optimization module is once every 24 hours, and the optimization objects include the policy network weights and parameters of each mathematical calculation formula of the multi-objective deep deterministic policy gradient algorithm.

[0020] Compared with existing technologies, this reinforcement learning-based intelligent traceability system for the entire laboratory testing process has the following advantages:

[0021] I. This invention introduces a multi-objective deep deterministic strategy gradient algorithm, combined with the system's comprehensive resource load rate formula, dynamic multi-objective reward function formula, and multi-objective action value function formula, to achieve dual optimization of resource utilization and traceability accuracy. The system can dynamically adjust the traceability strategy according to the real-time resource load. Under high load, it automatically switches to fingerprint storage mode to reduce resource consumption, while under low load, it adopts full storage mode to ensure the integrity of traceability data. This intelligent scheduling mechanism significantly improves the efficiency and reliability of the entire laboratory testing process.

[0022] Second, this invention establishes a model dynamic optimization module, which enables continuous iteration and improvement of the traceability decision algorithm. This module can periodically analyze historical calculation data and system operation results, and automatically adjust the policy network weights and key mathematical formula parameters in the multi-objective deep deterministic policy gradient algorithm. This dynamic optimization mechanism ensures that the traceability system can continuously evolve with changes in the laboratory environment and testing needs, and always maintain the optimal operating state. This mechanism significantly improves the adaptability and stability of the system, providing the laboratory with a reliable and efficient traceability solution.

[0023] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0025] Figure 1 This is a flowchart illustrating the overall system architecture and data interaction process.

[0026] Figure 2 Flowchart of core steps for enhancing intelligent decision-making through learning. Detailed Implementation

[0027] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0028] Example 1: End-to-End Traceability Scenario for Tumor Marker Detection in Medical Testing Laboratories

[0029] This embodiment is applied to the clinical laboratory department of a tertiary hospital. This department processes a large number of tumor marker tests daily from inpatients and outpatients, involving multiple indicators such as carcinoembryonic antigen (CEA) and alpha-fetoprotein (AFP). Based on the system of this invention, end-to-end traceability management can be achieved from sample reception and instrument testing to data output. This ensures the traceability of test results to meet the needs of clinical diagnosis and treatment evaluation, while dynamically balancing system resource consumption to avoid data loss or delay due to resource overload during peak periods. Specific steps are as follows... Figure 1 As shown.

[0030] 1. System operation during the sample receiving phase

[0031] The peak sample receiving period is from 8:00 AM to 10:00 AM daily. During this time, the system first activates the status data acquisition module. This module uses a sample progress sensor to collect real-time data on the coverage of traceability nodes, specifically including patient basic information registration, sample number allocation, sample-patient information association, overall system resource load rate of receiving personnel, sample information entry, and real-time CPU utilization for data format conversion. Simultaneously, this module also collects the original hash value and data generation timestamp for each sample traceability data set, including the patient information entry time and sample number generation time. All collected status data is transmitted to the reinforcement learning intelligent decision-making module via the MQTT protocol, with transmission latency strictly controlled within 100ms to ensure data real-time performance.

[0032] After receiving data, the reinforcement learning intelligent decision-making module processes the data using a multi-objective deep deterministic policy gradient algorithm. First, it quantifies the overall occupancy of storage and computing resources using the system's comprehensive resource load rate formula. At this point, due to the concentrated receipt of samples, the computing resource load rate has increased but has not exceeded the threshold, while the storage resource load rate is at a moderate level. The system's comprehensive resource load rate formula is: ,in, for The overall resource load rate of the real-time system; for Real-time storage resource load rate, i.e. The ratio of the current used storage capacity to the total storage capacity; for Calculate resource load rate in real time, i.e. The module measures the real-time utilization of CPU or GPU; 0.6 is the weighting coefficient for storage resource load rate, and 0.4 is the weighting coefficient for computational resource load rate. Subsequently, the module adjusts the target weights for traceability integrity and resource consumption control using a dynamic weighting coefficient formula. Considering that tumor marker detection results directly affect subsequent treatment plans, ensuring the integrity of the traceability chain for each sample is prioritized; therefore, the target weight for traceability integrity is set at a high level. Next, the module evaluates the current traceability node coverage, data accuracy, and timeliness using a multi-dimensional traceability integrity formula, and then analyzes the impact of current resource consumption on the system using a resource consumption elasticity penalty formula. The resource consumption elasticity penalty formula is as follows: ,in, The resource consumption elasticity penalty value at time t, with a value ranging from 0 to 1; This is the storage resource elasticity coefficient, with a value of 0.4. To calculate the resource elasticity coefficient, a value of 0.6 is used; This represents the maximum load rate of storage resources, with a value of 100%, which is the load rate when storage resources are fully utilized. This is the maximum load rate of computing resources, set to 100%, which is the load rate when the CPU or GPU is fully utilized. Let t be the storage resource load rate; The resource load at time t is calculated, and the final traceability scheduling decision is generated: a full-data storage mode is adopted to completely store all traceability data in the sample receiving stage. Simultaneously, the traceability nodes corresponding to suspected tumor patient samples are set as high-priority scheduling objects, prioritizing the completion of traceability data entry and storage association for this type of sample. Figure 2 As shown.

[0033] Upon receiving the aforementioned decision, the traceability scheduling execution module immediately performs a full-scale evidence storage mode switch operation, completely recording all traceability data from the receipt stage of all patient samples into the system, including patient name, sample number, receipt time, and receiving personnel. Simultaneously, according to high-priority scheduling rules, it prioritizes the traceability nodes of suspected tumor patient samples, ensuring that the traceability data and sample association are completed within 15 minutes of receipt for these samples. After the operation is complete, the module feeds back the execution results to the reinforcement learning intelligent decision-making module, including whether full-scale evidence storage was successful and whether high-priority node scheduling was completed on time. Simultaneously, it transmits the evidence storage index of each sample's traceability data to the hybrid architecture storage and verification module.

[0034] After receiving the evidence index, the hybrid architecture storage and verification module performs storage and verification based on a blockchain-cloud hybrid architecture. Specifically, the full traceability data for suspected cancer patient samples is stored on the Hyperledger Fabric blockchain, leveraging the blockchain's immutability to ensure the security of critical sample traceability data. For ordinary patient samples, the full traceability data is stored on Alibaba Cloud OSS to meet high-capacity storage needs and reduce storage costs. Subsequently, the module performs hash verification on all stored traceability data. By comparing the original hash value obtained during the collection phase with the hash value of the stored data, it confirms that the data has not been tampered with. After completing the accuracy verification, the verification result is fed back to the reinforcement learning intelligent decision-making module. The verification result includes information such as data accuracy, specific storage location, and verification time, providing accurate traceability basis for subsequent detection processes.

[0035] 2. System Operation During Instrument Testing Phase

[0036] After sample reception, the instrument detection phase begins, primarily using a chemiluminescence immunoassay analyzer to detect tumor markers in the samples. During this time, the status data acquisition module continuously operates, monitoring the computing resource load rate of the chemiluminescence immunoassay analyzer's supporting computing equipment via the edge computing node resource monitoring submodule—that is, the real-time GPU utilization rate for processing real-time optical signal data and generating detection curves. It also monitors the storage resource load rate of the detection process data via the cloud storage server resource monitoring submodule—that is, the percentage of server capacity used for storing raw optical signal data, reagent batch numbers, instrument operating status, etc. Simultaneously, the module collects traceability node coverage data via a sample progress sensor, including whether the detection start time record, instrument number registration, reagent addition record, and detection signal acquisition nodes are completely covered. It also collects the raw hash value and data generation timestamp for each step of the detection process, such as the detection start timestamp and reagent addition completion timestamp, and transmits all data in real-time to the reinforcement learning intelligent decision-making module via the MQTT protocol.

[0037] The reinforcement learning intelligent decision-making module is based on a multi-objective deep deterministic policy gradient algorithm. It analyzes the comprehensive reward value of the current scheduling action using a dynamic multi-objective reward function formula, which is as follows: ,in, Let be the dynamic multi-objective reward value at time t, with a value range from -1 to 1; Let t be the weight of the traceability integrity target. The weight of the resource consumption control target at time t; The score for the completeness of multi-dimensional source tracing at time t; Let t be the resource consumption elasticity penalty value; The value is the action impact coefficient, which is 0.1 when performing full evidence storage, 0.3 when performing fingerprint evidence storage, and 0.2 when performing high-priority node scheduling. The reward value represents the incremental improvement in traceability integrity at time t, taking into account both the improvement in traceability integrity and the control of resource consumption. The optimal value of different candidate actions is then evaluated using a multi-objective action value function formula. Candidate actions include full data storage, fingerprint storage, high-priority node scheduling, and low-priority node delayed scheduling. The multi-objective action value function formula is as follows: ,in, for The optimal action value of performing an action at a given time state; for System status at all times; for Actions that are performed at all times; for Real-time dynamic multi-objective reward value; This is a discount factor with a value of 0.9, used to weigh the importance of immediate rewards against future rewards; The candidate actions for time t+1 include four types of actions: full evidence storage, fingerprint evidence storage, high-priority node scheduling, and low-priority node delayed scheduling. State at time t+1 Next, execute candidate actions The optimal action value; for The stability coefficient of the state at any given time. The value is 1 when the fluctuation range between the system state at time t and time t is ≤5%, and 0.8 when the fluctuation range is >5%. At this time, due to the simultaneous operation of multiple chemiluminescence immunoassay analyzers entering peak testing periods, the computational resource load rate increases significantly. The overall resource load rate calculated using the system comprehensive resource load rate formula exceeds the medium threshold. The module then adjusts the weights using a dynamic weight coefficient formula, appropriately increasing the weight of the resource consumption control target. Simultaneously, it analyzes the impact of the timeliness of the test data using the traceability data timeliness decay factor formula to ensure that the time difference between the generation and storage of test data meets the requirements of clinical testing standards. Finally, a scheduling decision is generated. The traceability data timeliness decay factor formula is: ,in, The timeliness decay factor of the source data at time t, with a value ranging from 0 to 1; The difference between the time of generation of the traceability data and the time of evidence storage at time t; This is a time-sensitive threshold, ranging from 2 to 5 seconds. The attenuation coefficient ranges from 1.0 to 1.5: Switch to fingerprint storage mode, only store key fingerprint information and core indexes of the detection process data, such as the fingerprint corresponding to the peak of the optical signal and the time index of the key detection stage, to reduce storage resource consumption; set the traceability node corresponding to the sample in the key detection stage as a high-priority scheduling object, and set the traceability node of the sample that has completed sample preprocessing and is waiting for detection as a low-priority delayed scheduling object.

[0038] Upon receiving the decision, the traceability scheduling execution module immediately performs a fingerprint evidence storage mode switch operation, storing only the key data fingerprints and core indexes from each instrument's testing process, omitting the complete optical signal curves and reagent addition process logs. Simultaneously, it prioritizes the scheduling of traceability nodes for samples at critical testing stages, ensuring fingerprint evidence storage and node association are completed within 10 minutes of data generation at that stage. Low-priority nodes are processed centrally during instrument downtime, such as processing node data only after an instrument completes a batch of tests. After the operation is complete, the module feeds back the execution results to the reinforcement learning intelligent decision-making module, including the amount of fingerprint evidence storage data and the completion rate of high-priority node scheduling. It also transmits the evidence storage index of the testing process's traceability data to the hybrid architecture storage and verification module.

[0039] After receiving the evidence index, the hybrid architecture storage and verification module stores the fingerprint data and core index from the detection process in Alibaba Cloud OSS. Simultaneously, it additionally backs up the traceability fingerprint data of samples from key detection stages to the Hyperledger Fabric blockchain, ensuring data security in critical detection stages. Subsequently, the module performs hash verification, comparing the original hash value of the detection process data with the associated hash value of the stored fingerprint data to confirm data consistency. The verification result is fed back to the reinforcement learning intelligent decision-making module, providing accurate detection and traceability evidence for subsequent data output stages.

[0040] 3. System operation during the data output phase

[0041] After the test is completed, the system generates a tumor marker test report, which includes information such as the name of the test indicator, the test result value, the reference range, the instrument number, and the personnel who performed the test, and then enters the data output stage. At this time, the status data acquisition module continues to collect data. It collects the storage resource load rate of the report storage through the cloud storage server resource monitoring submodule, that is, the percentage of the server capacity used for storing the full text of the report and the report review records; and it collects the computing resource load rate of report generation and push through the edge computing node resource monitoring submodule, that is, the real-time CPU usage rate for processing report format conversion and pushing to the hospital HIS system and electronic health record system. At the same time, the module collects the coverage of traceability nodes through the sample progress sensor, including whether the nodes such as report generation time record, reviewer signature record, report modification record, and HIS system push record are completely covered. It also collects the original hash value of the report data and the report generation timestamp, and transmits all data to the reinforcement learning intelligent decision-making module through the MQTT protocol.

[0042] The reinforcement learning intelligent decision-making module is based on a multi-objective deep deterministic policy gradient algorithm, combined with a multi-dimensional source traceability integrity formula to evaluate the comprehensive situation of the source traceability data in the report. The multi-dimensional source traceability integrity formula is as follows: ; ,in, The score for the completeness of multi-dimensional source tracing at time t, with a value ranging from 0 to 1; The weight for the coverage of traceability nodes ranges from 0.4 to 0.6. The weight for the product of accuracy and timeliness of traceability data ranges from 0.4 to 0.6. + =1; Let t be the coverage of source nodes; To ensure the accuracy of the source data at time t; The timeliness decay factor for traceability data at time t includes whether the traceability coverage includes all nodes from sample receipt to testing to reporting, whether the data accuracy is consistent with the testing data, and whether the timeliness meets the requirements for clinical report issuance. The system's comprehensive resource load rate formula calculates that the peak testing period has passed, and both storage and computing resource load rates have decreased to low levels. Subsequently, the module adjusts the weights using a dynamic weight coefficient formula, which is: ,in, The weight of the traceability integrity target at time t, with a value ranging from 0 to 1; The weight of the resource consumption control target at time t, with a value ranging from 0 to 1; This is the weighting adjustment coefficient, with a value ranging from 2.0 to 3.0; This represents the system's maximum tolerable load rate, set at 90%. Given the overall system resource load rate at time t, the weight of the traceability integrity target is adjusted to the highest level, and the final scheduling decision is generated: switch back to full evidence storage mode to fully store the full text of the test report and all traceability data related to sample reception and instrument testing; prioritize scheduling the traceability data of the report review node and the HIS system push node to ensure that the report review process is traceable and that clinicians can quickly query complete traceability information.

[0043] The traceability scheduling and execution module performs a full-data storage mode switch operation, completely storing the full text of each test report and the corresponding traceability data from sample receipt and instrument testing. Simultaneously, it prioritizes the traceability storage of the report review node, recording the reviewer's name, review time, review comments, and the traceability association with the HIS system push node. It links the report ID with the patient's HIS system ID and the traceability data storage index, ensuring that the report is pushed to the HIS system within 20 minutes of approval. After the operation is complete, the module feeds back the execution results to the reinforcement learning intelligent decision-making module, and simultaneously transmits the storage index of the report's traceability data to the hybrid architecture storage and verification module.

[0044] After receiving the evidence index, the hybrid architecture storage and verification module stores the full report data and related traceability data in Alibaba Cloud OSS, and stores the report review records and reviewers' electronic signature data in the Hyperledger Fabric blockchain to ensure the compliance and traceability of the report review process. Subsequently, the module confirms the consistency between the report data and the related traceability data through hash verification, and the verification result is fed back to the reinforcement learning intelligent decision-making module to complete the traceability closed loop of the entire tumor marker detection process.

[0045] 4. Periodic optimization of the model dynamic optimization module

[0046] Throughout the entire process described above, the reinforcement learning intelligent decision-making module transmits historical computational data from each stage to the model dynamic optimization module in real time. This historical data includes resource load changes during peak and off-peak periods of sample reception, traceability integrity scores under different scheduling decisions, hash verification pass rates, and other information. This module iteratively optimizes the parameters of mathematical formulas such as the policy network weights of the multi-objective deep deterministic policy gradient algorithm, the system comprehensive resource load rate formula, the dynamic weight coefficient formula, the dynamic multi-objective reward function formula, and the multi-dimensional traceability integrity formula, based on historical data and an iterative optimization cycle of once every 24 hours. For example, it adjusts the weight adaptation of storage and computing resources in the system comprehensive resource load rate formula according to the resource load patterns at different times of the day; and it optimizes the weight adjustment logic of traceability integrity and resource consumption in the dynamic weight coefficient formula based on changes in the number of suspected tumor patients. After optimization, the module pushes the new parameters to the reinforcement learning intelligent decision-making module to update its data processing logic, ensuring that the traceability scheduling of the subsequent tumor marker detection process is more adapted to the resource fluctuations and changes in clinical testing needs of the laboratory.

[0047] In summary, in the entire process of tumor marker detection in medical testing laboratories, the system of this invention achieves efficient traceability management through the collaborative operation of five modules. The status data acquisition module collects resource load and traceability-related data in real time at each stage of sample reception, instrument testing, and data output, ensuring data real-time performance. The reinforcement learning intelligent decision-making module, based on a multi-objective deep deterministic strategy gradient algorithm and combined with the system's comprehensive resource load rate formula, generates scheduling decisions adapted to specific scenarios, prioritizing the integrity of traceability for critical clinical samples. The traceability scheduling and execution module switches storage modes and schedules node priorities as needed. The hybrid architecture storage and verification module ensures data security and accuracy through a blockchain-cloud architecture. The model dynamic optimization module iterates parameters periodically. The system not only meets the clinical need for traceable test results but also dynamically balances resource consumption, effectively adapting to peak sample fluctuations, and significantly improving traceability efficiency and reliability.

[0048] Example 2: Scenario of full-process traceability for heavy metal detection in soil in an environmental monitoring laboratory

[0049] This embodiment is applied to a provincial environmental monitoring station laboratory. The laboratory is mainly responsible for heavy metal testing of soil in farmland and industrial park areas within its jurisdiction. The testing indicators include lead, mercury, cadmium, chromium, etc. It needs to achieve full-process traceability from sample reception and instrument testing to data output in order to meet the needs of environmental protection departments for soil pollution control assessment and supervision, while adapting to the system resource scheduling needs under different seasonal sample volume fluctuations.

[0050] 1. System operation during the sample receiving phase

[0051] During the annual rainy season, the volume of soil samples received by the laboratory increases by 60% compared to the dry season due to the potential migration of heavy metals caused by rainwater erosion. At this time, the system activates the status data acquisition module. This module collects real-time data on the coverage of traceability nodes through a sample progress sensor, including the coverage of nodes such as sampling point name registration, sampling time recording, sampling personnel information entry, sampling depth recording, transport vehicle number registration, temperature and humidity recording during transportation, and laboratory reception registration. It also collects the storage resource load rate through a cloud storage server resource monitoring submodule, i.e., the percentage of server capacity currently used for storing sampling point map data, transportation records, and sampling personnel qualification information. Furthermore, it collects the computing resource load rate through an edge computing node resource monitoring submodule, i.e., the real-time CPU usage rate for processing sample information entry and associating sampling point data with sample numbers. Simultaneously, the module collects the original hash value and data generation timestamp for each soil sample traceability data, such as the sampling completion timestamp and transportation departure and arrival timestamps, and transmits all status data to the reinforcement learning intelligent decision-making module via the MQTT protocol, with transmission latency controlled within 100ms.

[0052] The reinforcement learning intelligent decision-making module processes data based on a multi-objective deep deterministic policy gradient algorithm. First, it quantifies the overall occupancy of storage and computing resources using a formula for the system's comprehensive resource load rate. Due to a surge in sample volume during the rainy season, both storage and computing resource load rates are at high levels. Next, the module adjusts the weights of the traceability integrity target and the resource consumption control target using a dynamic weighting coefficient formula. Considering the need to avoid system lag due to resource overload, the resource consumption control target weight is set higher than the traceability integrity target weight. Subsequently, the module analyzes the timeliness of sampled data using a traceability data timeliness decay factor formula, ensuring that the time difference between sampling completion and laboratory reception and evidence storage meets soil testing standards. Then, it analyzes the impact of current resource consumption on the system using a resource consumption elasticity penalty formula, ultimately generating a traceability scheduling decision: adopting a fingerprint evidence storage mode, storing only key information fingerprints and core indexes from the sample reception stage, such as sampling point number, sampling time, and transport vehicle ID fingerprints, reducing the amount of stored data; and setting the traceability nodes corresponding to soil samples around the industrial park as high-priority scheduling objects, as the soil in this area has a high risk of heavy metal pollution and the data requires key monitoring.

[0053] Upon receiving the decision, the traceability scheduling and execution module immediately performs a fingerprint evidence storage mode switch operation, inputting the fingerprints and core indexes of key information from all soil sample receiving stages into the system. Simultaneously, following high-priority scheduling rules, it prioritizes processing traceability nodes for soil samples from the vicinity of industrial parks, ensuring that traceability data and sample association are completed within 30 minutes of laboratory receipt. Traceability nodes for ordinary farmland soil samples are processed centrally during periods of low resource availability, such as midday when sample reception volume decreases. After completion, the module feeds back the execution results to the reinforcement learning intelligent decision-making module, including the integrity of the fingerprint evidence storage data and the completion status of high-priority node scheduling. It also transmits the traceability data evidence index from the sample receiving stage to the hybrid architecture storage and verification module.

[0054] After receiving the evidence index, the hybrid architecture storage and verification module stores the fingerprint data of soil samples from the surrounding industrial park on the Hyperledger Fabric blockchain, ensuring the immutability of traceability data for samples in high-risk areas and meeting environmental regulatory requirements. The fingerprint data of ordinary farmland soil samples is stored on Alibaba Cloud OSS. Subsequently, the module performs hash verification on all stored traceability data. By comparing the original hash value from the collection stage with the associated hash value of the stored fingerprint data, the module confirms the data accuracy. The verification result is fed back to the reinforcement learning intelligent decision-making module, providing a reliable traceability foundation for subsequent testing.

[0055] 2. System Operation During Instrument Testing Phase

[0056] After sample reception, the instrument testing phase begins, involving sample pretreatment and instrument testing. Sample pretreatment includes digestion and extraction to remove interference from impurities. Instrument testing utilizes inductively coupled plasma mass spectrometry (ICP-MS) to detect heavy metal content. During this process, the status data acquisition module continuously operates, monitoring the computing resource load of the ICP-MS-compatible computing equipment (i.e., the real-time GPU usage for processing elemental signal intensity data and generating concentration calibration curves) through the edge computing node resource monitoring submodule. It also monitors the storage resource load of the testing process data (i.e., the percentage of server capacity used for storing elemental signal spectra, pretreatment reagent batch numbers, and instrument calibration records) through the cloud storage server resource monitoring submodule. Simultaneously, the module uses a sample progress sensor to collect traceability node coverage data, including pretreatment start and end times, pretreatment equipment registration numbers, instrument calibration time records, and whether the detection signal acquisition nodes are fully covered. It also collects the raw hash values ​​and data generation timestamps of the testing process data, such as digestion completion timestamps and instrument calibration completion timestamps, and transmits all data in real-time to the reinforcement learning intelligent decision-making module via the MQTT protocol.

[0057] The reinforcement learning intelligent decision-making module, based on a multi-objective deep deterministic policy gradient algorithm, analyzes the comprehensive reward value of the current scheduling action using a dynamic multi-objective reward function formula, comprehensively considering the effect of improving traceability integrity and controlling resource consumption. It then evaluates the optimal value of different candidate actions using a multi-objective action value function formula. At this point, due to the completion of calibration for some ICP-MS instruments and the onset of peak testing periods, the computational resource load rate increases, but storage resources remain at a moderate level due to the use of fingerprint evidence storage. The overall resource load rate is calculated to be moderate using the system's comprehensive resource load rate formula. The module then uses a dynamic weighting coefficient formula to balance the weights of the traceability integrity target and the resource consumption control target. Combining this with a multi-dimensional traceability integrity formula, it evaluates the coverage and accuracy of traceability data in the testing process, ultimately generating the following scheduling decision: maintain the fingerprint evidence storage mode and continue to control storage resource usage; set the traceability nodes corresponding to instrument calibration records as high-priority scheduling objects, as calibration data directly affects the accuracy of testing results and requires priority evidence storage; and set the traceability nodes of samples that have completed preprocessing and are awaiting testing as regular priority.

[0058] After receiving the decision, the traceability scheduling and execution module continues to maintain the fingerprint evidence storage mode, storing key fingerprints of element signals and indexes of key preprocessing parameters during the testing process. Simultaneously, it prioritizes processing traceability nodes for instrument calibration records, associating the fingerprints of each instrument's calibration time, calibration standard material batch number, calibration results, and other data with the instrument number for evidence storage, ensuring evidence storage is completed within 20 minutes of calibration data generation. After the operation is complete, the module feeds back the execution results to the reinforcement learning intelligent decision-making module, including the calibration node evidence storage success rate and the integrity of the test data fingerprints. Simultaneously, it transmits the evidence storage index of the traceability data in the testing process to the hybrid architecture storage and verification module.

[0059] After receiving the evidence index, the hybrid architecture storage and verification module stores the fingerprint data of the testing process in Alibaba Cloud OSS and the fingerprint data of the instrument calibration record in the Hyperledger Fabric blockchain, ensuring the compliance of the calibration data and meeting the testing qualification requirements. Subsequently, the module confirms the consistency between the test fingerprint data and the original test signal data, and the consistency between the calibration record fingerprint and the original calibration data through hash verification. The verification results are fed back to the reinforcement learning intelligent decision-making module to ensure the accuracy of the test data traceability chain.

[0060] 3. System operation during the data output phase

[0061] After the testing is completed, the system generates a soil heavy metal testing report. The report includes the heavy metal concentration values ​​of samples from each sampling point, whether they exceed the "Soil Environmental Quality Standard for Agricultural Land Soil Pollution Risk Control", the testing method standard number, and other information, and then enters the data output stage. At this time, the status data acquisition module collects relevant data, monitors the storage resource load rate of the report storage through the cloud storage server resource monitoring submodule, and monitors the computing resource load rate of report generation and push through the edge computing node resource monitoring submodule. The push target is the provincial environmental protection department's regulatory platform. Simultaneously, the module collects the coverage of traceability nodes through the sample progress sensor, including whether the nodes such as report generation time records, data reviewer information entry, review opinion records, and regulatory platform push records are completely covered. It also collects the original hash value of the report data and the report generation timestamp, and transmits all data to the reinforcement learning intelligent decision-making module through the MQTT protocol.

[0062] The reinforcement learning intelligent decision-making module, based on a multi-objective deep deterministic policy gradient algorithm and combined with a multi-dimensional traceability integrity formula, evaluates the comprehensive situation of the traceability data in the report to confirm whether it covers the entire process of sampling, transportation, testing, and calibration. The module calculates using the system's comprehensive resource load rate formula that the peak testing period has passed and the resource load rate has dropped to a low level. Subsequently, the module adjusts the traceability integrity target weight to the highest level using a dynamic weight coefficient formula. Considering that the report needs to be submitted to the environmental protection department for soil pollution remediation decisions, and that full-chain traceability must be guaranteed, the module ultimately generates the following scheduling decision: switch to full-data storage mode, completely storing the full report and all related traceability data from the sampling, transportation, testing, and calibration stages; prioritize scheduling traceability data from the report review node and the environmental supervision platform push node.

[0063] The source tracing scheduling and execution module performs a full-data storage mode switch operation, completely storing the full text of each report along with associated sampling point information, transportation records, testing process data, instrument calibration records, and other source tracing data. Simultaneously, it prioritizes the source tracing and storage of the report review node, recording the reviewer's name, review time, review conclusion, and the source tracing association with the environmental supervision platform's push node. It associates the report ID with the supervision platform's sampling point ID and source tracing data storage index, ensuring that the report is pushed to the environmental protection department's supervision platform within one hour of approval. After the operation is complete, the module feeds back the execution results to the reinforcement learning intelligent decision-making module, and simultaneously transmits the source tracing data storage index from the reporting stage to the hybrid architecture storage and verification module.

[0064] After receiving the evidence index, the hybrid architecture storage and verification module stores the full report data and related traceability data in Alibaba Cloud OSS, and stores the report review records and reviewer signatures in the Hyperledger Fabric blockchain to ensure the compliance of the report review and facilitate verification by environmental protection departments. Subsequently, the module confirms the consistency between the report data and the related traceability data through hash verification, and the verification result is fed back to the reinforcement learning intelligent decision-making module to complete the traceability management of the entire process of soil heavy metal testing.

[0065] 4. Periodic optimization of the model dynamic optimization module

[0066] Throughout the entire soil heavy metal testing process, the reinforcement learning intelligent decision-making module transmits historical computational data to the model dynamic optimization module. This historical data includes information such as resource load differences between rainy and dry seasons, the scheduling effectiveness of samples from different pollution risk areas, changes in source tracing integrity scores, and hash verification pass rates. This module iterates and optimizes the parameters of the policy network weights of the multi-objective deep deterministic strategy gradient algorithm, as well as mathematical formulas such as the system's comprehensive resource load rate formula, dynamic weight coefficient formula, multi-dimensional source tracing integrity formula, and dynamic multi-objective reward function formula, based on historical data and following a 24-hour iteration optimization cycle. For example, considering the larger sample size during the rainy season and the smaller sample size during the dry season, the parameters of the system's comprehensive resource load rate formula are optimized to adapt to resource fluctuations in different seasons; based on the difference in regulatory priorities between industrial park and ordinary farmland samples, the weight logic of source tracing integrity in the dynamic weight coefficient formula is adjusted. The optimized parameters are then pushed to the reinforcement learning intelligent decision-making module, enabling it to better adapt to the seasonal work needs of environmental monitoring laboratories and improve the source tracing efficiency and resource utilization rationality of the entire soil testing process.

[0067] In summary, in the context of heavy metal testing in soil samples at environmental monitoring laboratories, this invention's system precisely adapts to seasonal work requirements. The status data acquisition module captures resource load and source tracing data during periods of surge in samples during the rainy season, providing a foundation for decision-making. The reinforcement learning intelligent decision-making module, relying on core algorithms and related formulas, formulates fingerprint evidence storage and high-priority node scheduling strategies, prioritizing the source tracing of high-pollution-risk samples. The source tracing scheduling and execution module flexibly switches evidence storage modes, balancing resource consumption and data recording needs. The hybrid architecture storage and verification module stores data hierarchically, balancing security and cost. The model dynamic optimization module adjusts parameters periodically to adapt to resource fluctuations during rainy and dry seasons. The system not only meets the environmental regulatory requirements for full-process source tracing of soil testing data but also achieves efficient resource utilization, ensuring compliance and efficiency in source tracing.

[0068] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. Laboratory testing full-process intelligent traceability system based on reinforcement learning, characterized in that, The system comprises: a state data acquisition module for acquiring resource load data and traceability integrity associated data in the sample receiving stage, instrument detection stage and data output stage, and transmitting the acquired state data to the enhanced learning intelligent decision module; the enhanced learning intelligent decision module processes the received state data based on a multi-objective deep deterministic policy gradient algorithm, a system comprehensive resource load rate formula, a dynamic multi-objective reward function formula, a multi-objective action value function formula, and a multi-dimensional traceability integrity formula and a traceability integrity increment formula to generate a traceability scheduling decision; simultaneously, the traceability scheduling decision is transmitted to the traceability scheduling execution module, and the historical calculation data generated in the data processing process is transmitted to the model dynamic optimization module; the traceability scheduling execution module performs traceability data full-amount storage or fingerprint storage mode switching operation and traceability node priority scheduling operation according to the received traceability scheduling decision; after the operation is completed, the execution result is fed back to the enhanced learning intelligent decision module, and the storage index of the traceability data is transmitted to the hybrid architecture storage and verification module; the hybrid architecture storage and verification module stores the traceability data corresponding to the storage index based on a blockchain-cloud hybrid architecture, and performs hash verification on the traceability data to complete accuracy verification; the hybrid architecture storage and verification module feeds back the verification result to the enhanced learning intelligent decision module; the model dynamic optimization module receives the historical calculation data transmitted by the enhanced learning intelligent decision module, iteratively optimizes the policy network weight of the multi-objective deep deterministic policy gradient algorithm and the parameters of the mathematical calculation formula, and pushes the optimized parameters to the enhanced learning intelligent decision module to update the data processing logic thereof.

2. The reinforcement learning based intelligent lab test full-process traceability system according to claim 1, characterized in that, The resource load data acquired by the state data acquisition module includes storage resource load rate and calculation resource load rate, and the traceability integrity associated data includes traceability node coverage, traceability data original hash value and data generation timestamp; The state data acquisition module acquires the traceability node coverage through a sample progress sensor, acquires the storage resource load rate through a cloud storage server resource monitoring sub-module, acquires the calculation resource load rate through an edge computing node resource monitoring sub-module, and transmits the state data by using MQTT protocol, with a transmission delay of ≤100 ms.

3. The reinforcement learning based intelligent lab test full-process traceability system according to claim 2, wherein, The system comprehensive resource load rate formula in the reinforcement learning intelligent decision module is: Wherein, is the system comprehensive resource load rate at the moment; is the storage resource load rate at the moment, that is, the ratio of the used capacity of the storage resource to the total capacity of the storage resource at the moment; is the computing resource load rate at the moment, that is, the real-time usage rate of the CPU or GPU at the moment; 0.6 is the weight coefficient of the storage resource load rate, and 0.4 is the weight coefficient of the computing resource load rate.

4. The reinforcement learning based intelligent lab test full flow traceability system according to claim 3, wherein, The multi-dimensional traceability integrity formula and the traceability integrity increment formula in the enhanced learning intelligent decision module are respectively: ; , wherein, is a multi-dimensional traceability integrity score at time t, and the value range is 0 to 1; is a traceability node coverage weight, and the value range is 0.4 to 0.6; is a weight of the product of traceability data accuracy and timeliness, and the value range is 0.4 to 0.6, and + =1; is a traceability node coverage at time t; is a traceability data accuracy at time t; is a traceability data timeliness decay factor at time t; is a traceability integrity increment at time t, and the value range is -1 to 1; is a multi-dimensional traceability integrity score at time t.

5. The reinforcement learning based intelligent laboratory testing full-process traceability system according to claim 4, characterized in that, The traceability data timeliness decay factor formula in the reinforcement learning intelligent decision module is: Wherein, The traceability data timeliness decay factor at time t, the value range is 0 to 1; The difference between the traceability data generation time and the storage time at time t; The timeliness threshold, the value range is 2 seconds to 5 seconds; The decay coefficient, the value range is 1.0 to 1.

5.

6. The reinforcement learning based intelligent laboratory testing full-process traceability system according to claim 5, characterized in that, The dynamic multi-target reward function formula in the reinforcement learning intelligent decision module is: Wherein, is a dynamic multi-target reward value at time t, and the value range is -1 to 1; is a trace integrity target weight at time t; is a resource consumption control target weight at time t; is a multi-dimensional trace integrity score at time t; is a resource consumption elasticity penalty value at time t; is an action influence coefficient, and the value is 0.1 when a full-amount evidence action is executed, the value is 0.3 when a fingerprint evidence action is executed, and the value is 0.2 when a high-priority node scheduling action is executed; is a trace integrity increment at time t.

7. The reinforcement learning based intelligent laboratory testing full-process traceability system according to claim 6, characterized in that, The dynamic weight coefficient formula in the reinforcement learning intelligent decision module is: wherein, is a trace integrity target weight at time t, and the value range is 0 to 1; is a resource consumption control target weight at time t, and the value range is 0 to 1; is a weight adjustment coefficient, and the value range is 2.0 to 3.0; is a maximum tolerable load rate of the system, and the value is 90%; is a system comprehensive resource load rate at time t.

8. The reinforcement learning based intelligent laboratory testing full-process traceability system according to claim 7, characterized in that, The resource consumption elasticity penalty formula in the reinforcement learning intelligent decision module is: Wherein, is the resource consumption elasticity penalty value at time t, and the value range is 0 to 1; is the storage resource elasticity coefficient, and the value is 0.4; is the computing resource elasticity coefficient, and the value is 0.6; is the maximum load rate of the storage resource, and the value is 100%, that is, the load rate when the storage resource is completely occupied; is the maximum load rate of the computing resource, and the value is 100%, that is, the load rate when the CPU or GPU is completely occupied; is the storage resource load rate at time t; is the computing resource load rate at time t.

9. The reinforcement learning based intelligent laboratory testing full-process traceability system according to claim 8, characterized in that, The multi-objective action value function formula in the reinforcement learning intelligent decision module is: Wherein, is the optimal action value of the action performed at the moment state; is the system state at the moment; is the action performed at the moment; is the dynamic multi-objective reward value at the moment; is a discount factor, taking a value of 0.9, used to weigh the importance of immediate rewards and future rewards; is the candidate action at the moment t+1, including four types of actions: full quantity storage, fingerprint storage, high priority node scheduling, and low priority node delay scheduling; is the state at the moment t+1 is the optimal action value of the candidate action performed at the moment t+1; is the state stability coefficient at the moment t+1, is 1 when the fluctuation amplitude of the system state at the moment t+1 and the system state at the moment t is ≤5%, and is 0.8 when the fluctuation amplitude is >5%.​​​​​​ 10. The reinforcement learning based intelligent lab testing full-process traceability system according to claim 2, wherein, When the traceability scheduling execution module performs full-amount storage / fingerprint storage mode switching operation, when the storage resource load rate > 80% or the calculation resource load rate > 80%, the fingerprint storage mode is switched; when the storage resource load rate < 40% and the calculation resource load rate < 40%, the full-amount storage mode is switched; the blockchain used by the hybrid architecture storage and verification module is Hyperledger Fabric blockchain, and the cloud storage used is Aliyun OSS; the iteration optimization period of the model dynamic optimization module is once every 24 hours, and the optimization objects include the policy network weight of the multi-objective deep deterministic policy gradient algorithm and the parameters of each mathematical calculation formula.

Citation Information

Patent Citations

  • Block chain-based traceability processing method and block chain distributed traceability system

    CN116362772A

  • Cloud-edge collaborative intelligent storage node dynamic deployment method and system

    CN120614249A