Laboratory detection whole-process intelligent traceability system based on reinforcement learning

By dynamically scheduling resources through reinforcement learning algorithms, the problem of data loss in laboratory testing traceability systems under resource competition is solved, achieving efficient and reliable traceability management and adapting to the needs of complex scenarios.

CN121073409AActive Publication Date: 2025-12-05连云港海关综合技术中心
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511630656.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2025-12-05
Estimated Expiration
2045-11-10

AI Technical Summary

Technical Problem

Existing laboratory testing traceability systems are prone to data loss or delayed evidence storage when resource competition is fierce. They lack intelligent decision-making mechanisms, cannot adapt to complex scenarios, and result in low traceability efficiency and waste of resources.

Method used

A laboratory testing intelligent traceability system based on reinforcement learning is adopted. It dynamically balances resource load and traceability integrity through a multi-objective deep deterministic strategy gradient algorithm, combined with dynamic weight adjustment and model iterative optimization, to achieve real-time resource scheduling and data storage mode switching.

Benefits of technology

It significantly improves the efficiency and reliability of laboratory testing processes, can adapt to changes in the laboratory environment, dynamically optimizes traceability strategies, and ensures the traceability of test results and the efficient use of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121073409A_ABST
    Figure CN121073409A_ABST
Patent Text Reader

Abstract

The invention discloses a laboratory detection whole-process intelligent traceability system based on reinforcement learning, and relates to the technical field of laboratory detection process intelligent management. A reinforcement learning intelligent decision module; a traceability scheduling execution module; a hybrid architecture storage and verification module; a model dynamic optimization module; the multi-target depth deterministic strategy gradient algorithm is introduced, the system comprehensive resource load rate formula, the dynamic multi-target reward function formula and the multi-target action value function formula are combined, dual optimization of resource utilization and traceability accuracy is achieved, the system can dynamically adjust the traceability strategy according to the real-time resource load condition, and the traceability accuracy is improved. The fingerprint evidence storage mode is automatically switched to reduce resource consumption when the load is high, the full-quantity evidence storage mode is adopted to ensure the integrity of the traceability data when the load is low, and the efficiency and the reliability of the whole process of laboratory detection are remarkably improved through the intelligent scheduling mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent management of laboratory detection process, in particular to a laboratory detection whole-process intelligent traceability system based on reinforcement learning. BACKGROUND

[0002] With the development of digitalization and automation of laboratory detection process, traceability management of detection data has become a key link to ensure the reliability of results. Traditional traceability systems rely on static rules or single indicators to trigger data archiving, which is difficult to dynamically balance resource consumption and traceability integrity. Especially when the sample volume surges or the instrument runs at high load, the system is prone to resource competition, resulting in loss or delayed archiving of traceability data, affecting the reliability and compliance of the detection process. In addition, the variability of the laboratory environment requires the traceability strategy to have real-time adaptability, and the existing technology lacks intelligent decision-making mechanisms to cope with complex scenarios.

[0003] The existing traceability technology has three defects: first, the fixed threshold value is used to trigger the mode switching of archiving, without considering the dynamic coupling relationship between computing resources and storage resources, which is prone to resource mismatch; second, the traceability priority scheduling relies on artificial preset rules and cannot be dynamically adjusted according to the real-time state of the node, which may miss critical traceability links; third, it lacks self-optimization ability driven by algorithms, and system parameters need to be manually adjusted, which is difficult to adapt to environmental changes in long-term laboratory operation. These deficiencies result in low traceability efficiency, serious resource waste, and inability to meet the needs of high-precision detection scenarios. SUMMARY

[0004] The purpose of the present application is to overcome the shortcomings of the prior art and provide a laboratory detection whole-process intelligent traceability system based on reinforcement learning. The present application dynamically balances resource load and traceability integrity through a multi-objective deep deterministic policy gradient algorithm, and the system automatically generates archiving mode switching and node scheduling decisions by real-time collection of resource and traceability data in the sample receiving, detection, and output stages. Combined with dynamic weight adjustment and model iteration optimization mechanism, the system can adapt to changes in the laboratory environment, significantly improve traceability efficiency and reliability, and is suitable for high-precision scenarios such as medical detection and environmental monitoring.

[0005] To solve the above technical problems, the present application provides the following technical solution: a laboratory detection whole-process intelligent traceability system based on reinforcement learning, which comprises:

[0006] A state data acquisition module for acquiring resource load data and traceability integrity associated data in the sample receiving stage, instrument detection stage, and data output stage, and transmitting the acquired state data to the reinforcement learning intelligent decision-making module;

[0007] The reinforcement learning intelligent decision-making module is based on a multi-objective deep deterministic policy gradient algorithm. It processes the received state data by combining the system's comprehensive resource load rate formula, dynamic multi-objective reward function formula, and multi-objective action value function formula to generate source tracing scheduling decisions. At the same time, the source tracing scheduling decisions are transmitted to the source tracing scheduling execution module, and the historical calculation data generated during the data processing is transmitted to the model dynamic optimization module.

[0008] Source tracing scheduling and execution module: Based on the received source tracing scheduling decision, it performs the full-data storage or fingerprint storage mode switching operation for source tracing data, as well as the source tracing node priority scheduling operation; after the operation is completed, it feeds back the execution result to the reinforcement learning intelligent decision module, and transmits the storage index of the source tracing data to the hybrid architecture storage and verification module;

[0009] Hybrid architecture storage and verification module: Based on the blockchain-cloud hybrid architecture, the module stores the traceability data corresponding to the evidence index and performs hash verification on the traceability data to complete the accuracy verification. The hybrid architecture storage and verification module feeds back the verification results to the reinforcement learning intelligent decision-making module.

[0010] Model Dynamic Optimization Module: This module receives historical computational data from the reinforcement learning intelligent decision-making module, iteratively optimizes the policy network weights and mathematical calculation formulas of the multi-objective deep deterministic policy gradient algorithm, and pushes the optimized parameters to the reinforcement learning intelligent decision-making module to update its data processing logic.

[0011] Furthermore, the resource load data collected by the status data acquisition module includes storage resource load rate and computing resource load rate, and the collected traceability integrity association data includes traceability node coverage, original hash value of traceability data, and data generation timestamp; the status data acquisition module collects traceability node coverage through a sample progress sensor, collects storage resource load rate through a cloud storage server resource monitoring submodule, and collects computing resource load rate through an edge computing node resource monitoring submodule, and the status data acquisition module uses the MQTT protocol to transmit status data with a transmission latency of ≤100ms.

[0012] Furthermore, the formula for the overall system resource load rate in the reinforcement learning intelligent decision-making module is: ,in, for The system's overall resource load rate at any given time is used to quantify the overall utilization of the system's storage and computing resources. for Real-time storage resource load rate, i.e. The ratio of the used storage capacity to the total storage capacity at any given time is obtained by the status data acquisition module through the cloud storage server resource monitoring submodule. for The resource load rate is calculated at the moment, that is, The real-time usage rate of CPU or GPU at the moment is collected by the edge computing node resource monitoring submodule through the state data collection module; 0.6 is the weight coefficient of the storage resource load rate, and 0.4 is the weight coefficient of the computing resource load rate.

[0013] Further, the traceability data timeliness decay factor formula in the reinforcement learning intelligent decision module is: , wherein, is the traceability data timeliness decay factor at the moment t, the value range is 0 to 1, and is used to quantify the influence of the time difference from generation to storage of the traceability data on the integrity; is the difference between the generation time and the storage time of the traceability data at the moment t, the unit is second, the generation time is automatically labeled by the detection instrument when the data is generated, and the storage time is automatically recorded by the hybrid architecture storage and verification module when the data is written; is the timeliness threshold, the value range is 2 seconds to 5 seconds; is the decay coefficient, the value range is 1.0 to 1.5, which is determined by the model dynamic optimization module based on the timeliness and traceability value correlation analysis of 50,000 different delay data.

[0014] Further, the dynamic weight coefficient formula in the reinforcement learning intelligent decision module is: , wherein, is the traceability integrity target weight at the moment t, the value range is 0 to 1, and is used to represent the degree of priority to ensure traceability integrity in the current scenario; is the resource consumption control target weight at the moment t, the value range is 0 to 1, and is complementary to , which is used to represent the degree of priority to control resource consumption in the current scenario; is the weight adjustment coefficient, the value range is 2.0 to 3.0; is the maximum tolerated load rate of the system, the value is 90%; is the system comprehensive resource load rate at the moment t.

[0015] Further, the multi-dimensional traceability integrity formula and the traceability integrity increment formula in the reinforcement learning intelligent decision module are respectively: ; , wherein, is the multi-dimensional traceability integrity score at the moment t, the value range is 0 to 1, and is used to comprehensively evaluate the coverage, accuracy and timeliness of the traceability data; is the traceability node coverage weight, the value range is 0.4 to 0.6; is the weight of the product of traceability data accuracy and timeliness, the value range is 0.4 to 0.6, and + = 1; is the traceability of the traceability node at time t, that is, the ratio of the number of stored traceability nodes to the total number of traceability nodes required to be stored at time t; is the trace data accuracy at time t, that is, the ratio of the number of data sets matched by the original hash value of the trace data and the stored hash value to the total number of data sets, obtained by comparing the trace data original hash value and the stored hash value by the hybrid architecture storage and verification module; is the trace data timeliness decay factor at time t; is the trace integrity comprehensive score at time t, with a value range of 0-1; is is the multi-dimensional trace integrity score at time t, stored by the historical cache unit of the enhanced learning intelligent decision module, with a storage period consistent with the system data processing period.

[0016] Further, the resource consumption elasticity penalty formula in the enhanced learning intelligent decision module is: wherein, is the resource consumption elasticity penalty value at time t, with a value range of 0-1, used to quantify the influence of the current resource consumption on the system; is the storage resource elasticity coefficient, with a value of 0.4, used to represent the influence weight of storage resource consumption on the system; is the computing resource elasticity coefficient, with a value of 0.6, used to represent the influence weight of computing resource consumption on the system; is the maximum load rate of storage resources, with a value of 100%, that is, the load rate when the storage resources are fully occupied; is the maximum load rate of computing resources, with a value of 100%, that is, the load rate when the CPU or GPU is fully occupied; is the storage resource load rate at time t; is the computing resource load rate at time t.

[0017] Further, the dynamic multi-target reward function formula in the enhanced learning intelligent decision module is: wherein, is the dynamic multi-target reward value at time t, with a value range of -1 to 1, used to provide a feedback signal for the multi-target deep deterministic policy gradient algorithm; is the trace integrity target weight at time t; is the resource consumption control target weight at time t; is the multi-dimensional trace integrity score at time t; is the resource consumption elasticity penalty value at time t; is the action influence coefficient, with a value of 0.1 when the full amount of trace evidence action is executed, a value of 0.3 when the fingerprint trace evidence action is executed, and a value of 0.2 when the high-priority node scheduling action is executed; The traceability integrity increment at time t is traced.

[0018] Further, the multi-objective action value function formula in the reinforcement learning intelligent decision module is: , wherein is the optimal action value of the action performed at time t. is the optimal action value of the action performed at time t. is the system state at time t. is the system state at time t. is the action performed at time t. is the action performed at time t. is the dynamic multi-objective reward value at time t. is the dynamic multi-objective reward value at time t. is a discount factor, which is 0.9, used to weigh the importance of immediate rewards and future rewards. is the candidate action at time t+1, including full storage, fingerprint storage, high-priority node scheduling, and low-priority node delay scheduling. is the state at time t+1 is the optimal action value of the candidate action performed at time t+1. is the optimal action value of the candidate action performed at time t+1. is the state stability coefficient at time t+1, wherein when the fluctuation amplitude of the system state at time t+1 and the system state at time t is ≤5%, the value is 1, and when the fluctuation amplitude is >5%, the value is 0.8. The fluctuation amplitude is calculated by comparing the change proportions of resource load rate and traceability node coverage in the system states at the two times.

[0019] Further, when the reinforcement learning intelligent decision module performs full storage / fingerprint storage mode switching operation, when the storage resource load rate >80% or the computing resource load rate >80%, it switches to the fingerprint storage mode; when the storage resource load rate <40% and the computing resource load rate <40%, it switches to the full storage mode. The blockchain used by the hybrid architecture storage and verification module is Hyperledger Fabric blockchain, which is used to store full data of high-priority traceability data. The cloud storage used is Aliyun OSS, which is used to store SHA-256 fingerprints and key indexes of low-priority traceability data. The iteration optimization period of the model dynamic optimization module is once every 24 hours, and the optimization objects include the policy network weights of the multi-objective deep deterministic policy gradient algorithm and the parameters of each mathematical calculation formula.

[0020] Compared with the prior art, the laboratory detection full-process intelligent traceability system based on reinforcement learning has the following beneficial effects:

[0021] One, the present application realizes the double optimization of resource utilization and traceability accuracy by introducing a multi-objective deep deterministic policy gradient algorithm, combining a system comprehensive resource load rate formula, a dynamic multi-objective reward function formula and a multi-objective action value function formula. The system can dynamically adjust the traceability strategy according to the real-time resource load condition, automatically switch to the fingerprint storage mode to reduce resource consumption under high load, and use the full storage mode to ensure the integrity of the traceability data under low load. This intelligent scheduling mechanism significantly improves the efficiency and reliability of the whole process of laboratory detection.

[0022] Two, the present application realizes the continuous iteration and improvement of the traceability decision algorithm by setting up a model dynamic optimization module. This module can regularly analyze historical calculation data and system operation effect, automatically adjust the strategy network weight in the multi-objective deep deterministic policy gradient algorithm and the parameters of the key mathematical formula. This dynamic optimization mechanism ensures that the traceability system can evolve with the changes of laboratory environment and detection demand, and always maintains the optimal operation state. This mechanism significantly improves the adaptability and stability of the system, and provides a reliable and efficient traceability solution for the laboratory.

[0023] Other advantages, objects and features of the present application will be in part apparent and in part pointed out hereinafter in the specification, and it is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application, as claimed. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creating any creative labor.

[0025] Figure 1 The system overall architecture and data interaction flowchart;

[0026] Figure 2 The reinforcement learning intelligent decision core step flowchart. DETAILED DESCRIPTION

[0027] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the specific embodiments, structures, features and effects according to the present application will be described in detail below with reference to the drawings and preferred embodiments.

[0028] Embodiment one: whole process traceability scene of tumor marker detection in medical detection laboratory

[0029] The embodiment is applied to a clinical laboratory of a certain third-grade class-A hospital. The department needs to process a large number of tumor marker detection requirements of inpatient and outpatient patients every day, and involves multiple index detection such as carcinoembryonic antigen and alpha-fetal protein. Relying on the system, the whole-process traceability management from sample receiving, instrument detection to data output can be realized, which not only guarantees the traceability of detection results to meet the needs of clinical diagnosis and treatment evaluation, but also dynamically balances the consumption of system resources, avoids the loss or delay of traceability data due to resource overload during peak period, and the specific steps are as shown in Figure 1 .

[0030] 1. Sample receiving stage system operation

[0031] The sample receiving peak period is from 8:00 to 10:00 every morning. The system first starts the state data acquisition module. The module acquires the traceability node coverage in real time through the sample progress sensor, including patient basic information registration, sample number allocation, sample and patient information association, sample information input, CPU real-time usage rate of data format conversion, sample information input, data format conversion, CPU real-time usage rate. At the same time, the module also acquires the original hash value and data generation timestamp of each sample traceability data, wherein the data generation timestamp includes patient information input time and sample number generation time. All collected state data is transmitted to the enhanced learning intelligent decision module through the MQTT protocol, and the transmission delay is strictly controlled within 100 ms to ensure real-time data.

[0032] After receiving the data, the enhanced learning intelligent decision module processes the data with the multi-objective deep deterministic policy gradient algorithm as the core. First, the overall occupation of the current storage and calculation resources is quantified by combining the system comprehensive resource load rate formula. At this time, due to the centralized reception of samples, the calculation resource load rate increases but does not exceed the threshold, the storage resource load rate is at a medium level, and the system comprehensive resource load rate formula is: wherein, is the system comprehensive resource load rate at time t; is the storage resource load rate at time t, that is, is the ratio of the storage resource used capacity to the total storage resource capacity at time t; is the calculation resource load rate at time t, that is, ​​​Real-time usage rate of CPU or GPU at time t; 0.6 is the weight coefficient of storage resource load rate, and 0.4 is the weight coefficient of computing resource load rate. Subsequently, the module adjusts the weight of the traceability integrity target and the weight of the resource consumption control target through a dynamic weight coefficient formula. Considering that the tumor marker detection result directly affects the subsequent treatment plan, the traceability chain of each sample needs to be prioritized to ensure integrity, so the traceability integrity target weight is set to a high level. Next, the module combines a multi-dimensional traceability integrity formula to evaluate the comprehensive situation of the current traceability node coverage, data accuracy and timeliness, and then analyzes the impact of current resource consumption on the system through a resource consumption elastic penalty formula. The resource consumption elastic penalty formula is: wherein, is the resource consumption elastic penalty value at time t, with a value range of 0 to 1; is the storage resource elasticity coefficient, with a value of 0.4; is the computing resource elasticity coefficient, with a value of 0.6; is the maximum load rate of storage resource, with a value of 100%, i.e. the load rate when the storage resource is fully occupied; is the maximum load rate of computing resource, with a value of 100%, i.e. the load rate when the CPU or GPU is fully occupied; is the storage resource load rate at time t; is the computing resource load at time t, and finally generates a traceability scheduling decision: adopt the full amount of evidence storage mode, completely store all traceability data of the sample receiving link, and set the traceability node corresponding to the suspected tumor patient sample as a high-priority scheduling object, and prioritize the traceability data entry and evidence storage association of this type of sample, as shown in Figure 2 .

[0033] After receiving the above decision, the traceability scheduling execution module immediately performs the full amount of evidence storage mode switching operation, and completely enters all patient sample receiving link traceability data into the system, including patient name, sample number, receiving time, receiving personnel, etc. At the same time, according to the high-priority scheduling rule, the traceability node of the suspected tumor patient sample is prioritized to ensure that the traceability data of this type of sample is associated with the sample within 15 minutes after receiving. After the operation is completed, the module feeds back the execution result to the enhanced learning intelligent decision-making module, including whether the full amount of evidence storage is successful, whether the high-priority node scheduling is completed on time, and at the same time, the evidence storage index of each sample traceability data is transmitted to the hybrid architecture storage and verification module.

[0034] The hybrid architecture storage and verification module receives the storage index and performs storage and verification based on the blockchain-cloud hybrid architecture. Among them, the traceability full data of the receiving link of suspected tumor patient samples is stored in the Hyperledger Fabric blockchain, and the traceability data of the key samples is ensured to be safe by using the tamper-proof characteristics of the blockchain; the traceability full data of the receiving link of ordinary patient samples is stored in the Aliyun OSS to meet the large-capacity storage demand and reduce the storage cost. Subsequently, the module performs hash verification on all stored traceability data, compares the original hash value obtained in the acquisition stage with the hash value of the stored data, confirms that the data has not been tampered with, completes the accuracy verification, and feeds back the verification result to the reinforcement learning intelligent decision-making module. The verification result includes whether the data is accurate, the specific storage location, the verification time, etc., providing accurate traceability basis for the subsequent detection link.

[0035] 2. Instrument detection stage system operation

[0036] After the sample is received, it enters the instrument detection link, mainly using the chemiluminescence immunoassay analyzer to detect the tumor markers in the sample. At this time, the state data acquisition module continues to work, and the computing resource load rate of the chemiluminescence immunoassay analyzer supporting computing device is acquired through the edge computing node resource monitoring submodule, that is, the real-time use rate of the GPU for processing real-time optical signal data and generating detection curves; the storage resource load rate of the detection process data is acquired through the cloud storage server resource monitoring submodule, that is, the used capacity proportion of the server for storing optical signal raw data, detection reagent batch number, instrument running state and other data. At the same time, the module acquires the traceability node coverage through the sample progress sensor, including whether the detection start time record, detection instrument number registration, reagent addition record, detection signal acquisition node are completely covered, and also acquires the original hash value and data generation timestamp of each detection process data, such as detection start timestamp and reagent addition completion timestamp, and transmits all data to the reinforcement learning intelligent decision-making module in real time through the MQTT protocol.

[0037] The reinforcement learning intelligent decision-making module is based on the multi-objective deep deterministic policy gradient algorithm, analyzes the comprehensive reward value of the current scheduling action combined with the dynamic multi-objective reward function formula, and the dynamic multi-objective reward function formula is: wherein, is the dynamic multi-objective reward value at time t, the value range is -1 to 1; is the traceability integrity target weight at time t; is the resource consumption control target weight at time t; is the multi-dimensional traceability integrity score at time t; is the resource consumption elasticity penalty value at time t; is an action influence coefficient, the value is 0.1 when the full quantity storage action is performed, the value is 0.3 when the fingerprint storage action is performed, and the value is 0.2 when the high-priority node scheduling action is performed; is a traceability integrity increment at time t, the reward value comprehensively considers the traceability integrity improvement effect and the resource consumption control effect; and then the optimal value of different candidate actions is evaluated through a multi-objective action value function formula, the candidate actions include full quantity storage, fingerprint storage, high-priority node scheduling, and low-priority node delay scheduling, and the multi-objective action value function formula is: , wherein, is the optimal action value of the action performed at time t; is the optimal action value of the action performed at time t; is the system state at time t; is the system state at time t; is the action performed at time t; is the action performed at time t; is the dynamic multi-objective reward value at time t; is a discount factor, the value is 0.9, and the importance of immediate reward and future reward is weighted; is a candidate action at time t+1, including full quantity storage, fingerprint storage, high-priority node scheduling, and low-priority node delay scheduling; is the state at time t+1 is the optimal action value of the candidate action performed at time t+1; is the system state at time t; is the system state at time t; is a state stability coefficient at time t, is 1 when the fluctuation amplitude of the system state at time t+1 and the system state at time t is less than or equal to 5%, and is 0.8 when the fluctuation amplitude is greater than 5%. At this time, multiple chemical luminescence immunoassay analyzers are simultaneously running to enter a detection peak, the resource load rate is greatly increased, the overall resource load rate is calculated through the system comprehensive resource load rate formula to be greater than the medium threshold. The module then adjusts the weight through the dynamic weight coefficient formula, appropriately increases the weight of the resource consumption control target, and analyzes the timeliness influence of the detection data in combination with a traceability data timeliness decay factor formula to ensure that the time difference from generation to storage of the detection data meets the requirements of the clinical detection standard, and finally generates a scheduling decision. The traceability data timeliness decay factor formula is: , wherein, is a traceability data timeliness decay factor at time t, the value range is 0 to 1; is the difference between the generation time and the storage time of the traceability data at time t; is a timeliness threshold, the value range is 2 seconds to 5 seconds; For the attenuation coefficient, the value range is 1.0 to 1.5: switch to the fingerprint storage mode, only store the key fingerprint information and core index of the detection process data, such as the fingerprint corresponding to the peak of the optical signal, the time index of the key detection stage, and reduce the storage resource occupation; set the traceability node corresponding to the sample in the key detection stage as a high-priority scheduling object, and set the traceability node of the sample that has completed sample pretreatment and is waiting for detection as a low-priority delayed scheduling object.

[0038] After the traceability scheduling execution module receives the decision, it immediately performs the fingerprint storage mode switching operation, only stores the key data fingerprint and core index in the detection process of each instrument, and no longer stores the complete optical signal curve and reagent addition process log; at the same time, the traceability nodes of the samples in the key detection stage are preferentially scheduled to ensure that the fingerprint storage and node association are completed within 10 minutes after the generation of the stage data, and the low-priority nodes are processed in the instrument idle gap set, such as processing the node data after a batch of detection is completed. After the operation is completed, the module feeds back the execution result to the enhanced learning intelligent decision module, and the feedback content includes the fingerprint storage data volume, the high-priority node scheduling completion rate, and the storage index of the traceability data in the detection link is transmitted to the hybrid architecture storage and verification module.

[0039] After the hybrid architecture storage and verification module receives the storage index, it stores the fingerprint data and core index of the detection link to Aliyun OSS, and additionally backs up the traceability fingerprint data of the sample in the key detection stage to HyperledgerFabric blockchain, to ensure the safety of the key detection link data. Subsequently, the module carries out hash verification, compares the original hash value of the detection process data with the associated hash value of the stored fingerprint data, confirms the data consistency, and feeds back the verification result to the enhanced learning intelligent decision module, to provide accurate detection traceability basis for the subsequent data output link.

[0040] 3. Data output stage system operation

[0041] After the detection is completed, the system generates a tumor marker detection report containing information such as detection index name, detection result value, reference range, detection instrument number, and detection personnel, and enters the data output stage. At this time, the state data acquisition module continues to collect data, the cloud storage server resource monitoring submodule collects the storage resource load rate of the report storage, that is, the server used capacity proportion of storing the full text of the report and the audit record of the report; the edge computing node resource monitoring submodule collects the computing resource load rate of the report generation and pushing, that is, the real-time CPU usage rate of processing the report format conversion and pushing to the hospital HIS system and electronic health record system. At the same time, the module collects the traceability node coverage degree through the sample progress sensor, including whether the report generation time record, the audit personnel signature record, the report modification record, and the HIS system pushing record are completely covered, and also collects the original hash value of the report data and the report generation timestamp, and transmits all the data to the enhanced learning intelligent decision module through the MQTT protocol.

[0042] The enhanced learning intelligent decision module is based on a multi-objective deep deterministic policy gradient algorithm, and combines a multi-dimensional traceability integrity formula to evaluate the comprehensive situation of the report traceability data, and the multi-dimensional traceability integrity formula is: , wherein, is the multi-dimensional traceability integrity score at time t, and the value range is 0 to 1; is the traceability node coverage degree weight, and the value range is 0.4 to 0.6; is the weight of the product of traceability data accuracy and timeliness, and the value range is 0.4 to 0.6, and + =1; is the traceability node coverage degree at time t; is the traceability data accuracy at time t; is the traceability data timeliness decay factor at time t, including whether the traceability coverage degree contains the full node from sample receiving to report, whether the data accuracy is consistent with the detection data, and whether the timeliness meets the clinical report delivery requirements; the system comprehensive resource load rate formula is used to calculate that at this time the detection peak has passed, and the storage and computing resource load rates have decreased to a low level. Subsequently, the module adjusts the weight through a dynamic weight coefficient formula, and the dynamic weight coefficient formula is: , wherein, is the traceability integrity target weight at time t, and the value range is 0 to 1; is the resource consumption control target weight at time t, and the value range is 0 to 1; is the weight adjustment coefficient, and the value range is 2.0 to 3.0; is the maximum tolerated load rate of the system, and the value is 90%; ​For the system comprehensive resource load rate at time t, the traceability integrity target weight is adjusted to the highest, and the final scheduling decision is generated: switch back to the full volume evidence storage mode, store the full text of the detection report and the associated sample receiving, instrument detection link all traceability data; prioritize scheduling the traceability data of the report audit node and the HIS system push node to ensure that the report audit process is traceable and the clinician can quickly query the complete traceability information.

[0043] The traceability scheduling execution module performs the full volume evidence storage mode switching operation to store the full text of each detection report and the corresponding sample receiving, instrument detection link traceability data; at the same time, it prioritizes the traceability storage of the report audit node, records the auditor's name, audit time, audit opinion, and the traceability association of the HIS system push node, associates the report ID with the patient HIS system ID and the traceability data storage index, and ensures that the report is pushed to the HIS system within 20 minutes after the audit is passed. After the operation is completed, the module feeds back the execution result to the reinforcement learning intelligent decision-making module, and transmits the traceability data storage index of the report link to the hybrid architecture storage and verification module.

[0044] The hybrid architecture storage and verification module receives the storage index, stores the full volume data of the report and the associated traceability data to Aliyun OSS, and stores the report audit record and the electronic signature data of the auditor to Hyperledger Fabric blockchain, to ensure the compliance and traceability of the report audit process. Subsequently, the module confirms the consistency of the report data and the associated traceability data through hash verification, and feeds back the verification result to the reinforcement learning intelligent decision-making module to complete the traceability closed loop of the tumor marker detection full process.

[0045] 4. Periodic optimization of the model dynamic optimization module

[0046] During the above-mentioned whole process, the reinforcement learning intelligent decision module transmits the historical calculation data of each stage to the model dynamic optimization module in real time, including the resource load change during the sample receiving peak and trough period, the traceability integrity score under different scheduling decisions, the hash verification pass rate and other information. The module iteratively optimizes the policy network weight of the multi-objective deep deterministic policy gradient algorithm and the parameters of the mathematical calculation formula such as the system comprehensive resource load rate formula, the dynamic weight coefficient formula, the dynamic multi-objective reward function formula and the multi-dimensional traceability integrity formula according to an iteration optimization period of once every 24 hours. For example, according to the resource load law of different time periods every day, the weight adaptability of the storage and computing resources in the system comprehensive resource load rate formula is adjusted; according to the change of the sample amount of suspected tumor patients, the weight adjustment logic of traceability integrity and resource consumption in the dynamic weight coefficient formula is optimized. After optimization, the module pushes the new parameters to the reinforcement learning intelligent decision module to update its data processing logic, ensuring that the subsequent traceability scheduling of the tumor marker detection whole process is more adaptive to the resource fluctuation and clinical detection demand change of the laboratory.

[0047] In summary, in the tumor marker detection whole process of the medical detection laboratory, the system of the application realizes efficient traceability management through the cooperative operation of the five modules. The state data acquisition module acquires the resource load and traceability correlation data of each stage of sample receiving, instrument detection and data output in real time, ensuring data real-time; the reinforcement learning intelligent decision module takes the multi-objective deep deterministic policy gradient algorithm as the core, generates scheduling decisions that adapt to the scene in combination with the system comprehensive resource load rate formula, and prioritizes the traceability integrity of critical samples for diagnosis and treatment; the traceability scheduling execution module switches the storage mode and schedules the node priority on demand; the hybrid architecture storage and verification module ensures data security and accuracy through the blockchain-cloud architecture; the model dynamic optimization module iterates the parameters regularly. The system not only meets the clinical demand for traceable detection results, but also dynamically balances resource consumption, effectively adapts to sample peak fluctuation, and significantly improves traceability efficiency and reliability.

[0048] Embodiment Two: Whole Process Traceability Scene of Soil Heavy Metal Detection in Environmental Monitoring Laboratory

[0049] This embodiment is applied to a provincial environmental monitoring station laboratory, which is mainly responsible for heavy metal detection of farmland and industrial park surrounding soil in the jurisdiction, including lead, mercury, cadmium, chromium, etc. The whole process traceability from sample receiving, instrument detection to data output needs to be realized to meet the needs of environmental protection departments for soil pollution control evaluation and supervision, while adapting to the system resource scheduling demand under the fluctuation of sample amount in different seasons.

[0050] 1. System operation in sample receiving stage

[0051] During the rainy season, the amount of soil samples received by the laboratory increases by 60% due to the migration of heavy metals in the soil caused by rain erosion. At this time, the system starts the state data collection module. This module collects the coverage of the traceability node in real time through the sample progress sensor, including the coverage of nodes such as sample point name registration, sampling time record, sampling personnel information input, sampling depth record, transportation vehicle number registration, transportation process temperature and humidity record, and laboratory reception registration; through the cloud storage server resource monitoring sub-module, the storage resource load rate is collected, that is, the current storage of sample point map data, transportation record, and sampling personnel qualification information data server used capacity proportion; through the edge computing node resource monitoring sub-module, the calculation resource load rate is collected, that is, the CPU real-time usage rate of processing sample information input and sample point data associated with sample number. At the same time, the module collects the original hash value and data generation timestamp of each soil sample traceability data, such as sampling completion timestamp, transportation departure and arrival timestamp, and transmits all state data to the enhanced learning intelligent decision-making module through the MQTT protocol, with a transmission delay controlled within 100ms.

[0052] The enhanced learning intelligent decision-making module processes data based on the multi-objective deep deterministic policy gradient algorithm. First, the overall occupation of current storage and computing resources is quantified by combining the system comprehensive resource load rate formula. Due to the surge in sample quantity during the rainy season, the storage and computing resource load rate is at a high level. Next, the module adjusts the traceability integrity target weight and resource consumption control target weight through the dynamic weight coefficient formula. Considering the need to avoid resource overload leading to system lag, the resource consumption control target weight is set higher than the traceability integrity target weight. Then, the module analyzes the timeliness of sampling data by combining the traceability data timeliness decay factor formula, ensuring that the time difference from sampling completion to laboratory reception and evidence storage meets the requirements of soil detection standards. Finally, the module analyzes the impact of current resource consumption on the system through the resource consumption elastic penalty formula, and generates a traceability scheduling decision: adopt a fingerprint evidence storage mode, only store key information fingerprints and core indexes such as sample point number, sampling time, and transportation vehicle ID fingerprint, reducing the amount of stored data; set the traceability nodes corresponding to the soil samples around the industrial park as high-priority scheduling objects, as the risk of heavy metal pollution in this area is higher, and the data needs to be closely monitored.

[0053] After receiving the decision, the traceability scheduling execution module immediately performs the fingerprint storage mode switching operation, and enters the key information fingerprints of all soil samples in the receiving link into the core index recording system; at the same time, according to the high-priority scheduling rule, the traceability nodes of the soil samples in the industrial park are preferentially processed to ensure that the association storage of the traceability data and the samples of such samples is completed within 30 minutes after the samples are received in the laboratory, and the traceability nodes of the ordinary farmland soil samples are processed in the resource idle period, such as when the sample receiving amount is reduced at noon. After the operation is completed, the module feeds back the execution result to the enhanced learning intelligent decision module, and the feedback content includes the fingerprint storage data integrity, the high-priority node scheduling completion condition, and the traceability data storage index of the sample receiving link is transmitted to the hybrid architecture storage and verification module.

[0054] After receiving the storage index, the hybrid architecture storage and verification module stores the fingerprint data of the soil samples in the receiving link in the industrial park into the Hyperledger Fabric block chain to ensure the non-tamperable characteristics of the high-risk area sample traceability data and meet the environmental protection supervision requirements; the fingerprint data of the receiving link of the ordinary farmland soil samples is stored in the Aliyun OSS. Subsequently, the module performs hash verification on all stored traceability data, confirms the data accuracy by comparing the original hash value in the collection stage with the associated hash value of the stored fingerprint data, and feeds back the verification result to the enhanced learning intelligent decision module to provide a reliable traceability basis for the subsequent detection link.

[0055] 2. System operation in the instrument detection stage

[0056] After the sample receiving is completed, the sample enters the instrument detection link, and the sample pretreatment and instrument detection are performed in turn. The sample pretreatment operation is digestion and extraction to remove impurities and interference, and the instrument detection uses an inductively coupled plasma mass spectrometer ICP-MS to detect the heavy metal content. At this time, the state data collection module continues to work, and the edge computing node resource monitoring sub-module collects the computing resource load rate of the ICP-MS supporting computing device, i.e. the real-time usage rate of the GPU for processing the element signal intensity data in the detection process and generating the concentration calibration curve; the cloud storage server resource monitoring sub-module collects the storage resource load rate of the detection process data, i.e. the server used capacity proportion for storing the element signal spectrum, pretreatment reagent batch number, instrument calibration record and other data. At the same time, the module collects the traceability node coverage degree through the sample progress sensor, including whether the pretreatment start and end time record, pretreatment device number registration, instrument calibration time record, detection signal collection node are completely covered, and also collects the original hash value and data generation timestamp of the detection process data, such as digestion completion timestamp and instrument calibration completion timestamp, and transmits all data in real time to the enhanced learning intelligent decision module through the MQTT protocol.

[0057] The reinforcement learning intelligent decision module is based on a multi-objective deep deterministic policy gradient algorithm, analyzes the comprehensive reward value of the current scheduling action in combination with a dynamic multi-objective reward function formula, comprehensively considers the effect of improving traceability integrity and the effect of controlling resource consumption, and then evaluates the optimal value of different candidate actions through a multi-objective action value function formula. At this time, because part of the ICP-MS instrument completes calibration and enters the detection peak, the computing resource load rate rises, but the storage resource is still at a medium level because of the use of fingerprint storage. The overall resource load is moderate, which is calculated by a system comprehensive resource load rate formula. The module then sets the traceability integrity target weight and the resource consumption control target weight to a balanced level through a dynamic weight coefficient formula, evaluates the coverage and accuracy of the traceability data in the detection link through a multi-dimensional traceability integrity formula, and finally generates a scheduling decision: maintain the fingerprint storage mode and continue to control the storage resource occupation; set the traceability nodes corresponding to the instrument calibration records as high-priority scheduling objects, because the calibration data directly affect the accuracy of the detection results and need to be stored preferentially; set the sample traceability nodes that have completed pretreatment and are waiting for detection as regular priority.

[0058] After receiving the decision, the traceability scheduling execution module continues to maintain the fingerprint storage mode, stores the key fingerprints of the element signals in the detection process and the indexes of the key parameters in the pretreatment, and at the same time, preferentially processes the traceability nodes of the instrument calibration records, associates and stores the fingerprints of the calibration time, calibration standard substance batch number, and calibration result data of each instrument with the instrument number for storage, and ensures that the calibration data is stored within 20 minutes after being generated. After the operation is completed, the module feeds back the execution result to the reinforcement learning intelligent decision module, the feedback content includes the storage success rate of the calibration nodes and the integrity of the detection data fingerprints, and at the same time, the storage indexes of the traceability data in the detection link are transmitted to the hybrid architecture storage and verification module.

[0059] After receiving the storage indexes, the hybrid architecture storage and verification module stores the fingerprint data in the detection process to the Aliyun OSS and stores the fingerprint data of the instrument calibration records to the Hyperledger Fabric blockchain, guarantees the compliance of the calibration data, and meets the detection qualification requirements. Then, the module confirms the consistency of the detection fingerprint data and the original detection signal data and the consistency of the calibration record fingerprint and the original calibration data through a hash verification, and feeds back the verification result to the reinforcement learning intelligent decision module, to ensure the accuracy of the traceability chain of the detection data.

[0060] 3. Data output stage system operation

[0061] After the detection is completed, the system generates a soil heavy metal detection report, which contains the heavy metal concentration values of each sampling point sample, whether it exceeds the Soil Environmental Quality Risk Control Standard for Agricultural Land Soil Pollution, the detection method standard number, and other information, and enters the data output stage. At this time, the state data acquisition module acquires relevant data, the cloud storage server resource monitoring submodule acquires the storage resource load rate of the report storage, and the edge computing node resource monitoring submodule acquires the computing resource load rate of the report generation and pushing. The pushing object is the provincial environmental protection department supervision platform. At the same time, the module acquires the traceability node coverage degree through the sample progress sensor, including whether the report generation time record, data audit personnel information input, audit opinion record, and supervision platform pushing record nodes are completely covered, and also acquires the original hash value of the report data and the report generation timestamp, and transmits all the data to the enhanced learning intelligent decision module through the MQTT protocol.

[0062] The enhanced learning intelligent decision module is based on a multi-objective deep deterministic policy gradient algorithm, and combines a multi-dimensional traceability integrity formula to evaluate the comprehensive situation of the report traceability data, and confirms whether it covers the whole link of sampling, transportation, detection, and calibration. Through a system comprehensive resource load rate formula, it is calculated that the detection peak has passed at this time, and the resource load rate has decreased to a low level. Subsequently, the module adjusts the traceability integrity target weight to the highest through a dynamic weight coefficient formula, considering that the report needs to be submitted to the environmental protection department for soil pollution control decision-making, and needs to guarantee the whole chain traceability, and finally generates a scheduling decision: switching to a full amount of evidence storage mode, and completely storing the full text of the report and all traceability data of the associated sampling, transportation, detection, and calibration links; prioritizing the scheduling of traceability data of the report audit node and the environmental protection supervision platform pushing node.

[0063] The traceability scheduling execution module executes the full amount of evidence storage mode switching operation, and completely stores each report full text and associated sampling point information, transportation records, detection process data, instrument calibration records, and other traceability data; at the same time, it preferentially completes the traceability storage of the report audit node, records the auditor's name, audit time, and audit conclusion, and associates the report ID with the sampling point ID of the supervision platform, and the traceability data storage index, to ensure that the report is pushed to the environmental protection department supervision platform within 1 hour after passing the audit. After the operation is completed, the module feeds back the execution result to the enhanced learning intelligent decision module, and transmits the traceability data storage index of the report link to the hybrid architecture storage and verification module.

[0064] After receiving the evidence index, the hybrid architecture storage and verification module stores the full-amount data and associated traceability data to Aliyun OSS, stores the report audit record and the audit personnel signature data to the Hyperledger Fabric blockchain, ensures the report audit compliance, and facilitates the environmental protection department to check. Subsequently, the module confirms the consistency of the report data and the associated traceability data through hash verification, and the verification result is fed back to the reinforcement learning intelligent decision-making module, completing the traceability management of the whole process of soil heavy metal detection.

[0065] 4. Periodic optimization of the model dynamic optimization module

[0066] In the whole process of soil heavy metal detection, the reinforcement learning intelligent decision-making module transmits historical calculation data to the model dynamic optimization module, including the resource load difference between the rainy season and the dry season, the scheduling effect of samples in different pollution risk areas, the traceability integrity score change, the hash verification pass rate and other information. The module adjusts the policy network weight of the multi-objective deep deterministic policy gradient algorithm and the parameters of the mathematical calculation formulas such as the system comprehensive resource load rate formula, the dynamic weight coefficient formula, the multi-dimensional traceability integrity formula and the dynamic multi-objective reward function formula based on the historical data according to an iteration optimization period of every 24 hours. For example, according to the law that the sample amount is large in the rainy season and small in the dry season, the parameters of the system comprehensive resource load rate formula are optimized to adapt to the resource fluctuation in different seasons; according to the difference in supervision priority between industrial parks and ordinary farmland samples, the weight logic of traceability integrity in the dynamic weight coefficient formula is adjusted. The optimized parameters are pushed to the reinforcement learning intelligent decision-making module, so that it can better adapt to the seasonal work demand of the environmental monitoring laboratory and improve the traceability efficiency and resource utilization rationality of the whole process of soil detection.

[0067] In summary, in the soil heavy metal detection scenario of the environmental monitoring laboratory, the system of the present application precisely adapts to the seasonal work demand. The state data acquisition module captures the resource load and traceability data when the sample amount increases sharply in the rainy season, providing a basis for decision-making; the reinforcement learning intelligent decision-making module relies on the core algorithm and related formulas to develop fingerprint evidence and high-priority node scheduling strategies, focusing on ensuring the traceability of high-pollution-risk samples; the traceability scheduling execution module flexibly switches the evidence storage mode to balance resource occupation and data recording needs; the hybrid architecture storage and verification module stores data in stages, taking into account safety and cost; the model dynamic optimization module adjusts parameters periodically to adapt to resource fluctuations in the rainy season and the dry season. The system not only meets the requirements of environmental protection supervision for the whole process of soil detection data traceability, but also realizes efficient use of resources, ensuring traceability compliance and efficiency.

[0068] The above merely describes the preferred embodiments of the present application, and is not intended to limit the present application in any form. Although the present application has been disclosed with the preferred embodiments as above, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content to obtain equivalent embodiments with equivalent changes, as long as the changes or modifications do not deviate from the technical solution of the present application. Any modification, change, equivalent change and modification of the above embodiments made according to the technical essence of the present application still belong to the scope of the technical solution of the present application.

Claims

1. Laboratory testing full-process intelligent traceability system based on reinforcement learning, characterized in that, The system comprises: a state data acquisition module for acquiring resource load data and traceability integrity associated data in the sample receiving stage, instrument detection stage and data output stage, and transmitting the acquired state data to the enhanced learning intelligent decision module; the enhanced learning intelligent decision module processes the received state data based on a multi-objective deep deterministic policy gradient algorithm, a system comprehensive resource load rate formula, a dynamic multi-objective reward function formula, a multi-objective action value function formula, and a multi-dimensional traceability integrity formula and a traceability integrity increment formula to generate a traceability scheduling decision; simultaneously, the traceability scheduling decision is transmitted to the traceability scheduling execution module, and the historical calculation data generated in the data processing process is transmitted to the model dynamic optimization module; the traceability scheduling execution module performs traceability data full-amount storage or fingerprint storage mode switching operation and traceability node priority scheduling operation according to the received traceability scheduling decision; after the operation is completed, the execution result is fed back to the enhanced learning intelligent decision module, and the storage index of the traceability data is transmitted to the hybrid architecture storage and verification module; the hybrid architecture storage and verification module stores the traceability data corresponding to the storage index based on a blockchain-cloud hybrid architecture, and performs hash verification on the traceability data to complete accuracy verification; the hybrid architecture storage and verification module feeds back the verification result to the enhanced learning intelligent decision module; the model dynamic optimization module receives the historical calculation data transmitted by the enhanced learning intelligent decision module, iteratively optimizes the policy network weight of the multi-objective deep deterministic policy gradient algorithm and the parameters of the mathematical calculation formula, and pushes the optimized parameters to the enhanced learning intelligent decision module to update the data processing logic thereof.

2. The reinforcement learning based intelligent lab test full-process traceability system according to claim 1, characterized in that, The resource load data acquired by the state data acquisition module includes storage resource load rate and calculation resource load rate, and the traceability integrity associated data includes traceability node coverage, traceability data original hash value and data generation timestamp; The state data acquisition module acquires the traceability node coverage through a sample progress sensor, acquires the storage resource load rate through a cloud storage server resource monitoring sub-module, acquires the calculation resource load rate through an edge computing node resource monitoring sub-module, and transmits the state data by using MQTT protocol, with a transmission delay of ≤100 ms.

3. The reinforcement learning based intelligent lab test full flow traceability system according to claim 1, wherein, The formula for the overall system resource load rate in the reinforcement learning intelligent decision-making module is: ,in, for The overall resource load rate of the real-time system; for Real-time storage resource load rate, i.e. The ratio of the current used storage capacity to the total storage capacity; for Calculate resource load rate in real time, i.e. Real-time CPU or GPU utilization; 0.6 is the weighting factor for storage resource load rate, and 0.4 is the weighting factor for computing resource load rate.

4. The reinforcement learning based intelligent lab test full flow traceability system according to claim 1, wherein, The traceability data timeliness decay factor formula in the reinforcement learning intelligent decision module is: Wherein, t is the traceability data timeliness decay factor at time t, and the value range is 0 to 1; t is the difference between the traceability data generation time and the storage time at time t; t is the timeliness threshold, and the value range is 2 seconds to 5 seconds; t is the decay coefficient, and the value range is 1.0 to 1.

5.

5. The reinforcement learning based intelligent laboratory testing full-process traceability system according to claim 1, wherein, The dynamic weight coefficient formula in the reinforcement learning intelligent decision module is: wherein, is a traceability integrity target weight at time t, and the value range is 0 to 1; is a resource consumption control target weight at time t, and the value range is 0 to 1; is a weight adjustment coefficient, and the value range is 2.0 to 3.0; is a maximum tolerable load rate of the system, and the value is 90%; is a system comprehensive resource load rate at time t.

6. The reinforcement learning based intelligent lab test full flow traceability system according to claim 1, wherein, The multi-dimensional traceability integrity formula and the traceability integrity increment formula in the enhanced learning intelligent decision module are respectively: ; , wherein is a multi-dimensional traceability integrity score at time t, and the value range is 0 to 1; is a traceability node coverage weight, and the value range is 0.4 to 0.6; is a weight of the product of traceability data accuracy and timeliness, and the value range is 0.4 to 0.6, and + =1; is a traceability node coverage at time t; is a traceability data accuracy at time t; is a traceability data timeliness decay factor at time t; is a traceability integrity increment at time t, and the value range is -1 to 1; is a multi-dimensional traceability integrity score at time t.

7. The reinforcement learning based intelligent lab testing full-process traceability system according to claim 1, wherein, The resource consumption elasticity penalty formula in the reinforcement learning intelligent decision module is: Wherein, is the resource consumption elasticity penalty value at time t, and the value range is 0 to 1; is the storage resource elasticity coefficient, and the value is 0.4; is the computing resource elasticity coefficient, and the value is 0.6; is the maximum load rate of the storage resource, and the value is 100%, that is, the load rate when the storage resource is completely occupied; is the maximum load rate of the computing resource, and the value is 100%, that is, the load rate when the CPU or GPU is completely occupied; is the storage resource load rate at time t; is the computing resource load rate at time t.

8. The reinforcement learning based intelligent lab testing full-process traceability system according to claim 1, wherein, The dynamic multi-target reward function formula in the reinforcement learning intelligent decision module is: Wherein, is a dynamic multi-target reward value at time t, and the value range is -1 to 1; is a trace integrity target weight at time t; is a resource consumption control target weight at time t; is a multi-dimensional trace integrity score at time t; is a resource consumption elasticity penalty value at time t; is an action influence coefficient, which is 0.1 when a full-amount notarization action is executed, 0.3 when a fingerprint notarization action is executed, and 0.2 when a high-priority node scheduling action is executed; is a trace integrity increment at time t.

9. The reinforcement learning based intelligent lab testing full-process traceability system according to claim 1, wherein, The multi-objective action value function formula in the reinforcement learning intelligent decision module is: Wherein, is the optimal action value of the action performed at the moment state; is the system state at the moment; is the action performed at the moment; is the dynamic multi-objective reward value at the moment; is a discount factor, taking a value of 0.9, used to weigh the importance of immediate rewards and future rewards; is the candidate action at the moment t+1, including four types of actions: full quantity storage, fingerprint storage, high priority node scheduling, and low priority node delay scheduling; is the state at the moment t+1 is the optimal action value of the candidate action performed at the moment t+1; is the state stability coefficient at the moment t+1, is 1 when the fluctuation amplitude of the system state at the moment t+1 and the system state at the moment t is ≤5%, and is 0.8 when the fluctuation amplitude is >5%.​​​​​​ 10. The reinforcement learning based intelligent lab testing full-process traceability system according to claim 1, wherein, When the traceability scheduling execution module performs full-amount storage / fingerprint storage mode switching operation, when the storage resource load rate > 80% or the calculation resource load rate > 80%, the fingerprint storage mode is switched; when the storage resource load rate < 40% and the calculation resource load rate < 40%, the full-amount storage mode is switched; the blockchain used by the hybrid architecture storage and verification module is Hyperledger Fabric blockchain, and the cloud storage used is Aliyun OSS; the iteration optimization period of the model dynamic optimization module is once every 24 hours, and the optimization objects include the policy network weight of the multi-objective deep deterministic policy gradient algorithm and the parameters of each mathematical calculation formula.

Citation Information

Patent Citations

  • Block chain-based traceability processing method and block chain distributed traceability system

    CN116362772A

  • Scheduling automation system application state management method

    CN119292745A

  • Asset management system based on full life cycle

    CN120525447A

  • Competition management system and method based on cloud computing

    CN120596254A

  • Cloud-edge collaborative intelligent storage node dynamic deployment method and system

    CN120614249A