Unmanned laboratory sample automatic regulation and control method and system based on reinforcement learning
By constructing a directed demand chain and a distributed reinforcement control model, the problem of full-process collaborative optimization in the automated sample configuration of unmanned laboratories was solved, realizing dynamic adjustment of the experimental process and equipment load management, and improving the accuracy and efficiency of sample configuration.
Patent Information
- Application Number
- CN202511430687.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-09
AI Technical Summary
Existing technologies lack the ability to achieve full-process collaborative optimization and dynamic adjustment in the automated sample configuration process of unmanned laboratories. This leads to resource configuration conflicts, chaotic execution sequences, slow anomaly localization, and a failure to dynamically correlate the status of experimental equipment with standard requirements, affecting the accuracy, efficiency, and reliability of the experiment.
By constructing a directed demand chain based on reinforcement learning, combined with a distributed reinforcement control model and blockchain technology, the logical relationship and priority mapping of experimental atomic experiments are realized, a sequence of pre-control instructions is generated, the experimental process is dynamically optimized, and real-time anomaly tracking and equipment load management are performed.
It significantly improves the accuracy and efficiency of automated sample preparation, reduces operational risks, and ensures the traceability and compliance of the experimental process with standard requirements.
Smart Images

Figure CN120911935A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of experiment control, and particularly relates to an unmanned laboratory sample automatic regulation method and system based on reinforcement learning. BACKGROUND
[0002] In the current automated experiment technical field, especially in the sample automation configuration process of unmanned laboratories, traditional methods usually rely on predefined static processes or isolated equipment control strategies. These methods lack the ability of global coordination and dynamic adjustment of the whole experiment process, resulting in the highlighting of multiple key technical problems: first, the logical dependency relationship between experiment steps and the priority conflict across experiments are difficult to be globally coordinated, which easily causes resource configuration conflicts or execution sequence confusion; second, due to the lack of real-time verification and tracing mechanism for the execution effect and efficiency of sub-experiments, when an exception occurs, it is often slow to locate, and the potential impact of the exception on subsequent steps cannot be accurately evaluated; third, the compliance of the state of the experimental equipment with the requirements of the experimental standard cannot be dynamically associated, so that the control process cannot be adaptively adjusted according to the real-time state, thereby affecting the accuracy, efficiency and reliability of the overall experiment. SUMMARY
[0003] In view of the deficiencies of the prior art, the present application provides an unmanned laboratory sample automatic regulation method and system based on reinforcement learning. The method first constructs a directed demand chain reflecting atomic sub-experiments, logical relationships and priorities of experiments by clustering algorithm and blockchain technology; then maps it to a process control verification chain composed of experimental equipment priority and real-time load by using smart contract, and generates a pre-control instruction sequence in combination with a distributed reinforcement control model; the model integrates sub-experiment accuracy, verification pass probability, implementation efficiency and interaction risk transfer matrix, and dynamically optimizes decisions relying on reinforcement learning and experimental standard verification chain; finally, sample automation configuration is realized through instruction execution; the present application realizes global coordination and real-time exception tracking across experiments, significantly improves configuration accuracy and efficiency, and reduces operational risk.
[0004] To achieve the above-mentioned purpose, the present application provides the following technical solutions:
[0005] The unmanned laboratory sample automatic regulation method based on reinforcement learning comprises:
[0006] Obtaining a directed demand chain corresponding to different experiment demands; the directed demand chain is obtained by atomic sub-experiments corresponding to each experiment, implementation logical relationship between sub-experiments, sub-experiment implementation priority of different experiments, and clustering algorithm and blockchain construction;
[0007] A preset flow control verification chain is mapped to each sub-experiment in the directed demand chain through a smart contract in the directed demand chain, and a pre-control instruction sequence is obtained in combination with a configured distributed reinforcement control model; the distributed reinforcement control model is constructed in combination with a reinforcement learning algorithm and an experiment standard verification chain by a per-sub-experiment sample configuration accuracy, a verification pass probability, a corresponding sub-experiment implementation efficiency, and an interaction risk transfer matrix; the interaction risk transfer matrix is used to measure the influence degree of the verification pass probability and the implementation efficiency of the corresponding sub-experiment of the current experiment on the verification pass probability and the implementation efficiency of the next sub-experiment of the same experiment or different experiments; the experiment standard verification chain is constructed by a standard operation step of each sub-experiment corresponding to the sub-experiment of each experiment and a standard evaluation index of each sub-experiment in combination with an evaluation algorithm; and the flow control verification chain is constructed by an experimental equipment corresponding to each experiment and an operation priority of the corresponding experimental equipment in each experiment and a maximum real-time load coefficient of each experimental equipment.
[0008] The pre-control instruction sequence is executed.
[0009] Specifically, the construction process of the directed demand chain includes:
[0010] Different experiment flow information is obtained, and each experiment flow information is decomposed into an atomic sub-experiment sequence of the corresponding experiment according to the operation indivisibility and the equipment singularity double rules, wherein the operation indivisibility means that each sub-experiment is the smallest operation unit in the corresponding experiment flow, and the equipment singularity means that the complete execution of the sub-experiment depends on only one experimental equipment; the atomic sub-experiment includes an operation trigger condition, an execution time sequence priority, a standard execution effect, and a standard execution efficiency of the corresponding sub-experiment; the execution time sequence priority includes a time sequence priority of the corresponding sub-experiment in the experiment and an operation priority of different experiments facing the same experimental equipment;
[0011] The atomic sub-experiment sequence corresponding to the same experiment is used to construct a first operation node sequence, and the operation trigger condition and the time sequence priority of the corresponding sub-experiment in the experiment between different sub-experiments under the same experiment are used to construct a first connection relationship;
[0012] Based on the first operation node sequence and the first connection relationship, a horizontal demand chain corresponding to each experiment is obtained, and the standard execution effect and the standard execution efficiency of the corresponding sub-experiment are stored in the corresponding first operation node for saving;
[0013] Based on the experimental equipment label of the corresponding sub-experiment of the horizontal demand chain, a horizontal mapping between each sub-experiment and the experimental equipment is constructed.
[0014] Specifically, the construction process of the directed demand chain further includes:
[0015] Clustering the experimental equipment unique identifiers corresponding to different experiments as parameters to obtain initial sub-experiment clustering groups;
[0016] Based on each corresponding sub-experiment in the initial sub-experiment clustering group, the operation type quantitative features and the operation object attribute quantitative features of each sub-experiment are extracted by combining the convolutional neural network to obtain the operation complexity of the corresponding sub-experiment; the operation complexity is constructed by the number of operation steps of the corresponding sub-experiment, the control parameter quantity of each operation step, the validation index parameter size, and the size of the corresponding sub-experiment implementation time under the standard execution flow;
[0017] Based on the operation priority of different experiments facing the same experimental step, combining the operation complexity of the corresponding sub-experiment, taking the minimum influence degree of the current sub-experiment on the implementation efficiency and validation probability of the next sub-experiment as the target, combining the adaptive clustering algorithm, the secondary clustering sub-group set is obtained by clustering each initial sub-experiment clustering group twice.
[0018] Specifically, the construction process of the directed demand chain also includes:
[0019] The implementation sequence corresponding to the minimum influence degree of the current initial sub-experiment clustering group on the implementation efficiency and validation probability of the next sub-experiment is constructed as a vertical connection relationship sequence;
[0020] The main vertical node is constructed by the label corresponding to each experimental equipment, the vertical sub-node sequence is constructed by the secondary clustering sub-group set corresponding to each main vertical node, and the vertical cross-experiment node chain is constructed based on the vertical connection relationship sequence and the vertical sub-node sequence.
[0021] Based on the vertical cross-experiment node chain corresponding to each experimental equipment, the real-time state evaluation score of the corresponding experimental equipment, and the maximum load coefficient corresponding to the current real-time state evaluation score, a second load mapping connection is constructed.
[0022] The horizontal demand chain and the horizontal mapping, and the vertical cross-experiment node chain and the second load mapping connection are stored in the blockchain for on-chain storage to obtain a directed demand chain.
[0023] Specifically, the distributed reinforcement control model includes a federal control sub-model and M sub-device control models; each experimental equipment in the process control verification chain corresponds to each sub-experiment of the corresponding horizontal demand chain one by one and is sorted by the sub-experiment implementation time corresponding to each experiment; the construction process of the distributed reinforcement control model includes:
[0024] Based on the standard evaluation index of the corresponding sub-experiment in the experimental standard verification chain, sample configuration detection data after the actual execution of the corresponding sub-experiment are collected, the degree of agreement between the actual detection data and the standard evaluation index is calculated, and the ratio of the number of samples with an agreement greater than or equal to the standard threshold to the total number of detected samples is calculated and used as the sample configuration accuracy of the corresponding sub-experiment; the detection data includes component detection values, measured concentration values, and purity detection results.
[0025] Extract the number of valid executions of each sub-experiment during its N executions in history, where the sample configuration accuracy is greater than or equal to the preset pass threshold and the operation steps meet the standard operation steps of the sub-experiment in the experimental standard verification chain. Use the ratio of the number of valid executions to the total number of N executions as the prior value of the verification pass probability of the corresponding sub-experiment.
[0026] Based on the prior value of the verification success probability of the current sub-experiment as the weight, the verification evaluation index of the corresponding sub-experiment and the real-time index detection data, combined with the comprehensive evaluation algorithm, the verification success probability of the corresponding sub-experiment is obtained.
[0027] If the current sub-experiment is a new sub-experiment, the initial verification pass probability of the corresponding sub-experiment is obtained by multiplying the verification pass probability of similar sub-experiments by the performance coefficient of the corresponding experimental equipment at the current moment; the performance coefficient at the current moment is constructed by the ratio of the maximum real-time load coefficient of the experimental equipment at the current moment to the rated load coefficient of the equipment.
[0028] Specifically, the construction process of the distributed reinforcement control model also includes:
[0029] Extract the standard execution efficiency index and the actual execution time T1 of the corresponding sub-experiment in the experimental standard validation chain, and then... , , The product of the real-time load coefficient and the actual load coefficient is used as the implementation efficiency of the corresponding sub-experiment.
[0030] Where T0 represents the standard execution time of the corresponding sub-experiment. This represents the implementation efficiency of the (k-1)th sub-experiment implemented on the (j-1)th experimental device in the i-th experiment. The implementation efficiency of the k-th sub-experiment implemented on the j-th experimental device for the i-th experiment. The probability risk transmission value; This represents the implementation efficiency of the g-th sub-experiment under the p-th experiment on the vertical cross-experiment node chain corresponding to the j-th experimental device. The efficiency of implementing the s-th sub-experiment on the j-th experimental device for the i-th experiment. The probability risk propagation value, the kth sub-experiment and the gth sub-experiment are the same sub-experiment, and the operation priority implemented by the sth sub-experiment is lower than that of the gth sub-experiment.
[0031] Specifically, the construction process of the distributed reinforcement control model further comprises:
[0032] a first horizontal sub-experiment pair continuously performed within the same horizontal demand chain and a second vertical sub-experiment pair corresponding to the same vertical cross-experiment node chain As a sample pair, through a Bayesian algorithm, the following are obtained: , and , ; wherein represents the verification passing probability of the k-1th sub-experiment of the ith experiment performed on the j-1th experimental equipment on the horizontal demand chain , the probability risk transfer value of the verification passing probability of the kth sub-experiment of the ith experiment performed on the jth experimental equipment ; represents the verification passing probability of the gth sub-experiment of the pth experiment on the vertical cross-experiment node chain corresponding to the jth experimental equipment , the probability risk transfer value of the verification passing probability of the s th sub-experiment of the ith experiment performed on the jth experimental equipment ; represents the second vertical sub-experiment pair composed of the s th sub-experiment of the ith experiment and the gth sub-experiment of the pth experiment adjacent to the jth experimental equipment ; represents the first horizontal sub-experiment pair constructed by the k-1th sub-experiment of the ith experiment performed on the j-1th experimental equipment and the kth sub-experiment of the ith experiment performed on the jth experimental equipment ;
[0033] Based on the product of and and the product of and , the risk transfer pair of the gth sub-experiment of the ith experiment performed on the jth experimental equipment is constructed.
[0034] Based on the combination of the horizontal demand chain and the vertical cross-experiment node chain corresponding to all experiments and the hidden Markov algorithm, the risk transfer pair acquisition process is repeated to obtain the real-time risk transfer matrix corresponding to all experiments.
[0035] Specifically, the construction process of the distributed reinforcement control model further comprises:
[0036] Based on the vertical cross-experiment node chain of each experimental device, the second load mapping connection data of each experimental device, the performance evaluation data of the corresponding sub-experiment, and the probability risk transfer value of each experimental device in the real-time risk transfer matrix, the input state space of the sub-device control model is constructed. Based on the horizontal demand chain data of all experiments, the horizontal mapping data of all experiments, the output state of all sub-device control models, and the real-time risk transfer matrix, the input state space of the federated control sub-model is constructed.
[0037] Specifically, the construction process of the distributed reinforcement control model also includes:
[0038] Based on the input state space of the sub-device control model and the input state space of the federated control sub-model, the execution action trigger information is constructed.
[0039] When the real-time evaluation result and implementation efficiency of any sub-experiment in the horizontal demand chain do not meet the corresponding standard execution effect and standard execution efficiency, or when the overall performance evaluation value and implementation efficiency of the corresponding experiment do not meet the preset overall standard execution effect and overall standard execution efficiency, the anomaly detection and location information is triggered. That is, through the horizontal mapping and second load mapping connection between the federated control sub-model and the M sub-device control models, bidirectional anomaly location is performed on the horizontal demand chain and the vertical cross-experiment node chain, and the anomaly location is identified by the anomaly identification model constructed by random forest.
[0040] The corresponding location and type parameters of the anomalies are fed back to the federated control sub-model, and an anomaly adjustment strategy is generated in combination with the preset anomaly adjustment strategy library.
[0041] The generated anomaly adjustment strategy is connected to the second load mapping through the horizontal mapping and distributed to the corresponding anomaly device and associated anomaly device for adjustment until the evaluation results and implementation efficiency of the sub-experiment corresponding to each horizontal demand chain meet the corresponding standard execution effect and standard execution efficiency.
[0042] An automated sample control system for unmanned laboratories based on reinforcement learning includes: a demand chain module, an analysis mapping module, and an execution module;
[0043] The demand chain module is used to obtain directed demand chains corresponding to different experimental needs; the directed demand chain is constructed by the atomic sub-experiments corresponding to each experiment and the implementation logic relationship between the sub-experiments, the implementation priority of the sub-experiments of different experiments, clustering algorithms, and blockchain.
[0044] The analysis and mapping module is used to preset the process control verification chain, and to map each sub-experiment in the directed demand chain to the process control verification chain through the smart contract in the directed demand chain, and to obtain the pre-control instruction sequence by combining the configured distributed reinforcement control model.
[0045] The distributed reinforcement control model is constructed by the configuration accuracy of each sub-experiment sample, the verification passing probability, the corresponding sub-experiment implementation efficiency and the interaction risk transfer matrix, combined with a reinforcement learning algorithm and an experiment standard verification chain; the interaction risk transfer matrix is used to measure the influence degree of the verification passing probability and the implementation efficiency of the corresponding sub-experiment of the current experiment on the verification passing probability and the implementation efficiency of the next sub-experiment of the same experiment or different experiments; the experiment standard verification chain is constructed by the standard operation steps of the corresponding sub-experiment of each experiment and the standard evaluation index of each sub-experiment combined with an evaluation algorithm; and the flow control verification chain is constructed by the corresponding experimental equipment of each experiment, the operation priority of the corresponding experimental equipment in each experiment and the maximum real-time load coefficient of each experimental equipment.
[0046] The execution module is used for executing the pre-control instruction sequence.
[0047] Compared with the prior art, the present application has the following beneficial effects:
[0048] The present application aims at the deficiencies of the prior art, and by constructing a directed demand chain that integrates experimental atomic sub-experiments, logical relationships, implementation priorities, clustering results and blockchain storage, the precise disassembly and tamper-proof storage of experimental requirements can be realized, and the sub-experiments executable by similar devices can be merged through a clustering algorithm to improve the device reuse efficiency; then by means of a smart contract, the sub-experiments in the directed demand chain are mapped to a flow control verification chain constructed by experimental equipment, operation priority and maximum real-time load coefficient of the equipment, which can effectively avoid device overload and optimize the rationality of device scheduling; at the same time, based on the configuration accuracy of sub-experiment sample, the verification passing probability, the implementation efficiency and the interaction risk transfer matrix, combined with a reinforcement learning algorithm and an experiment standard verification chain, a distributed reinforcement control model is constructed, which can accurately generate a pre-control instruction sequence, not only can dynamically improve the accuracy and stability of sub-experiment execution and reduce the abnormal transmission risk between sub-experiments, but also can significantly improve the overall efficiency, accuracy and reliability of sample automatic configuration in an unmanned laboratory through full-process standardization and intelligent control, reduce the demand for manual intervention and the probability of experimental abnormalities, and ensure that the experimental process is traceable, verifiable and meets the standard requirements. BRIEF DESCRIPTION OF DRAWINGS
[0049] Figure 1 It is a flow chart of the unmanned laboratory sample automatic regulation and control method based on reinforcement learning of the present application embodiment 1;
[0050] Figure 2 It is an interaction schematic diagram of the directed demand chain of the present application embodiment 1;
[0051] Figure 3 It is a structure diagram of the unmanned laboratory sample automatic regulation and control system based on reinforcement learning of the present application embodiment 2. DETAILED DESCRIPTION
[0052] Embodiment 1
[0053] For example, the present application provides an embodiment of a method for automatic regulation of unmanned laboratory samples based on reinforcement learning, comprising the following steps: Figure 1
[0054] S1, obtaining a directed demand chain corresponding to different experimental requirements; the directed demand chain is obtained by atomic sub-experiments corresponding to each experiment, the logical relationship between the sub-experiments, the priority of the sub-experiments of different experiments, the clustering algorithm and the construction of the block chain;
[0055] It should be further explained that the construction process of the directed demand chain in the embodiment includes:
[0056] Obtaining different experimental process information, and decomposing each experimental process information into an atomic sub-experiment sequence corresponding to the experiment according to the double rules of operation indivisibility and equipment singularity, wherein the operation indivisibility means that each sub-experiment is the smallest operation unit that cannot be further divided in the corresponding experimental process, and the equipment singularity means that the complete execution of the sub-experiment depends only on a unique experimental equipment, without the need for multiple equipment auxiliary operations; It should be further explained that the atomic sub-experiment includes the operation trigger condition, the execution time sequence priority, the standard execution effect and the standard execution efficiency of the corresponding sub-experiment; It should be further explained that the execution time sequence priority includes the time sequence priority of the corresponding sub-experiment in the experiment and the operation priority of different experiments facing the same experimental equipment; It should be further explained that the time sequence priority here refers to the time sequence of the execution of the corresponding sub-experiment under the same experiment, and it should be further explained that the operation priority of different experiments facing the same experimental equipment refers to the operation sequence of the sub-experiments of different experiments operating on the same experimental equipment;
[0057] Based on the atomic sub-experiment sequence corresponding to the same experiment, a first operation node sequence is constructed, and a first connection relationship is constructed based on the operation trigger condition between different sub-experiments under the same experiment and the time sequence priority of the corresponding sub-experiment in the experiment;
[0058] Based on the first operation node sequence and the first connection relationship, a horizontal demand chain corresponding to each experiment is obtained, and the standard execution effect and the standard execution efficiency of the corresponding sub-experiment are stored in the corresponding first operation node for saving;
[0059] It needs to be further explained that the standard execution effect and the standard execution efficiency described in the embodiment are based on historical experimental data of the same type, and specifically, in historical experiments, after strictly following the standard execution process to complete the operation, the experimental batches with the lowest overall abnormal rate and the highest execution efficiency are selected, and the average of the performance evaluation results and the efficiency evaluation results obtained by each sub-experiment in the detection stage is calculated respectively as the standard execution effect and the standard execution efficiency of each sub-experiment, which is used to evaluate and optimize the execution results and execution efficiency of the corresponding sub-experiment.
[0060] Based on the experimental equipment label of the corresponding sub-experiment of the horizontal demand chain, a horizontal mapping between each sub-experiment and the experimental equipment is constructed.
[0061] A primary clustering of initial sub-experiment clustering groups is obtained by clustering the unique identifiers of the experimental equipment corresponding to different experiments as parameters;
[0062] Based on the corresponding sub-experiment in each initial sub-experiment clustering group, the operation complexity of the corresponding sub-experiment is obtained by extracting the operation type quantitative features and operation object attribute quantitative features of each sub-experiment combined with a convolutional neural network; the operation complexity is constructed by the number of operation steps of the corresponding sub-experiment, the number of control parameters of each operation step, the size of the validation index parameter, and the size of the implementation time of the corresponding sub-experiment under the standard execution process; wherein the operation type quantitative features include operation type coding, operation precision requirement, automation operation level, and operation action step number; the operation object attribute quantitative features include but are not limited to sample state coding, sample concentration / purity, and consumable type coding;
[0063] Based on the operation priority of different experiments facing the same experimental step combined with the operation complexity of the corresponding sub-experiment, the influence degree of the implementation efficiency and the validation passing probability of the current sub-experiment on the next sub-experiment is minimized as the target, and the adaptive clustering algorithm is combined to perform secondary clustering on each initial sub-experiment clustering group to obtain a secondary clustering sub-group set;
[0064] A vertical connection relationship sequence is constructed based on the implementation order corresponding to the minimum influence degree of the implementation efficiency and the validation passing probability of all previous sub-experiments on the next sub-experiment in the current initial sub-experiment clustering group;
[0065] It needs to be further explained that one construction process of the vertical connection relationship sequence in the embodiment includes:
[0066] Based on the secondary clustering sub-group set, the feature data of all sub-experiments in each sub-group is extracted, including operation priority quantitative value, operation complexity quantitative value, implementation efficiency influence coefficient, and validation passing probability influence coefficient;
[0067] Based on the characteristic data of all sub-experiments in each sub-group, the one-way influence degree value between any two sub-experiments in each sub-group is obtained through an association algorithm;
[0068] Through a topological sorting algorithm, a preliminary longitudinal connection relationship sequence is generated with the sum of one-way influence degree values between all pairs of sub-experiments as the optimization objective;
[0069] Based on historical experimental data, the actual influence degree value of each previous sub-experiment on the next sub-experiment in the preliminary longitudinal connection relationship sequence is verified;
[0070] If the actual influence degree value of a continuous pair of sub-experiments exceeds a preset influence threshold, the relative position of the pair of sub-experiments in the sequence is adjusted;
[0071] The topological sorting and verification and adjustment steps are repeatedly executed until the actual influence degree values of all continuous pairs of sub-experiments in the entire sequence are lower than the preset influence threshold, and a final longitudinal connection relationship sequence is obtained;
[0072] The final longitudinal connection relationship sequence is associatedly stored with the corresponding secondary clustering sub-group set, and each sub-experiment is labeled with its position number in the sequence and the influence degree value with the predecessor and successor sub-experiments.
[0073] A main longitudinal node is constructed with the corresponding label of each experimental device, a longitudinal sub-node sequence is constructed with the corresponding secondary clustering sub-group set of each main longitudinal node, and a longitudinal cross-experimental node chain is constructed based on the longitudinal connection relationship sequence and the longitudinal sub-node sequence;
[0074] Based on the longitudinal cross-experimental node chain corresponding to each experimental device, the real-time state evaluation score of the corresponding experimental device, and the maximum load coefficient corresponding to the current real-time state evaluation score, a second load mapping connection is constructed.
[0075] It should be further noted that the second load mapping connection in the embodiment is used to establish a dynamic matching relationship between the experimental device and its load capacity, and its core role is to dynamically allocate the sub-experiment sequence planned in the longitudinal cross-experimental node chain to the most suitable physical device for execution according to the real-time performance state and the maximum load capacity of the device, so as to realize load balancing, prevent device overload, ensure the execution of sub-experiments in the optimal order, and finally improve the overall experimental efficiency and resource utilization.
[0076] The horizontal demand chain and the horizontal mapping, and the longitudinal cross-experimental node chain and the second load mapping connection are stored in the blockchain for on-chain storage, and a directed demand chain is obtained.
[0077] Please refer to Figure 2, assuming that n1, n2, n3 to nm are m sub-experiments corresponding to the nth experiment, a1, a2, a3 to am are process control verification chains constructed by one-to-one corresponding experimental equipment combined with corresponding control models, the connection between n1, n2, n3 to nm is a first connection relationship, the connection between n1, n2, n3 to nm and a1, a2, a3 to am is a horizontal mapping, g1, g2, g3 to gq are a sequence of vertical nodes with a3 as the main vertical node, and the directed connection between g1, g2, g3 to gq is a second load mapping connection corresponding to a1, a2, a3 to am;
[0078] Taking a biological sample concentration detection experiment as an example, the experiment includes 5 sub-experiments n1-n5, wherein n1: consumable barcode scanning confirmation, n2: sample sampling, n3: liquid addition and dilution, n4: mixing and shaking, and n5: spectral detection, corresponding to experimental equipment a1-a5, a1: Internet of Things terminal scanning equipment, a2: six-axis intelligent collaborative robot, a3: self-developed liquid addition system, a4: self-developed mixing system, and a5: self-developed spectral system; the first connection relationship between n1-n5 is n1→n2→n3→n4→n5 (the consumables need to be scanned and confirmed first, and then the spectral detection is performed after sampling, liquid addition, and mixing), and the horizontal mapping of n1-n5 and a1-a5 is n1 binding a1 (sample area scanning), n2 binding a2 (sample area sampling), n3 binding a3 (liquid addition area), n4 binding a4 (mixing and shaking area), and n5 binding a5 (finished product area); a3 (self-developed liquid addition system) is the main vertical node, and its vertical node sequence is g1-g3, wherein g1 (experimental A sample liquid addition and dilution), g2 (experimental B sample liquid addition and dilution), and g3 (quality control sample liquid addition and calibration), and the second load mapping connection between g1-g3 is g3→g1→g2 (according to the real-time state evaluation score and the maximum load coefficient of a3, the quality control calibration is preferentially performed to ensure the subsequent liquid addition accuracy and reduce the influence of the efficiency and pass rate between sub-experiments); the design of this embodiment not only realizes the standardization of the experimental process, the accurate matching of sub-experiments and self-developed equipment / laboratory area through horizontal mapping and the first connection relationship, and meets the characteristics of dataization of standard operation process in documents and reduction of personnel dependence, but also optimizes the cross-experiment scheduling of core equipment such as liquid addition system through the vertical node sequence and the second load mapping connection, avoids equipment overload, and realizes unmanned operation by combining Internet of Things scanning, intelligent robots, and self-developed equipment, improves experimental replicability and control, effectively guarantees sample configuration accuracy and experimental efficiency.
[0079] S2, a preset process control verification chain is mapped to each sub-experiment in the directed demand chain through a smart contract in the directed demand chain, and a pre-control instruction sequence is obtained in combination with a configured distributed reinforcement control model; the distributed reinforcement control model is constructed in combination with a reinforcement learning algorithm and an experimental standard verification chain by a sample configuration accuracy rate of each sub-experiment, a verification passing probability, a corresponding sub-experiment implementation efficiency, and an interaction risk transfer matrix; the interaction risk transfer matrix is used to measure the influence degree of the verification passing probability and the implementation efficiency of the corresponding sub-experiment of the current experiment on the verification passing probability and the implementation efficiency of the next sub-experiment of the same experiment or different experiments; the experimental standard verification chain is constructed by a standard operation step of each sub-experiment corresponding to each experiment and a standard evaluation index of each sub-experiment in combination with an evaluation algorithm; and the process control verification chain is constructed by an experimental equipment corresponding to each experiment and an operation priority of the corresponding experimental equipment in each experiment and a maximum real-time load coefficient of each experimental equipment;
[0080] The distributed reinforcement control model includes a federal control sub-model and M sub-device control models; each experimental equipment in the process control verification chain corresponds to each sub-experiment of the corresponding horizontal demand chain one by one and is sequentially labeled according to the implementation time of each sub-experiment corresponding to each experiment;
[0081] It should be further explained that the construction process of the distributed reinforcement control model in the embodiment includes:
[0082] Based on the standard evaluation index of the corresponding sub-experiment in the experimental standard verification chain, sample configuration detection data after actual execution of the corresponding sub-experiment is collected, the degree of coincidence between the actual detection data and the standard evaluation index is calculated, and the ratio of the number of samples with a coincidence degree greater than or equal to a standard threshold to the total number of detection samples is calculated and taken as the sample configuration accuracy rate of the corresponding sub-experiment; the detection data includes component detection values, concentration measured values, and purity detection results;
[0083] The effective execution number of each sub-experiment history N times of execution process that meets the sample configuration accuracy rate greater than or equal to a preset qualified threshold and the operation step meets the standard operation step of the sub-experiment in the experimental standard verification chain is extracted, and the ratio of the effective execution number to the total execution number N times is taken as the verification passing probability prior value of the corresponding sub-experiment;
[0084] Based on the verification passing probability prior value of the current sub-experiment as a weight, the verification evaluation index and real-time index detection data of the corresponding sub-experiment, and in combination with a comprehensive evaluation algorithm, the verification passing probability of the corresponding sub-experiment is obtained;
[0085] If the current sub-experiment is a new sub-experiment, the initial verification pass probability of the corresponding sub-experiment is obtained by multiplying the verification pass probability of the same type of sub-experiment by the performance coefficient of the corresponding experimental equipment at the current time; the performance coefficient at the current time is constructed by the ratio of the maximum real-time load coefficient at the current time to the rated load coefficient of the experimental equipment; the maximum real-time load coefficient is constructed by the maximum load experiment number corresponding to the life evaluation score of the corresponding experimental equipment obtained in real time;
[0086] The standard execution efficiency index of the corresponding sub-experiment in the experimental standard verification chain and the actual execution time T1 of the corresponding sub-experiment are extracted, and the product of the standard execution efficiency index, the actual execution time T1, the standard execution time T0, and the real-time load coefficient is taken as the implementation efficiency of the corresponding sub-experiment. 、 、
[0087] Wherein, T0 represents the standard execution time of the corresponding sub-experiment, represents the implementation efficiency of the k-1th sub-experiment implemented by the i th experiment on the j-1th experimental equipment , the probability risk transfer value of the implementation efficiency of the kth sub-experiment implemented by the i th experiment on the jth experimental equipment represents the implementation efficiency of the gth sub-experiment under the pth experiment on the longitudinal cross-experiment node chain corresponding to the jth experimental equipment , the probability risk transfer value of the implementation efficiency of the s th sub-experiment implemented by the i th experiment on the jth experimental equipment , the kth sub-experiment and the gth sub-experiment are the same sub-experiment, and the operation priority corresponding to the implementation of the s th sub-experiment is less than that of the gth sub-experiment.
[0088] The first horizontal sub-experiment pair continuously executed in the same horizontal demand chain and the second longitudinal sub-experiment pair corresponding to the same longitudinal cross-experiment node chain are taken as sample pairs, and the Bayesian algorithm is used to obtain 、 and 、 ; wherein represents the verification pass probability of the k-1th sub-experiment executed by the i th experiment on the j-1th experimental equipment on the horizontal demand chain , the probability risk transfer value of the corresponding verification pass probability of the kth sub-experiment executed by the i th experiment on the jth experimental equipment represents the verification pass probability of the gth sub-experiment under the pth experiment on the longitudinal cross-experiment node chain corresponding to the jth experimental equipment , the verification pass probability of the s th sub-experiment executed by the i th experiment on the jth experimental equipment The probability risk transmission value; This represents the s-th sub-experiment under the i-th experiment on the vertical cross-experiment node chain corresponding to the j-th experimental device. and the g-th sub-experiment under the adjacent p-th experiment The second longitudinal sub-experiment pair was formed; This indicates that in the horizontal demand chain, the i-th experiment executes the k-1-th sub-experiment on the j-1-th experimental device. The k-th sub-experiment performed on the j-th experimental device in relation to the i-th experiment. The first horizontal sub-experiment pair was constructed; for the sub-experiment implementation efficiency, the ratio of standard execution time T0 to actual execution time T1 reflects the time matching degree, and conditional probability is used. To quantify the lateral risk transmission of sub-experiment efficiency from preceding equipment (j-1) to subsequent equipment j within the same experiment, we use... On the same device j, the vertical risk transmission of the implementation efficiency of high-priority sub-experiment g to the implementation efficiency of low-priority sub-experiment s is quantified, and then combined with the real-time load coefficient to obtain the sub-experiment implementation efficiency.
[0089] , Its purpose is to quantify and manage the transmission relationship between two key efficiency risks in the experimental process, among which Used to assess the risk of longitudinal impact of the execution efficiency of sub-experiments on preceding devices on the execution efficiency of sub-experiments on subsequent devices within the same experiment, aiming to ensure the consistency and stability of the experimental process; This is used to evaluate the risk of horizontal interference between the execution efficiency of high-priority sub-experiments and the execution efficiency of low-priority sub-experiments on different experiments on the same device. It aims to achieve cross-experiment task scheduling optimization and resource contention management, thereby jointly ensuring the efficient and reliable operation of the overall experimental system. , They serve the same purpose, measuring the mutual influence between the probabilities of passing verification for different sub-experiments.
[0090] based on and The product and sum and The product is used to construct a risk transfer pair in which the i-th experiment implements the g-th sub-experiment on the j-th experimental device;
[0091] Based on the horizontal demand chain and vertical cross-experiment node chain corresponding to all experiments, and combined with the hidden Markov algorithm, the process of obtaining risk transfer pairs is repeated to obtain the real-time risk transfer matrix corresponding to all current experiments.
[0092] constructing an input state space of the sub-device control model based on the longitudinal cross-experimental node chain of each experimental device, the second load mapping connection data of each experimental device, the performance evaluation data of the corresponding sub-experiment, and the probability risk transfer value corresponding to each experimental device in the real-time risk transfer matrix, constructing an input state space of the federal control sub-model based on the horizontal demand chain data of all experiments, the horizontal mapping data of all experiments, and the output state of all sub-device control models and the real-time risk transfer matrix;
[0093] Based on the input state space of the sub-device control model and the input state space of the federal control sub-model, an execution action trigger information is constructed, specifically:
[0094] When the real-time evaluation result of any sub-experiment in the horizontal demand chain and the implementation efficiency do not meet the corresponding standard execution effect and standard execution efficiency, or the overall performance evaluation value of the corresponding experiment and the implementation efficiency do not meet the preset overall standard execution effect and overall standard execution efficiency, an abnormal detection positioning information is triggered, that is, through the horizontal mapping and the second load mapping connection between the federal control sub-model and the M sub-device control models, the horizontal demand chain and the longitudinal cross-experimental node chain are bidirectionally positioned, and the positioned abnormalities are identified through the abnormal identification model constructed by the random forest;
[0095] The corresponding positioning and identification of the abnormal position parameters and type parameters are fed back to the federal control sub-model, and an abnormal adjustment strategy is generated in combination with a preset abnormal adjustment strategy library;
[0096] The generated abnormal adjustment strategy is issued to the corresponding positioned abnormal device and associated abnormal device through the horizontal mapping and the second load mapping connection for adjustment until the evaluation result and the implementation efficiency of each sub-experiment corresponding to each horizontal demand chain meet the corresponding standard execution effect and standard execution efficiency;
[0097] It should be further explained that the execution action trigger information constructed in the embodiment includes sub-device control model trigger execution logic and federal control sub-model trigger execution logic, wherein the sub-device control model trigger execution logic is specifically:
[0098] When the real-time load coefficient of the trigger device exceeds the preset load safety threshold, the current real-time load coefficient, the maximum real-time load coefficient of the device, and the operation complexity of each sub-experiment in the longitudinal cross-experimental node chain of the device are collected;
[0099] Adjust the sub-experiment execution order based on the order from low to high operation complexity, and pause the execution of the sub-experiment with high operation complexity; calculate the difference between the current real-time load coefficient of the device and the preset load safety threshold to obtain the load excess amount;
[0100] The load exceeding amount and the adjusted sub-experiment execution sequence are fed back to the federal control sub-model; at the same time, the operation speed of the current sub-experiment is reduced to reduce the real-time load of the device, and the real-time load coefficient of the device is continuously monitored until the real-time load coefficient of the device is below the preset load safety threshold.
[0101] When the sub-experiment verification pass probability is triggered to be lower than the preset qualified threshold, the historical execution data of the sub-experiment is collected, the historical execution data including the effective execution times and the total execution times of the sub-experiment; the actual detection data after the actual execution of the sub-experiment is collected, the actual detection data including the detection information corresponding to the standard evaluation index, and the evaluation result information obtained by configuring the evaluation algorithm; it needs to be further explained that the evaluation algorithm corresponding to each sub-experiment is configured by a person skilled in the art according to the properties of the corresponding detection object;
[0102] The standard evaluation index of the sub-experiment in the experiment standard verification chain is obtained, and the actual detection data is compared with the standard evaluation index to locate the deviation item that does not match;
[0103] The operation parameters of the corresponding sub-experiment are adjusted according to the deviation item, the operation parameters including the precision control parameters and the execution time; the sub-experiment is re-executed and new verification data is collected;
[0104] The verification pass probability of the sub-experiment is updated based on the new verification data; if the updated verification pass probability is still lower than the preset qualified threshold, the standard operation parameters of the same type of sub-experiment are called to make a second adjustment to the operation parameters of the sub-experiment, and the steps of re-execution and verification pass probability updating are repeated until the verification pass probability of the sub-experiment reaches the preset qualified threshold.
[0105] When the sub-experiment implementation efficiency is triggered to be lower than the preset efficiency threshold, the actual execution time or the actual operation speed of the corresponding sub-experiment is collected; the standard execution efficiency index of the corresponding sub-experiment in the experiment standard verification chain is obtained, the standard execution efficiency index including the standard execution time and the standard operation average speed;
[0106] The real-time load coefficient of the corresponding device of the corresponding sub-experiment is collected, the difference between the preset efficiency threshold and the actual implementation efficiency of the corresponding sub-experiment is calculated to obtain the efficiency deviation; the causes of the efficiency deviation are analyzed, the causes including but not limited to high device load and unreasonable sub-experiment operation parameters; if the cause of the efficiency deviation is high device load, the information that the device load is high is fed back to the federal control sub-model, and the federal control sub-model is requested to coordinate the device load distribution;
[0107] If the efficiency deviation is caused by unreasonable operation parameters of the sub-experiment, the operation rate of the sub-experiment is adjusted according to the matching relationship between the standard operation average rate and the real-time load coefficient of the equipment; the sub-experiment is re-executed and the new actual implementation efficiency is calculated, and the efficiency deviation analysis and operation adjustment steps are repeated until the actual implementation efficiency of the sub-experiment reaches the preset efficiency threshold.
[0108] When the probability risk transfer value corresponding to the equipment in the real-time risk transfer matrix exceeds the preset risk threshold, the probability risk transfer value exceeding the preset risk threshold is extracted from the real-time risk transfer matrix;
[0109] According to the extracted probability risk transfer value, a corresponding sub-experiment pair is determined, which includes a sub-experiment pair in a horizontal demand chain and a sub-experiment pair in a vertical cross-experiment node chain;
[0110] The historical interaction data of the determined sub-experiment pair is collected, and the historical interaction data includes the verification passing probability of the previous sub-experiment, the implementation efficiency of the previous sub-experiment, and the verification passing probability change rate of the next sub-experiment, and the implementation efficiency change rate of the next sub-experiment;
[0111] Based on the collected historical interaction data of the sub-experiment pair, the probability risk transfer value exceeding the preset risk threshold is corrected combined with the Bayes formula;
[0112] Based on the correction result, it is judged whether the corrected probability risk transfer value still exceeds the preset risk threshold;
[0113] If the corrected probability risk transfer value still exceeds the preset risk threshold, the execution parameters of the previous sub-experiment in the sub-experiment pair are adjusted, the execution parameters include prolonging the execution time of the previous sub-experiment to improve the verification passing probability of the previous sub-experiment, and reducing the risk impact of the previous sub-experiment on the next sub-experiment; and the probability risk transfer value in the real-time risk transfer matrix is updated.
[0114] When a new sub-experiment in the vertical cross-experiment node chain or the operation priority adjustment of the sub-experiment is triggered, the specific triggering scenario is first judged, if it is a new sub-experiment in the vertical cross-experiment node chain, the sub-equipment control model first collects the operation complexity, operation priority and standard execution parameter of the new sub-experiment, the standard execution parameter contains the standard execution effect and the standard execution efficiency, and then the new sub-experiment is classified into the secondary clustering sub-group set corresponding to the equipment;
[0115] If it is the operation priority adjustment of the sub-experiment in the vertical cross-experiment node chain, the sub-equipment control model first collects the new operation priority of the adjusted sub-experiment, and then reorders the execution order of all sub-experiments in the vertical cross-experiment node chain based on the new operation priority;
[0116] According to the load demand of the added sub-experiment or the adjusted load change value of the sub-experiment operation priority, the maximum real-time load coefficient matching relationship of the corresponding equipment is adjusted through the sub-equipment control model corresponding to the current experimental equipment, and the second load mapping connection of the corresponding equipment is updated; the updated longitudinal cross-experimental node chain data is fed back to the federal control sub-model.
[0117] It should be further pointed out that the federal control sub-model triggering execution logic in the embodiment is as follows:
[0118] When triggering any sub-experiment in the horizontal demand chain that does not meet the operation trigger condition, locate the sub-experiment in the horizontal demand chain that does not meet the operation trigger condition, and mark it as the current sub-experiment;
[0119] Determine the pre-sub-experiment of the current sub-experiment, and the operation trigger condition includes that the pre-sub-experiment is not completed or the pre-sub-experiment does not meet the standard;
[0120] Collect the execution state of the pre-sub-experiment, and the execution state includes the completion progress and verification result of the pre-sub-experiment;
[0121] Determine the execution state of the pre-sub-experiment, if the pre-sub-experiment is not completed, send an acceleration execution instruction to the sub-equipment control model corresponding to the pre-sub-experiment, and the acceleration execution instruction is formulated based on the operation complexity of the pre-sub-experiment to adjust the execution rate of the pre-sub-experiment; if the pre-sub-experiment does not meet the standard, send a re-execution instruction to the sub-equipment control model corresponding to the pre-sub-experiment, and simultaneously issue an operation parameter optimization suggestion, instruct the sub-equipment control model to re-execute the pre-sub-experiment and optimize the operation parameter; continuously monitor the execution state of the pre-sub-experiment, and when the pre-sub-experiment meets the operation trigger condition, send an execution instruction to the sub-equipment control model corresponding to the current sub-experiment; at the same time, adjust the execution time sequence of all subsequent sub-experiments in the horizontal demand chain according to the execution instruction of the current sub-experiment.
[0122] When the timing deviation of the execution of the sub-experiments in the transverse demand chain exceeds the preset timing threshold, the first connection relationship of the transverse demand chain is obtained, the first connection relationship includes the logical and constraint conditions between the sub-experiments, and the timing deviation is that the actual execution sequence does not conform to the first connection relationship; the actual execution sequence and the execution progress of each sub-experiment in the current transverse demand chain are collected; the step difference between the actual execution sequence and the logical sequence of the sub-experiments in the first connection relationship is calculated by comparing the actual execution sequence with the logical sequence of the sub-experiments in the first connection relationship, and the timing deviation amount is obtained; the sub-experiments involved in the timing deviation are determined according to the timing deviation amount; a timing adjustment instruction is sent to the sub-device control model corresponding to the sub-experiments involved in the timing deviation; the timing adjustment instruction includes pausing the sub-experiments that deviate from the sequence, and simultaneously instructing the sub-experiments that should be in front in the first connection relationship to be executed preferentially; the actual execution sequence of the sub-experiments is continuously monitored, and after the actual execution sequence conforms to the logical sequence of the experiments in the first connection relationship, an instruction to resume the execution of the subsequent sub-experiments is sent to the sub-device control model; and the execution timing record of the transverse demand chain is updated to record the adjusted execution sequence and execution progress of the sub-experiments.
[0123] When the output state of any sub-device control model triggers its own trigger information, the trigger information includes that the device load exceeds the preset load safety threshold, and the sub-experiment verification pass probability is lower than the preset qualified threshold;
[0124] The trigger information fed back by the sub-device control model that triggers the own trigger information is received; the current device state data of the device corresponding to the sub-device control model is received, and the device state data includes the real-time load coefficient of the device, the sub-experiment performance data, and the probability risk transfer value;
[0125] The associated devices that interact with the sub-experiment of the device corresponding to the sub-device control model are determined;
[0126] The device state data of all associated devices is collected; the trigger information type of the sub-device control model is judged, if the trigger information type is that the device load exceeds the preset load safety threshold, then based on the corresponding relationship between the sub-experiments and the experimental devices in the transverse mapping, the low-priority sub-experiments on the device are screened; the associated devices with idle load are coordinated, and the screened low-priority sub-experiments are transferred to the associated devices with idle load for execution;
[0127] If the trigger information type is that the sub-experiment verification pass probability is lower than the preset qualified threshold, the global standard parameter template in the experimental standard verification chain is called;
[0128] The operation parameter adjustment strategy of the sub-device control model is formulated based on the global standard parameter template, and is delivered to the sub-device control model; at the same time, the operation parameter adjustment strategy is broadcast to other similar sub-device control models through the blockchain for reference and adjustment by the similar sub-device control models.
[0129] extracting, from the real-time risk transfer matrix, a cross-device or cross-experiment risk transfer pair when a mean value of implementation efficiency and verification pass probability of a cross-device or cross-experiment risk transfer pair in the real-time risk transfer matrix exceeds a preset global risk threshold, the cross-device or cross-experiment risk transfer pair specifically including a product of a probability risk transfer value of an implementation efficiency of a sub-experiment of an i-th experiment on a j-th experimental device in a horizontal demand chain and a probability risk transfer value of an implementation efficiency of a corresponding sub-experiment of the j-th experimental device on a vertical cross-experiment node chain, and a product of a probability risk transfer value of a verification pass probability of the sub-experiment of the i-th experiment on the j-th experimental device in the horizontal demand chain and a probability risk transfer value of a verification pass probability of the corresponding sub-experiment of the j-th experimental device on the vertical cross-experiment node chain; and determining, according to the extracted risk transfer pair, a corresponding sub-experiment pair (including a same-device cross-sub-experiment pair, a different-device same-experiment corresponding sub-experiment pair) and a device pair (including a cross-experiment associated device pair, a same-experiment multi-device pair);
[0130] collecting real-time interaction risk data of the determined sub-experiment pair and device pair, the data including an efficiency influence value, a pass rate influence value, and device load change data of a preceding sub-experiment pair on a following sub-experiment in real-time execution;
[0131] calculating, based on a hidden Markov algorithm, the collected real-time interaction risk data to obtain a risk transfer pair of a current experiment scenario; correcting, using the obtained risk transfer pair, risk transfer pair data at a corresponding position in the real-time risk transfer matrix; calculating a mean value of all cross-device or cross-experiment risk transfer pairs in the corrected real-time risk transfer matrix; determining whether the mean value still exceeds the preset global risk threshold; if the mean value still exceeds the preset global risk threshold, adjusting an execution time period of a cross-device sub-experiment, so that execution time periods of sub-experiments with a high risk transfer relationship are staggered to avoid parallel execution, or the operation priority of the cross-sub-experiment is reduced (low-risk cross-sub-experiments are preferentially executed); and issuing, by a federal control sub-model, a risk control instruction to related sub-device control models involved in the risk transfer pair, instructing the sub-device control models to execute the adjusted sub-experiment execution time period or operation priority, and synchronizing the adjustment result to the real-time risk transfer matrix to complete updating.
[0132] when a newly added experiment demand triggers an update of a horizontal demand chain or a vertical cross-experiment node chain, collecting flow information of the newly added experiment; and decomposing, according to an operation indivisibility rule and a device singularity rule, the flow information of the newly added experiment into an atomic sub-experiment sequence;
[0133] constructing a first operation node sequence of the newly added experiment based on the atomic sub-experiment sequence;
[0134] constructing a first connection relationship of the newly added experiment based on an operation trigger condition between atomic sub-experiments and an execution operation priority of a sub-experiment in the experiment; and integrating the first operation node sequence and the first connection relationship to obtain a horizontal demand chain of the newly added experiment.
[0135] constructing a horizontal mapping of the new experiment based on the equipment labels of each sub-experiment of the new experiment, the horizontal mapping including the correspondence between each sub-experiment of the new experiment and the corresponding experimental equipment;
[0136] extracting the unique identification of the equipment of each sub-experiment of the new experiment; and distributing the sub-experiments of the new experiment to the vertical cross-experiment node chain of the corresponding equipment according to the unique identification of the equipment;
[0137] sending an instruction to the sub-equipment control model corresponding to the equipment of the sub-experiment, instructing the sub-equipment control model to update the secondary clustering sub-group set and the sub-experiment operation priority ranking; analyzing the risk transmission relationship between the new sub-experiment and the existing sub-experiment; incorporating the risk transmission relationship into the real-time risk transfer matrix, and updating the global real-time risk transfer matrix; synchronizing the updated horizontal demand chain and vertical cross-experiment node chain data to all sub-equipment control models;
[0138] re-generating the pre-control instruction sequence based on the updated horizontal demand chain, vertical cross-experiment node chain data, and real-time risk transfer matrix.
[0139] taking the difference between the sub-experiment sample configuration accuracy and the preset accuracy threshold, the difference between the sub-experiment verification pass probability and the preset qualified threshold, and the difference between the sub-experiment implementation efficiency and the preset efficiency threshold as positive reward items;
[0140] Based on the directed demand chain and the process control verification chain, a distributed reinforcement control model is constructed by deploying a federated learning framework, which sets the federated control sub-model (Model_T) as the parameter aggregation node, and each sub-equipment control model (Model_1~Model_M) as the local training node; based on the vertical cross-experiment node chain data, real-time load data, and device trigger information related data (including but not limited to load data when the real-time load coefficient of the equipment exceeds the preset load safety threshold, and execution data when the sub-experiment verification pass probability is lower than the qualified threshold) of the corresponding experimental equipment of each sub-equipment control model, local training is performed without sharing the original data; through a secure aggregation algorithm such as federated averaging algorithm, the encrypted model parameters uploaded by each local training node are weighted and averaged, wherein the weight is positively correlated with the sub-experiment execution frequency of the corresponding equipment and the response success rate in the trigger scenario, and abnormal parameters deviating from the parameter mean value by more than 2 times the standard deviation are excluded, to generate global update parameters; by distributing the global update parameters to each local training node to achieve parameter synchronization, and preferentially retaining parameters with good training effect in the trigger scenario during the aggregation process;
[0141] Model initialization is achieved by inputting the initial sample configuration accuracy, initial validation pass probability, initial implementation efficiency, initial interaction risk transfer matrix, and initial threshold corresponding to the device trigger information of each sub-experiment in the vertical cross-experimental node chain of the corresponding experimental equipment into the sub-device control model. The initial sample configuration accuracy is obtained based on the average of historical data from similar sub-experiments; the initial validation pass probability is obtained by combining the historical validation pass probability of similar sub-experiments with device performance coefficient correction; the initial implementation efficiency is obtained by assigning a standard execution efficiency value; and the initial interaction risk transfer matrix is obtained based on the historical influence coefficient matrix of similar sub-experiments. Hyperparameter initialization is achieved by initializing the model weights of each sub-device control model to random normal distribution values and setting the total number of training iterations N and the learning rate η. Finally, the federated control sub-model initialization is achieved by inputting the horizontal demand chain, horizontal mapping, initial output of each sub-device control model, and initial threshold corresponding to the trigger information of the federated control sub-model into Model_T, and initializing the global parameters of Model_T to be consistent with the initial weights of each local training node.
[0142] The reward design is based on the reward function of the sub-device control model (R_score) and the reward function of the federated control sub-model (R_total). R_score is obtained by weighting the difference between the sub-experiment sample configuration accuracy and the accuracy threshold, the difference between the sub-experiment verification pass probability and the preset qualified threshold, the difference between the sub-experiment implementation efficiency and the preset efficiency threshold, and the absolute deviation between the real-time load coefficient of the device and the preset load safety threshold. The weight coefficients are dynamically adjusted according to the triggering scenario. R_total is obtained by weighting the reward of each sub-device control model, the compliance rate of the horizontal demand chain sub-experiment timing, and the deviation of the cross-device or cross-sub-experiment interaction risk value from the preset global risk threshold. The weight coefficients are also dynamically adjusted according to the triggering scenario of the federated control sub-model.
[0143] Federated training is achieved through iterative execution of local training, parameter aggregation, and model update steps: In the local training phase, each sub-device control model trains its local model using reinforcement learning algorithms, based on real-time data from the corresponding device's sub-experiments, dynamically updated data across the vertical experimental node chain, and actual response data to device trigger information, with the goal of maximizing R_score. Local parameters are iteratively updated. In the parameter aggregation phase, each local training node uploads the trained encrypted model parameters to Model_T. Model_T uses a federated averaging algorithm to calculate the weighted average and remove outlier parameters, generating globally updated parameters. In the model update phase, Model_T distributes the global parameters to each sub-device control model to update its local parameters. Based on the globally updated parameters and the horizontal demand chain, it executes data-optimized global control strategies and trigger information judgment logic in real time.
[0144] The trained distributed reinforcement control model and the operation parameters thereof are obtained through continuous iteration until a model convergence condition is met, i.e., the R_ sub-fluctuation amplitude of each sub-device control model and the R_total fluctuation amplitude of the federal control sub-model are continuously lower than a set fluctuation threshold, and the success rate of abnormal adjustment of a sub-experiment under a triggered scenario is not lower than 95%, or the number of iterations reaches the total number of iterations N, and the trained distributed reinforcement control model and the operation parameters thereof are used to generate a pre-control instruction sequence.
[0145] S3, executing the pre-control instruction sequence, and detecting real-time state data of each experiment corresponding sub-experiment in real time, and feeding back the real-time detected state data to the distributed reinforcement control model to verify the sample configuration and the experimental operation process of the corresponding experiment in combination with the experimental standard verification chain of the corresponding experiment, until the corresponding experiment is completed.
[0146] The embodiment transforms complex experiments into minimal and unique atomic sub-experiments by operating the inseparability and device singularity double rule disassembly process, and clearly defines the operation trigger condition, execution time priority, standard execution effect and standard execution efficiency of each sub-experiment, effectively avoiding the process confusion problem caused by operation splitting ambiguity and multi-device cross assistance in traditional experiments; at the same time, based on the atomic sub-experiment, a horizontal demand chain and a horizontal mapping are constructed, so that the sub-experiment is accurately bound with the experimental equipment and the laboratory area, and the sub-experiment execution sequence is combined with the first connection relationship to realize the standardization landing of the experimental process, which is consistent with the characteristics of the data of the standard operation process in the document, greatly reduces the dependence on the experience of the operator, and even a novice or an automatic device can perform the experiment according to the unified standard; in addition, the directed demand chain is stored on the blockchain, and the data of the horizontal demand chain, the vertical cross-experiment node chain and the load mapping are tamper-proof, which can trace the execution equipment, operation parameters and result data of each sub-experiment in real time, which is convenient for experimental problem backtracking and quality control, and is especially suitable for the scene with high experimental repeatability requirement in the field of life science. Secondly, the embodiment reasonably allocates the sub-experiments of different experiments to the vertical cross-experiment node chain of the corresponding device through primary clustering (according to the unique identifier of the device) and secondary clustering (according to the operation priority, complexity and minimization of the influence between sub-experiments), and dynamically matches the real-time state and load capacity of the device by combining the second load mapping connection, for example, the self-developed liquid adding system a3 is sorted as g3→g1→g2 according to the real-time state evaluation score and the maximum load coefficient, and the quality control calibration is preferentially executed, so that the single device is not overloaded or idle and wasted due to factor experiment accumulation, and the device utilization rate is significantly improved; at the same time, the federal control sub-model in the distributed reinforcement control model can coordinate the associated devices to share the load, for example, when the load of a certain device exceeds the threshold, the low-priority sub-experiment is transferred to the idle associated device, which further optimizes the cross-device resource scheduling, reduces the idle time and overload loss of the device, prolongs the service life of the device, and reduces the operation and maintenance cost of the laboratory equipment; and the construction of the vertical cross-experiment node chain enables the same device to efficiently process similar sub-experiments of multiple experiments, for example, the liquid adding system can add and dilute and perform quality control calibration for experiments A and B, which avoids the efficiency loss caused by frequent switching of the operation type of the device and improves the processing efficiency of the device.Third, the distributed reinforcement control model in this embodiment constructs an evaluation system through core indicators such as sub-experiment configuration accuracy rate, verification passing probability, and implementation efficiency, and combines an experimental standard verification chain (including sub-experiment standard operation steps and standard evaluation indicators) to measure experimental quality in real time. For example, based on the spectral system, the accuracy rate is calculated by collecting component, concentration, and purity data, and the standard execution effect is determined by comparing the mean value of the historical optimal experimental batch to ensure that the sub-experiment results are accurate and meet the standards. The design of the interaction risk transfer matrix and the real-time risk transfer matrix in the model can quantify the efficiency and passing probability influence between sub-experiments. For example, the vertical influence of the previous equipment sub-experiment on the subsequent equipment in the same experiment, and the horizontal influence of different sub-experiments of the same equipment. The risk transfer value is corrected by the Bayesian algorithm and the hidden Markov algorithm to avoid quality problems caused by high-risk sub-experiment interaction in advance. Fourth, this embodiment constructs a double-layer trigger execution logic of the sub-equipment control model and the federal control sub-model, covering all-scenario abnormalities such as device load, sub-experiment performance, timing deviation, and cross-device risk. When the device load exceeds the safety threshold, the sub-equipment control model can adjust the execution order and reduce the operation speed according to the operation complexity of the sub-experiment, and at the same time feedback the federal control sub-model to coordinate the load. When the sub-experiment passing probability is not up to standard, the deviation item is located by comparing the actual detection data with the standard evaluation indicators, the operation parameters are adjusted, and the execution is restarted. When the cross-device risk exceeds the global risk threshold, the federal control sub-model can adjust the sub-experiment execution period or priority to avoid risk transfer expansion. This abnormality processing mechanism of local response-global coordination can quickly locate and solve problems in the experimental process. For example, when there is a timing deviation in the horizontal demand chain of the sub-experiment, the federal control sub-model can pause the deviated sub-experiment and execute the preceding sub-experiment preferentially to ensure that the process returns to the right track and avoid the failure of the entire experiment due to local abnormalities, thereby significantly improving the experimental fault tolerance and stability. At the same time, real-time feedback state data is provided when executing the pre-control instruction, and the operation is dynamically adjusted in combination with the experimental standard verification chain to further ensure the controllability of the quality in the experimental process. Even if a sudden situation such as temporary performance fluctuation of the equipment occurs, it can be corrected in time to ensure the completion quality of the experiment. Fifth, the deployment of the federal learning framework enables each sub-equipment control model to train based on only local sub-experiment data without sharing original data (such as device load data and sub-experiment detection data), effectively protecting sensitive laboratory experimental data such as sample components and experimental schemes in the field of life sciences, and avoiding the risk of data leakage. The federal control sub-model aggregates node parameters to generate global update parameters through a secure aggregation algorithm, realizes synchronous optimization of model parameters, utilizes the full experimental data to improve the model generalization ability, and also considers data privacy protection, which meets the requirements of laboratory data security management. In addition, simulation trigger scenarios are provided in the model training process to optimize the response strategy of the model to abnormal scenarios. With the accumulation of experimental data, the model can be iteratively updated to further improve the control accuracy and abnormality processing capability.In summary, the embodiment realizes the transformation of sample configuration in an unmanned laboratory from manual dependence to automation, standardization and precision through the standardized construction of a directed demand chain, precise regulation of a distributed reinforcement control model, safe optimization of federated learning and exception protection of a double-layer trigger execution logic, has significant advantages in improving experimental efficiency and quality, optimizing equipment resource utilization, ensuring data security and reducing personnel risk, and provides strong support for the actual landing and large-scale application of unmanned laboratories.
[0147] Embodiment 2:
[0148] Please refer to Figure 3 Another embodiment provided by the application is an unmanned laboratory sample automatic regulation system based on reinforcement learning, comprising a demand chain module, an analysis mapping module and an execution module.
[0149] The demand chain module is configured to obtain a directed demand chain corresponding to different experimental demands, wherein the directed demand chain is constructed by atomic sub-experiments corresponding to each experiment, logical relationships between the sub-experiments, priority of the sub-experiments of different experiments, a clustering algorithm and a blockchain.
[0150] The analysis mapping module is configured to preset a process control verification chain, map each sub-experiment in the directed demand chain to the process control verification chain through a smart contract in the directed demand chain, and obtain a pre-control instruction sequence in combination with a configured distributed reinforcement control model.
[0151] The distributed reinforcement control model is constructed by a sub-experiment sample configuration accuracy, a verification passing probability, a corresponding sub-experiment implementation efficiency and an interaction risk transfer matrix, in combination with a reinforcement learning algorithm and an experimental standard verification chain. The interaction risk transfer matrix is configured to measure the influence degree of the verification passing probability and the implementation efficiency of a corresponding sub-experiment on the verification passing probability and the implementation efficiency of a next sub-experiment of the same experiment or different experiments. The experimental standard verification chain is constructed by standard operation steps of a sub-experiment corresponding to each experiment and standard evaluation indexes of each sub-experiment in combination with an evaluation algorithm. The process control verification chain is constructed by experimental equipment corresponding to each experiment, operation priority of the corresponding experimental equipment in each experiment and a maximum real-time load coefficient of each experimental equipment.
[0152] The execution module is configured to execute the pre-control instruction sequence.
[0153] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make changes, modifications, replacements and variations to the above-described embodiments without departing from the purpose of the present application and the scope protected by the claims, and these are all within the protection of the present application.
Claims
1. A method for automated sample conditioning in a laboratory without human intervention based on reinforcement learning, characterized in that, The method comprises the following steps: obtaining a directed demand chain corresponding to different experimental requirements; the directed demand chain is constructed by atomic sub-experiments corresponding to each experiment, logical relationships between the sub-experiments, priority of the sub-experiments of different experiments, a clustering algorithm and a blockchain; presetting a process control verification chain, mapping each sub-experiment in the directed demand chain to the process control verification chain through a smart contract in the directed demand chain, and obtaining a pre-control instruction sequence by combining a configured distributed reinforcement control model; the distributed reinforcement control model is constructed by combining a reinforcement learning algorithm and an experimental standard verification chain, and comprises sample configuration accuracy of each sub-experiment, verification passing probability, corresponding sub-experiment implementation efficiency and an interaction risk transfer matrix; the interaction risk transfer matrix is used to measure the influence degree of the verification passing probability and the implementation efficiency of the sub-experiment corresponding to the current experiment on the verification passing probability and the implementation efficiency of the next sub-experiment of the same experiment or different experiments; the experimental standard verification chain is constructed by combining an evaluation algorithm with standard operation steps of the sub-experiment corresponding to each experiment and standard evaluation indexes of each sub-experiment; the process control verification chain is constructed by combining an operation priority of the experimental equipment corresponding to each experiment in each experiment and a maximum real-time load coefficient of each experimental equipment; executing the pre-control instruction sequence.
2. The reinforcement learning based unattended laboratory sample robotic method of claim 1, wherein, The construction process of the directed demand chain comprises the following steps: obtaining different experimental process information, and decomposing each experimental process information into an atomic sub-experiment sequence of the corresponding experiment according to operation indivisibility and equipment singularity; the operation indivisibility means that each sub-experiment is the smallest operation unit in the corresponding experimental process, and the equipment singularity means that complete execution of the sub-experiment depends on only one experimental equipment; the atomic sub-experiment comprises an operation trigger condition, an execution time sequence priority, a standard execution effect and a standard execution efficiency of the corresponding sub-experiment; the execution time sequence priority comprises a time sequence priority of the corresponding sub-experiment in the experiment and an operation priority of different experiments facing the same experimental equipment; constructing a first operation node sequence by using the atomic sub-experiment sequence corresponding to the same experiment, and constructing a first connection relationship by using the operation trigger condition between different sub-experiments of the same experiment and the time sequence priority of the corresponding sub-experiment in the experiment; obtaining a horizontal demand chain corresponding to each experiment based on the first operation node sequence and the first connection relationship, and storing the standard execution effect and the standard execution efficiency of the corresponding sub-experiment into the corresponding first operation node for storage; constructing a horizontal mapping between each sub-experiment and the experimental equipment based on the experimental equipment label of the corresponding sub-experiment of the horizontal demand chain.
3. The reinforcement learning based unattended laboratory sample robotic method of claim 2, wherein, The construction process of the directed demand chain further comprises the following steps: performing one-level clustering by taking the unique identifier of the experimental equipment of the sub-experiment corresponding to different experiments as a parameter, and obtaining an initial sub-experiment clustering group; Based on the corresponding sub-experiment in each initial sub-experiment clustering group, the operation type quantitative features and the operation object attribute quantitative features of each sub-experiment are extracted by combining a convolutional neural network to obtain the operation complexity of the corresponding sub-experiment; the operation complexity is constructed by the number of operation steps of the corresponding sub-experiment, the control parameter quantity of each operation step, the verification index parameter size, and the size of the implementation duration of the corresponding sub-experiment under the standard execution flow; Based on the operation priority of the same experiment step of different experiments and the operation complexity of the corresponding sub-experiment, the influence degree of the current sub-experiment on the implementation efficiency and the verification passing probability of the next sub-experiment is minimized as the target, and the adaptive clustering algorithm is combined to perform secondary clustering on each initial sub-experiment clustering group to obtain a secondary clustering sub-group set.
4. The reinforcement learning based unattended laboratory sample robotic method of claim 3, wherein, The construction process of the directed demand chain also includes: A vertical connection relationship sequence is constructed according to the implementation sequence corresponding to the minimum influence degree of all previous sub-experiments on the implementation efficiency and the verification passing probability of the next sub-experiment in the current initial sub-experiment clustering group; A main vertical node is constructed according to the corresponding label of each experimental equipment, a vertical sub-node sequence is constructed according to the corresponding secondary clustering sub-group set of each main vertical node, and a vertical cross-experiment node chain is constructed based on the vertical connection relationship sequence and the vertical sub-node sequence; Based on the vertical cross-experiment node chain corresponding to each experimental equipment, the real-time state evaluation score of the corresponding experimental equipment, and the maximum load coefficient corresponding to the current real-time state evaluation score, a second load mapping connection is constructed; The horizontal demand chain and the horizontal mapping and the vertical cross-experiment node chain and the second load mapping connection are stored in the blockchain for on-chain storage to obtain a directed demand chain.
5. The reinforcement learning based unattended laboratory sample robotic method of claim 4, wherein, The distributed reinforcement control model includes a federal control sub-model and M sub-device control models; each experimental equipment in the process control verification chain corresponds to each sub-experiment of the corresponding horizontal demand chain and is sorted and labeled in the order of the implementation time of each experiment corresponding to the sub-experiment; the construction process of the distributed reinforcement control model includes: Based on the standard evaluation index of the corresponding sub-experiment in the experimental standard verification chain, sample configuration detection data after actual execution of the corresponding sub-experiment are collected, the degree of coincidence between the actual detection data and the standard evaluation index is calculated, and the ratio of the number of samples with a coincidence degree greater than or equal to a standard threshold to the total number of detection samples is calculated and taken as the sample configuration accuracy of the corresponding sub-experiment; the detection data includes component detection values, concentration measured values, and purity detection results; The effective execution times of each sub-experiment in the historical N execution processes that meet the sample configuration accuracy greater than or equal to a preset qualified threshold and the operation steps that meet the standard operation steps of the sub-experiment in the experimental standard verification chain are extracted, and the ratio of the effective execution times to the total execution times N is taken as the verification passing probability prior value of the corresponding sub-experiment; Based on the verification passing probability prior value of the current sub-experiment as a weight, the verification evaluation index and the real-time index detection data of the corresponding sub-experiment, and a comprehensive evaluation algorithm, the verification passing probability of the corresponding sub-experiment is obtained. If the current sub-experiment is a new sub-experiment, an initial verification passing probability of the corresponding sub-experiment is obtained by multiplying the verification passing probability of the same type of sub-experiment by a performance coefficient of the corresponding experimental equipment at a current time; the performance coefficient at the current time is obtained by constructing a ratio of a maximum real-time load coefficient of the experimental equipment at the current time to a rated load coefficient of the equipment.
6. The reinforcement learning based unattended laboratory sample robotic method of claim 5, wherein, The construction process of the distributed reinforcement control model further includes: The standard execution efficiency index of the corresponding sub-experiment in the extraction experiment standard verification chain, the actual execution time T1 of the corresponding sub-experiment, and the product of the real-time load coefficient are multiplied as the implementation efficiency of the corresponding sub-experiment. , , wherein T0 represents a standard execution duration of the corresponding sub-experiment, represents the implementation efficiency of the kth sub-experiment of the ith experiment implemented on the jth experimental device , the implementation efficiency of the kth sub-experiment of the ith experiment implemented on the jth experimental device , the probability risk transfer value of the implementation efficiency of the kth sub-experiment of the ith experiment implemented on the jth experimental device represents the implementation efficiency of the gth sub-experiment of the pth experiment on the longitudinal cross-experimental node chain corresponding to the jth experimental device , the implementation efficiency of the s th sub-experiment of the ith experiment implemented on the jth experimental device , the probability risk transfer value of the implementation efficiency of the s th sub-experiment of the ith experiment implemented on the jth experimental device, the kth sub-experiment and the gth sub-experiment are the same sub-experiment, and the operation priority corresponding to the s th sub-experiment is less than that of the gth sub-experiment.
7. The reinforcement learning based unattended laboratory sample robotic method of claim 6, wherein, The construction process of the distributed reinforcement control model further includes: A first horizontal sub-experiment pair continuously performed within the same horizontal demand chain A second longitudinal sub-experiment pair corresponding to the same longitudinal cross-experiment node chain As a sample pair, through the Bayesian algorithm, the following is obtained , and , ; wherein represents the verification passing probability of the i-th experiment performing the k-1-th sub-experiment on the j-1-th experiment equipment on the horizontal demand chain , the probability risk transfer value of the verification passing probability of the i-th experiment performing the k-th sub-experiment on the j-th experiment equipment ; represents the verification passing probability of the g-th sub-experiment under the p-th experiment on the longitudinal cross-experiment node chain corresponding to the j-th experiment equipment , the probability risk transfer value of the verification passing probability of the i-th experiment performing the s-th sub-experiment on the j-th experiment equipment ; represents the second longitudinal sub-experiment pair composed of the s-th sub-experiment under the i-th experiment and the g-th sub-experiment under the adjacent p-th experiment on the longitudinal cross-experiment node chain corresponding to the j-th experiment equipment ; represents the first horizontal sub-experiment pair constructed by the k-1-th sub-experiment of the i-th experiment performed on the j-1-th experiment equipment and the k-th sub-experiment of the i-th experiment performed on the j-th experiment equipment based on and The product and sum and The product is used to construct a risk transfer pair in which the i-th experiment implements the g-th sub-experiment on the j-th experimental device; Based on the transverse demand chain and the longitudinal cross-experiment node chain corresponding to all experiments, and in combination with a hidden Markov algorithm, a process of repeatedly obtaining risk transfer pairs is repeated to obtain a real-time risk transfer matrix corresponding to all experiments.
8. The reinforcement learning based unattended laboratory sample robotic method of claim 7, wherein, The construction process of the distributed reinforcement control model further includes: Based on the longitudinal cross-experiment node chain of each experimental equipment, the second load mapping connection data of each experimental equipment, the performance evaluation data of the corresponding sub-experiment, and the probability risk transfer value corresponding to each experimental equipment in the real-time risk transfer matrix, an input state space of a sub-equipment control model is constructed, and based on the transverse demand chain data of all experiments, the transverse mapping data of all experiments, the output state of all sub-equipment control models, and the input state space of a federal control sub-model, an input state space of a federal control sub-model is constructed.
9. The reinforcement learning based unattended laboratory sample robotic method of claim 8, wherein, The construction process of the distributed reinforcement control model further includes: Based on the input state space of the sub-equipment control model and the input state space of the federal control sub-model, an execution action trigger information is constructed; When the real-time evaluation result and the implementation efficiency of any sub-experiment in the transverse demand chain do not satisfy the corresponding standard execution effect and standard execution efficiency, or the overall performance evaluation value and the implementation efficiency of the corresponding experiment do not satisfy the preset overall standard execution effect and overall standard execution efficiency, an abnormality detection positioning information is triggered, that is, the transverse demand chain and the longitudinal cross-experiment node chain are bidirectionally positioned for abnormalities through the transverse mapping and the second load mapping connection between the federal control sub-model and the M sub-equipment control models, and the positioned abnormalities are identified through an abnormality identification model constructed by a random forest; The corresponding positioned abnormal position parameters and type parameters are fed back to the federal control sub-model, and an abnormality adjustment strategy is generated in combination with a preset abnormality adjustment strategy library; The generated abnormality adjustment strategy is issued to the corresponding positioned abnormal equipment and associated abnormal equipment through the transverse mapping and the second load mapping connection for adjustment until the evaluation result and the implementation efficiency of each sub-experiment corresponding to each transverse demand chain satisfy the corresponding standard execution effect and standard execution efficiency.
10. The unmanned laboratory sample automatic regulation system based on reinforcement learning for implementing the unmanned laboratory sample automatic regulation method based on reinforcement learning in any one of claims 1-9, characterized in that, It includes: a demand chain module, an analysis mapping module, and an execution module; The demand chain module is configured to obtain a directed demand chain corresponding to different experimental demands, and the directed demand chain is constructed by atomic sub-experiments corresponding to each experiment, implementation logical relationships between the sub-experiments, implementation priorities of the sub-experiments of different experiments, a clustering algorithm, and a block chain. The analysis mapping module is configured to preset a process control verification chain, map each sub-experiment in the directed demand chain to the process control verification chain through a smart contract in the directed demand chain, and obtain a pre-control instruction sequence in combination with a configured distributed reinforcement control model. The distributed reinforcement control model is constructed by combining a reinforcement learning algorithm and an experiment standard verification chain according to a per-sub-experiment configuration accuracy rate, a verification passing probability, a corresponding sub-experiment implementation efficiency and an interactive risk transfer matrix; the interactive risk transfer matrix is used for measuring the influence degree of the verification passing probability and the implementation efficiency of the corresponding sub-experiment of the current experiment on the verification passing probability and the implementation efficiency of the next sub-experiment of the same experiment or different experiments; the experiment standard verification chain is constructed by combining a standard operation step of each sub-experiment corresponding to each experiment and a standard evaluation index of each sub-experiment with an evaluation algorithm; and the flow control verification chain is constructed by each experiment corresponding experiment equipment, an operation priority of the corresponding experiment equipment in each experiment and a maximum real-time load coefficient of each experiment equipment. The execution module is configured to execute the pre-control instruction sequence.
Citation Information
Patent Citations
Chemical experiment monitoring optimization method and device based on digital twinning
CN116682500A
Multi-machine multi-task scheduling system for automatic chemistry laboratory
CN119168341A
Laboratory automatic data processing and management method, system, equipment and medium
CN120278484A
Digital twin-based chemical experiment monitoring and optimization method and apparatus
WO2025010771A1