A method and system for generating precise control actions for a dual-system medical laboratory robot inspector

By using a dual-system medical laboratory robot system, multi-view visual fusion and multi-frame temporal aggregation technologies, the challenges of visual perception and long-sequence operations in medical laboratories have been solved, achieving high-precision, standardized, and safe automated operation.

CN122077634APending Publication Date: 2026-05-26CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610315852.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-16
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Medical laboratory automation faces challenges such as visual perception difficulties, accumulation of errors in long-sequence operations, adaptability deficiencies of traditional automation solutions, and limitations of end-to-end learning models, making it difficult to meet the high precision and standardization requirements of unstructured dynamic scenarios.

Method used

A dual-system medical laboratory robot system is adopted, including a top-level LLM planner and a low-level motion generator. Through multi-view visual fusion, multi-frame temporal aggregation and a two-stage semantic alignment retrieval mechanism, atomic instructions that conform to medical standards are generated to achieve precise control.

Benefits of technology

It significantly improves the robot's visual perception capabilities in unstructured scenarios, reduces errors in long-sequence operations, enhances the system's adaptability and generalization ability, and ensures the standardization and safety of medical operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122077634A_ABST
    Figure CN122077634A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for generating precise control actions for a dual-system medical laboratory robotic technician. The method comprises the following steps: a high-order semantic planner performs semantic analysis on multi-frame temporal visual data from a visual observation system and a structured subtask list from a top-level LLM planner using a two-stage cosine similarity semantic alignment retrieval mechanism to generate atomic instructions conforming to medical standards; a low-order action generator receives and processes the atomic instructions to generate a sequence of robotic arm action trajectories; the robotic arm executes the operation according to the generated action trajectory, delivering the sample container to be analyzed into the clinical testing equipment. The system includes an execution system equipped with dual robotic arms, a visual observation system, and a control system. Through the synergistic effect of the two-layer architecture, this invention minimizes the risk of error accumulation in long-sequence operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot operation, specifically a method and system for generating precise control actions for a dual-system medical laboratory robot inspector. Background Technology

[0002] With the continuous advancement of intelligent medical technology, medical laboratory automation has become a core development direction for improving testing throughput, ensuring operator safety, and reducing exposure to biochemical hazards. In a highly automated medical laboratory environment, robotic medical technicians need to complete a series of critical operations such as pipetting, test tube processing, sample transfer, and docking with blood gas analyzers. The accuracy and stability of these operations directly determine the reliability and consistency of laboratory test results, such as blood gas analysis.

[0003] Unlike highly structured industrial production lines, medical laboratory settings are significantly unstructured and dynamic, placing higher demands on the perception, reasoning, and decision-making capabilities of robotic medical technicians. The core challenges currently facing the field of medical laboratory automation are mainly reflected in the following four aspects:

[0004] First, there is the challenge of visual perception of the objects being manipulated. Objects in the laboratory are often transparent or semi-transparent containers (such as test tubes and reagent bottles). Due to refraction, reflection, and occlusion effects, these objects exhibit complex appearance features, leading to the degradation of depth information and boundary cues, making it difficult for robots to accurately identify and locate target objects.

[0005] Second, there is the problem of error accumulation in long-sequence operations. Clinical testing procedures are essentially long-cycle chains of atomic operations, and errors accumulate exponentially with the length of the operation sequence. Even a small deviation in any sub-step can lead to serious consequences such as sample contamination, distorted test results, or even failure of the entire process.

[0006] Third, the adaptability of traditional automation solutions. Existing traditional automation methods rely on hard-coded pipelines, which can maintain a certain degree of stability in fixed scenarios, but lack the ability to adapt flexibly to changes in laboratory layout, changes in the specifications of the objects being operated on, or sudden interference, making it difficult to meet the needs of dynamic scenarios.

[0007] Fourth, the limitations of existing end-to-end learning models. In recent years, end-to-end imitation learning (IL) and vision-language-action (VLA) models have provided new paradigms for robot operation, directly learning the mapping relationship between multimodal observations and actions through expert demonstrations, showing certain potential in general tasks. However, in scenarios with long sequences and high standardization requirements, such as medical testing, these models reveal significant shortcomings: on the one hand, the models are prone to "semantic illusions," making it difficult to reliably parse the complex temporal logic contained in high-level instructions; on the other hand, treating the entire workflow as a single task for learning leads to an over-reliance on finite optimal demonstration paths. When the system enters an out-of-distribution (OOD) state, the lack of an effective error correction mechanism causes continuous error accumulation, ultimately leading to task failure.

[0008] Furthermore, while existing hierarchical learning methods attempt to improve robustness by decoupling semantic planning and low-level control, they are mostly geared towards general object manipulation scenarios and do not fully consider the strict temporal logic and standardization requirements of medical testing procedures, thus failing to meet the safety regulations and operational standards in medical settings. Therefore, developing a medical laboratory robot operating scheme that can balance long-sequence inference with millisecond-level fine control, while also considering adaptability and standardization, has become an urgent technical problem to be solved in this field. Summary of the Invention

[0009] The purpose of this invention is to provide a method for generating precise control actions for a dual-system medical laboratory robot technician, comprising the following steps:

[0010] Step 1) Build a laboratory robot control system, including an execution system equipped with at least one robotic arm and a vision observation system;

[0011] Step 2) Receive global natural language instructions from the user, and then use the top-level LLM planner to decompose the global natural language instructions into a structured list of subtasks;

[0012] Step 3) The higher-order semantic planner performs semantic analysis on multi-frame temporal visual data from the visual observation system and the structured subtask list from the top-level LLM planner through a two-stage cosine similarity semantic alignment retrieval mechanism to generate atomic instructions that conform to medical standards.

[0013] Step 4) The low-order motion generator receives and processes atomic instructions to generate a sequence of robotic arm motion trajectories;

[0014] Step 5) The robotic arm performs operations according to the generated motion trajectory, sending the sample container to be analyzed into the clinical testing equipment so that the clinical testing equipment can complete the subsequent clinical testing and analysis.

[0015] Step 6) Repeat steps 3)-5) until all subtasks in the structured subtask list have been executed, thus realizing the medical testing task corresponding to the global natural language command.

[0016] Furthermore, the dual robotic arms of the execution system are used to perform operations such as grasping, transporting, and aligning the test sample containers with the detection modules of the clinical testing equipment;

[0017] The visual observation system includes a global scene image acquisition camera and at least one local visual detail image acquisition camera; the number of local visual detail image acquisition cameras is the same as the number of robotic arms.

[0018] The global scene image acquisition camera is used to acquire global scene images during the operation of the robotic arm;

[0019] A local visual detail image acquisition camera is mounted on the wrist of the robotic arm to acquire local visual detail images.

[0020] Furthermore, global natural language commands correspond to the detection functions of clinical testing equipment;

[0021] The structured subtask list includes sample grabbing, sample transport, sample alignment, sample retrieval, sample return, and pose reset;

[0022] Sample grasping refers to the process by which a robotic arm grasps a container of samples to be tested from a sample rack.

[0023] Sample delivery refers to the process by which a robotic arm transports sample containers to the area where clinical testing equipment is located.

[0024] Sample alignment refers to aligning the sample container with the detection module of the clinical testing equipment;

[0025] Sample retrieval refers to removing the sample container from the clinical testing equipment after the test is completed;

[0026] Sample return refers to putting the sample container back into the designated location;

[0027] Posture reset refers to returning the robotic arm to its initial posture.

[0028] Furthermore, in step 2), the top-level LLM planner receives global natural language instructions and invokes pre-stored instrument operation standards to generate a structured list of subtasks. ,Right now:

[0029] (1)

[0030] in, This represents the total number of structured subtasks. Indicates the first Natural language descriptions of structured subtasks.

[0031] Furthermore, the high-order semantic planner and the low-order action generator were trained on a historical medical testing dataset.

[0032] The historical medical testing dataset is obtained in the following ways:

[0033] S1) Control the robotic arm to perform operations starting from random initial postures of different users, and mark the transition nodes of the robotic arm's atomic skills in real time during the operation, generate text labels in the time dimension of the robotic arm's operation trajectory, generate long sequence task demonstration data, and short sequence demonstration data corresponding to each sub-task;

[0034] S2) Train the low-order action generator using long sequence task demonstration data and short sequence demonstration data, and then use the trained low-order action generator to perform tasks.

[0035] During the execution of the task, the user can switch atomic instructions through intervention to guide or correct the operation strategy of the robotic arm, collect data in the distributed scenario, and write it into the long sequence task demonstration data and short sequence demonstration data to form a historical medical test dataset.

[0036] S3) Use historical medical test datasets to train the high-order semantic planner and fine-tune the low-order action generator.

[0037] Furthermore, in step 3), the steps for generating atomic instructions that conform to medical standards include:

[0038] Step 3.1) At each time step Obtained through time series observation window Frame-by-frame temporal visual data;

[0039] Among them, the time series observation window As shown below:

[0040] (2)

[0041] in, The time interval between adjacent image frames. Indicates the first Image frames at time points; k = 0, 1, ..., H; This represents the image frame at time t;

[0042] Step 3.2) Employ a pre-trained visual encoder For each image frame Feature extraction is performed to obtain frame-level visual features;

[0043] Step 3.3) Input frame-level visual features into the temporal aggregation Transformer It captures inter-frame dynamic information through temporal aggregation using a Transformer, and then integrates the inter-frame dynamic information using an MLP layer to generate a visual-semantic embedding. ,Right now:

[0044] (3)

[0045] In the formula, I represents the image frame;

[0046] Using a frozen text encoder Natural language description of each atomic skill in the expert skill retrieval database Encode to obtain text embedding vectors ;

[0047] Step 3.4) Computational Vision - Semantic Embedding With text embedding vectors The cosine similarity is used to select the instruction with the highest semantic matching degree as the output instruction, that is:

[0048] (4)

[0049] Step 3.5) Describe the atomic skills in natural language. Mapped to a predefined set of canonical atomic skills M represents the predefined number of canonical atomic skills.

[0050] Step 3.6) Use the same frozen text encoder as in Step 3.3). For each predefined specification atomic skill Encode to obtain text embedding vectors ;

[0051] Step 3.7) Calculate the text embedding vector Embedded vectors with each predefined canonical skill The cosine similarity is used to select the predefined specification skills with the highest matching degree as the final atomic instructions that conform to medical standards. ,Right now:

[0052] (5)

[0053] In the formula, A predefined set of canonical atomic skills.

[0054] Furthermore, in step 4), the step of generating the robotic arm motion trajectory sequence includes:

[0055] Step 4.1) Acquire multi-view visual observation images from the visual observation system. and atomic instructions from the higher-order semantic planner And extract visual features and text features respectively;

[0056] Step 4.2) Utilize the Qwen3VL-2B backbone network Visual and textual features are processed to generate semantic feature sequences. ,Right now:

[0057] (6)

[0058] in, The VLM outputs the sequence length of the tokens. R is the embedding dimension; R is the set of real numbers;

[0059] Obtain the proprioceptive state of the robotic arm Projection transformation is performed through the MLP layer to obtain the proprioceptive embedding. ,Right now:

[0060] (7)

[0061] Step 4.4) Concatenate the proprioceptive embedding with the semantic feature sequence to construct a condition variable. ,Right now:

[0062] (8)

[0063] Step 4.5) Define the conditional action distribution and the tensor of the actual action sequence within the time window The motion generation process is modeled as time-varying with the virtual flow. Evolutionary Probability Flow ,Right now:

[0064] (9)

[0065] in, Represents a real sequence of actions; Standard Gaussian noise;

[0066] Step 4.6) Apply conditional flow with respect to virtual time Taking the derivative, we obtain the closed-form expression for the conditional objective vector field. ,Right now:

[0067] (10)

[0068] Step 4.7) The training objective of the Flow Matching motion head is to minimize the mean square error between the predicted velocity field and the target velocity. Starting from standard Gaussian noise, the learned vector field is... By performing numerical integration, we obtain Solution of time , will solve As a sequence of robotic arm motion trajectories that satisfies the current semantic and physical constraints;

[0069] Among them, the loss function during training As shown below:

[0070] (11)

[0071] in, For cross-attention movement heads, Representing an interval A uniform distribution on the surface; E represents the expectation.

[0072] A dual-system medical laboratory robot system using the aforementioned manipulation action generation method includes an execution system equipped with at least one robotic arm, a vision observation system, and a control system.

[0073] The execution system is used to execute global natural language commands to send the sample container to be analyzed into the clinical testing equipment so that the clinical testing equipment can complete the subsequent clinical testing and analysis.

[0074] The visual observation system is used to acquire multi-view visual observation images;

[0075] The control system includes a top-level LLM planner, a high-order semantic planner, and a low-order action generator.

[0076] The top-level LLM planner decomposes global natural language instructions into a list of structured subtasks;

[0077] The higher-order semantic planner performs semantic analysis on multi-frame temporal visual data from the visual observation system and the structured subtask list from the top-level LLM planner through a two-stage cosine similarity semantic alignment retrieval mechanism to generate atomic instructions that conform to medical standards.

[0078] The low-order motion generator receives and processes atomic instructions to generate a sequence of robotic arm motion trajectories.

[0079] Furthermore, the high-order semantic planner includes a SigLIP2 visual encoder, a temporal Transformer, an MLP layer, a frozen text encoder, a two-order semantic alignment retrieval module, an LLM generation instruction library, and a predefined security skills library.

[0080] Among them, the SigLIP2 visual encoder extracts frame-level features from multi-frame temporal visual data;

[0081] The temporal Transformer captures inter-frame dynamic information based on frame-level features;

[0082] The MLP layer integrates inter-frame dynamic information into visual-semantic embeddings;

[0083] The frozen text encoder converts atomic skill descriptions in the LLM-generated instruction library and canonical operation instructions in the predefined safety skill library into text embedding vectors.

[0084] The two-level semantic alignment retrieval module implements visual embedding and LLM instruction embedding respectively, and then performs semantic alignment between LLM instruction embedding and predefined skill embedding to output the final atomic instruction.

[0085] The low-order action generator includes a multi-view visual encoder, a tokenizer, a pre-trained VLM backbone network, an MLP layer, a flow matching action generation head, and a proprioceptive state processing module.

[0086] The multi-view visual encoder is used to extract visual features from multi-view visual data;

[0087] The tokenizer converts atomic instruction text into text features;

[0088] The pre-trained VLM backbone network processes visual and textual features to generate semantic feature sequences;

[0089] The proprioception state processing module acquires the proprioception state of the robotic arm. ;

[0090] The MLP layer performs a projection transformation on the proprioceptive state to obtain the proprioceptive embedding, and then concatenates the proprioceptive embedding with the semantic feature sequence to construct a conditional variable.

[0091] Guided by conditional variables, the Flow Matching motion generation head converts Gaussian noise into a sequence of robotic arm motion trajectories through an iterative denoising process.

[0092] Furthermore, a two-stage training strategy is adopted when training the low-order action generator;

[0093] In the first training phase, all parameters of the VLM backbone network are frozen, and only randomly initialized cross-attention action heads are trained to align the embedding space of the action heads with the semantic representation of the VLM.

[0094] In the second training phase, the VLM backbone network is unfrozen, and the entire LG-Executor module is fine-tuned end-to-end.

[0095] The technical effects of this invention are undeniable, and its beneficial effects are as follows:

[0096] 1. Effectively solves the visual perception challenges in unstructured scenes.

[0097] This invention significantly enhances a robot's perception capabilities in complex visual scenes through multi-view visual fusion and multi-frame temporal aggregation strategies. Addressing the issues of blurred depth information and degraded boundary cues caused by refraction and reflection in transparent / semi-transparent containers, the multi-view visual input from the top global camera and the wrist local camera captures visual information from different dimensions, enabling comprehensive observation of the manipulated object. Simultaneously, the multi-frame temporal aggregation Transformer employed by HS-Planner integrates dynamic information from consecutive image frames to generate a unified visual-semantic embedding, effectively compensating for the shortcomings of single-frame images in depth perception. This allows the robot to accurately identify and locate the manipulated object, providing a reliable visual foundation for subsequent fine-grained operations.

[0098] 2. Significantly reduces the risk of error accumulation in long sequence operations.

[0099] The layered architecture design of this invention fundamentally solves the problem of error accumulation in long sequence operations. HS-Planner achieves semantic reasoning and canonicality verification through a two-stage cosine similarity semantic alignment retrieval mechanism, effectively avoiding semantic illusions and unsafe operational decisions. Simultaneously, the introduced prediction offset strategy endows the model with forward-looking reasoning capabilities, anticipating skill transition needs and reducing switching deviations between sub-steps. LG-Executor generates continuous, smooth, high-precision motion trajectories based on Flow Matching, meeting millimeter-level (approximately 5mm) operational accuracy requirements and ensuring that the execution error of each sub-step is controlled within acceptable limits. Through the synergistic effect of the two-layer architecture, this invention minimizes the risk of error accumulation in long sequence operations.

[0100] 3. Significantly improves the system's adaptability and generalization ability to complex scenarios.

[0101] This invention significantly improves the system's scene adaptability and instruction generalization ability through retrieval-based semantic alignment and a two-stage data acquisition strategy. The HS-Planner uses retrieval-based semantic matching to replace the autoregressive text generation of traditional end-to-end models, effectively avoiding semantic illusions. The introduction of a predefined specification skill library enables the system to handle dynamic scenarios such as layout adjustments and object specification changes. The DAgger-based enhancement stage in the two-stage data acquisition strategy focuses on collecting data from out-of-distribution (OOD) scenarios, enabling the model to handle various sudden disturbances. Furthermore, for non-standard, open-domain text instructions, the retrieval-based semantic alignment mechanism of this invention exhibits strong generalization ability. Experimental verification shows that when using Gemini 3 Pro to generate 20 semantic variant instructions, the retrieval accuracy of this invention decreases by less than 10% compared to the standard instruction set, while traditional fixed-label classification heads lack this generalization ability and cannot effectively handle semantic drift problems.

[0102] 4. Strictly ensure the standardization and safety of medical procedures.

[0103] This invention ensures the standardization and safety of medical operations through multiple mechanisms, meeting the stringent requirements of medical laboratories. The top-level skill text planning module integrates instrument operation standards into the generation process of atomic skill descriptions, laying the foundation for standardized operations. The second stage of HS-Planner, LLM-predefined skill alignment, ensures that the output atomic instructions conform to medical operation standards and safety criteria through semantic matching with a predefined standardized skill library, eliminating unsafe and unexecutable operational decisions. The motion trajector generated by LG-Executor strictly adheres to geometric constraints and dynamic smoothness requirements, avoiding safety risks such as sample contamination and instrument damage caused by abrupt actions. The synergistic effect of these multiple standardization and safety assurance mechanisms enables this invention to fully adapt to the stringent requirements of medical scenarios, providing reliable safety support for clinical applications.

[0104] 5. The performance of the HS-Planner core architecture has been experimentally verified and is superior to existing solutions.

[0105] The performance of the core architecture of the HS-Planner (high-level semantic inference engine) of this invention has been fully verified through targeted ablation studies, and it is significantly better than existing similar schemes, which fully proves the scientific nature and rationality of the architecture design. The ablation experiments were carried out on the core architecture of the HS-Planner (high-level semantic inference engine), and three model variants were designed for comparison: (1) Classification Head: a single-frame image encoder + Softmax classification head is used, which can only predict fixed atomic skill tags and lacks temporal context awareness and open vocabulary semantic capabilities; (2) Single-Frame Retrieval Variant: based on retrieval matching, but only uses the current frame visual embedding for cosine similarity calculation, omitting the temporal aggregation module; (3) Complete Model (Ours): a complete HS-Planner architecture containing multi-frame temporal aggregation and two layers of cosine matching. The experimental results are as follows. Figure 6 ;

[0106] 1) In terms of retrieval performance: The Top-1 retrieval accuracy of the complete model reached 93.6%, which is 2.55 and 7.58 percentage points higher than the classification head baseline (91.05%) and the single-frame retrieval variant (86.02%), respectively. The model confidence remained at a high level of 95.3%, which verified the synergistic effect of multi-frame temporal aggregation and retrieval semantic alignment.

[0107] 2) Regarding the success rate of subtasks: In subtasks with strong sequence dependencies such as "sample insertion" and "sample retrieval", the complete model has the most significant advantage and can effectively distinguish spatially ambiguous operations such as insertion and retrieval that overlap visually, while single-frame retrieval variants are difficult to accurately distinguish due to the lack of dynamic information of historical frames.

[0108] 3) In terms of instruction generalization ability: When using Gemini 3 Pro to generate 20 non-standard semantic variant instructions (such as changing "Grasp tube" to "Pick up the transparent test tube"), the retrieval accuracy of the complete model only decreased by less than 10% compared to the standard instruction set, while the fixed label classification head cannot handle this kind of semantic drift problem.

[0109] 6. The blood gas analysis physical experimental verification system is highly efficient and reliable.

[0110] This invention conducted a physical experiment on blood gas analysis for core application scenarios, verifying the system's actual operational performance and stability. In the experiment, the dual-system method was continuously executed 50 times across the entire blood gas analysis process (including sample grabbing, transport, alignment, detection, return, and attitude reset). The results showed that the system achieved an average success rate of 92% and an average completion time of 87 ± 0.8 seconds. These results demonstrate that this invention can stably complete high-precision medical testing operations, effectively avoiding process interruptions or operational errors. Compared to traditional manual operations, it significantly shortens the single-sample testing cycle, and compared to existing automated solutions, it improves operational stability, fully meeting the dual requirements of medical laboratories for testing efficiency and result reliability. Attached Figure Description

[0111] Figure 1 RoboMLT Dual System Overall Framework Diagram;

[0112] Figure 2 Architecture diagram of the High-Level Semantic Inference Engine (HS-Planner);

[0113] Figure 3 : Low-level action executor (LG-Executor) architecture diagram;

[0114] Figure 4 Blood gas analysis task operation flowchart;

[0115] Figure 5 Experimental platform hardware layout diagram;

[0116] Figure 6 HS-Planner ablation experiment. Detailed Implementation

[0117] The present invention will be further described below with reference to embodiments, but it should not be construed that the scope of the present invention is limited to the following embodiments. Various substitutions and modifications made based on ordinary technical knowledge and common practices in the art without departing from the above-described technical concept of the present invention should be included within the scope of protection of the present invention.

[0118] Example 1:

[0119] See Figures 1 to 6 A method for generating precise control actions for a dual-system medical laboratory robot inspector includes the following steps:

[0120] Step 1) Build a laboratory robot control system, including an execution system equipped with at least one robotic arm (determined according to actual needs) and a vision observation system;

[0121] Step 2) Receive global natural language instructions from the user, and then use the top-level LLM planner to decompose the global natural language instructions into a structured list of subtasks;

[0122] Step 3) The higher-order semantic planner performs semantic analysis on multi-frame temporal visual data from the visual observation system and the structured subtask list from the top-level LLM planner through a two-stage cosine similarity semantic alignment retrieval mechanism to generate atomic instructions that conform to medical standards.

[0123] Step 4) The low-order motion generator receives and processes atomic instructions to generate a sequence of robotic arm motion trajectories;

[0124] Step 5) The robotic arm performs operations according to the generated motion trajectory, sending the sample container to be analyzed into the clinical testing equipment so that the clinical testing equipment can complete the subsequent clinical testing and analysis.

[0125] Step 6) Repeat steps 3)-5) until all subtasks in the structured subtask list have been executed, thus realizing the medical testing task corresponding to the global natural language command.

[0126] Example 2:

[0127] A method for generating precise control actions of a dual-system medical laboratory robot inspector, with the same technical content as in Embodiment 1, further wherein the dual robotic arms of the execution system are used to perform operations such as grasping, transporting, and aligning the test sample container with the test module of the clinical testing equipment.

[0128] The visual observation system includes a global scene image acquisition camera and at least one local visual detail image acquisition camera; the number of local visual detail image acquisition cameras is the same as the number of robotic arms.

[0129] The global scene image acquisition camera is used to acquire global scene images during the operation of the robotic arm;

[0130] A local visual detail image acquisition camera is mounted on the wrist of the robotic arm to acquire images of local visual details. One local visual detail image acquisition camera corresponds to one robotic arm.

[0131] Example 3:

[0132] A method for generating precise control actions of a dual-system medical laboratory robot inspector, with the same technical content as any one of embodiments 1-2, further wherein global natural language commands correspond to the detection functions of clinical testing equipment;

[0133] The structured subtask list includes sample grabbing, sample transport, sample alignment, sample retrieval, sample return, and pose reset;

[0134] Sample grasping refers to the process by which a robotic arm grasps a container of samples to be tested from a sample rack.

[0135] Sample delivery refers to the process by which a robotic arm transports sample containers to the area where clinical testing equipment is located.

[0136] Sample alignment refers to aligning the sample container with the detection module of the clinical testing equipment;

[0137] Sample retrieval refers to removing the sample container from the clinical testing equipment after the test is completed;

[0138] Sample return refers to putting the sample container back into the designated location;

[0139] Posture reset refers to returning the robotic arm to its initial posture.

[0140] Example 4:

[0141] A method for generating precise control actions for a dual-system medical laboratory robot inspector, with technical content identical to any one of embodiments 1-3, further comprising the following step 2): the top-level LLM planner receives global natural language instructions and invokes pre-stored instrument operation standards to generate a structured list of subtasks. ,Right now:

[0142] (1)

[0143] in, This represents the total number of structured subtasks. Indicates the first Natural language descriptions of structured subtasks.

[0144] Example 5:

[0145] A method for generating precise control actions for a dual-system medical laboratory robot inspector, with the same technical content as any one of embodiments 1-4, further wherein the high-order semantic planner and the low-order action generator are trained on a historical medical test dataset.

[0146] The historical medical testing dataset is obtained in the following ways:

[0147] S1) Control the robotic arm to perform operations starting from random initial postures of different users, and mark the transition nodes of the robotic arm's atomic skills in real time during the operation, generate text labels in the time dimension of the robotic arm's operation trajectory, generate long sequence task demonstration data, and short sequence demonstration data corresponding to each sub-task;

[0148] S2) Train the low-order action generator using long sequence task demonstration data and short sequence demonstration data, and then use the trained low-order action generator to perform tasks.

[0149] During the execution of the task, the user can switch atomic instructions through intervention to guide or correct the operation strategy of the robotic arm, collect data in the distributed scenario, and write it into the long sequence task demonstration data and short sequence demonstration data to form a historical medical test dataset.

[0150] S3) Use historical medical test datasets to train the high-order semantic planner and fine-tune the low-order action generator.

[0151] Example 6:

[0152] A method for generating precise control actions of a dual-system medical laboratory robot inspector, with technical content identical to any one of embodiments 1-5, further comprising, in step 3), generating atomic instructions conforming to medical standards, including:

[0153] Step 3.1) At each time step Obtained through time series observation window Frame-by-frame temporal visual data;

[0154] Among them, the time series observation window As shown below:

[0155] (2)

[0156] in, The time interval between adjacent image frames. Indicates the first Image frames at time points; k = 0, 1, ..., H; This represents the image frame at time t;

[0157] Step 3.2) Employ a pre-trained visual encoder For each image frame Feature extraction is performed to obtain frame-level visual features;

[0158] Step 3.3) Input frame-level visual features into the temporal aggregation Transformer It captures inter-frame dynamic information through temporal aggregation using a Transformer, and then integrates the inter-frame dynamic information using an MLP layer to generate a visual-semantic embedding. ,Right now:

[0159] (3)

[0160] In the formula, I represents the image frame;

[0161] Using a frozen text encoder Natural language description of each atomic skill in the expert skill retrieval database Encode to obtain text embedding vectors ;

[0162] Step 3.4) Computational Vision - Semantic Embedding With text embedding vectors The cosine similarity is used to select the instruction with the highest semantic matching degree as the output instruction, that is:

[0163] (4)

[0164] Step 3.5) Describe the atomic skills in natural language. Mapped to a predefined set of canonical atomic skills M represents the predefined number of canonical atomic skills.

[0165] Step 3.6) Use the same frozen text encoder as in Step 3.3). For each predefined specification atomic skill Encode to obtain text embedding vectors ;

[0166] Step 3.7) Calculate the text embedding vector Embedded vectors with each predefined canonical skill The cosine similarity is used to select the predefined specification skills with the highest matching degree as the final atomic instructions that conform to medical standards. ,Right now:

[0167] (5)

[0168] In the formula, A predefined set of canonical atomic skills.

[0169] Example 7:

[0170] A method for generating precise control movements of a dual-system medical laboratory robot inspector, with technical content identical to any one of embodiments 1-6, further comprising, in step 4), generating the robotic arm motion trajectory sequence, including:

[0171] Step 4.1) Acquire multi-view visual observation images from the visual observation system. and atomic instructions from the higher-order semantic planner And extract visual features and text features respectively;

[0172] Step 4.2) Utilize the Qwen3VL-2B backbone network Visual and textual features are processed to generate semantic feature sequences. ,Right now:

[0173] (6)

[0174] in, The VLM outputs the sequence length of the tokens. R is the embedding dimension; R is the set of real numbers;

[0175] Obtain the proprioceptive state of the robotic arm Projection transformation is performed through the MLP layer to obtain the proprioceptive embedding. ,Right now:

[0176] (7)

[0177] Step 4.4) Concatenate the proprioceptive embedding with the semantic feature sequence to construct a condition variable. ,Right now:

[0178] (8)

[0179] Step 4.5) Define the conditional action distribution and the tensor of the actual action sequence within the time window The motion generation process is modeled as time-varying with the virtual flow. Evolutionary Probability Flow ,Right now:

[0180] (9)

[0181] in, Represents a real sequence of actions; Standard Gaussian noise;

[0182] Step 4.6) Apply conditional flow with respect to virtual time Taking the derivative, we obtain the closed-form expression for the conditional objective vector field. ,Right now:

[0183] (10)

[0184] Step 4.7) The training objective of the Flow Matching motion head is to minimize the mean square error between the predicted velocity field and the target velocity. Starting from standard Gaussian noise, the learned vector field is... By performing numerical integration, we obtain Solution of time , will solve As a sequence of robotic arm motion trajectories that satisfies the current semantic and physical constraints;

[0185] Among them, the loss function during training As shown below:

[0186] (11)

[0187] in, For cross-attention movement heads, Representing an interval A uniform distribution on the surface; E represents the expectation.

[0188] Example 8:

[0189] A dual-system medical laboratory robot system using the manipulation action generation method described in any one of Examples 1-7 includes an execution system equipped with at least one robotic arm, a vision observation system, and a control system.

[0190] The execution system is used to execute global natural language commands to send the sample container to be analyzed into the clinical testing equipment so that the clinical testing equipment can complete the subsequent clinical testing and analysis.

[0191] The visual observation system is used to acquire multi-view visual observation images;

[0192] The control system includes a top-level LLM planner, a high-order semantic planner, and a low-order action generator.

[0193] The top-level LLM planner decomposes global natural language instructions into a list of structured subtasks;

[0194] The higher-order semantic planner performs semantic analysis on multi-frame temporal visual data from the visual observation system and the structured subtask list from the top-level LLM planner through a two-stage cosine similarity semantic alignment retrieval mechanism to generate atomic instructions that conform to medical standards.

[0195] The low-order motion generator receives and processes atomic instructions to generate a sequence of robotic arm motion trajectories.

[0196] Example 9:

[0197] A dual-system medical laboratory robot system using the manipulation action generation method described in any one of Embodiments 1-7, with the same technical content as Embodiment 8, further comprising the higher-order semantic planner including a SigLIP2 visual encoder, a temporal Transformer, an MLP layer, a frozen text encoder, a dual-order semantic alignment retrieval module, an LLM generation instruction library, and a predefined safety skill library.

[0198] Among them, the SigLIP2 visual encoder extracts frame-level features from multi-frame temporal visual data;

[0199] The temporal Transformer captures inter-frame dynamic information based on frame-level features;

[0200] The MLP layer integrates inter-frame dynamic information into visual-semantic embeddings;

[0201] The frozen text encoder converts atomic skill descriptions in the LLM-generated instruction library and canonical operation instructions in the predefined safety skill library into text embedding vectors.

[0202] The two-level semantic alignment retrieval module implements visual embedding and LLM instruction embedding respectively, and then performs semantic alignment between LLM instruction embedding and predefined skill embedding to output the final atomic instruction.

[0203] The low-order action generator includes a multi-view visual encoder, a tokenizer, a pre-trained VLM backbone network, an MLP layer, a flow matching action generation head, and a proprioceptive state processing module.

[0204] The multi-view visual encoder is used to extract visual features from multi-view visual data;

[0205] The tokenizer converts atomic instruction text into text features;

[0206] The pre-trained VLM backbone network processes visual and textual features to generate semantic feature sequences;

[0207] The proprioception state processing module acquires the proprioception state of the robotic arm. ;

[0208] The MLP layer performs a projection transformation on the proprioceptive state to obtain the proprioceptive embedding, and then concatenates the proprioceptive embedding with the semantic feature sequence to construct a conditional variable.

[0209] Guided by conditional variables, the Flow Matching motion generation head converts Gaussian noise into a sequence of robotic arm motion trajectories through an iterative denoising process.

[0210] Example 10:

[0211] A dual-system medical laboratory robot system using the manipulation action generation method described in any one of Embodiments 1-7, with the same technical content as any one of Embodiments 8-9, further wherein a two-stage training strategy is adopted during the training of the low-order action generator.

[0212] In the first training phase, all parameters of the VLM backbone network are frozen, and only randomly initialized cross-attention action heads are trained to align the embedding space of the action heads with the semantic representation of the VLM.

[0213] In the second training phase, the VLM backbone network is unfrozen, and the entire LG-Executor module is fine-tuned end-to-end.

[0214] Example 11:

[0215] This invention proposes a dual-system medical laboratory robot for fine operation (RoboMLT). Based on the cognitive dual-system theory (the two systems consist of System 2 and System 1, where System 2 is a low-frequency task planner and System 1 is a high-frequency action executor), a hierarchical architecture is designed. The core objective is to balance the dual requirements of long sequence reasoning and millisecond-level fine control, while ensuring the standardization and normalization of medical operation procedures.

[0216] This system achieves end-to-end automation from natural language commands to precise operations by the robotic medical technician through the collaborative work of a top-level skill text planning module, a high-order semantic planner (HS-Planner, System 2 in the dual system), and a low-order action generator (LG-Executor, System 1 in the dual system). The specific technical solution is as follows:

[0217] (a) Hardware Dependence and Data Acquisition

[0218] 1. Hardware platform configuration

[0219] The technical implementation of this invention depends on a specific hardware platform, the layout and composition of which are as follows: Figure 5 As shown: The core hardware includes the Agilex dual-arm Static-Aloha platform, two Piper robotic arms (left and right), three Intel RealSense D435i RGB cameras, and experimental objects such as a blood gas analyzer and test tube racks. The Agilex dual-arm Static-Aloha platform serves as the core execution vehicle, with the two Piper robotic arms forming a 14-dimensional motion space (defined by the target joint positions of the two robotic arms), responsible for performing key operations such as grasping, transporting, and alignment. The visual observation system consists of three cameras: one mounted on top to acquire global scene images, and the other two mounted on the wrists of the left and right robotic arms respectively, to capture local visual details during the operation, providing high-precision visual support for fine manipulation. The transparent test tubes in the experimental objects simulate sample containers in clinical testing, the professional blood gas analyzer simulates clinical testing equipment, and its narrow entrance requires millimeter-level (approximately 5mm gap) precision. The test tube racks are used for sample storage and return.

[0220] 2. Data Acquisition Methods

[0221] To enable the training and fine-tuning of the two core modules (LG-Executor and HS-Planner), this invention constructs a unified robot dataset. The data collection process is divided into two stages, supplemented by a prediction offset strategy to improve data effectiveness.

[0222] The first phase involves interactive annotation and initial data acquisition. A real-time interactive annotation method is employed. During expert demonstrations of long-sequence medical testing tasks (such as blood gas analysis), operators use keyboard commands to mark the transition nodes of atomic skills in real time, generating text labels that are precisely aligned with the robot's operational trajectory in the time dimension. To improve the LG-Executor's command-following generalization ability, an additional set of command-trajectory data based on random initial postures is recorded. This involves controlling the robot to execute operations from different random initial postures and collecting the corresponding command and trajectory information. This phase's dataset contains 100 sets of complete long-sequence task demonstration data, and 50 sets of short-sequence demonstration data for each sub-task (such as grasping, transporting, and aligning) under random initial postures. This dataset serves as the foundation for the initial fine-tuning of the LG-Executor.

[0223] The second phase involves DAgger-based augmented data acquisition. The LG-Executor module, fine-tuned in the first phase, executes medical testing tasks. During real-time inference, a human operator intervenes via keyboard to switch atomic commands, guiding or correcting the robot's operational strategy. The focus is on acquiring data from out-of-distribution (OOD) scenarios to compensate for insufficient coverage in the initial dataset. The data acquired in both phases are integrated to form the final training dataset, used for training the HS-Planner and further fine-tuning the LG-Executor.

[0224] 3. Predictive Migration Strategy

[0225] To endow HS-Planner with forward-looking reasoning capabilities based on historical observation trends and improve the smoothness of transitions between atomic skills, this invention introduces a prediction offset strategy. Specifically, during data processing, the HS-Planner's prediction target is shifted forward by 10 consecutive observation frames in the time dimension. This constraint forces the model to predict future states based on historical observation data, thereby anticipating the need for skill transitions and significantly reducing action stuttering or deviations during skill switching, thus improving the smoothness and stability of the entire operation process.

[0226] (II) Core Module Design and Workflow

[0227] The overall architecture of this invention adopts a layered design, with the core comprising a global instruction and context input module, a global LLM planner, an operation manual and prompt module, a high-order semantic planner (HS-Planner), a low-order action generator (LG-Executor), and a historical observation data interface. Specifically, the global LLM planner is responsible for decomposing user natural language instructions (such as "complete blood gas analysis of all samples") into a structured list of subtasks; the operation manual and prompt module provides standardized operational guidelines for the LLM planner; the historical observation data interface is used to input multi-frame temporal visual data; the HS-Planner performs slow, cautious semantic reasoning through a two-stage cosine similarity semantic alignment retrieval mechanism; and the LG-Executor, based on a VLM-constrained autoregressive flow matching controller, achieves fine-grained execution of atomic instructions.

[0228] Top-level skill text planning module

[0229] The core function of this module is to convert natural language instructions into a structured set of atomic skill descriptions, providing a foundation for subsequent semantic reasoning. The specific implementation process is as follows: First, two types of input information are integrated: one is the global natural language instruction input by laboratory operators (e.g., "Complete blood gas analysis of all samples"), and the other is the corresponding instrument operation standards (e.g., blood gas analyzer usage specifications, sample processing procedures, etc.). These two types of information are input as prompts into the Qwen3 Large Language Model (LLM). Based on these prompts, the Qwen3 LLM generates a structured set of atomic skill descriptions, mathematically expressed as:

[0230]

[0231] in, The total number of atomic skill descriptions. Indicates the first The natural language descriptions of each atomic skill are generated, and the resulting set of atomic skill descriptions will be stored in the expert skill retrieval library as a candidate instruction set for semantic matching by HS-Planner.

[0232] High-order semantic planner (HS-Planner): Semantic reasoning based on two-order semantic alignment retrieval

[0233] The core function of HS-Planner is to realize semantic reasoning from visual observation to standardized executable instructions. Internally, it includes a SigLIP2 visual encoder, a temporal Transformer, an MLP layer, a frozen text encoder, a two-level semantic alignment retrieval module, an LLM-generated instruction library, and a predefined safety skill library. Multi-frame historical observation data is processed by the SigLIP2 visual encoder to extract frame-level features. After the temporal Transformer captures inter-frame dynamic information, it is integrated into a unified visual-semantic embedding through the MLP layer. Atomic skill descriptions in the LLM-generated instruction library and standardized operation instructions in the predefined safety skill library are converted into text embedding vectors by the frozen text encoder. The two-level semantic alignment retrieval module performs semantic alignment between visual embedding and LLM instruction embedding, and between LLM instruction embedding and predefined skill embedding, ultimately outputting standardized executable atomic instructions.

[0234] Through a two-stage cosine similarity semantic alignment retrieval mechanism, HS-Planner maps multi-frame temporal visual observations into atomic instructions that conform to medical standards, while avoiding semantic illusions and unsafe operational decisions. The specific workflow is divided into two stages:

[0235] Phase 1: Visual-LM semantic alignment. At each time step... Captured through time series observation window For continuous image data frames, the mathematical expression for the time-series observation window is:

[0236]

[0237] in, The time interval between adjacent image frames. Indicates the first Image frames at any given moment.

[0238] To obtain the overall visual-semantic representation of this temporal window, a pre-trained SigLIP2-300M visual encoder was used. For each image frame Feature extraction is performed to obtain frame-level visual features; these frame-level features are then input into a temporal aggregation Transformer. Together with the MLP layer, the Transformer captures inter-frame dynamic information through temporal aggregation, and then integrates it through the MLP layer to generate a unified visual-semantic embedding. The mathematical expression is:

[0239]

[0240] At the same time, using a frozen text encoder Generate instructions for each LLM in the expert skills retrieval database. Encode the text to obtain the corresponding text embedding vector. .

[0241] Computer Vision - Semantic Embedding With each LLM instruction embedding vector The cosine similarity is used to select the instruction with the highest semantic matching degree as the output of this stage, which can be expressed mathematically as:

[0242]

[0243] This stage replaces autoregressive text generation with retrieval-based matching, effectively avoiding the semantic illusion problem that language models may produce in complex medical scenarios.

[0244] Phase Two: LLM - Predefined Skill Alignment. To ensure the executability, safety, and standardization of output instructions with medical processes, the LLM instructions retrieved in Phase One are further mapped to a predefined set of standardized atomic skills. The mathematical expression of the predefined set of standardized atomic skills is as follows:

[0245]

[0246] Uses the same frozen text encoder as the first stage. For each predefined specification atomic skill Encode the text to obtain the corresponding text embedding vector. .

[0247] Calculate the LLM instruction embedding vector Embedded vectors with each predefined canonical skill The cosine similarity is used to select the predefined specification skill with the highest matching degree as the final execution instruction. The mathematical expression is:

[0248]

[0249] This phase ensures that downstream lower-level control systems receive only validated atomic operation instructions that comply with safety standards and medical guidelines by aligning with a predefined skill set.

[0250] Low-order action generator (LG-Executor): Temporal semantic guidance for continuous action generation based on Qwen3VL and Flow Matching

[0251] LG-Executor is a high-frequency vision-language-action (VLA) strategy that employs a temporal semantic guided continuous action flow (TSG-CCAF) architecture. Its core function is to convert atomic instructions output by the HS-Planner into continuous action trajectories that satisfy geometric constraints and dynamic smoothness. Its modules include a multi-view visual encoder, a tokenizer, a pre-trained VLM backbone network (Qwen3VL-2B), a key-value cache, a DeepStack encoder, a flow-matching action generation head, an ActionDecoder, and a proprioceptive state processing module. Multi-view visual observations (primary view, left wrist view, right wrist view, etc.) extract visual features through the multi-view visual encoder. Atomic instruction text is converted into text features by the tokenizer. Both are input into the Qwen3VL-2B backbone network to generate conditional latent semantic prefixes. The proprioceptive state is projected through an MLP and fused with the semantic prefixes to form conditional variables for action generation. Guided by these conditional variables, the flow-matching action generation head converts Gaussian noise into precise, continuous action segments through an iterative denoising process. The specific implementation process is as follows:

[0252] 3.1 Multimodal Feature Extraction and Condition Variable Construction

[0253] Qwen3VL-2B was selected as the vision-language backbone network, possessing powerful open vocabulary perception and spatial localization capabilities, enabling precise alignment of visual features of medical instruments with text descriptions. Within each control cycle, multi-view visual observations were input. (Images from the main view, left wrist view, and right wrist view, respectively) and atomic command text output by HS-Planner Visual and textual features are extracted using a SigLIP2-Large visual encoder and a tokenizer, respectively. Conditional latent semantic prefixes are then generated via a Qwen3VL-2B backbone network, mathematically expressed as:

[0254]

[0255] in, The VLM outputs the sequence length of the tokens. As an embedding dimension, this semantic feature sequence encapsulates rich visual-linguistic contextual information.

[0256] Obtain the robot's proprioceptive state (Including physical state information such as joint angles and angular velocities), the projection transformation is performed through an MLP layer to match the dimensions with the semantic embedding:

[0257]

[0258] By concatenating proprioceptive embeddings with semantic feature sequences, the final conditional variables are constructed, providing semantic and state-aware guidance for action generation.

[0259]

[0260] 3.2 Continuous Action Generation Based on Flow Matching

[0261] Introducing Flow Matching technology to clarify the modeling conditions and action distribution ,definition The tensor formed by the actual action sequence within the time window ( The length of the action sequence. (The motion dimension is defined by the target joint positions of the dual robotic arms). The noise is standard Gaussian noise. Flow Matching models the motion generation process as time-varying noise with the virtual flow. The probability flow of evolution:

[0262]

[0263] in, It represents the real action sequence and achieves interpolation between the noise distribution and the target action distribution by constructing the optimal transmission (OT) probability path.

[0264] Conditional flow with respect to virtual time Taking the derivative, we obtain the closed-form expression for the conditional objective vector field:

[0265]

[0266] The training objective of the Flow Matching motion head is to minimize the mean square error between the predicted velocity field and the target velocity. The loss function is defined as:

[0267]

[0268] in, For cross-attention movement heads, Representing an interval A uniform distribution on the surface.

[0269] During the inference phase, the model starts with standard Gaussian noise and learns the vector field. By performing numerical integration, we obtain The solution at that time is the predicted action sequence that satisfies the current semantic and physical constraints:

[0270]

[0271] 3.3 Two-stage training strategy

[0272] To ensure stable convergence of the model and preserve the semantic alignment capability of the pre-trained VLM, LG-Executor employs a two-stage training strategy:

[0273] Phase 1: Action Head Alignment. Freeze all parameters of the VLM backbone network and train only randomly initialized cross-attention action heads to align the embedding space of the action heads with the semantic representation of the VLM, avoiding the destruction of pre-trained knowledge. This phase requires fewer training iterations and can effectively prevent catastrophic forgetting.

[0274] Phase Two: End-to-End Fine-Tuning. The VLM backbone network is unfrozen, and the entire LG-Executor module is fine-tuned end-to-end. To preserve pre-trained knowledge, the VLM backbone network uses a lower learning rate than the action head, achieving task-specific adaptation through joint optimization while maintaining semantic coherence to ensure the accuracy required for medical procedures.

[0275] 3.4 Asynchronous Inference Framework

[0276] To bridge the latency gap between high-order semantic reasoning and high-frequency control, a decoupled multi-process asynchronous framework is adopted. This framework coordinates three parallel processes (semantic planning process, inference client process, and future-state-aware execution process) through explicit locks on shared state. The specific mechanism is as follows:

[0277] Shared state definition: Action queue (block size is) ), historical multi-perspective buffer (capacity is) Current task instructions ;

[0278] Asynchronous planning and action flow: Semantic planning process saturates in the history buffer ( When the inference client process receives the inference action blocks from the VLA server in real time, it performs exponential weighted aggregation with the blocks to be executed in the current queue to ensure a smooth transition of continuous trajectories.

[0279] Future state-aware execution: The main control loop interacts with the hardware at a frequency of 30Hz, when the action queue length is below a threshold ( When a new inference request is initiated, this is done to mitigate inference latency. The resulting misalignment is addressed using a state deduction strategy, which involves "pre-reading" the next action in the action queue. Given a set of actions to be executed, estimate the robot's state when inference is complete:

[0280]

[0281] in, This is the current proprioceptive state. For the first One pending action. The request tuple... Sending the data to the inference client process enables the VLM model to generate actions based on future states, ensuring accurate tracking of dynamic tasks.

[0282] Overall System Workflow

[0283] The overall workflow of this invention follows a logical chain of "instruction input - planning decomposition - semantic reasoning - action generation - operation execution", and the specific steps are as follows:

[0284] The first step is to input instructions: receive global natural language instructions (such as "complete blood gas analysis of all samples") from laboratory operators.

[0285] The second step is planning decomposition: The top-level LLM planner (Qwen3) combines the instrument operation standards to decompose the global natural language instructions into a structured list of subtasks;

[0286] The third step is semantic reasoning: HS-Planner uses a two-stage cosine similarity semantic alignment retrieval mechanism to perform semantic analysis on multi-frame temporal visual observations in the historical multi-view buffer and map them into atomic instructions that conform to medical standards.

[0287] The fourth step is action generation: LG-Executor receives atomic instructions, extracts multimodal features, constructs condition variables, generates actions through Flow Matching, and outputs a continuous, smooth, and high-precision action trajectory. The action flow is generated and aggregated in real time through an asynchronous inference framework.

[0288] Step 5, Operation Execution: The Piper robotic arm performs operations according to the generated motion trajectory to complete the entire process of the blood gas analysis task, which includes 6 key operation stages: (1) Sample Grabbing: Grab the sample tube to be tested from the test tube rack; (2) Sample Transport: Transport the test tube to the area where the blood gas analyzer is located; (3) Fine Alignment: Align the test tube with the analyzer probe with millimeter-level precision; (4) Sample Removal: Remove the test tube from the analyzer after the test is completed; (5) Sample Return: Put the test tube back into the designated slot; (6) Posture Reset: The robotic arm returns to the initial posture to prepare for the next round of operation.

[0289] Step 6, Execute in a loop: Repeat steps 3 to 5 until all subtasks in the structured subtask list have been completed, thus realizing the medical testing task corresponding to the global natural language command.

[0290] (III) Model Training and Optimization

[0291] HS-Planner training: With retrieval accuracy as the core indicator, a two-order semantic alignment loss function is adopted. The parameters of modules such as visual encoder, temporal Transformer, and text encoder are optimized by gradient descent. The training data is visual-instruction aligned data collected in two stages. A prediction offset strategy is introduced during the training process to improve the model's look-ahead inference ability.

[0292] LG-Executor training: It is divided into two stages. The first stage (action head alignment) uses the dataset collected in the first stage as the training data and the mean square error of the action trajectory as the loss function to optimize the action head parameters. The second stage (end-to-end fine-tuning) uses the complete dataset after the two stages are integrated. Combining the geometric constraints and dynamic smoothness requirements of medical operations, the VLM backbone network and action head parameters are iteratively optimized to ensure the accuracy of action generation.

Claims

1. A method for generating precise control actions for a dual-system medical laboratory robot inspector, characterized in that, Includes the following steps: Step 1) Build a laboratory robot control system, including an execution system equipped with at least one robotic arm and a vision observation system; Step 2) Receive global natural language instructions from the user, and then use the top-level LLM planner to decompose the global natural language instructions into a structured list of subtasks; Step 3) The higher-order semantic planner performs semantic analysis on multi-frame temporal visual data from the visual observation system and the structured subtask list from the top-level LLM planner through a two-stage cosine similarity semantic alignment retrieval mechanism to generate atomic instructions that conform to medical standards. Step 4) The low-order motion generator receives and processes atomic instructions to generate a sequence of robotic arm motion trajectories; Step 5) The robotic arm performs the operation according to the generated motion trajectory, and sends the sample container to be analyzed into the clinical testing equipment; Step 6) Repeat steps 3)-5) until all subtasks in the structured subtask list have been executed, thus realizing the medical testing task corresponding to the global natural language command.

2. The method for generating precise control actions for a dual-system medical laboratory robot technician according to claim 1, characterized in that, The robotic arm of the execution system is used to perform operations such as grasping, transporting, and aligning the test sample container with the test module of the clinical testing equipment; The visual observation system includes a global scene image acquisition camera and at least one local visual detail image acquisition camera; the number of local visual detail image acquisition cameras is the same as the number of robotic arms. The global scene image acquisition camera is used to acquire global scene images during the operation of the robotic arm; A local visual detail image acquisition camera is mounted on the wrist of the robotic arm to acquire local visual detail images.

3. The method for generating precise control actions for a dual-system medical laboratory robot technician according to claim 1, characterized in that, Global natural language commands correspond to the detection functions of clinical testing equipment; The structured subtask list includes sample grabbing, sample transport, sample alignment, sample retrieval, sample return, and pose reset; Sample grasping refers to the process by which a robotic arm grasps a container of samples to be tested from a sample rack. Sample delivery refers to the process by which a robotic arm transports sample containers to the area where clinical testing equipment is located. Sample alignment refers to aligning the sample container with the detection module of the clinical testing equipment; Sample retrieval refers to removing the sample container from the clinical testing equipment after the test is completed; Sample return refers to putting the sample container back into the designated location; Posture reset refers to returning the robotic arm to its initial posture.

4. The method for generating precise control actions for a dual-system medical laboratory robot technician according to claim 1, characterized in that, In step 2), the top-level LLM planner receives global natural language instructions and invokes pre-stored instrument operation standards to generate a structured list of subtasks. ,Right now: (1) in, This represents the total number of structured subtasks. Indicates the first Natural language descriptions of structured subtasks.

5. The method for generating precise control actions for a dual-system medical laboratory robot technician according to claim 1, characterized in that, The high-order semantic planner and the low-order action generator were trained on a historical medical testing dataset. The historical medical testing dataset is obtained in the following ways: S1) Control the robotic arm to perform operations starting from random initial postures of different users, and mark the transition nodes of the robotic arm's atomic skills in real time during the operation, generate text labels in the time dimension of the robotic arm's operation trajectory, generate long sequence task demonstration data, and short sequence demonstration data corresponding to each sub-task; S2) Train the low-order action generator using long sequence task demonstration data and short sequence demonstration data, and then use the trained low-order action generator to perform tasks. During the execution of the task, the user can switch atomic instructions through intervention to guide or correct the operation strategy of the robotic arm, collect data in the distributed scenario, and write it into the long sequence task demonstration data and short sequence demonstration data to form a historical medical test dataset. S3) Use historical medical test datasets to train the high-order semantic planner and fine-tune the low-order action generator.

6. The method for generating precise control actions for a dual-system medical laboratory robot technician according to claim 1, characterized in that, Step 3), the steps for generating atomic instructions that conform to medical standards, include: Step 3.1) At each time step Obtained through time series observation window Frame-by-frame temporal visual data; Among them, the time series observation window As shown below: (2) in, The time interval between adjacent image frames. Indicates the first Image frames at time points; k = 0, 1, ..., H; This represents the image frame at time t; Step 3.2) Employ a pre-trained visual encoder For each image frame Feature extraction is performed to obtain frame-level visual features; Step 3.3) Input frame-level visual features into the temporal aggregation Transformer It captures inter-frame dynamic information through temporal aggregation using a Transformer, and then integrates the inter-frame dynamic information using an MLP layer to generate a visual-semantic embedding. ,Right now: (3) In the formula, I represents the image frame; Using a frozen text encoder Natural language description of each atomic skill in the expert skill retrieval database Encode to obtain text embedding vectors ; Step 3.4) Computational Vision-Semantic Embedding With text embedding vectors The cosine similarity is used to select the instruction with the highest semantic matching degree as the output instruction, that is: (4) Step 3.5) Describe the atomic skills in natural language. Mapped to a predefined set of canonical atomic skills M represents the predefined number of canonical atomic skills. Step 3.6) Use the same frozen text encoder as in Step 3.3). For each predefined specification atomic skill Encode to obtain text embedding vectors ; Step 3.7) Calculate the text embedding vector Embedded vectors with each predefined canonical skill The cosine similarity is used to select the predefined specification skills with the highest matching degree as the final atomic instructions that conform to medical standards. ,Right now: (5) In the formula, A predefined set of canonical atomic skills.

7. The method for generating precise control actions for a dual-system medical laboratory robot technician according to claim 1, characterized in that, Step 4), the steps for generating the robotic arm motion trajectory sequence include: Step 4.1) Acquire multi-view visual observation images from the visual observation system. and atomic instructions from the higher-order semantic planner And extract visual features and text features respectively; Step 4.2) Utilize the Qwen3VL-2B backbone network Visual and textual features are processed to generate semantic feature sequences. ,Right now: (6) in, The VLM outputs the sequence length of the tokens. R is the embedding dimension; R is the set of real numbers; Obtain the proprioceptive state of the robotic arm Projection transformation is performed through the MLP layer to obtain the proprioceptive embedding. ,Right now: (7) Step 4.4) Concatenate the proprioceptive embedding with the semantic feature sequence to construct a condition variable. ,Right now: (8) Step 4.5) Define the conditional action distribution and the tensor of the actual action sequence within the time window The motion generation process is modeled as time-varying with the virtual flow. Evolutionary Probability Flow ,Right now: (9) in, Represents a real sequence of actions; Standard Gaussian noise; Step 4.6) Apply conditional flow with respect to virtual time Taking the derivative, we obtain the closed-form expression for the conditional objective vector field. ,Right now: (10) Step 4.7) The training objective of the Flow Matching motion head is to minimize the mean square error between the predicted velocity field and the target velocity. Starting from standard Gaussian noise, the learned vector field is... By performing numerical integration, we obtain Solution of time , will solve As a sequence of robotic arm motion trajectories that satisfies the current semantic and physical constraints; Among them, the loss function during training As shown below: (11) in, For cross-attention movement heads, Representing an interval A uniform distribution on the surface; E represents the expectation.

8. A dual-system medical laboratory robot system using the manipulation action generation method according to any one of claims 1-7, characterized in that, It includes an execution system equipped with at least one robotic arm, a vision observation system, and a control system; The execution system is used to execute global natural language commands and send the sample container to be analyzed into the clinical testing equipment; The visual observation system is used to acquire multi-view visual observation images; The control system includes a top-level LLM planner, a high-order semantic planner, and a low-order action generator. The top-level LLM planner decomposes global natural language instructions into a list of structured subtasks; The higher-order semantic planner performs semantic analysis on multi-frame temporal visual data from the visual observation system and the structured subtask list from the top-level LLM planner through a two-stage cosine similarity semantic alignment retrieval mechanism to generate atomic instructions that conform to medical standards. The low-order motion generator receives and processes atomic instructions to generate a sequence of robotic arm motion trajectories.

9. The dual-system medical laboratory robot system according to claim 8, characterized in that, The high-order semantic planner includes a SigLIP2 visual encoder, a temporal Transformer, an MLP layer, a frozen text encoder, a two-order semantic alignment retrieval module, an LLM generation instruction library, and a predefined security skills library. Among them, the SigLIP2 visual encoder extracts frame-level features from multi-frame temporal visual data; The temporal Transformer captures inter-frame dynamic information based on frame-level features; The MLP layer integrates inter-frame dynamic information into visual-semantic embeddings; The frozen text encoder converts atomic skill descriptions in the LLM-generated instruction library and canonical operation instructions in the predefined safety skill library into text embedding vectors. The two-level semantic alignment retrieval module implements visual embedding and LLM instruction embedding respectively, and then performs semantic alignment between LLM instruction embedding and predefined skill embedding to output the final atomic instruction. The low-order action generator includes a multi-view visual encoder, a tokenizer, a pre-trained VLM backbone network, an MLP layer, a flow matching action generation head, and a proprioceptive state processing module. The multi-view visual encoder is used to extract visual features from multi-view visual data; The tokenizer converts atomic instruction text into text features; The pre-trained VLM backbone network processes visual and textual features to generate semantic feature sequences; The proprioception state processing module acquires the proprioception state of the robotic arm. ; The MLP layer performs a projection transformation on the proprioceptive state to obtain the proprioceptive embedding, and then concatenates the proprioceptive embedding with the semantic feature sequence to construct a conditional variable. Guided by conditional variables, the Flow Matching motion generation head converts Gaussian noise into a sequence of robotic arm motion trajectories through an iterative denoising process.

10. The dual-system medical laboratory robot system according to claim 8, characterized in that, A two-stage training strategy is used when training the low-order action generator; In the first training phase, all parameters of the VLM backbone network are frozen, and only randomly initialized cross-attention action heads are trained to align the embedding space of the action heads with the semantic representation of the VLM. In the second training phase, the VLM backbone network is unfrozen, and the entire LG-Executor module is fine-tuned end-to-end.