Methods and apparatus for multi-agent path finding within diagnostic laboratory systems

A neural network-based Multi-Agent Pathfinding approach addresses the challenges of sample container routing in diagnostic laboratories by optimizing paths and timings, reducing congestion and ensuring timely processing.

HK40135821APending Publication Date: 2026-07-31SIEMENS HEALTHCARE DIAGNOSTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
HK62026125131
Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-10
Filing Date
2026-06-23
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Diagnostic laboratory systems face challenges in determining the optimal path and timing for transporting and handling multiple sample containers while avoiding collisions and unnecessary delays, particularly due to non-instantaneous processing times and priority constraints.

Method used

Employing a neural network trained with Multi-Agent Pathfinding (MAPF) to determine the actions of agents, incorporating arrival, deadline, and priority constraints, along with non-instantaneous target processing times, to optimize sample container routing in diagnostic laboratory systems.

Benefits of technology

The solution effectively reduces congestion and ensures timely processing of samples by enforcing schedule and priority constraints, enhancing the efficiency and reliability of sample handling in diagnostic laboratories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In some embodiments, a method for path finding of sample containers in a diagnostic laboratory system includes receiving a batch of sample containers in a diagnostic laboratory system having a diagnostic laboratory device; obtaining a grid for the diagnostic laboratory system, the grid having units comprising diagnostic laboratory devices and one or more tracks connecting the diagnostic laboratory devices; distributing an intelligent agent to each sample carrier; determining a target for each agent assigned to the sample carrier holding the sample container, the target including an earliest arrival time constraint, a deadline constraint, and a target processing time constraint; and employing a multi-agent path finding (MAPF) trained neural network to determine the action of each agent, the neural network being trained with arrival, deadline and priority constraints, and non-instantaneous target processing times. Numerous other aspects are provided.
Need to check novelty before this filing date? Find Prior Art

Description

(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202480043989.6 (22) Application Date 2024.08.07 (30) Priority Data 63 / 518565 2023.08.10 US (85) PCT International Application Entering National Phase Date 2025.12.29 (86) PCT International Application Application Data PCT / US2024 / 041204 2024.08.07 (87) PCT International Application Publication Data WO2025 / 034800 EN 2025.02.13 (71) Applicant Siemens Medical Diagnostics Inc., USA Address New York, USA (72) Inventor R. Prasad A. Kapoor K. Abdulrahman (74) Patent Agency China Patent Agency (Hong Kong) Limited 72001 Patent Attorneys Zhang Tao and Liu Chunyuan (51) Int.Cl. G05D 1 / 644 (2006.01) G05D 1 / 698 (2006.01) G06N 3 / 0464 (2006.01) G01N 35 / 02 (2006.01) G05D 1 / 46 (2006.01) (54) Invention Title: Method and Apparatus for Multi-Agent Pathfinding in a Diagnostic Laboratory System (57) Abstract: In some embodiments, a method for pathfinding of sample containers in a diagnostic laboratory system includes: receiving a batch of sample containers in a diagnostic laboratory system having diagnostic laboratory equipment; obtaining a grid for the diagnostic laboratory system having cells including diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment; assigning agents to each sample carrier; determining a target for each agent assigned to a sample carrier holding a sample container, the target including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint; and using a neural network trained with Multi-Agent Pathfinding (MAPF) to determine the action of each agent, the neural network being trained using arrival, deadline, and priority constraints and a non-instantaneous target processing time. Many other aspects are provided.Claims 4 pages, Description 15 pages, Drawings 7 pages, CN 121464409 A 2026.02.03 CN 1 21 46 44 09 A 1. A method for pathfinding of sample containers in a diagnostic laboratory system, the method comprising: receiving a batch of sample containers in a diagnostic laboratory system having diagnostic laboratory equipment, one or more tracks connecting the diagnostic laboratory equipment, and a plurality of sample carriers configured to transport the sample containers within the diagnostic laboratory system, each sample container containing a sample to be processed; obtaining a grid for the diagnostic laboratory system having cells including the diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment; assigning agents to each of the sample carriers; determining a target for each agent assigned to a sample carrier holding a sample container, including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint; and determining the action of each agent using a neural network trained with Multi-Agent Pathfinding (MAPF), the neural network being trained using arrival, deadline, and priority constraints and a non-instantaneous target processing time. 2. The method of claim 1, further comprising employing a sample carrier controller of the diagnostic laboratory system to perform one or more of the determined actions, including transferring a sample container for processing. 3. The method of claim 1, wherein determining a target for each agent assigned to a sample carrier holding a sample container, including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint, comprises receiving the earliest arrival time constraint, the deadline time constraint, and the target processing time constraint from a scheduler program. 4. The method of claim 1, further comprising determining a series of targets for each agent, each target including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint. 5. The method of claim 1, wherein determining the action of each agent using the MAPF-trained neural network comprises inputting local field-of-view information of each agent into the neural network, the local field-of-view information representing the state of neighboring vertices within a predetermined distance of the agent's current position. 6. The method of claim 5, wherein the predetermined distance is a predetermined radius R, which defines the local field of view of each agent as a square of length 2*R+1. 7. The method of claim 5, wherein each unit has a unit type, and wherein the agent's local field of vision includes a central unit and a plurality of units surrounding the central unit, the central unit including the agent.8. The method of claim 7, wherein determining the action of each agent using the MAPF-trained neural network comprises inputting one or more masks for the local field of view into the neural network, the one or more masks comprising at least one of an obstacle mask, a mask of the positions of all agents within the local field of view, a mask of the target positions of other agents within the local field of view, and a mask of the target position of the agent. 9. The method of claim 1, wherein determining the action of each agent using the MAPF-trained neural network comprises inputting at least one of: distance information to the target of the agent, and the order of other agents if they share the target position with the agent. 10. The method of claim 1, wherein determining the action of each agent using the MAPF-trained neural network comprises inputting next target position information of each agent into the neural network. 11. The method of claim 10, wherein the next target position information of each agent comprises information about the difference between the agent's current position and the agent's next target position. 12. The method of claim 10, wherein the next target location information for each agent includes at least one of the following: the time remaining until the earliest arrival time of the agent, the time remaining until the target deadline, and whether the target deadline has passed. Claims 1 / 4 page 2 CN 121464409 A 13. The method of claim 1, wherein the neural network is trained using at least one of a global positive reward for all agents when any agent completes the target and a global negative reward for all agents when any agent is late for the deadline. 14. A diagnostic laboratory system comprising: a diagnostic laboratory device; one or more tracks connected to the diagnostic laboratory device; a processor; and a memory coupled to the processor, the memory including computer-executable instructions that, when executed by the processor, cause the processor to: obtain a grid for the diagnostic laboratory system, the grid having cells including the diagnostic laboratory device and one or more tracks connected to the diagnostic laboratory device; assign agents to each sample carrier within the diagnostic laboratory system; determine a target for each agent assigned to a sample carrier holding a sample container, including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint; and determine the action of each agent using a neural network trained with Multi-Agent Pathfinding (MAPF), the neural network being trained using arrival, deadline, and priority constraints and a non-instantaneous target processing time.15. The diagnostic laboratory system of claim 14, further comprising a sample carrier controller, wherein the memory includes computer-executable instructions, when executed by the processor, causing the processor to employ the sample carrier controller to perform one or more determined actions, including transferring a sample container for processing. 16. The diagnostic laboratory system of claim 14, wherein the memory includes computer-executable instructions, when executed by the processor, causing the processor to determine a set of objectives for each agent, each objective including an earliest arrival time constraint, a deadline time constraint, and an objective processing time constraint. 17. The diagnostic laboratory system of claim 14, wherein the memory includes computer-executable instructions, when executed by the processor, causing the processor to input local field-of-view information for each agent into the neural network, the local field-of-view information representing the state of neighboring vertices within a predetermined distance from the agent's current position. 18. The diagnostic laboratory system of claim 17, wherein the predetermined distance is a predetermined radius R, which defines the local field of view of each agent as a square of length 2*R+1. 19. The diagnostic laboratory system of claim 17, wherein each unit has a unit type, and wherein the agent's local field of view includes a central unit and a plurality of units surrounding the central unit, the central unit including the agent. 20. The diagnostic laboratory system of claim 19, wherein the memory includes computer-executable instructions, when executed by the processor, causing the processor to input one or more masks for the local field of view into the neural network, the one or more masks including at least one of an obstacle mask, a mask of the positions of all agents within the local field of view, a mask of the target positions of other agents within the local field of view, and a mask of the target positions of the agent. 21. The diagnostic laboratory system of claim 14, wherein the memory includes computer-executable instructions, when executed by the processor, causing the processor to: for each agent, input distance information to a target of the agent, and the order of other agents if they share a target position with the agent.22. The diagnostic laboratory system of claim 14, wherein the memory includes computer-executable instructions, which, when executed by the processor, cause the processor to: determine the action of each agent using the MAPF-trained neural network by inputting next target position information of each agent into the neural network. 23. The diagnostic laboratory system of claim 22, wherein the next target position information of each agent includes information about the difference between the agent's current position and the agent's next target position. 24. The diagnostic laboratory system of claim 22, wherein the next target position information of each agent includes at least one of: the time remaining until the earliest arrival time, the time remaining until the target deadline, and whether the target deadline has passed. 25. The diagnostic laboratory system of claim 14, wherein the neural network is trained using at least one of a global positive reward for all agents when any agent completes the target and a global negative reward for all agents when any agent is late for the deadline. 26. A method for pathfinding of a sample container in a diagnostic laboratory system, the method comprising: creating a grid for the diagnostic laboratory system, the grid having cells including diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment; creating a plurality of agents, each agent representing a sample carrier within the diagnostic laboratory system; determining a target for each agent assigned to a sample carrier holding a sample container, the target including an earliest arrival time constraint, a deadline constraint, and a target processing time constraint; creating a neural network that takes the earliest arrival time constraint, the deadline constraint, and the target processing time constraint of the agent as input, and takes an action or a probability of an action of the agent as output; training the neural network using arrival, deadline, and priority constraints and a non-instantaneous target processing time; and creating a pathfinding program that uses the trained neural network to generate actions for each agent within the diagnostic laboratory system. 27. The method of claim 26, further comprising deploying the trained neural network and the pathfinding program for use by the diagnostic laboratory system. 28. The method of claim 27, further comprising using the pathfinding program to determine one or more actions, including transferring the sample container for processing. 29. The method of claim 26, wherein the neural network is configured to input local field-of-view information of each agent into the neural network, the local field-of-view information representing the state of neighboring vertices within a predetermined distance of the agent's current position.30. The method of claim 29, wherein the neural network is configured to input one or more masks for the local field of view into the neural network, the one or more masks comprising at least one of an obstacle mask, a mask of the positions of all agents within the local field of view, a mask of the target positions of other agents within the local field of view, and a mask of the target position of the agent. 31. The method of claim 26, wherein the neural network is configured to: for each agent, input at least one of the following: distance information to the target of the agent, and the order of the other agents if they share a target position with the agent. 32. The method of claim 26, wherein the neural network is configured to input next target position information for each agent. 33. The method of claim 32, wherein the next target position information for each agent includes information about the difference between the agent's current position and the agent's next target position. 34. The method of claim 32, wherein the next target location information for each agent includes at least one of the following: the time remaining until the earliest arrival time of the agent, the time remaining until the target deadline, and whether the target deadline has passed. 35. The method of claim 26, further comprising: training the neural network using at least one of a global positive reward for all agents when any agent completes the target and a global negative reward for all agents when any agent is late for the deadline. Claims 4 / 4 Page 5 CN 121464409 A Method and apparatus for multi-agent pathfinding in a diagnostic laboratory system

[0001] Cross-Reference to Related Applications This application claims the benefit of U.S. Provisional Application No. 63 / 518,565, filed August 10, 2023, pursuant to 35 USC § 119(e). The entire contents of the above-cited patent application(s) are expressly incorporated herein by reference. Technical Field

[0002] This disclosure relates to diagnostic laboratories, and more specifically, to methods and apparatus for multi-agent pathfinding within a diagnostic laboratory system. Background Art

[0003] Diagnostic laboratory systems perform tests or examinations to identify analytes or other components in biological samples such as serum, plasma, urine, interstitial fluid, cerebrospinal fluid, etc. Medical technicians collect biological samples in sample containers such as test tubes and then bring the biological samples to a diagnostic laboratory system for testing. Each biological sample may have multiple testing requirements, each of which requires a different test to be performed by the diagnostic laboratory system.

[0004] A diagnostic laboratory system may include multiple instruments, each configured to perform one or more tests. When a biological sample with a single test requirement is received in the diagnostic laboratory system, the biological sample is transferred to an instrument configured to perform the test. Similarly, when a biological sample with multiple test requirements is received in the diagnostic laboratory system, the biological sample is directed to one or more instruments collectively configured to perform multiple tests.

[0005] A scheduler creates a timetable for directing sample containers to specific instruments for testing the samples contained therein. However, determining the optimal path and timing for transporting and handling multiple sample containers while avoiding collisions and unnecessary delays is a challenge. Therefore, there is a need for improved methods and apparatus for sample container routing within a diagnostic laboratory system.

[0006] In some embodiments, a method for pathfinding of sample containers in a diagnostic laboratory system is provided, the method comprising: receiving a batch of sample containers in a diagnostic laboratory system, the diagnostic laboratory system having a diagnostic laboratory device, one or more tracks connected to the diagnostic laboratory device, and a plurality of sample carriers configured to transport the sample containers within the diagnostic laboratory system, each of the sample containers containing a sample to be processed; obtaining a grid of the diagnostic laboratory system having cells including the diagnostic laboratory device and one or more tracks connected to the diagnostic laboratory device; assigning agents to each of the sample carriers; determining a target for each agent assigned to a sample carrier holding a sample container, the target including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint; and employing a neural network trained with multi-agent pathfinding (MAPF) to determine the action of each agent, the neural network being trained using arrival, deadline, and priority constraints and a non-instantaneous target processing time.

[0007] In some embodiments, a diagnostic laboratory system is provided, the diagnostic laboratory system including a diagnostic laboratory device; one or more tracks connected to the diagnostic laboratory device; a processor; and a memory coupled to the processor.The memory includes computer-executable instructions that, when executed by the processor, cause the processor to: obtain a grid of the diagnostic laboratory system having cells including the diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment; assign agents to each sample carrier within the diagnostic laboratory system; determine a target for each agent assigned to a sample carrier holding a sample container, the target including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint; and determine the action of each agent using a MAPF-trained neural network trained with arrival, deadline, and priority constraints, as well as non-instantaneous target processing time.

[0008] In some embodiments, a method for pathfinding of sample containers in a diagnostic laboratory system is provided, the method comprising: creating a grid for the diagnostic laboratory system, the grid having cells including diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment; creating a plurality of agents, each agent representing a sample carrier within the diagnostic laboratory system; determining a target for each agent assigned to a sample carrier holding a sample container, the target including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint; creating a neural network that takes the earliest arrival time constraint, the deadline time constraint, and the target processing time constraint of the agent as input, and takes an action or probability of an action of the agent as output; training the neural network using arrival, deadline, and priority constraints and a non-instantaneous target processing time; and creating a pathfinding program that uses the trained neural network to generate actions for each agent within the diagnostic laboratory system.

[0009] Other aspects, features, and advantages of this disclosure may be apparent from the following description and illustration of exemplary embodiments, including the best mode contemplated for carrying out this disclosure. This disclosure may also have other embodiments and different embodiments, and certain details thereof may be modified in various ways without departing from the scope of this disclosure. Brief Description of the Drawings

[0010] The drawings described below are provided for illustrative purposes and are not necessarily drawn to scale. Therefore, the drawings and description are to be considered illustrative in nature and not restrictive. The drawings are not intended to limit the scope of this disclosure in any way.

[0011] FIG1 illustrates an embodiment of a diagnostic laboratory system according to embodiments provided herein.

[0012] FIG2A illustrates an example embodiment of the computer of FIG1 according to some embodiments provided herein.

[0013] FIG2B illustrates an example flowchart according to some embodiments provided herein, depicting the information flow from the scheduler program to the pathfinding program, from the pathfinding program to the neural network and from the neural network to the pathfinding program, and from the pathfinding program to the sample carrier controller.

[0014] FIG3A illustrates an example embodiment of another diagnostic laboratory system according to embodiments provided herein.

[0015] FIG3B illustrates a portion of the diagnostic laboratory system of FIG3A with a first agent and a second agent according to embodiments provided herein.

[0016] FIG3C illustrates a portion of the diagnostic laboratory system of FIG3B according to embodiments provided herein, showing vertices connected by edges for cells that can be reached by the sample carrier.

[0017] FIG3D illustrates a portion of the diagnostic laboratory of FIG3B, showing vertices and edges without cells for clarity.

[0018] FIG3E illustrates several example masks representing the states of cells within the diagnostic laboratory system of FIG3A according to embodiments provided herein.

[0019] Figure 4 is a flowchart of an example method for pathfinding of sample containers in a diagnostic laboratory system according to some embodiments provided herein (page 2 / 15, CN 121464409 A).

[0020] Figure 5 is a flowchart of another method for finding a path for sample containers in a diagnostic laboratory system according to some embodiments provided herein. Detailed Description

[0021] Regardless of grammatical usage, individuals with male, female, or other gender identities are included in the terminology.

[0022] Scheduling and / or routing sample containers to specific instruments in a diagnostic laboratory system is complex because diagnostic laboratory systems have many different types of components and may receive many different types of samples. For example, a diagnostic laboratory system may include many instruments, each of which may be configured with different test menus, test procedures, and test durations. Furthermore, samples to be analyzed within a diagnostic laboratory system may have priorities, such as routine or stat samples requiring immediate processing.

[0023] The embodiments provided herein employ multi-agent pathfinding techniques to determine actions for transporting and processing sample containers within a diagnostic laboratory system. Pathfinding plays an important role in many areas of automation. Many automated setups involve multiple interacting components that need to act in concert. The Multi-Agent Pathfinding Problem (MAPF) illustrates important aspects of the target domain, such as collision avoidance and minimization of travel time. However, a key assumption in current MAPF is that task processing is instantaneous. This discrepancy is amplified in areas of laboratory or factory automation where tasks, as part of an automation pipeline, require non-trivial processing.Machines may take different amounts of time to process different tasks. Without a corresponding time slack for machine availability, agent congestion may occur because such methods cannot capture priority constraints and queuing effects caused by non-instantaneous task processing time.

[0024] According to one or more embodiments provided herein, the above-described MAPF problem is solved by enforcing a schedule and priority constraints. For example, the schedule can be enforced by automated pipeline capabilities such as machine operating characteristics. In some embodiments, a time-constrained MAPF (MAPF-TC) approach is provided, which can be modeled as a sequential decision problem.

[0025] In one or more embodiments, a multi-agent reinforcement learning (MARL) model is employed to solve a time-constrained multi-agent pathfinding problem. In some embodiments, a time-constrained version of the MAPF problem is provided, which introduces arrival, deadline, and priority constraints with non-instantaneous target processing time. Furthermore, in some embodiments, novel rewards and observation spaces are provided for modeling time aspects and priority constraints.

[0026] These and other embodiments provided herein are described below with reference to Figures 1-5.

[0027] Referring to Figure 1, Figure 1 illustrates an embodiment of a diagnostic laboratory system 100 according to embodiments provided herein. Diagnostic laboratory system 100 may include multiple instruments 102 (e.g., diagnostic laboratory devices) configured to process samples stored in sample containers 104 (some labeled) and to test the samples (e.g., trials or other tests). Performing a test may include performing one or more operations on the sample. For example, each operation may include one or more measurements. One or more of the instruments 102 may include multiple different modules configured to perform operations.

[0028] Samples may be various biological samples collected from individuals such as patients being evaluated by a medical professional. Samples may be collected in sample containers 104 and delivered to diagnostic laboratory system 100, and subsequently transported throughout the diagnostic laboratory system 100 via track 108, such as to instruments 102. For example, sample containers 104 may be transported by sample carriers 110 (some labeled). In the embodiment of FIG1, diagnostic laboratory system 100 has three instruments 102, which include a sample handler 114, a first analyzer 116, and a second analyzer 118. Diagnostic laboratory system 100 may include fewer or more instruments than shown in FIG1.

[0029] In some embodiments, the track 108 may be close to or extend around the instrument 102, as shown in FIG1.Instrument manual 3 / 15 pages 8 CN 121464409 A Parts or modules of instrument 102 may have devices for transferring sample containers 104 to sample carriers 110 and transferring sample containers 104 from sample carriers 110, such as automated mechanical devices (robots) (not shown in FIG. 1). Track 108 may include multiple segments 120 (some labeled) that can be interconnected. Sample carriers 110 can move as shown by dashed lines 126 in the segments 120. In some embodiments, some segments 120 may be integrated with one or more of the instruments 102.

[0030] Diagnostic laboratory systems such as laboratory system 100 may have many instruments and may have tracks linked to other laboratory systems. Such laboratory systems (including diagnostic laboratory system 100) can move and process multiple sample carriers 110 and their respective sample containers 104 simultaneously. In some embodiments, diagnostic laboratory system 100 can move and process hundreds or thousands of sample carriers 110 and their respective sample containers 104 simultaneously.

[0031] The diagnostic laboratory system 100 may include or be coupled to a computer 130, which is configured to execute one or more programs and control the operation of the diagnostic laboratory system 100. The computer 130 may be configured to communicate with the instrument 102 and other components of the diagnostic laboratory system 100, such as components in a transport system (e.g., track 108 and components controlling its operation). The transport system may include some or all of the components (e.g., motors, sensors, power supplies, etc.) configured to transport samples throughout the diagnostic laboratory system 100. The computer 130 may include a processor 132 configured to execute programs, including programs other than those described herein. The programs may be implemented as computer code, computer-executable instructions, etc.

[0032] The computer 130 may include or access a memory 134, which may store one or more programs 136 and / or data. The memory 134 may be any suitable type of memory, such as, but not limited to, one or more of volatile and / or non-volatile memory. In one or more embodiments, memory 134 may be non-transitory memory (e.g., hard disk drive, solid-state drive, flash drive, another non-transitory computer-readable medium, etc.). Program 136 may be computer code and / or instructions executable on or by processor 132. As described below, such a program can control all or part of the operation of the diagnostic laboratory system 100.

[0033] Computer 130 may be coupled to workstation 138, which is configured to allow a user to interface with the diagnostic laboratory system 100. Workstation 138 may include a display 140, keyboard 142, and other peripheral devices.The program within memory 134 enables display 140 to display information about the handling of samples within diagnostic laboratory system 100.

[0034] When a sample is received in diagnostic laboratory system 100 for testing, the sample undergoes a complex sequence of operations or processes within a specific workflow defined by the specific test. Each type of test can have a unique sequence of operations. A workflow sequence may begin with a sample handling or sample container handling operation, followed by the aspiration and dispersion of the sample and / or reagents into cuvettes. The mixture in the cuvettes may undergo other operations required for the test. A workflow sequence may end with a measurement operation, such as a photometric measurement, to determine the chemical properties of the sample. Each operation in the workflow sequence can be performed using one or more of instruments 102.

[0035] FIG. 2A illustrates an example embodiment of computer 130 of FIG. 1 according to some embodiments. In the embodiment of FIG. 2A, memory 134 of computer 130 includes scheduler program 202 coupled to pathfinding program 204. As described below, the pathfinding program 204 can receive schedule information from the scheduler program 202 and use a neural network 206 trained with Multi-Agent Pathfinding (MAPF) to determine the actions (such as carrier movement of each sample carrier 110) within the diagnostic laboratory system 100. These actions can be communicated from the pathfinding program 204 to the sample carrier controller 208 for execution. For example, the sample carrier controller 208 can control the operation of the motors, sensors, power supplies, etc. used to transport the sample carriers 110 on the track 108. Although shown in the memory 134 of the computer 130, it will be understood that the scheduler program 202 and / or the sample carrier controller 208 may reside in different memories and / or be executed by different processors.

[0036] FIG2B illustrates an example flowchart according to some embodiments, depicting the information flow from scheduler program 202 to pathfinding program 204, from pathfinding program 204 to neural network 206 and from neural network 206 to pathfinding program 204, and from pathfinding program 204 to carrier controller 208.

[0037] Scheduler program 202 may be configured to determine a schedule or workflow for processing all sample containers 104 within the diagnostic laboratory system 100 (e.g., where sample carrier 110 should travel to pick up and transport sample containers 104 for processing). For example, scheduler program 202 may determine that sample containers 104 within sample disposal unit 114 should be retrieved and delivered to one or more of first analyzer 116 and second analyzer 118 for processing.In some embodiments, scheduler program 202 may generate a workflow indicating the diagnostic laboratory equipment to be accessed by each sample container and their order (e.g., based on information from scanning or imaging tags attached to the sample container regarding one or more analyses to be performed on the sample in the sample container). In one or more embodiments, processor 132 (or another processor) may simulate the workflow for each sample container to estimate workflow completion time information for each sample container.

[0038] Schedule information may be provided to pathfinding program 204, which in turn may employ a MAPF-trained neural network 206 to determine the actions of each sample carrier 110 within the diagnostic laboratory system 100, such that the schedule determined by scheduler program 202 is effectively implemented (e.g., collision-free delivery of sample containers for processing within any required deadlines). These actions may be fed to sample carrier controller 208 for execution. Thus, scheduler program 202, pathfinding program 204, MAPF-trained neural network 206, and sample carrier controller 208 may be deployed within the diagnostic laboratory system for use to facilitate sample processing.

[0039] In one or more embodiments, the time-constrained MAPF problem can be defined by an undirected graph G=(V,E), a set of m agents, and a timetable S. The vertex set V corresponds to the positions on the graph, and the edge set E represents the motion constraints for each vertex. At each step, an agent can move or wait at its current position. Each agent is assigned a series of target positions according to the timetable S. The timetable S represents time windows that include the earliest arrival time (Ati(j)), the deadline (Dti(j)), and the target processing time (Pti(j)) for agent i and target j. For example, the timetable S, along with the earliest arrival time, deadline, and target processing time constraints, can be provided by a scheduler program 202.

[0040] Ati(j) constrains when an agent can begin its target processing and where the agent might arrive before starting, but the target will not be processed before the scheduled time. Dti(j) constrains the latest time an agent is allowed to arrive at its target for processing. Pt i(j) represents (e.g., and constrains) non-instantaneous target processing that agents are not allowed to move. This schedule also enforces priority constraints on target locations shared among multiple agents (e.g., factory machines such as sample analyzers). While soft violations with penalized deadlines are permissible, priority constraints are still enforced, effectively inducing a queue of agents sharing target locations. In some embodiments, agents are always present on the graph from the beginning and do not disappear after completing their targets.The persistence of all agents on spatially constrained orbits is a key challenge.

[0041] In one or more embodiments, the interaction between the agent and the environment can be modeled as a partially observable Markov decision process (POMDP) ​​(S, O, A, P, R, ), where S is the set of environmental states, O is the set of partial observations, A is the set of actions, P is a function representing the transition probability, R is the reward function, and is a discount factor. In the embodiments described herein, the environment is constrained to a 2D grid, where each agent is restricted to local field of view observations. In other embodiments, a 3D grid can be employed. Homogeneous policies can be learned, which can be deployed to different numbers of agents and executed in a distributed manner.

[0042] In some embodiments, the observation space can include two components. First, the local field of view can represent the state of neighboring vertices relative to the agent's current position. Second, information about the next target position can be directly encoded using vectors that represent the spatial and temporal information of the target in three dimensions, as described below.

[0043] FIG. 3A illustrates an example embodiment of another diagnostic laboratory system 300 according to embodiments provided herein. Referring to FIG. 3A, the diagnostic laboratory system 300 is represented as a grid 302 that defines a plurality of cells 304 (only some are labeled) within the diagnostic laboratory system 300. As indicated by the key in FIG. 3A, cells 304 without shading represent sample carrier movement locations, such as track segments, track connections, machine padding locations, etc. Cells 304 with medium shading represent diagnostic laboratory equipment, such as sample analyzers, sample handlers, etc. Cells 304 with dark shading represent empty or unusable areas of the diagnostic laboratory system 300. In the embodiment of FIG. 3A, carrier movement includes movement in the north (N), south (S), east (E), and west (W) directions. In some embodiments, the size of a cell 304 may be determined based on one or more of the size of the sample carrier 110 (e.g., large enough to accommodate the sample carrier), the movement resolution of the track segment used, etc. For example, in some embodiments, the size of each unit 304 may be determined to hold a single sample carrier 110. Other unit sizes may be used.

[0044] FIG3B illustrates a portion 300a of the diagnostic laboratory system 300 of FIG3A according to an embodiment provided herein, having a first agent 308a and a second agent 308b. A first local FOV 310a is shown for the first agent 308a, and a second local FOV 310b is shown for the second agent 308b, each FOV including the nearest neighbor unit 304. Other FOV sizes may be employed.The first agent 308a is located on a track segment close to the first diagnostic laboratory device 312a (e.g., the first sample analyzer). The second agent 308b is located within the second diagnostic laboratory device 312b (e.g., the second sample analyzer).

[0045] FIG3C illustrates a portion 300a of a diagnostic laboratory system 300 according to an embodiment provided herein, showing vertices 314 (connected by edges 316) of units 304 that can be reached for a sample carrier. FIG3D illustrates a portion 300a, which for clarity shows vertices 314 and edges 316, but not units 304.

[0046] As shown in FIG3D, a local field of view 310a can represent the state of neighboring vertices 314 relative to the current position of the first agent 308a within a radius R. For example, the local field of view (FOV) can be a square of length 2*R+1 around the agent.

[0047] In some embodiments, the state of each unit 304 within the field of view can be represented by a plurality of masks 318a-d, as shown in FIG3E. Although four masks are shown in FIG3E, it should be understood that fewer or more masks may be used, as well as larger masks (e.g., for a larger field of view). For example, each unit state can be represented by one or more of the following: (1) a binary mask of obstacles; (2) a binary mask of the positions of all agents within the FOV; (3) a binary mask of the target positions of other agents within the FOV; and (4) a binary mask of the target position of an agent within the unit, projected onto the boundary if it is located outside the FOV. Furthermore, in some embodiments, if other agents share the same target position with an agent, the state of each unit of the unit with that agent can also be represented by the normalized distance from each vertex in the FOV to the target of that agent and / or the order of other agents in the queue.

[0048] As stated above, in some embodiments, a vector can be used to directly encode information about the agent's next target location, representing the spatial and temporal information of the target in three dimensions. For example, the spatial vector Gs can have a first vector component Gs1 representing the difference between the agent's current x-coordinate and the target x-coordinate. A second vector component Gs2 can represent the difference between the agent's current y-coordinate and the target y-coordinate. A third vector component Gs3 can represent the magnitude of the vector from the agent's current position to its target position, which is limited to an absolute value (e.g., 60 in some embodiments).

[0049] The temporal vector Gt can represent the temporal information of the agent's target in three dimensions.For example, if the current time is less than At, the first vector component Gt1 can represent the time remaining until the earliest arrival time At; otherwise, the first vector component Gt1 can represent 0.0. If the current time is between At and Dt, the second vector Gt2 can represent the time remaining until the deadline Dt; otherwise, the second vector Gt2 can represent 0.0. If the deadline Dt has passed, indicating that the agent is late, the third vector Gt3 can be set to 1; otherwise, the third vector Gt3 can be set to 0.0. Other spatial vectors and / or time vectors may be used.

[0050] Returning to FIG2A, in some embodiments, the pathfinding procedure 204 may include (stored in memory 134) computer-executable instructions that, when executed by processor 132, cause processor 132 to obtain a grid for the diagnostic laboratory system (e.g., grid 302 for the diagnostic laboratory system 300 in FIG3A). The grid defines multiple cells, including cells within the diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment (e.g., as shown in Figures 3A and 3B).

[0051] The pathfinding procedure 204 can also assign an agent to each sample carrier within the diagnostic laboratory system. For example, a first agent 308a and a second agent 308b are assigned to sample carriers 320a and 320b, respectively (Figure 3B).

[0052] The pathfinding procedure 204 can further determine a target for each agent assigned to a sample carrier holding a sample container, the target including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint. As stated, in some embodiments, such information can be provided from a scheduler procedure 202 (Figure 2A).

[0053] The pathfinding procedure 204 can employ a MAPF-trained neural network (e.g., a MAPF-trained neural network 206) to determine the action of each agent, which is trained using arrival, deadline, and priority constraints, as well as non-instantaneous target processing times. In the example of Figure 3A, the actions of the agent (e.g., the sample carrier) may include waiting or moving north, south, east, or west. Other actions may be provided.

[0054] In some embodiments, the sample carrier controller 208 may be employed to perform one or more of the determined actions (via computer program instructions within memory 134). For example, the sample carrier controller 208 may cause the transfer of the sample container for processing within the diagnostic laboratory system 300.

[0055] Training the neural network MAPF neural network 206 can be any suitable neural network, such as a convolutional neural network, a transformer neural network, etc.In some embodiments, each agent can receive a positive reward when it reaches its goal within the scheduled time window. A negative penalty can be given for each time step in which an agent occupies its goal before the earliest arrival time. This approach aims to reduce congestion, as agents waiting for their goals may block the path of other agents. A negative penalty can be given for each time step in which any agent has not reached its goal after the deadline.

[0056] In the reinforcement learning (RL) exposition of MAPF, penalties are typically given when an agent collides with static obstacles or other agents in the environment. The embodiments provided herein aim to reduce unnecessary subproblems for more efficient learning. One subproblem is static obstacle avoidance. Accordingly, invalid actions can be masked (e.g., ignored) in cases where an agent will collide with an obstacle. Another subproblem is collision avoidance with other agents. The inventors have observed that, without collision penalties, neural network policies tend to converge faster on goal completion behavior while implicitly avoiding collisions. Explicit collision penalties can slow down training because agents are more likely to collide with each other early in the learning process. Collision avoidance is still enforced through the environment. Furthermore, in some embodiments, no penalty is imposed on agent movement. Additionally, in one or more embodiments, an agent may need to wait or reroute its path without being later than its deadline.

[0057] Unlike some existing methods, a global positive reward can be distributed to all agents when any agent completes the goal. Similarly, a global negative reward can be distributed to all agents when any agent is late for its deadline. The idea is to incentivize cooperation among agents, especially those who slack off until their deadline and can afford to take longer paths.

[0058] Table 1 provides examples of reward values ​​for different events. Other reward values ​​may be used.

[0059] Table 1 Event Reward Achieved Target in Time Window +1.0 Achieved Target Before Time Window -0.1 Achieved Target After Time Window (Late) -0.1 Achieved Target During Processing 0.0 Environmental Collision Masked Agent Collision 0.0 Agent Action (Move / Wait) 0.0 Shared Target Reward +0.2 Shared Delay Reward -0.04

[0060] In some embodiments, the local view is processed by a convolutional neural network (CNN) (such as a 3-layer CNN), where each layer except the last layer is followed by a max pooling operation, and the last layer is followed by a global average to produce a view representation vector. The position and time window vectors are then fed through fully connected layers to produce a combined target representation vector.Then, both the vision and target representation vectors are passed through recursive modules to incorporate information about past states as a mitigation of the partial observability problem. The model is optimized using proximal policy optimization (PPO) as described in Schulman, John, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. “Proximal Policy Optimization Algorithms” arXiv. http: / / arxiv.org / abs / 1707.06347.

[0061] Selective scenarios can be employed to investigate the impact of different components on model performance. In some embodiments, the time window can be parameterized to alter the resulting scheduling relaxation.

[0062] Unlike the warehouse domain, automated factories are constrained by empty space, machine dimensional dimensions, and security constraints such as power and ventilation outlets. To capture these various constraints, a parametric layout planning generator can be employed to generate a parametric layout for a diagnostic laboratory system. Training on various such layouts can generate a model that can satisfy nearly 100% of the objectives within their respective time windows. In such layouts, agents may be able to learn traffic-like behaviors, such as parallel lanes / corridors for reverse flow. While scaling to a larger number of agents within the same fixed layout is an important aspect of MAPF, it may not maintain the same value under time constraints where the throughput of the entire system is limited by (automated) scheduling capabilities. Adding more agents may simply result in agents waiting idly and having ample time to navigate tracks. As the number of agents in MAPF increases under time constraints, the bottleneck becomes scheduling rather than pathfinding.

[0063] In some embodiments, a more restrictive grid layout (e.g., a single lane) may still be able to accomplish nearly 100% of the agent objectives. In well-connected layouts where tracks form continuous paths, agents can learn to circle around the tracks, thus maintaining a flow state until their deadlines expire without blocking other agents (who may then deviate from their objectives). In a more constrained layout with forked dead ends, the agent can learn to reverse through fill segments around the decision point (e.g., the location of a sample carrier adjacent to a diagnostic laboratory device).

[0064] In some embodiments, a layout design can be created to identify the impact of different factors on performance, as well as to establish constraints on the pathfinding procedure 204. For example, the layout can be modified in three design dimensions: redundancy, size, and fill. Redundancy can be modified by connecting tracks to remove forked dead ends. Size can be tested by changing the number of targets / channels.In some embodiments, this may include producing two variations: a small variation with 3 targets / channels and a large variation with 6 targets / channels, respectively. This factor can affect the number of agents (e.g., the number of agents can be set to equal the number of targets plus 2). Finally, different padding can be added around the decision nodes.

[0065] The scheduling distribution can be varied across multiple design dimensions, such as the single agent shortest distance (A* factor) (which is the minimum distance an agent travels assuming no other agents are present in the target) and the congestion estimate for multi-agents (where the average of all single agent shortest distances is added as a congestion estimate). These represent scheduling constraints that affect the earliest arrival time of each agent, thus providing the minimum necessary conditions for scheduling feasibility. Furthermore, the size of the time window from the earliest arrival time to the deadline and the task processing runtime before completion can be varied.

[0066] One objective of the MAPF-trained neural network 206 is to maximize the number of targets completed within its scheduled time window. The relative throughput of targets completed within the time window can be measured. This standardizes the comparison to understand the impact of different time window distributions on reinforcement learning models. Example metrics that can be considered include the percentage of objectives completed within the lateness tolerance, thus highlighting the suboptimal nature of the model (if any), and the time taken to complete a predetermined number of objectives (e.g., 100), thus standardizing across different timeframes. A coefficient of variation, defined as the ratio between the mean and standard deviation of objective completion times, can also be used to capture the reliability of the model when producing similar performance.

[0067] Figure 4 is a flowchart of an example method 400 for pathfinding of sample containers in a diagnostic laboratory system according to some embodiments. In some implementations, one or more process blocks of Figure 4 may be executed by a pathfinding procedure (such as pathfinding procedure 204 of Figure 2A). For example, memory 134 may include computer-executable instructions that, when executed by processor 132, cause processor 132 to perform one or more of the process steps described with reference to Figure 4. In at least some embodiments, such computer-executable instructions may be stored in non-transitory memory and / or included in a non-transitory computer-readable medium.

[0068] As shown in FIG4, method 400 may include receiving a batch of sample containers in a diagnostic laboratory system having diagnostic laboratory equipment, one or more tracks connected to the diagnostic laboratory equipment, and a plurality of sample carriers configured to transport the sample containers within the diagnostic laboratory system, each sample container containing a sample to be processed (box 402).For example, a sample handler (such as sample handler 114 of FIG. 1) may receive a batch of sample containers (e.g., sample container 104) in a diagnostic laboratory system (e.g., diagnostic laboratory system 100 of FIG. 1 or diagnostic laboratory system 300 of FIG. 3A), the diagnostic laboratory system having diagnostic laboratory equipment (e.g., sample handler, sample analyzer, etc.), one or more tracks connecting the diagnostic laboratory equipment, and a plurality of sample carriers configured to transport sample containers within the diagnostic laboratory system, each sample container containing a sample to be processed, as described above (e.g., sample handler 114 of FIG. 1, first sample analyzer 116, second analyzer 118, track 108, sample container 104, and sample carrier 110). Note that some sample carriers may be empty until they receive sample containers at the input module of the diagnostic laboratory system (e.g., at sample handler 114 of diagnostic laboratory system 100).

[0069] Also as shown in FIG. 4, method 400 may include obtaining a grid for the diagnostic laboratory system having cells (box 404) including diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment. For example, the pathfinding procedure 204 of FIG2A can obtain a grid 302 for the diagnostic laboratory system 300 (FIG. 3A), the grid 302 having cells 304 including diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment, as described above. In some embodiments, the grid can be stored in memory 134, or automatically generated by scheduler procedure 202 and / or pathfinding procedure 204 and / or another procedure.

[0070] As further shown in FIG4, method 400 may include assigning an agent to each sample carrier (block 406). Specification 9 / 15 pages 14 CN 121464409 A For example, pathfinding procedure 204 may assign an agent to each sample carrier 110 within the diagnostic laboratory system 100, as described above.

[0071] Also as shown in FIG4, method 400 may include determining a target for each agent assigned to a sample carrier holding a sample container, the target including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint (block 408). For example, pathfinding procedure 204 (e.g., via scheduler procedure 202) can determine a target for each agent assigned to the sample carrier 110 holding the sample container 104, the target including an earliest arrival time constraint, a deadline constraint, and a target processing time constraint, as described above. In some embodiments, this may include receiving the earliest arrival time, deadline, and target processing time constraint from a procedure such as scheduler procedure 202 (FIG. 2A).Furthermore, in some embodiments, the pathfinding procedure 204 (and / or scheduler procedure 202) may determine a set of objectives for each agent, each objective including an earliest arrival time constraint, a deadline constraint, and an objective processing time constraint.

[0072] As further shown in FIG4, method 400 may include employing a MAPF-trained neural network to determine the action of each agent, the neural network being trained using arrival, deadline, and priority constraints, as well as non-instantaneous objective processing time (box 410). For example, pathfinding procedure 204 may employ a MAPF-trained neural network 206 to determine the action of each agent, the neural network 206 being trained using arrival, deadline, and priority constraints, as well as non-instantaneous objective processing time, as described above.

[0073] In some embodiments, employing a MAPF-trained neural network to determine the action of each agent may include inputting local field-of-view information of each agent into the neural network, the local field-of-view information representing the state of neighboring vertices within a predetermined distance of the agent's current position. For example, the predetermined distance may be a predetermined radius R, which defines the local field of view of each agent as a square of length 2*R+1, as shown in local field of view 310a in FIG3D.

[0074] Each unit may have a unit type, and the agent's local field of view may include a central unit having the agent and multiple units surrounding the central unit, such as the local fields of view 310a and 310b shown for FIG3B-3D.

[0075] In one or more embodiments, using a MAPF-trained neural network to determine the action of each agent may include inputting one or more masks for the local field of view (e.g., masks 318a-d of FIG3E) into the neural network, the one or more masks including at least one of an obstacle mask, a mask of the positions of all agents within the local field of view, a mask of the target positions of other agents within the local field of view, and a mask of the target positions of the agent, as described above. Furthermore, using a MAPF-trained neural network to determine the action of each agent may include inputting at least one of distance information to the agent's target and the order of other agents if they share the target position with the agent.

[0076] Using a MAPF-trained neural network to determine the action of each agent may also include inputting the next target position information of each agent into the neural network. For example, the next target location information for each agent may include the difference between the agent's current location and the agent's next target location, the time remaining until the agent's earliest arrival time, the time remaining until the target deadline, and / or information on whether the target deadline has passed.

[0077] In some embodiments, neural network 206 may be trained using at least one of a global positive reward for all agents when any agent completes the objective and a global negative reward for all agents when any agent is late for the deadline.

[0078] Although FIG4 shows an example block of method 400, in some embodiments, method 400 may include additional blocks, fewer blocks, different blocks, or blocks with different arrangements compared to those blocks depicted in FIG4. Additionally or alternatively, two or more blocks of method 400 may be executed in parallel.

[0079] In some embodiments, a sample carrier controller 208 (FIG. 2A) of a diagnostic laboratory system 100 or 300 may be employed to perform one or more of the actions determined by pathfinding procedure 204 and / or neural network 206, including transferring sample containers for processing (e.g., transferring sample container 104 to sample analyzer 116 or 118).

[0080] FIG5 is a flowchart of another method 500 for path finding of sample containers in a diagnostic laboratory system according to some embodiments. In some implementations, one or more process blocks of FIG5 may be executed by a processor (such as processor 132 of FIG2A). For example, memory 134 may include computer-executable instructions that, when executed by processor 132, cause processor 132 to perform one or more of the process steps described with reference to FIG5. In at least some embodiments, such computer-executable instructions may be stored in non-transitory memory and / or included in a non-transitory computer-readable medium.

[0081] As shown in FIG5, method 500 may include creating a grid for a diagnostic laboratory system having units (block 502) including diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment. For example, processor 132 executing one or more programs may create a grid for a diagnostic laboratory system having units including diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment, as described above.

[0082] Also as shown in FIG5, method 500 may include creating a plurality of agents, each agent representing a sample carrier within a diagnostic laboratory system (block 504). For example, processor 132 executing one or more programs may create a plurality of agents, each agent representing a sample carrier within a diagnostic laboratory system, as described above.

[0083] As further shown in FIG5, method 500 may include determining a target for each agent assigned to a sample carrier holding a sample container, the target including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint (block 506).For example, processor 132 executing scheduler program 202 and / or another program (e.g., pathfinding program 204) can determine a target for each agent assigned to a sample carrier holding a sample container, including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint, as described above.

[0084] Also as shown in FIG5, method 500 may include creating a neural network that takes the agent's earliest arrival time constraint, deadline time constraint, and target processing time constraint as input and takes the agent's action or the probability of the action as output (box 508). As described above, in some embodiments, neural network 206 may be any suitable neural network, such as a convolutional neural network, a transformer neural network, etc.

[0085] As further shown in FIG5, method 500 may include training the neural network using arrival, deadline, and priority constraints, as well as non-instantaneous target processing time (box 510). For example, processor 132 may train neural network 206 using arrival, deadline, and priority constraints, as well as non-instantaneous target processing time, as described above.

[0086] Also as shown in FIG5, method 500 may include creating a pathfinding procedure that employs a trained neural network to generate actions for each agent within the diagnostic laboratory system (box 212). For example, processor 132 may be used to create pathfinding procedure 204 that employs a trained neural network 206 to generate actions for each agent within the diagnostic laboratory system, as described above.

[0087] Although FIG5 shows example boxes of method 500, in some implementations, method 500 may include additional boxes, fewer boxes, different boxes, or boxes with different arrangements compared to those depicted in FIG5. Additionally or alternatively, two or more boxes of method 500 may be executed in parallel.

[0088] Non-limiting illustrative embodiments The following provides a non-limiting list of illustrative embodiments of the present disclosure: An illustrative method for pathfinding of sample containers in a diagnostic laboratory system, the method comprising: receiving a batch of sample containers in a diagnostic laboratory system having diagnostic laboratory equipment, one or more tracks connecting the diagnostic laboratory equipment, and a plurality of sample carriers configured to transport the sample containers within the diagnostic laboratory system, each of the sample containers containing a sample to be processed; obtaining a grid for the diagnostic laboratory system having cells including the diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment; assigning agents to each sample carrier within the diagnostic laboratory system; determining a target for each agent assigned to a sample carrier holding a sample container, including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint; and employing a neural network trained with multi-agent pathfinding (MAPF) to determine the action of each agent, the neural network being trained using arrival, deadline, and priority constraints and a non-instantaneous target processing time.

[0089] The illustrative method of any of the foregoing illustrative embodiments further includes employing a sample carrier controller of the diagnostic laboratory system to perform one or more of the determined actions, including transferring a sample container for processing.

[0090] The illustrative method of any of the foregoing illustrative embodiments, wherein determining a target for each agent assigned to a sample carrier holding a sample container, including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint, includes receiving the earliest arrival time constraint, the deadline time constraint, and the target processing time constraint from a scheduler program.

[0091] The illustrative method of any of the foregoing illustrative embodiments, further includes determining a series of targets for each agent, each target including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint.

[0092] The illustrative method of any of the foregoing illustrative embodiments, wherein determining the action of each agent using the MAPF-trained neural network includes inputting local field-of-view information of each agent into the neural network, the local field-of-view information representing the state of neighboring vertices within a predetermined distance of the agent's current position.

[0093] The illustrative method of any of the foregoing illustrative embodiments, wherein the predetermined distance is a predetermined radius R, which defines the local field of view of each agent as a square of length 2*R+1.

[0094] An illustrative method in any of the foregoing illustrative embodiments, wherein each unit has a unit type, and wherein the agent's local field of view includes a central unit and a plurality of units surrounding the central unit, the central unit including the agent.

[0095] An illustrative method in any of the foregoing illustrative embodiments, wherein determining the action of each agent using the MAPF-trained neural network includes inputting one or more masks for the local field of view into the neural network, the one or more masks including at least one of an obstacle mask, a mask of the positions of all agents within the local field of view, a mask of the target positions of other agents within the local field of view, and a mask of the target position of the agent.

[0096] An illustrative method in any of the foregoing illustrative embodiments, wherein determining the action of each agent using the MAPF-trained neural network includes inputting at least one of distance information to the target of the agent and the order of other agents if they share the target position with the agent.

[0097] In any of the foregoing illustrative embodiments, the illustrative method of determining the action of each agent using the MAPF-trained neural network includes inputting the next target position information of each agent into the neural network.

[0098] In any of the foregoing illustrative embodiments, the next target position information of each agent includes information about the difference between the agent's current position and the agent's next target position. Specification 12 / 15 pages 17 CN 121464409 A

[0099] In any of the foregoing illustrative embodiments, the next target position information of each agent includes at least one of the time remaining until the agent's earliest arrival time, the time remaining until the target deadline, and whether the target deadline has passed.

[0100] In any of the foregoing illustrative embodiments, the illustrative method of training the neural network uses at least one of a global positive reward for all agents when any agent completes the target and a global negative reward for all agents when any agent is late for the deadline.

[0101] An illustrative diagnostic laboratory system includes: a diagnostic laboratory device; one or more tracks connected to the diagnostic laboratory device; a processor; and a memory coupled to the processor, the memory including computer-executable instructions that, when executed by the processor, cause the processor to: obtain a grid for the diagnostic laboratory system, the grid having cells including the diagnostic laboratory device and one or more tracks connected to the diagnostic laboratory device; assign agents to each sample carrier within the diagnostic laboratory system; determine a target for each agent assigned to a sample carrier holding a sample container, including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint; and determine an action for each agent using a neural network trained with Multi-Agent Pathfinding (MAPF), the neural network being trained using arrival, deadline, and priority constraints, and a non-instantaneous target processing time.

[0102] The illustrative diagnostic laboratory system of any of the foregoing illustrative embodiments further includes a sample carrier controller, and wherein, when executed by the processor, the computer-executable instructions cause the processor to use the sample carrier controller to perform one or more of the determined actions, including transferring a sample container for processing.

[0103] In any of the foregoing illustrative embodiments of the illustrative diagnostic laboratory system, the memory includes computer-executable instructions that, when executed by the processor, cause the processor to determine a set of objectives for each agent, each objective including an earliest arrival time constraint, a deadline time constraint, and an objective processing time constraint.

[0104] In any of the foregoing illustrative embodiments of the illustrative diagnostic laboratory system, the memory includes computer-executable instructions that, when executed by the processor, cause the processor to input local field-of-view information of each agent into the neural network, the local field-of-view information representing the state of neighboring vertices within a predetermined distance of the agent's current position.

[0105] In any of the foregoing illustrative embodiments of the illustrative diagnostic laboratory system, the predetermined distance is a predetermined radius R, which defines the local field of view of each agent as a square of length 2*R+1.

[0106] An illustrative diagnostic laboratory system of any of the foregoing illustrative embodiments, wherein each unit has a unit type, and wherein the local field of view of the agent includes a central unit and a plurality of units surrounding the central unit, the central unit including the agent.

[0107] An illustrative diagnostic laboratory system in any of the foregoing illustrative embodiments, wherein the memory includes computer-executable instructions, when executed by the processor, causing the processor to input one or more masks for the local field of view into the neural network, the one or more masks including at least one of an obstacle mask, a mask of the positions of all agents within the local field of view, a mask of the target position of other agents within the local field of view, and a mask of the target position of the agent.

[0108] An illustrative diagnostic laboratory system in any of the foregoing illustrative embodiments, wherein the memory includes computer-executable instructions, when executed by the processor, causing the processor to: for each agent, input distance information to the target of the agent, and the order of the other agents if the other agents to be processed share the target position with the agent described in the specification (pages 13 / 15, CN 121464409 A).

[0109] An illustrative diagnostic laboratory system of any of the foregoing illustrative embodiments, wherein the memory includes computer-executable instructions that, when executed by the processor, cause the processor to: determine the action of each agent using the MAPF-trained neural network by inputting the next target position information of each agent into the neural network.

[0110] An illustrative diagnostic laboratory system of any of the foregoing illustrative embodiments, wherein the next target position information of each agent includes information about the difference between the agent's current position and the agent's next target position.

[0111] An illustrative diagnostic laboratory system of any of the foregoing illustrative embodiments, wherein the next target position information of each agent includes at least one of the time remaining until the earliest arrival time, the time remaining until the target deadline, and whether the target deadline has passed.

[0112] An illustrative diagnostic laboratory system of any of the foregoing illustrative embodiments, wherein the neural network is trained using at least one of a global positive reward for all agents when any agent completes the target and a global negative reward for all agents when any agent is late for the deadline.

[0113] An illustrative method for pathfinding of sample containers in a diagnostic laboratory system, the method comprising: creating a grid for the diagnostic laboratory system, the grid having cells including diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment; creating a plurality of agents, each agent representing a sample carrier within the diagnostic laboratory system; determining a target for each agent assigned to a sample carrier holding a sample container, the target including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint; creating a neural network that takes the earliest arrival time constraint, the deadline time constraint, and the target processing time constraint of the agent as input, and takes an action or probability of an action of the agent as output; training the neural network using arrival, deadline, and priority constraints and a non-instantaneous target processing time; and creating a pathfinding program that uses the trained neural network to generate actions for each agent within the diagnostic laboratory system.

[0114] The illustrative method of any of the foregoing illustrative embodiments further includes deploying the trained neural network and the pathfinding program for use by the diagnostic laboratory system.

[0115] The illustrative method of any of the foregoing illustrative embodiments further includes employing the pathfinding procedure to determine one or more actions, including transferring the sample container for processing.

[0116] The illustrative method of any of the foregoing illustrative embodiments, wherein the neural network is configured to input local field-of-view information for each agent into the neural network, the local field-of-view information representing the state of neighboring vertices within a predetermined distance of the agent's current position.

[0117] The illustrative method of any of the foregoing illustrative embodiments, wherein the neural network is configured to input one or more masks for the local field of view into the neural network, the one or more masks including at least one of an obstacle mask, a mask of the positions of all agents within the local field of view, a mask of the target position of other agents within the local field of view, and a mask of the target position of the agent.

[0118] An illustrative method in any of the foregoing illustrative embodiments, wherein the neural network is configured to: for each agent, input distance information to the target of the agent, and at least one of the following: if other agents to be processed share the target location with the agent.

[0119] An illustrative method in any of the foregoing illustrative embodiments, wherein the neural network is configured to input the next target location information for each agent.Specification page 14 / 15 19 CN 121464409 A

[0120] An illustrative method in any of the foregoing illustrative embodiments, wherein the next target location information for each agent includes information about the difference between the agent's current location and the agent's next target location.

[0121] An illustrative method in any of the foregoing illustrative embodiments, wherein the next target location information for each agent includes at least one of the time remaining until the agent's earliest arrival time, the time remaining until the target deadline, and whether the target deadline has passed.

[0122] An illustrative method in any of the foregoing illustrative embodiments further includes: training the neural network with at least one of a global positive reward for all agents when any agent completes the target and a global negative reward for all agents when any agent is late for the deadline.

[0123] While this disclosure is susceptible to various modifications and alternatives, specific method and apparatus embodiments have been shown by way of example in the accompanying drawings and have been described in detail herein. However, it should be understood that the specific methods and apparatus disclosed herein are not intended to limit this disclosure. Instruction manual, page 15 / 15, 20 CN 121464409 A, Figure 1; Instruction manual, Figure 1 / 7, page 21 CN 121464409 A, Figure 2A, Figure 2B; Instruction manual, Figure 2 / 7, page 22 CN 121464409 A, Figure 3A; Instruction manual, Figure 3 / 7, page 23 CN 121464409 A, Figure 3B, Figure 3C; Instruction manual, Figure 4 / 7, page 24 CN 121464409 A, Figure 3D, Figure 3E; Instruction manual, Figure 5 / 7, page 25 CN 121464409 A, Figure 4; Instruction manual, Figure 6 / 7, page 26 CN 121464409 A, Figure 5; Instruction manual, Figure 7 / 7, page 27 CN 121464409 A.

Claims

1. A method for pathfinding of sample containers in a diagnostic laboratory system, the method comprising: A batch of sample containers is received in a diagnostic laboratory system, the diagnostic laboratory system having diagnostic laboratory equipment, one or more tracks connected to the diagnostic laboratory equipment, and a plurality of sample carriers configured to transport the sample containers within the diagnostic laboratory system, each of the sample containers containing a sample to be processed. Obtain a grid for the diagnostic laboratory system, the grid having cells including the diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment; The intelligent agent is assigned to each of the sample carriers; For each agent assigned to a sample carrier holding a sample container, a target is determined, including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint. as well as The action of each agent is determined by a neural network trained with Multi-Agent Pathfinding (MAPF), which is trained using arrival, deadline, and priority constraints as well as non-instantaneous target processing time.

2. The method of claim 1, further comprising using a sample carrier controller of the diagnostic laboratory system to perform one or more of the determined actions, including transferring a sample container for processing.

3. The method according to claim 1, wherein, For each agent assigned to a sample carrier holding a sample container, a target is determined, including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint, which are received from the scheduler program.

4. The method according to claim 1, further comprising determining a series of objectives for each agent, each objective including an earliest arrival time constraint, a deadline constraint, and an objective processing time constraint.

5. The method according to claim 1, wherein, Determining the action of each agent using the MAPF-trained neural network involves inputting local field-of-view information of each agent into the neural network, wherein the local field-of-view information represents the state of neighboring vertices within a predetermined distance of the agent's current position.

6. The method according to claim 5, wherein, The predetermined distance is a predetermined radius R, which defines the local field of view of each agent as a square with a length of 2*R+1.

7. The method according to claim 5, wherein, Each unit has a unit type, and wherein the agent's local field of vision includes a central unit and a plurality of units surrounding the central unit, the central unit including the agent.

8. The method according to claim 7, wherein, Determining the action of each agent using the MAPF-trained neural network includes inputting one or more masks for the local field of view into the neural network, the one or more masks including at least one of an obstacle mask, a mask of the positions of all agents within the local field of view, a mask of the target position of other agents within the local field of view, and a mask of the target position of the agent.

9. The method according to claim 1, wherein, Determining the action of each agent using the MAPF-trained neural network includes inputting at least one of the following: distance information to the target of the agent, and the order of the other agents if they share the target location with the agent.

10. The method according to claim 1, wherein, Determining the action of each agent using the MAPF-trained neural network involves inputting the next target position information of each agent into the neural network.

11. The method according to claim 10, wherein, The next target location information for each agent includes information about the difference between the agent's current location and the agent's next target location.

12. The method according to claim 10, wherein, The next target location information for each agent includes at least one of the following: the time remaining until the agent's earliest arrival time, the time remaining until the target deadline, and whether the target deadline has passed.

13. The method according to claim 1, wherein, The neural network is trained using at least one of the following: a global positive reward for all agents when any agent completes the objective and a global negative reward for all agents when any agent is late for the deadline.

14. A diagnostic laboratory system, the diagnostic laboratory system comprising: Diagnostic laboratory equipment; One or more tracks connecting the diagnostic laboratory equipment; processor; as well as A memory coupled to the processor, the memory including computer-executable instructions, which, when executed by the processor, cause the processor to: Obtain a grid for the diagnostic laboratory system, the grid having cells including the diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment; Assign an agent to each sample carrier within the diagnostic laboratory system; For each agent assigned to a sample carrier holding a sample container, a target is determined, including an earliest arrival time constraint, a deadline time constraint, and a target processing time constraint. as well as The action of each agent is determined by a neural network trained with Multi-Agent Pathfinding (MAPF), which is trained using arrival, deadline, and priority constraints as well as non-instantaneous target processing time.

15. The diagnostic laboratory system of claim 14, further comprising a sample carrier controller, and wherein, The memory includes computer-executable instructions that, when executed by the processor, cause the processor to employ the sample carrier controller to perform one or more defined actions, including transferring a sample container for processing.

16. The diagnostic laboratory system of claim 14, wherein, The memory includes computer-executable instructions that, when executed by the processor, cause the processor to determine a set of objectives for each agent, each objective including an earliest arrival time constraint, a deadline constraint, and an objective processing time constraint.

17. The diagnostic laboratory system of claim 14, wherein, The memory includes computer-executable instructions that, when executed by the processor, cause the processor to input local field-of-view information for each agent into the neural network, the local field-of-view information representing the state of neighboring vertices within a predetermined distance of the agent's current position.

18. The diagnostic laboratory system of claim 17, wherein, The predetermined distance is a predetermined radius R, which defines the local field of view of each agent as a square with a length of 2*R+1.

19. The diagnostic laboratory system of claim 17, wherein, Each unit has a unit type, and wherein the agent's local field of vision includes a central unit and a plurality of units surrounding the central unit, the central unit including the agent.

20. The diagnostic laboratory system of claim 19, wherein, The memory includes computer-executable instructions that, when executed by the processor, cause the processor to input one or more masks for the local field of view into the neural network, the one or more masks including at least one of an obstacle mask, a mask of the positions of all agents within the local field of view, a mask of the target positions of other agents within the local field of view, and a mask of the target positions of the agents.

21. The diagnostic laboratory system of claim 14, wherein, The memory includes computer-executable instructions that, when executed by the processor, cause the processor to: for each agent, input distance information to a target of the agent, and the order of other agents if they share the target location with the agent.

22. The diagnostic laboratory system of claim 14, wherein, The memory includes computer-executable instructions that, when executed by the processor, cause the processor to: determine the action of each agent using the MAPF-trained neural network by inputting the next target position information of each agent into the neural network.

23. The diagnostic laboratory system according to claim 22, wherein, The next target location information for each agent includes information about the difference between the agent's current location and the agent's next target location.

24. The diagnostic laboratory system of claim 22, wherein, The next target location information for each agent includes at least one of the following: the time remaining until the earliest arrival time, the time remaining until the target deadline, and whether the target deadline has passed.

25. The diagnostic laboratory system according to claim 14, wherein, The neural network is trained using at least one of the following: a global positive reward for all agents when any agent completes the objective and a global negative reward for all agents when any agent is late for the deadline.

26. A method for pathfinding of sample containers in a diagnostic laboratory system, the method comprising: Create a grid for a diagnostic laboratory system, the grid having cells including diagnostic laboratory equipment and one or more tracks connecting the diagnostic laboratory equipment; Create multiple intelligent agents, each representing a sample carrier within the diagnostic laboratory system; For each agent assigned to a sample carrier holding a sample container, a target is determined, the target including an earliest arrival time constraint, a deadline constraint, and a target processing time constraint; Create a neural network that takes the agent's earliest arrival time constraint, deadline constraint, and target processing time constraint as input, and takes the agent's action or the probability of the action as output. The neural network is trained using arrival and deadline times, priority constraints, and non-instantaneous target processing time. as well as A pathfinding program is created, which uses a trained neural network to generate actions for each agent within the diagnostic laboratory system.

27. The method of claim 26, further comprising deploying the trained neural network and the pathfinding procedure for use by the diagnostic laboratory system.

28. The method of claim 27, further comprising employing the pathfinding procedure to determine one or more actions, including transferring the sample container for processing.

29. The method according to claim 26, wherein, The neural network is configured to input local field-of-view information for each agent into the neural network, the local field-of-view information representing the state of neighboring vertices within a predetermined distance of the agent's current position.

30. The method according to claim 29, wherein, The neural network is configured to input one or more masks for the local field of view into the neural network, the one or more masks including at least one of an obstacle mask, a mask of the positions of all agents within the local field of view, a mask of the target positions of other agents within the local field of view, and a mask of the target positions of the agents.

31. The method according to claim 26, wherein, The neural network is configured such that, for each agent, at least one of the following is input: distance information to the target of the agent, and the order of the other agents if they share the target location with the agent.

32. The method according to claim 26, wherein, The neural network is configured to take into account the next target location information for each agent.

33. The method according to claim 32, wherein, The next target location information for each agent includes information about the difference between the agent's current location and the agent's next target location.

34. The method according to claim 32, wherein, The next target location information for each agent includes at least one of the following: the time remaining until the agent's earliest arrival time, the time remaining until the target deadline, and whether the target deadline has passed.

35. The method of claim 26, further comprising: The neural network is trained using at least one of a global positive reward for all agents when any agent completes the objective and a global negative reward for all agents when any agent is late for the deadline.