Air-ground integrated medical rescue command and decision support method

By using the CTDE architecture and multi-agent reinforcement learning, combined with spatiotemporal optimization algorithms, the problems of data heterogeneity and lack of decision-making coordination in existing medical rescue systems have been solved. Autonomous collaborative decision-making and scheduling in dynamic environments have been achieved, improving rescue efficiency and accuracy.

CN121983256APending Publication Date: 2026-05-05CSSC HAISHEN MEDICAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CSSC HAISHEN MEDICAL TECH CO LTD
Filing Date
2025-12-29
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In existing medical rescue systems, the data formats and communication protocols of heterogeneous units such as drones, ambulances, and hospitals vary, lacking the ability to fuse and represent data in real time under a unified spatiotemporal framework. Existing scheduling systems cannot adapt to the highly dynamic, multi-objective, and strongly coupled complex decision-making environment in emergency rescue scenarios, lack effective task-level collaboration mechanisms, and traditional optimization algorithms struggle to handle real-world uncertainties and some observability.

Method used

The CTDE architecture, which combines centralized training and distributed execution, combines multi-agent reinforcement learning and spatiotemporal optimization algorithms. It simulates rescue scenarios in a high-fidelity rescue simulation environment to generate a global rescue situation. It uses a hybrid decision engine to generate strategies for each rescue agent and optimizes the strategies through a composite reward mechanism to achieve autonomous collaborative decision-making by the agents.

Benefits of technology

It enables autonomous collaborative decision-making and scheduling of multiple types of rescue resources in dynamic and uncertain environments. The intelligent agent can spontaneously learn complex collaborative strategies and dynamically adjust strategies to adapt to uncertainty. The centralized evaluator can implicitly consider the overall benefits of the team, thereby improving rescue efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983256A_ABST
    Figure CN121983256A_ABST
Patent Text Reader

Abstract

The invention discloses an air-ground integrated medical rescue command and decision support method and system. The method comprises the following steps: constructing and operating a high-fidelity rescue simulation environment; in the multi-agent decision center, obtaining a global rescue situation state and local observation of each rescue agent; based on a global rescue situation state and local observation of each rescue agent, a strategy is generated for each type of rescue agent through a trained hybrid decision engine, a combined action instruction is output, and the combined action instruction is issued to the corresponding rescue agent in the high-fidelity rescue simulation environment to be executed. Driving an environment state to transfer, and obtaining a composite reward including a team reward and an individual reward; and carrying out iterative optimization on strategies in the multi-agent decision center based on environmental state transition and composite rewards. The strategy can be dynamically adjusted, the overall benefit of the team is considered, and the final strategy tends to be globally better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the intersection of intelligent emergency response and smart healthcare technologies, specifically to an integrated air-ground medical rescue collaborative decision-making system based on multi-agent reinforcement learning (MARL) and spatiotemporal optimization algorithms, and the implementation of its core algorithms. Background Technology

[0002] The existing medical rescue system suffers from the following major technical bottlenecks: the data formats and communication protocols of heterogeneous units such as drones, ambulances, and hospitals vary, lacking the ability to integrate and represent data in real time within a unified spatiotemporal framework; existing dispatch systems are mostly based on rules or simple heuristic algorithms, which cannot adapt to the highly dynamic, multi-objective, and strongly coupled complex decision-making environment in emergency rescue scenarios; there is a lack of effective task-level collaboration mechanisms between air and ground rescue units, making it difficult to achieve efficient collaboration modes such as "drone reconnaissance-ambulance precise transfer" or "drone emergency material delivery-ambulance en route treatment"; and traditional optimization algorithms rely on precise mathematical models and the assumption of complete information, making it difficult to handle uncertainties and partial observability in reality. Summary of the Invention

[0003] The main technical problems to be solved by this invention include: how to overcome the main technical problems existing in the above-mentioned medical rescue systems.

[0004] In a first aspect, embodiments of the present invention provide an integrated air-ground medical rescue command and decision support method, based on a centralized training and distributed execution CTDE architecture, implemented through interaction between a multi-agent decision-making center and a high-fidelity rescue simulation environment, comprising the following steps: The high-fidelity rescue simulation environment is constructed and run. The high-fidelity rescue simulation environment simulates rescue scenarios that include casualty events, rescue agents, geospatial conditions, and dynamic uncertainties, and generates a global rescue situation status. In the multi-agent decision-making center, the global rescue situation status and the local observations of each rescue agent are obtained; Based on the global rescue situation and the local observations of each rescue agent, a hybrid decision engine, after training, generates a strategy for each type of rescue agent and outputs joint action instructions. The hybrid decision engine integrates a multi-agent reinforcement learning module and a spatiotemporal optimization algorithm module. The joint action command is sent to the corresponding rescue agent in the high-fidelity rescue simulation environment to drive the environmental state transition and obtain a composite reward including team rewards and individual rewards. Based on the environmental state transition and the composite reward, the strategy in the multi-agent decision-making center is iteratively optimized.

[0005] In a specific embodiment of the present invention, the centralized training, distributed execution CTDE architecture includes: a centralized evaluator network that uses the global rescue situation state and the actions of all rescue agents to estimate the value of joint actions during the training phase; Multiple distributed actuator networks, each corresponding to a different type or individual rescue agent, are used to output action strategies based on their respective local observations; The parameter updates of the actuator network are guided by the policy gradients provided by the evaluator network.

[0006] In a specific embodiment of the present invention, the step of generating a strategy for each type of rescue agent based on the global rescue situation and the local observations of each rescue agent, and outputting joint action instructions through a trained hybrid decision engine, includes: The strategy output by the multi-agent reinforcement learning module is used to assign appropriate rescue agents to new rescue missions. For the rescue agent that has been assigned a specific rescue mission, the spatiotemporal optimization algorithm module is invoked to calculate the optimized path from its current location to the mission target point based on the instantaneous state of the current environment. The key node information of the optimized path is used as prior knowledge and input into the policy network of the corresponding rescue agent to generate the final executable action instructions.

[0007] In a specific embodiment of the present invention, the composite reward includes: Global team rewards: When a rescue mission is successfully completed, all participating rescue agents receive a positive reward; when a mission fails, they receive a negative reward. Individual efficiency rewards are given or penalized based on the rescue agent's mobility efficiency or task execution progress. Collaborative rewards: When two or more rescue agents complete a preset collaborative action pattern, an additional positive reward is given. Constraints and penalties are imposed on actions that violate preset operating rules or physical constraints.

[0008] In a specific embodiment of the present invention, the actuator network and / or the evaluator network are deep neural networks. The input layer of the deep neural network is designed as a module for encoding the global rescue situation and / or the local observations. This encoding module fuses multi-source heterogeneous data into a unified feature vector.

[0009] Secondly, embodiments of the present invention provide an integrated air-ground intelligent medical rescue collaborative decision-making system, comprising: The high-fidelity rescue simulation environment module is used to simulate and generate dynamic rescue scenarios, physical interactions of rescue intelligent agents, and global states. The multi-agent decision-making central module includes a hybrid decision engine based on the CTDE architecture, which is used to receive state and observation, calculate and output joint action instructions; The training and optimization module is used to manage the interaction process between the multi-agent decision-making center module and the high-fidelity rescue simulation environment module, collect experience data, and update the parameters of the hybrid decision engine.

[0010] In a specific embodiment of the present invention, the high-fidelity rescue simulation environment module includes: The scene generation unit is used to configure or randomly generate the location, type, and number of casualty events; The physics simulation unit is used to simulate the kinematics, dynamics, and environmental interaction of rescue intelligent agents; Uncertainty injection unit, used to introduce communication delays, equipment failures or environmental abrupt changes into the simulation; The evaluation unit is used to calculate key performance indicators such as rescue response time, success rate, and resource utilization.

[0011] In a specific embodiment of the present invention, the hybrid decision engine includes a multi-agent reinforcement learning module whose algorithm is based on the multi-agent proximal policy optimization (MAPPO) algorithm or the multi-agent deep deterministic policy gradient (MADDPG) algorithm.

[0012] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the air-ground integrated intelligent medical rescue collaborative decision-making method.

[0013] Fourthly, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the air-ground integrated intelligent medical rescue collaborative decision-making method described above.

[0014] Compared with existing technologies, it has the following outstanding advantages: Compared with existing technologies, this invention enables intelligent agents to spontaneously learn complex collaborative strategies, such as relaying, containment, and division of labor, through training, without requiring hard-coded rules. The RL-based decision model dynamically adjusts its strategies in the face of uncertainties injected into the simulation environment (such as unit destruction or sudden new tasks), exhibiting robustness. The centralized Critic mechanism allows agents to implicitly consider the overall team benefit while pursuing individual rewards, ultimately leading to a globally optimal strategy. The designed CTDE framework, hybrid decision engine, and reward mechanism lay the foundation for integrating more types of intelligent agents (such as medical robots and logistics vehicles). Attached Figure Description

[0015] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating an integrated air-ground medical rescue command and decision support method provided in an embodiment of the present invention; Figure 2 A schematic diagram of the framework of an integrated air-ground medical rescue command and decision support system provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the computer hardware of the present invention. Detailed Implementation

[0016] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0017] It should also be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0018] It should also be understood that, in various embodiments of the present invention, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0019] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0020] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0021] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0022] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0023] To make the above-mentioned features and effects of the present invention clearer and easier to understand, specific embodiments are described below in conjunction with the accompanying drawings. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are merely illustrative. The scope of protection of the present invention is not limited to the disclosed embodiments, but is defined by the appended claims.

[0024] The following are system embodiments corresponding to the above method embodiments. This embodiment can be implemented in conjunction with the above embodiments. The relevant technical details mentioned in the above embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.

[0025] The technological bottlenecks in existing medical rescue systems are the root cause of their inefficiency, slow response, and difficulty in coordination. These include: data flows from different rescue units are isolated, heterogeneous, and asynchronous, making real-time fusion and sharing impossible on a unified "battle map." This results in a fragmented, delayed, and incomplete understanding of the overall situation from the command center. For example, drones might transmit H.264-encoded real-time video streams and metadata with GPS coordinates; ambulances might transmit vital sign data (such as ECG and blood oxygen saturation) from onboard monitors via 4G / 5G, following HL7 or proprietary protocols; and hospital information systems might manage bed, operating room, and blood bank inventory information based on SQL databases. The lack of a unified data model and semantic standards among these three systems is akin to communicating using different languages ​​and units of measurement, making direct integration extremely difficult. Furthermore, each system employs different communication protocols (such as RTSP, MQTT, and HTTP API) and security authentication mechanisms. Establishing a temporary data channel requires complex adaptation and integration development, making real-time connectivity impossible in the time-sensitive rescue scenarios. Drones use GPS timestamps and WGS-84 coordinates, while ambulance locations may be derived from cellular network positioning. The hospital's information system time may also deviate from the network time protocol by milliseconds. Without a unified spatiotemporal framework, precise correlation and trajectory extrapolation are virtually impossible, making it impossible to accurately predict convergence times. Commanders are like "blind men touching an elephant," needing to switch between multiple isolated screens and piece together information from experience. Decisions based on incomplete and inconsistent images are highly prone to misjudgment. For example, an ambulance that is stuck in traffic or has low battery might be dispatched to the scene.

[0026] The core decision-making of existing systems relies on a pre-defined, linear "if-then" rule base or simple algorithm, which cannot handle the multi-objective, multi-constraint, tightly coupled, and continuously changing "non-linear" optimization problems in rescue scenarios. A typical rule is "dispatch the nearest available ambulance." This ignores the differences in the severity of injuries (a severely injured patient with internal bleeding may need an ambulance further away but equipped with specialists and blood), dynamic road congestion, and the real-time load on hospital capacity. The rules are rigid and cannot handle multi-factor trade-offs. Some systems may employ simple heuristics such as greedy algorithms. For example, always prioritizing the most severely injured. This may be effective in a single event, but in multi-point concurrent disaster scenarios, resources may be occupied for a long time by a few severely injured patients, causing more moderately and lightly injured patients to deteriorate due to excessive waiting time, ultimately reducing the overall survival rate. It lacks long-term planning and a global perspective. Decision variables are highly coupled. For example, choosing which hospital to take for injured patient A (decision X) directly affects whether that hospital has the resources to take in subsequent injured patient B (decision Y). This sequential decision-making and resource competition relationship cannot be properly resolved by static rules or single-step optimization algorithms. The system often makes locally optimal but globally suboptimal, or even short-sighted decisions. The response strategy is rigid and unable to adapt to sudden changes. When facing complex, large-scale rescue events, the scheduling efficiency drops sharply, and resource utilization suffers from severe uneven distribution and mismatch.

[0027] There is only basic status reporting between rescue units, lacking direct driving, feedback, and closed-loop coordination mechanisms at the mission level, thus failing to create a synergy where "1+1>2". The ideal "drone-led reconnaissance - ambulance precise transfer" process typically involves: the drone arriving at the scene, the operator reporting the location to the command center via voice or text; the command center then notifies the ambulance driver of a general address via radio. This information transmission chain is long and inefficient, and the drone cannot directly and in real-time push its high-definition video footage or precise coordinates to the ambulance's navigation terminal, forcing the ambulance to blindly search for the last few hundred meters. Units are "passively responding to instructions" rather than "actively seeking cooperation." For example, if an ambulance discovers severe traffic congestion en route, it cannot proactively publish this information as a "cooperative service" for nearby drones or other ambulances planning their routes. Similarly, if a drone discovers a wounded person needs specific medication, it cannot directly initiate a "supply delivery request" to the nearest ambulance carrying that medication. For tasks requiring precise time and space coordination, such as "relay rescue" (e.g., helicopters transporting the wounded to an assembly point, followed by ambulance transfers to hospitals), existing systems rely on manual estimation and frequent communication for coordination. This makes them highly susceptible to failure or delays in any link (e.g., traffic, weather) that could cause the entire chain to fail or stall, preventing seamless coordination. Rescue operations are a mechanical aggregation of multiple individual tasks, rather than a smooth collaborative process. Efficiency is significantly reduced at the "seams" between units, response delays increase, and secondary risks may arise from coordination errors.

[0028] Traditional operations research optimization algorithms (such as linear programming and dynamic programming) rely on a well-defined, information-complete, and highly deterministic mathematical model. However, real-world rescue environments are filled with noise, missing information, and uncertainty, causing the algorithm's premises to fail and the solutions to deviate from reality. The idealized assumption of complete information: The algorithm assumes that the dispatch center knows all the exact variables: the precise travel time for each road segment, the specific injury condition of each injured person, and the real-time availability of beds in each hospital. However, in reality, information is partially observable, delayed, and prone to error: traffic conditions change rapidly, initial assessments of injuries may be inaccurate, and updates to hospital bed availability are delayed. Uncertainties that are difficult to model: Sudden weather changes forcing drones to change routes, ambulance malfunctions, temporary communication interruptions, and deterioration of a patient's condition during transport… these random events are difficult to incorporate into traditional deterministic optimization models. Even using stochastic programming or robust optimization, the complexity explodes with the dimension of uncertainty, making real-time solutions difficult. Slow adaptation to dynamics: Traditional models are typically "one-off optimizations." When new events occur (such as additional casualties) or existing information is updated (such as a road becoming accessible), it is often necessary to solve the entire model from scratch, which is computationally time-consuming and cannot meet the needs of critical decision-making at the minute or even second level. Optimization algorithms that perform well in the laboratory or simulation often become infeasible or significantly less effective once deployed in the real world because the preconditions are not met. The system appears "fragile," only able to work under ideal conditions and unable to cope with the chaos and uncertainty of the real world.

[0029] In summary, existing systems are a collection of data silos, rigid rules, weak collaborative interfaces, and fragile models. They are better suited for routine pre-hospital emergency care with fixed procedures and transparent information, but cannot cope with large-scale, sudden, and highly dynamic complex disaster relief scenarios. The multi-agent collaborative decision-making system proposed in this invention aims to fundamentally break through these bottlenecks and build a new generation of rescue system characterized by data fusion, intelligent decision-making, proactive collaboration, and robustness against disturbances.

[0030] This invention aims to solve the problem of autonomous collaborative decision-making and scheduling of multiple types of rescue resources (agents) under dynamic and uncertain environments. The core objective is to verify the effectiveness of a hybrid framework combining multi-agent reinforcement learning and spatiotemporal optimization algorithms in simulated rescue environments. This invention proposes a multi-agent decision-making hub under a "Centralized Training, Distributed Execution" (CTDE) architecture, coupled with a high-fidelity rescue simulation environment for technology verification and core algorithm development. The core components are detailed below: The high-fidelity rescue simulation environment (technology verification platform) can replace the real world, providing a controllable, repeatable, and quantifiable scenario for algorithm training and testing. This includes a scenario generator: randomly or according to a preset script to generate casualty events of different locations, injuries, and numbers; a physics engine: used to simulate the kinematics and dynamic constraints of drones / ambulances, as well as the influence of airspace and traffic networks; an uncertainty module: used to introduce random factors, such as communication delays, temporary vehicle malfunctions, and worsening of casualty conditions; and an evaluator: used to calculate key performance indicators (KPIs), such as weighted average response time, rescue success rate, and resource utilization rate.

[0031] The multi-agent decision-making hub (core algorithm) can use MAPPO (Multi-Agent Proximal Policy Optimization) or MADDPG (Multi-Agent Deep Deterministic Policy Gradient) as its basic framework.

[0032] Specifically, each drone and each ambulance is considered an intelligent agent. The observation space includes each agent's local observations, encompassing its own state (position, battery / fuel level, mission status), information about other units within its communication range, and local information about the assigned casualty / target hospital. The action space includes discrete action sets or continuous action spaces. For example: {go to reconnaissance point A, go to rendezvous point B, return to recharge, hover and wait}, or continuous action vectors representing the target's heading and speed. The reward function includes a global team reward: successfully rescuing a casualty earns a large positive reward to all participating agents; failure to rescue a casualty within the allotted time results in a negative reward. Individual efficiency rewards: agents receive small rewards for efficient movement (reducing distance to the target per unit time), and small penalties for ineffective movement or inactivity. Collaborative rewards: specially designed rewards encourage collaborative behavior. For example, when a drone scouts and accurately guides an ambulance to the casualty's location, both parties receive additional collaborative rewards. Constraint penalties: violations of constraints (such as entering a no-fly zone or speeding) are penalized.

[0033] Specifically, the network structure includes an Actor network (distributed execution): each agent has an independent policy network that outputs actions based on local observations. A Centralized Critic network (centralized training): during the training phase, the Critic network can access the global state S (a panoramic view provided by the simulation environment) and the actions of all agents to more accurately evaluate the value Q(S, A) of the joint action A, thereby guiding the updates of each Actor network.

[0034] In this embodiment, to address the low exploration efficiency of pure reinforcement learning in the early stages of training and its difficulty in strictly satisfying complex spatiotemporal constraints, a hybrid architecture of "RL-guided optimization" is adopted. The trained MARL policy quickly assigns suitable drones and ambulances (agents) to each newly emerging rescue task. For an agent assigned a specific task, before its action is executed, a lightweight spatiotemporal optimization algorithm (such as a variant considering real-time constraints) calculates the optimal feasible path from its current location to the task objective point based on the specific environment at the current instant, and inputs the critical waypoints of this path as "suggested actions" into the agent's Actor network for final decision-making. This effectively provides prior knowledge for RL, accelerating training and ensuring basic feasibility.

[0035] Specifically, the CTDE architecture's workflow in two key phases is as follows: During centralized training, all agents are taught to collaborate. During training, the system has a centralized "brain," the Critic network. This brain can see the global state S (e.g., the location of all wounded, the location and status of all drones and ambulances, and the overall traffic situation). The Critic evaluates the long-term benefits of each agent's action combination (aA, aB) to the team as a whole. Feedback generated from this "global perspective" simultaneously guides and optimizes each agent's own policy network. During training, agents understand how their actions affect teammates and the global objective, thus learning highly collaborative strategies. This avoids "going it alone" and sabotage.

[0036] In distributed execution, the centralized "God brain" is removed during execution. Each agent relies solely on its local sensors and communication to obtain local observations oA / oB (e.g., its own portion of the map, nearby teammates, and assigned tasks). Using a pre-trained individual policy network, it makes decisions independently and in parallel based on local information, generating actions aA' / aB'. No single point of failure: It does not depend on a potentially faulty central command node. Low communication requirements: There is no need to continuously transmit large amounts of data back to the center during execution. Fast response: Local decision-making and real-time response. High scalability: Adding new agents only requires deploying their individual networks.

[0037] In this embodiment, a "global perspective" is needed during training to learn collaboration: How can drones and ambulances learn to "relay rescue"? Only by allowing the algorithm to see the whole picture during training can it understand the global value of the action sequence of "the drone first scouting and guiding." During execution, "local autonomy" is needed to cope with reality: In real rescue operations, communication may be interrupted, and the central server may be overloaded. Drones and ambulances must be able to rely on their own "cerebellum" (the trained policy network) to make reasonable decisions that align with the team's interests, even with only partial information.

[0038] In short, CTDE allows the system to function like a "coaching team with a holistic view" during training, instructing each player. During the game (execution), each player becomes a "star player" capable of independent judgment and seamless teamwork, eliminating the need for real-time adjustments from the coach. This is an ideal architecture for achieving the dual goals of "intelligent collaboration" and "efficient execution."

[0039] like Figure 1 As shown, this application includes the following steps: 110: Simulation Environment Setup and Benchmarking In the embodiments of this application, two-dimensional / three-dimensional meshed or continuous space rescue scenarios are constructed using Python and simulation libraries (such as PyGame, Unity ML-Agents, or OpenAIGym custom environments).

[0040] Specifically, the simulation environment elements include a map: a two-dimensional grid containing: ordinary traversable areas, obstacles (such as buildings and mountains), multiple hospitals with different capacities and locations, and roads (optional, with different travel speeds). Dynamic entities include wounded soldiers, who appear randomly or according to a script at various locations on the map. Each wounded soldier has a unique ID and an initial injury level (e.g., minor injury, serious injury), and the injury may worsen over time. There are multiple ambulances, with initial locations that can be set (e.g., parked at a hospital). Each ambulance has capacity, speed, and current status (idle, heading to a wounded soldier, transporting a wounded soldier to a hospital). There are multiple drones, which are faster but cannot transport wounded soldiers; they can only conduct reconnaissance. They can report the location and condition of wounded soldiers in real time and can guide ambulances. Time and movement: discrete time steps. At each time step, all dynamic entities move one grid according to their speed (or continuously update their positions). The decision cycle can be every time step or triggered by an event. Drones and ambulances can only observe information within a certain range.

[0041] The rescue process includes: An injured person is found, but their location and condition are unknown to the system (unless detected by a drone). A drone can be dispatched for reconnaissance; once the injured person is located, their location and condition become known. The command center (dispatcher) allocates an ambulance to the injured person's location based on this information. Upon arrival, the ambulance loads the injured person and then proceeds to a hospital. After arrival at the hospital, the injured person receives treatment, and the ambulance becomes available for use.

[0042] Therefore, the evaluation indicators include: weighted rescue time: the time from the injury to the arrival at the hospital, weighted according to the severity of the injury (serious injuries have a higher weight); mortality rate: the proportion of injured persons who are not rescued within the specified time; resource utilization rate: the percentage of busy time for ambulances and drones; coordination efficiency: such as the success rate of drones guiding ambulances, the connection time of relay rescue, etc.

[0043] In this application, the environment is constructed according to the following steps, creating a custom environment using the OpenAI Gym framework. A rule-based benchmark scheduler (e.g., nearest-first) is implemented. A scheduler based on classical optimization algorithms (e.g., genetic algorithms) is implemented for comparison. An evaluation metric set is defined and computed within the environment. Specifically, the following steps are included: Step 1: Create a custom environment. This involves creating a class that inherits from gym.Env and implementing the following methods: - init: Initializes the map, entities, parameters, etc.

[0044] - reset: Resets the environment state to start a new round.

[0045] - step: Execute a time step, including updating the state of all entities and returning observations, rewards, completion flags, and information.

[0046] - render: Optional, visualizes the current state.

[0047] The state representation (observations) will use the global state as the observation, including the location and status of all wounded, ambulances, and drones. For simplicity, a multidimensional array or dictionary will be used to represent this.

[0048] Action Space: For multi-agent systems, the action space can be discrete: each agent (ambulance, drone) can choose a direction of movement (up, down, left, right, stay) or perform a special action (such as loading or unloading wounded personnel, reconnaissance, etc.) at each time step. However, centralized scheduling is used in the baseline scheduler, so actions may be sets of instructions.

[0049] Step 2: Rule-based baseline scheduler, based on nearest-distance priority.

[0050] Specifically, for each newly identified known casualty, the nearest available ambulance is selected and assigned to the casualty's location. If the ambulance is already transporting a casualty, its destination is the hospital, and this cannot be changed.

[0051] Step 3: Scheduler based on genetic algorithm In a genetic algorithm, a scheduling scheme is encoded as a chromosome, representing the sequence of tasks assigned to each ambulance. The objective function (such as minimizing weighted rescue time) is then optimized through operations like selection, crossover, and mutation. Because the environment is dynamic, the genetic algorithm can be rerun at each decision point for rescheduling.

[0052] Step 4: Evaluation metrics will be calculated at the end of each episode: Weighted rescue time for all injured persons (average); mortality rate (percentage of injured persons exceeding the maximum waiting time); average utilization rate of ambulances and drones (busy time steps / total time steps).

[0053] Since this environment will be used for subsequent MARL training, it is also necessary to design observation and action spaces suitable for MARL, as well as a reward function. However, for the baseline scheduler part, we will only focus on scheduling decisions for now, without involving rewards.

[0054] 120: Development and Training of the MARL Core Algorithm; Specifically, in this embodiment, the MAPPO or MADDPG algorithm is implemented using a deep learning framework (such as PyTorch). The agent's observation and action spaces, as well as the composite reward function as described in Part III, are designed and encoded. Large-scale distributed training is performed in a simulation environment, utilizing an experience replay pool and parallel environment sampling to accelerate training. Changes in policy performance during training are saved and compared with benchmark algorithms.

[0055] Specifically, the following steps are included: First, the rescue simulation environment needs to be packaged into a format that meets the training requirements of MAPPO. Each agent (ambulance or drone) has its own observation space and action space. Observation Space: For each agent, observations include its own state: position, speed, fuel / battery, current status (idle, heading to the injured, heading to the hospital, etc.); if a task has been assigned, it includes the location of the task target, the injury level of the injured, etc.; surrounding environmental information, the relative positions and states of other agents, injured persons, hospitals, and other entities within a certain range.

[0056] Action space, for ambulances, might include movements such as moving up, down, left, and right, as well as loading and unloading injured persons. For drones, actions might include movement (in eight directions), reconnaissance, and communication. It can also be designed as a continuous action space, such as the direction and speed of movement.

[0057] The embodiments of this application use a discrete action space.

[0058] Next, PyTorch is used to build the policy network (Actor) and the value network (Critic). Due to the CTDE architecture, each agent has its own Actor network, while the Critic network can access global information during training.

[0059] The Actor network takes the local observations of the agents as input and outputs the probability distribution of actions; the Critic network takes the concatenation of observations and actions of all agents (or the global state) as input and outputs the state-value function.

[0060] The reward function includes: a global team reward: successful rescue of the wounded earns a positive reward to all participating agents; death of the wounded results in a negative reward; individual efficiency rewards: agents receive a small reward for efficient movement (reducing the distance to the target per unit time), and a small penalty for ineffective movement or inactivity; and a collaborative reward: when the drone scouts and accurately guides the ambulance to the wounded's location, both parties receive an additional collaborative reward. Constraint penalties: violations of constraints (such as entering a no-fly zone or speeding) incur penalties.

[0061] Calculate the reward for each agent in the environment and return it at each step (or each round).

[0062] Finally, the MAPPO algorithm is used for training. The main steps are as follows: Initialize the environment and acquire initial observations. Each agent selects an action based on the current observation through the Actor network. The environment executes the actions of all agents, obtaining the next observation and reward. The experience (observation, action, reward, next observation, termination) is stored in the experience replay pool. When the experience replay pool is large enough, a batch of data is sampled from it to update the Actor and Critic networks. The above steps are repeated until the preset number of training steps is reached. To accelerate training, multiple environment instances can be used to collect experience in parallel.

[0063] Use PyTorch's distributed training capabilities, or use frameworks like Ray RLlib to simplify distributed training. Ray RLlib is chosen here because it includes the MAPPO algorithm and can be easily extended to multiple environments.

[0064] During training, the model and training metrics (such as cumulative reward, success rate, etc.) are saved periodically. Simultaneously, benchmark algorithms (such as nearest-neighbor priority rules and genetic algorithms) are run in the same test scenario to compare performance.

[0065] 130: Hybrid Architecture Integration and Optimization; Specifically, a spatiotemporal optimization algorithm is integrated with the MARL policy, and an interactive interface is designed. Comparative experiments are conducted to compare the performance differences of pure MARL policy, pure optimization algorithm, and hybrid policy in the same test scenario. The advantages and disadvantages of different policies are analyzed, such as computational latency, optimality of the solution, and anti-interference ability.

[0066] Specifically, the upper layer uses MARL (Multi-Agent Reinforcement Learning) for high-level task allocation and decision-making, while the lower layer uses spatiotemporal optimization algorithms (such as A*, Dijkstra's algorithm, genetic algorithms, etc.) for specific path planning. The purpose of this design is to combine the advantages of both: MARL can handle complex multi-agent cooperation and dynamic environments, while traditional optimization algorithms can guarantee the optimality (or near-optimality) of the path under a given task and satisfy various constraints.

[0067] When the MARL policy selects a target for an agent (e.g., to go to a wounded person's location or a hospital), it passes this target and the current environmental state (including the map, obstacles, the positions of other agents, etc.) to the spatiotemporal optimization algorithm.

[0068] The spatiotemporal optimization algorithm calculates the optimal path from the agent's current position to the target position based on the current state (considering time, distance, constraints, etc.) and returns it to the MARL policy.

[0069] The MARL policy uses this path as a suggestion, and the agent executes actions along this path (e.g., moving along the path). Simultaneously, the MARL policy can also reallocate tasks based on environmental changes (such as new obstacles or agent malfunctions) and invoke the spatiotemporal optimization algorithm again to replan the path.

[0070] Compare the three strategies in the same test scenario: Pure MARL policy: The agent's actions are entirely output by the MARL policy, including movement direction, speed, etc. The MARL policy needs to learn how to avoid obstacles and plan paths on its own.

[0071] Pure optimization algorithms: These use classic optimization algorithms (such as genetic algorithms) for task allocation and path planning. Note that pure optimization algorithms may not be able to handle highly dynamic environments, so we assume that they replan at each time step, but task allocation is fixed or follows some rule.

[0072] Hybrid strategy: The upper layer MARL performs task allocation, and the lower layer optimization algorithm performs path planning.

[0073] Evaluation metrics: Computational latency: the time from receiving the environmental state to outputting an action. Solution optimality: measured by cumulative reward, task completion time, path length, etc. Integrity: the policy's robustness to random events introduced into the environment (such as agent failure or new task insertion).

[0074] The implementation steps are described in detail below.

[0075] Step 1: Implement a pure MARL policy; using the previously designed MAPPO algorithm, enable the agent to learn to move, rescue, and coordinate in the environment. The agent's observations include environmental information, the states of other agents, and the states of the wounded, etc., and the action space consists of continuous or discrete movement actions.

[0076] Step 2: Implement a pure optimization algorithm; select a benchmark optimization algorithm, such as task allocation and path planning based on a genetic algorithm. The genetic algorithm encodes task allocation and path planning as chromosomes, optimizing them through selection, crossover, and mutation. The fitness function can be the total task completion time, total path length, etc.

[0077] Step 3: Implement a hybrid strategy; combine MAPPO's task allocation (target selection) with a fast path planning algorithm (such as the A algorithm). MAPPO outputs the target point for each agent, then the A algorithm calculates the path to the target point, and the agent moves according to the path.

[0078] Step 4: Comparative Experiments; Design multiple test scenarios, from simple to complex, for example: Scenario 1: A small number of wounded soldiers, a small number of intelligent agents, and a static environment.

[0079] Scenario 2: A large number of wounded, insufficient number of intelligent agents, static environment.

[0080] Scenario 3: Dynamic environment with moving obstacles, and the location of the wounded appears randomly.

[0081] Three strategies are run in each scenario, and evaluation metrics are recorded.

[0082] Step 5: Analyze the advantages and disadvantages; based on the experimental results, analyze the advantages and disadvantages of the three strategies in terms of computational delay, optimality, and anti-interference ability.

[0083] 140: Verification and Analysis; Specifically, on a reserved set of test scenarios, statistical evaluation metrics are used to demonstrate the superiority of the hybrid MARL framework over benchmark methods. The decision-making trajectories of agents in typical rescue cases are visualized to observe and analyze the rationality and intelligence of their collaborative behavior. A technical verification report is written, including the core algorithm code, experimental data, analytical conclusions, and suggestions for future physical system integration.

[0084] Furthermore, the MARL reward function design for air-to-ground collaborative rescue utilizes an innovative composite reward structure to effectively balance individual efficiency and team collaboration, guiding agents to learn highly practical strategies. A hybrid decision-making paradigm combining RL and classical optimization algorithms is employed: lower-level optimization ensures the instantaneous feasibility of actions, while upper-level RL handles high-level task allocation and sequential decision-making, combining flexibility and reliability. A global state representation method for the rescue situation in the CTDE architecture can transform multi-source heterogeneous rescue situation information into feature vectors that can be processed by a centralized Critic network through a compact and efficient encoding method.

[0085] like Figure 1 As shown, this embodiment of the invention provides an integrated air-ground medical rescue command and decision support method, based on centralized training and distributed execution CTDE architecture, which is realized by the interaction between a multi-agent decision center and a high-fidelity rescue simulation environment, including the following steps: 210: Construct and run the high-fidelity rescue simulation environment, which simulates rescue scenarios including casualty events, rescue agents, geospatial conditions and dynamic uncertainties, and generates a global rescue situation status; 220: In the multi-agent decision-making center, acquire the global rescue situation status and the local observations of each rescue agent; 230: Based on the global rescue situation and the local observations of each of the rescue agents, a strategy is generated for each type of rescue agent through a trained hybrid decision engine, and joint action instructions are output. The hybrid decision engine integrates a multi-agent reinforcement learning module and a spatiotemporal optimization algorithm module. 240: The joint action command is sent to the corresponding rescue agent in the high-fidelity rescue simulation environment to drive the environmental state transition and obtain a composite reward including team rewards and individual rewards; 250: Based on the environmental state transition and the composite reward, the strategy in the multi-agent decision center is iteratively optimized.

[0086] In a specific embodiment of the present invention, the centralized training, distributed execution CTDE architecture includes: a centralized evaluator network that uses the global rescue situation state and the actions of all rescue agents to estimate the value of joint actions during the training phase; Multiple distributed actuator networks, each corresponding to a different type or individual rescue agent, are used to output action strategies based on their respective local observations; The parameter updates of the actuator network are guided by the policy gradients provided by the evaluator network.

[0087] In a specific embodiment of the present invention, the step of generating a strategy for each type of rescue agent based on the global rescue situation and the local observations of each rescue agent, and outputting joint action instructions through a trained hybrid decision engine, includes: The strategy output by the multi-agent reinforcement learning module is used to assign appropriate rescue agents to new rescue missions. For the rescue agent that has been assigned a specific rescue mission, the spatiotemporal optimization algorithm module is invoked to calculate the optimized path from its current location to the mission target point based on the instantaneous state of the current environment. The key node information of the optimized path is used as prior knowledge and input into the policy network of the corresponding rescue agent to generate the final executable action instructions.

[0088] In a specific embodiment of the present invention, the composite reward includes: Global team rewards: When a rescue mission is successfully completed, all participating rescue agents receive a positive reward; when a mission fails, they receive a negative reward. Individual efficiency rewards are given or penalized based on the rescue agent's mobility efficiency or task execution progress. Collaborative rewards: When two or more rescue agents complete a preset collaborative action pattern, an additional positive reward is given. Constraints and penalties are imposed on actions that violate preset operating rules or physical constraints.

[0089] In a specific embodiment of the present invention, the actuator network and / or the evaluator network are deep neural networks. The input layer of the deep neural network is designed as a module for encoding the global rescue situation and / or the local observations. This encoding module fuses multi-source heterogeneous data into a unified feature vector.

[0090] like Figure 2 As shown, this embodiment of the invention provides an integrated air-ground intelligent medical rescue collaborative decision-making system, comprising: The high-fidelity rescue simulation environment module is used to simulate and generate dynamic rescue scenarios, physical interactions of rescue intelligent agents, and global states. The multi-agent decision-making central module includes a hybrid decision engine based on the CTDE architecture, which is used to receive state and observation, calculate and output joint action instructions; The training and optimization module is used to manage the interaction process between the multi-agent decision-making center module and the high-fidelity rescue simulation environment module, collect experience data, and update the parameters of the hybrid decision engine.

[0091] In a specific embodiment of the present invention, the high-fidelity rescue simulation environment module includes: The scene generation unit is used to configure or randomly generate the location, type, and number of casualty events; The physics simulation unit is used to simulate the kinematics, dynamics, and environmental interaction of rescue intelligent agents; Uncertainty injection unit, used to introduce communication delays, equipment failures or environmental abrupt changes into the simulation; The evaluation unit is used to calculate key performance indicators such as rescue response time, success rate, and resource utilization.

[0092] In a specific embodiment of the present invention, the hybrid decision engine includes a multi-agent reinforcement learning module whose algorithm is based on the multi-agent proximal policy optimization (MAPPO) algorithm or the multi-agent deep deterministic policy gradient (MADDPG) algorithm.

[0093] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a multimodal image fusion method based on multi-level alignment and task awareness.

[0094] This invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the air-ground integrated intelligent medical rescue collaborative decision-making method as described above.

[0095] In addition, combined Figure 1 The air-ground integrated intelligent medical rescue collaborative decision-making method described in this embodiment of the invention can be implemented by electronic devices, such as computer devices.

[0096] Figure 3 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention.

[0097] In some embodiments, the computer device may further include a communication interface 83 and a bus 80. For example, Figure 3 As shown, the processor 81, memory 82, and communication interface 83 are connected through bus 80 and complete communication with each other.

[0098] Specifically, the processor 81 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of the present invention.

[0099] The memory 82 can be used to store or cache various data files that need to be processed and / or communicated, as well as possible computer program instructions executed by the processor 81.

[0100] The processor 81 reads and executes computer program instructions stored in the memory 82 to implement any of the air-ground integrated medical rescue command and decision support methods in the above embodiments.

[0101] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0102] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for integrated air-ground medical rescue command and decision support, characterized in that, Based on centralized training and distributed execution, the CTDE architecture, which is implemented through the interaction between a multi-agent decision-making center and a high-fidelity rescue simulation environment, includes the following steps: The high-fidelity rescue simulation environment is constructed and run. The high-fidelity rescue simulation environment simulates rescue scenarios that include casualty events, rescue agents, geospatial conditions, and dynamic uncertainties, and generates a global rescue situation status. In the multi-agent decision-making center, the global rescue situation status and the local observations of each rescue agent are obtained; Based on the global rescue situation and the local observations of each rescue agent, a hybrid decision engine, after training, generates a strategy for each type of rescue agent and outputs joint action instructions. The hybrid decision engine integrates a multi-agent reinforcement learning module and a spatiotemporal optimization algorithm module. The joint action command is sent to the corresponding rescue agent in the high-fidelity rescue simulation environment to drive the environmental state transition and obtain a composite reward including team rewards and individual rewards. Based on the environmental state transition and the composite reward, the strategy in the multi-agent decision-making center is iteratively optimized.

2. The method according to claim 1, characterized in that, The centralized training, distributed execution CTDE architecture includes: a centralized evaluator network that uses the global rescue situation state and the actions of all rescue agents to estimate the value of joint actions during the training phase; Multiple distributed actuator networks, each corresponding to a different type or individual rescue agent, are used to output action strategies based on their respective local observations; The parameter updates of the actuator network are guided by the policy gradients provided by the evaluator network.

3. The method according to claim 1 or 2, characterized in that, Based on the global rescue situation and the local observations of each rescue agent, a trained hybrid decision engine generates a strategy for each type of rescue agent and outputs joint action instructions, including: The strategy output by the multi-agent reinforcement learning module is used to assign appropriate rescue agents to new rescue missions. For the rescue agent that has been assigned a specific rescue mission, the spatiotemporal optimization algorithm module is invoked to calculate the optimized path from its current location to the mission target point based on the instantaneous state of the current environment. The key node information of the optimized path is used as prior knowledge and input into the policy network of the corresponding rescue agent to generate the final executable action instructions.

4. The method according to claim 1, characterized in that, The composite reward includes: Global team rewards: When a rescue mission is successfully completed, all participating rescue agents receive a positive reward; when a mission fails, they receive a negative reward. Individual efficiency rewards are given or penalized based on the rescue agent's mobility efficiency or task execution progress. Collaborative rewards: When two or more rescue agents complete a preset collaborative action pattern, an additional positive reward is given. Constraints and penalties are imposed on actions that violate preset operating rules or physical constraints.

5. The method according to claim 2, characterized in that, The actuator network and / or the evaluator network are deep neural networks. The input layer of the deep neural network is designed as a module that encodes the global rescue situation and / or the local observations. This encoding module fuses multi-source heterogeneous data into a unified feature vector.

6. A collaborative decision-making system for air-ground integrated intelligent medical rescue, used to implement the method of any one of claims 1-5, characterized in that, include: The high-fidelity rescue simulation environment module is used to simulate and generate dynamic rescue scenarios, physical interactions of rescue intelligent agents, and global states. The multi-agent decision-making central module includes a hybrid decision engine based on the CTDE architecture, which is used to receive state and observation, calculate and output joint action instructions; The training and optimization module is used to manage the interaction process between the multi-agent decision-making center module and the high-fidelity rescue simulation environment module, collect experience data, and update the parameters of the hybrid decision engine.

7. The system according to claim 6, characterized in that, The high-fidelity rescue simulation environment module includes: The scene generation unit is used to configure or randomly generate the location, type, and number of casualty events; The physics simulation unit is used to simulate the kinematics, dynamics, and environmental interaction of rescue intelligent agents; Uncertainty injection unit, used to introduce communication delays, equipment failures or environmental abrupt changes into the simulation; The evaluation unit is used to calculate key performance indicators such as rescue response time, success rate, and resource utilization.

8. The system according to claim 6, characterized in that, The hybrid decision engine includes a multi-agent reinforcement learning module, whose algorithm is based on the multi-agent proximal policy optimization (MAPPO) algorithm or the multi-agent deep deterministic policy gradient (MADDPG) algorithm.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 5.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.

Citation Information

Cited By

  • Accident rescue method, device and apparatus

    CN122266174A