An intelligent aviation medicine simulation rescue path planning method

By introducing risk-weighted sampling point generation strategy and guidance rewards based on congestion delay index in aeronautical medical simulated rescue path planning, the problems of weak adaptability to dynamic obstacles and complex environments in the prior art and increasing energy consumption and operational complexity are solved, and a more flexible, smooth and safe rescue path planning effect is achieved.

CN119781505BActive Publication Date: 2025-05-27SECOND MEDICAL CENT OF CHINESE PLA GENERAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510279903.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-27
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing aeronautical medical simulation rescue path planning method has weak adaptability to dynamic obstacles and complex environments, resulting in inflexible path selection, poor path planning effect, and neglecting environmental dynamics and complexity, resulting in sharp turn increase energy consumption and operational complexity and increase safety risks.

Method used

A risk-weighted sampling point generation strategy is adopted to construct dynamic adjustment of exploration factors based on environmental complexity indicators, and adaptively increase the exploration ratio to improve the global optimality of the path; at the same time, guide rewards are introduced based on the congestion delay index, and the path is ensured to be smoother and safer by defining the steering cost.

Benefits of technology

It improves the flexibility and effectiveness of rescue path planning, enhances the ability to adapt to dynamic obstacles and complex environments, reduces energy consumption and operational complexity caused by sharp turns, and reduces safety risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119781505B_ABST
    Figure CN119781505B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent aviation medical simulation rescue path planning method, which includes simulating a rescue scenario, training a Q-network, and aviation medical simulation rescue path planning. The present invention belongs to the field of path planning, specifically referring to an intelligent aviation medical simulation rescue path planning method. This solution adopts a sampling point generation strategy introducing risk weighting, constructs a dynamically adjusted exploration factor based on the environmental complexity index, adaptively increases the exploration ratio, improves the global optimality of the path, and further improves the rescue path planning effect; introduces a heuristic reward based on the congestion delay index, adapts to the dynamic and complex environment, and makes the rescue path smoother and safer by defining the turning cost and punishing sharp-turn paths, thereby improving the rescue path planning effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of path planning, and in particular to an intelligent aviation medical simulation rescue path planning method. Background Art

[0002] Aeromedical simulation rescue path planning is a path planning technology for actual aeromedical rescue missions in a simulated environment. It models the environmental factors in the rescue scenario and uses intelligent algorithms to design efficient, smooth, and safe flight paths to ensure that the rescue mission can be completed quickly. However, general aeromedical simulation rescue path planning methods have weak adaptability to dynamic obstacles and complex environments, and cannot dynamically balance exploration and utilization, which leads to inflexible path selection and poor path planning effects. General aeromedical simulation rescue path planning methods ignore the dynamics and complexity of the environment, increase energy consumption and increase operational complexity due to sharp turns, which in turn increases safety risks and poor rescue path planning effects. Summary of the invention

[0003] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides an intelligent aviation medical simulation rescue path planning method. In view of the fact that general aviation medical simulation rescue path planning methods have weak adaptability to dynamic obstacles and complex environments, and cannot dynamically balance exploration and utilization, which leads to inflexible path selection and poor path planning effects, this scheme adopts a risk-weighted sampling point generation strategy, constructs a dynamically adjusted exploration factor based on the environmental complexity index, adaptively increases the exploration ratio, improves the global optimality of the path, and thereby improves the rescue path planning effect; in view of the fact that general aviation medical simulation rescue path planning methods ignore the dynamics and complexity of the environment, increase energy consumption and increase operation complexity due to sharp turns, which leads to increased safety risks and poor rescue path planning effects, this scheme introduces guided rewards based on the congestion delay index, adapts to dynamic and complex environments, and punishes sharp turn paths by defining turning costs, making the rescue path smoother and safer, thereby improving the rescue path planning effect.

[0004] The technical solution adopted by the present invention is as follows: The present invention provides an intelligent aviation medical simulation rescue path planning method, the method comprising the following steps:

[0005] Step S1: simulate a rescue scenario;

[0006] Step S2: training the Q network;

[0007] Step S3: Aeromedical simulation rescue path planning.

[0008] Furthermore, in step S1, the simulated rescue scenario is to determine a three-dimensional rescue environment; the three-dimensional rescue environment includes a rescue starting point, a rescue target point, obstacles, weather conditions and aircraft density; and the environmental data is dynamically updated.

[0009] Furthermore, in step S2, the training of the Q network specifically includes the following steps:

[0010] Step S21: Adaptive sampling; optimizing the sampling area for dynamic obstacles and quickly generating path points; specifically: defining the sampling area, expressed as: ; ; ; Define the dynamic exploration factor, expressed as: ; Introduce the environmental complexity index H, expressed as: ; When the number of consecutive sampling failures reaches the threshold, the ellipse range is dynamically expanded, expressed as: ; ;in, It is the location of the rescue starting point; is the location of the rescue target point; is half of the major axis of the elliptical sampling area; It is half of the minor axis of the ellipse sampling area; C is the center position of the ellipse; is the dynamic adjustment factor at the i-th sampling; and are the minimum and maximum values ​​of the adjustment factor respectively; N is the maximum number of sampling times; i is the sampling number index; , , and is the environmental complexity coefficient; is the number of obstacles; is the area of ​​the region; is the rate of environmental change; is the current path length; It is a path reward; is the completeness of environmental information;

[0011] Step S22: Generate sampling points; introduce risk weights , expressed as: : ; ; ;in, is A random number in a range; is a random sampling position; is the risk factor;

[0012] Step S23: Dynamically adjust the exploration factor; the dynamic adjustment of the exploration factor is expressed as: ; The greedy strategy is expressed as: ; ;in, is the exploration factor for the kth iteration; is the initial exploration factor; is the decay coefficient of the exploration factor; and is the adjustment factor; is the probability of selecting action a in state s; m is the number of all candidate actions; is the optimal action in the current state; is the expected value of the cumulative reward; is the Q value difference between the optimal action and the current action; is the initial selection probability; is the selected attenuation coefficient;

[0013] Step S24: Guide reward calculation; introduce congestion delay index , the reward function The definition is expressed as: ; ; ; ;in, is to reach the target point, It is a corresponding reward; It is a collision obstacle. It is a corresponding reward; is the reward weight coefficient; d is the Euclidean distance from the agent to the target point; is the weather weight; , and They are normal weather, moderately bad weather and severely bad weather. and is the corresponding weight; is the task priority weight; and They are ordinary tasks and high priority tasks, 1 and is the corresponding weight; is the number of aircraft in the current airspace; is the airspace area;

[0014] Step S25: Define the steering cost; define the vector from the parent node to the current node as V1, the vector from the current node to the child node as V2, and define the steering angle as ,and ; Define the turn cost , expressed as: ; Evaluate the total cost of the path , expressed as: ; ; ; ;in, and is the steering weight coefficient; is the path length cost; is the guidance cost; d(·) is the Euclidean distance; N[·] is the node; j is the node position index; is the environmental weight; , , and is the environmental factor; is the position of the current node;

[0015] Step S26: Q value update; according to the reward and path evaluation results, the Q value is dynamically updated to optimize the path planning; if the change of the Q value converges, the single task training is completed; if the cumulative reward when completing the path planning task under the condition of the training task reaching the number of times tends to be stable, the training is completed and the Q network is obtained; the updated Q value is expressed as: ;in, and are the expected values ​​of the cumulative rewards after and before the update, respectively; is the learning rate; is the cost weight; is the discount factor; is in the next state Choose the Q value that has the most action.

[0016] Furthermore, in step S3, the aviation medical simulation rescue path planning is based on the Q network obtained after training. The Q network is deployed in the actual aircraft control system, and the rescue starting point and target point are collected in real time. The dynamic aviation medical rescue path planning is realized through sensors and real-time environmental data updates.

[0017] The beneficial effects achieved by the present invention using the above scheme are as follows:

[0018] (1) In view of the fact that general aviation medical simulation rescue path planning methods have weak adaptability to dynamic obstacles and complex environments, and cannot dynamically balance exploration and utilization, which leads to inflexible path selection and poor path planning effects, this scheme adopts a risk-weighted sampling point generation strategy, constructs a dynamically adjusted exploration factor based on the environmental complexity index, adaptively increases the exploration ratio, improves the global optimality of the path, and thus improves the rescue path planning effect.

[0019] (2) In view of the fact that general aviation medical simulation rescue path planning methods ignore the dynamics and complexity of the environment, increase energy consumption and operational complexity due to sharp turns, and thus increase safety risks and poor rescue path planning effects, this scheme introduces guided rewards based on the congestion delay index to adapt to dynamic and complex environments. By defining turning costs and punishing sharp turn paths, the rescue path is made smoother and safer, thereby improving the rescue path planning effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A schematic diagram of a flow chart of an intelligent aviation medical simulation rescue path planning method provided by the present invention;

[0021] Figure 2 It is a schematic diagram of the process of step S2.

[0022] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0024] In the description of the present invention, it is necessary to understand that terms such as “upper”, “lower”, “front”, “back”, “left”, “right”, “top”, “bottom”, “inside” and “outside” indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the referred system or element must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as a limitation on the present invention.

[0025] Example 1, see Figure 1 The present invention provides an intelligent aviation medical simulation rescue path planning method, which includes the following steps:

[0026] Step S1: simulate a rescue scenario; construct a three-dimensional rescue environment including dynamic updates;

[0027] Step S2: training the Q network; gradually training the Q network to achieve the optimal path planning strategy through adaptive sampling, dynamic adjustment of the exploration factor, guided reward calculation and steering cost optimization;

[0028] Step S3: Aeromedical simulation rescue path planning: deploy the trained Q network into the aircraft control system, and dynamically plan the optimal rescue path in combination with real-time environmental data.

[0029] Example 2, see Figure 1 This embodiment is based on the above embodiment. In step S1, the simulated rescue scenario is to determine a three-dimensional rescue environment; the three-dimensional rescue environment includes a rescue starting point, a rescue target point, obstacles, weather conditions and aircraft density; and the environmental data is dynamically updated; the rescue starting point is the initial position of the rescue aircraft; the rescue target point is the position where rescue is required; the obstacles are high-rise buildings, terrain, no-fly zones, drones and static obstacles; the weather conditions include normal, moderately severe and severely severe weather; the aircraft density is the distribution of other aircraft in the airspace.

[0030] Example 3, see Figure 1 and Figure 2 This embodiment is based on the above embodiment. In step S2, training the Q network includes the following steps:

[0031] Step S21: Adaptive sampling; optimize the sampling area for dynamic obstacles, avoid risk areas first, quickly generate high potential path points, and reduce blind sampling; specifically: define the sampling area, expressed as: ; ; ; Define the dynamic exploration factor, expressed as: ; Introduce the environmental complexity index H, expressed as: ; When the number of consecutive sampling failures reaches the threshold, the ellipse range is dynamically expanded, expressed as: ; ;in, It is the location of the rescue starting point; is the location of the rescue target point; is half of the major axis of the elliptical sampling area; It is half of the minor axis of the ellipse sampling area; C is the center position of the ellipse; is the dynamic adjustment factor at the i-th sampling; and are the minimum and maximum values ​​of the adjustment factor respectively; N is the maximum number of sampling times; i is the sampling number index; , , and is the environmental complexity coefficient; is the number of obstacles; is the area of ​​the region; is the rate of environmental change; is the current path length; It is a path reward; is the completeness of environmental information;

[0032] Step S22: Generate sampling points; introduce risk weights , expressed as: : ; ; ;in, is A random number in a range; is a random sampling position; is the risk factor.

[0033] By performing the above operations, the general aviation medical simulation rescue path planning method has weak adaptability to dynamic obstacles and complex environments, and cannot dynamically balance exploration and utilization, which leads to inflexible path selection and poor path planning effect. This scheme adopts a risk-weighted sampling point generation strategy, constructs a dynamically adjusted exploration factor based on the environmental complexity index, adaptively increases the exploration ratio, improves the global optimality of the path, and thus improves the rescue path planning effect.

[0034] Example 4, see Figure 1 and Figure 2 This embodiment is based on the above embodiment. In step S2, training the Q network further includes the following steps:

[0035] Step S23: Dynamically adjust the exploration factor; used to dynamically balance exploration and utilization, improve the quality of path planning, and avoid path blindness or local optimal problems caused by fixed exploration factors; dynamic adjustment of the exploration factor is expressed as: ; The greedy strategy is expressed as: ; ;in, is the exploration factor for the kth iteration; is the initial exploration factor; is the decay coefficient of the exploration factor; and is the adjustment factor; is the probability of selecting action a in state s; m is the number of all candidate actions; is the optimal action in the current state; is the expected value of the cumulative reward; is the Q value difference between the optimal action and the current action; is the initial selection probability; is the selected attenuation coefficient;

[0036] Step S24: Guide reward calculation; design continuous reward gradient to guide the agent to approach the target point first, avoid the agent from choosing congested areas, and improve the path quality; used to improve the agent's sensitivity to reward signals, accelerate convergence, dynamically optimize path selection, and adapt to real-time complex environments; introduce congestion delay index , the reward function The definition is expressed as: ; ; ; ;in, is to reach the target point, It is a corresponding reward; It is a collision obstacle. It is a corresponding reward; is the reward weight coefficient; d is the Euclidean distance from the agent to the target point; is the weather weight; , and They are normal weather, moderately bad weather and severely bad weather. and is the corresponding weight; is the task priority weight; and They are ordinary tasks and high priority tasks, 1 and is the corresponding weight; is the number of aircraft in the current airspace; is the airspace area;

[0037] Step S25: Define the steering cost; used to reduce the sharp turn problem in the rescue path, make the path smoother and safer, and meet the operational requirements of the rescue aircraft; specifically: define the vector from the parent node to the current node as V1, the vector from the current node to the child node as V2, and define the steering angle as ,and ; Define the turn cost , expressed as: ; Evaluate the total cost of the path , expressed as: ; ; ; ;in, and is the steering weight coefficient; is the path length cost; is the guidance cost; d(·) is the Euclidean distance; N[·] is the node; j is the node position index; is the environmental weight; , , and is the environmental factor; is the position of the current node;

[0038] Step S26: Q value update; according to the reward and path evaluation results, the Q value is dynamically updated to optimize the path planning; if the change of the Q value converges, the single task training is completed; if the cumulative reward when completing the path planning task under the condition of the training task reaching the number of times tends to be stable, the training is completed and the Q network is obtained; the updated Q value is expressed as: ;in, and are the expected values ​​of the cumulative rewards after and before the update, respectively; is the learning rate; is the cost weight; is the discount factor; is in the next state Choose the Q value that has the most action.

[0039] By performing the above operations, the general aviation medical simulation rescue path planning method ignores the dynamics and complexity of the environment, increases energy consumption and operation complexity due to sharp turns, and thus leads to increased safety risks and poor rescue path planning effects. This scheme introduces guided rewards based on the congestion delay index to adapt to dynamic and complex environments. By defining the turning cost and punishing the sharp turn path, the rescue path is made smoother and safer, thereby improving the rescue path planning effect.

[0040] Example 5, see Figure 1 This embodiment is based on the above embodiment. In step S3, the aviation medical simulation rescue path planning is based on the Q network obtained after training. The Q network is deployed to the actual aircraft control system, and the rescue starting point and target point are collected in real time. The dynamic aviation medical rescue path planning is realized through sensors and real-time environmental data updates.

[0041] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0042] While the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that many changes, modifications, substitutions and variations can be made to the embodiments without departing from the principles and spirit of the invention.

[0043] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.

Claims

1. An intelligent aviation medical simulation rescue path planning method, characterized by: The method comprises the following steps: Step S1: simulate a rescue scenario; construct a three-dimensional rescue environment including dynamic updates; Step S2: training the Q network; gradually training the Q network to achieve the optimal path planning strategy through adaptive sampling, sampling point generation, dynamic adjustment of exploration factor, guided reward calculation, steering cost optimization and Q value update; Step S3: Aeromedical simulation rescue path planning: deploy the trained Q network to the aircraft control system, and dynamically plan the optimal rescue path in combination with real-time environmental data; The adaptive sampling is to optimize the sampling area for dynamic obstacles and quickly generate path points; specifically, the sampling area is defined as: ; ; ; Define the dynamic adjustment factor, expressed as: ; Introduce the environmental complexity index H, expressed as: ; When the number of consecutive sampling failures reaches the threshold, the ellipse range is dynamically expanded, expressed as: ; ;in, It is the location of the rescue starting point; is the location of the rescue target point; is half of the major axis of the elliptical sampling area; It is half of the minor axis of the ellipse sampling area; C is the center position of the ellipse; is the dynamic adjustment factor at the i-th sampling; and are the minimum and maximum values ​​of the adjustment factor respectively; N is the maximum number of sampling times; i is the sampling number index; , , and is the environmental complexity coefficient; is the number of obstacles; is the area of ​​the region; is the rate of environmental change; is the current path length; It is a path reward; It is the completeness of environmental information.

2. The intelligent aviation medical simulation rescue path planning method according to claim 1 is characterized by: In step S2, the training of the Q network specifically includes the following steps: Step S21: adaptive sampling; Step S22: Generate sampling points; introduce risk weights , expressed as: : ; ; ;in, is A random number in a range; is a random sampling position; is the risk factor; Step S23: Dynamically adjust the exploration factor; the exploration factor is expressed as: ; The greedy strategy is expressed as: ; ;in, is the exploration factor for the kth iteration; is the initial exploration factor; is the decay coefficient of the exploration factor; and is the adjustment factor; is the probability of selecting action a in state s; m is the number of all candidate actions; is the optimal action in the current state; is the expected value of the cumulative reward; is the Q value difference between the optimal action and the current action; is the initial selection probability; is the selected attenuation coefficient; Step S24: guide reward calculation; Step S25: define the turning cost; Step S26: Q value update.

3. The intelligent aviation medical simulation rescue path planning method according to claim 2 is characterized by: In step S24, the guided reward calculation is to introduce the congestion delay index , the reward function The definition is expressed as: ; ; ; ;in, is to reach the target point, It is a corresponding reward; It is a collision obstacle. It is a corresponding reward; is the reward weight coefficient; d is the Euclidean distance from the agent to the target point; is the weather weight; , and They are normal weather, moderately bad weather and severely bad weather.

1. and is the corresponding weight; is the task priority weight; and They are ordinary tasks and high priority tasks, 1 and is the corresponding weight; is the number of aircraft in the current airspace; is the airspace area.

4. The intelligent aviation medical simulation rescue path planning method according to claim 3 is characterized by: In step S25, the turning cost is defined as: the vector from the parent node to the current node is defined as V1, the vector from the current node to the child node is defined as V2, and the turning angle is defined as ,and ; Define the turn cost , expressed as: ; Evaluate the total cost of the path , expressed as: ; ; ; ;in, and is the steering weight coefficient; is the path length cost; is the guidance cost; d(·) is the Euclidean distance; N[·] is the node; j is the node position index; is the environmental weight; , , and is the environmental factor; is the position of the current node.

5. The intelligent aviation medical simulation rescue path planning method according to claim 4 is characterized by: In step S26, the Q value update is to dynamically update the Q value and optimize the path planning according to the reward and path evaluation results; if the change of the Q value converges, the single task training is completed; if the cumulative reward when the path planning task is completed under the condition of the training task reaching the number of times tends to be stable, the training is completed and the Q network is obtained; the updated Q value is expressed as: ;in, and are the expected values ​​of the cumulative rewards after and before the update, respectively; is the learning rate; is the cost weight; is the discount factor; is in the next state Select the most actionable Q value; is the total cost of the path.

6. The intelligent aviation medical simulation rescue path planning method according to claim 5 is characterized by: In step S1, the simulated rescue scene is a three-dimensional rescue environment; the three-dimensional rescue environment includes a rescue starting point, a rescue target point, obstacles, weather conditions and aircraft density; and the environmental data is dynamically updated.

7. The intelligent aviation medical simulation rescue path planning method according to claim 6 is characterized by: In step S3, the aviation medical rescue simulation path planning is based on the Q network obtained after training, and the rescue starting point and target point are collected in real time. The dynamic aviation medical rescue path planning is realized through sensor and real-time environmental data update.

Citation Information

Patent Citations

  • Path planning method of cellular access type unmanned aerial vehicle in urban environment and related device

    CN116880549A

  • Systems and methods for adaptive path planning

    US20210103286A1