Hybrid path planning method and system for multi-vehicle path conflict avoidance

By establishing a risk assessment mechanism and reinforcement learning, combined with path conflict penalty items, and dynamically adjusting the action selection method, the problem of difficult balance between efficiency and safety in multi-vehicle path planning is solved, and more efficient and safe path selection is achieved.

CN120668164APending Publication Date: 2025-09-19GUANGDONG UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510701190.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing path planning methods have difficulty in dynamically balancing efficiency and safety in multi-vehicle environments with unmanned vehicles, fail to effectively predict path conflicts, and traditional methods lack real-time conflict penalty items, resulting in irrational path selection.

Method used

By establishing a two-dimensional grid map, combining the risk scoring mechanism with reinforcement learning, introducing the fusion of greedy strategy and DQN strategy, and using conflict penalty terms to optimize path selection, multi-vehicle path conflict avoidance is achieved.

Benefits of technology

It improves the safety of path planning and the coordination of multi-vehicle systems, dynamically adjusts strategies to adapt to complex environments, and improves the robustness and explainability of path selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120668164A_ABST
    Figure CN120668164A_ABST
Patent Text Reader

Abstract

The invention discloses a mixed path planning method and system for multi-vehicle path conflict avoidance, and belongs to the field of unmanned vehicle path planning. Comprising the following steps: (1) establishing a two-dimensional plane area map, dividing M * N grid units, and calculating a comprehensive risk assessment value; (2) for any unmanned vehicle, calculating a current planned action, and predicting path trajectories of other unmanned vehicles in a plurality of time steps in the future at each moment; judging whether the planned action conflicts with the path tracks of other unmanned vehicles or not, and if yes, calculating a conflict penalty term; (3) selecting a greedy strategy, a DQN strategy or a fusion strategy according to the comprehensive risk assessment value of the grid unit where the unmanned vehicle is located currently so as to determine an optimal action; and (4) executing the optimal action and updating the unmanned vehicle state information and the risk score. The method breaks through the disadvantage that the traditional greedy and reinforcement learning strategies are isolated or randomly mixed, and has the advantages of high real-time performance and high robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of unmanned vehicle path planning, and in particular to a hybrid path planning method and system for multi-vehicle path conflict avoidance. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] In recent years, with breakthroughs in artificial intelligence and multimodal sensor fusion, the automotive industry is undergoing a revolution centered on intelligent and connected technologies. As the most representative technology in this field, autonomous driving systems, through the collaborative application of deep neural network algorithms, high-precision map modeling, and communication technologies, have enabled a paradigm shift from traditional manual driving to autonomous machine decision-making. Currently, autonomous vehicles using autonomous driving systems are widely used in warehousing and logistics, industrial inspections, and intelligent transportation. This, in turn, places higher demands on their ability to achieve efficient and safe path planning in complex environments.

[0004] Existing path planning methods mainly include heuristic search-based algorithms (such as A* and Dijkstra), sampling methods (such as RRT), and the recently emerging path planning methods based on reinforcement learning. Although reinforcement learning has shown certain advantages in terms of not requiring explicit environment modeling and autonomous learning strategies, it still has many shortcomings in practical applications:

[0005] 1. Most reinforcement learning methods, such as local path planning methods based on Q-learning and combined optimization methods using neural networks and reinforcement learning, often rely solely on obstacle distance or path length as a reward function and fail to explicitly consider the risk level of each area in the environment. This can cause autonomous vehicles to mistakenly enter high-risk areas during path planning, posing a safety hazard.

[0006] 2. Traditional methods often use greedy and reinforcement learning strategies in isolation or randomly mix them, lacking adaptive adjustment mechanisms based on environmental risks. In complex traffic scenarios, path selection in high-risk areas requires a combination of global optimization and local obstacle avoidance, but existing methods struggle to dynamically balance efficiency and safety.

[0007] 3. In multi-vehicle path planning, existing methods generate coordinated trajectories using priority schemes but fail to incorporate real-time conflict penalties, leading to frequent path overlap and adjacent interference. Furthermore, while artificial potential field-based methods can quantify risk, they lack multi-vehicle trajectory prediction, making them inadequate for rapid response to path conflicts in dynamic missions. Summary of the Invention

[0008] To address the shortcomings of the existing technology, the present invention provides a hybrid path planning method and system for multi-vehicle path conflict avoidance. By combining reinforcement learning with a risk scoring mechanism and greedy hybrid path planning, the method and system improve the safety of the planning process and the coordination of the multi-vehicle system while ensuring path efficiency. This solves the problems of low efficiency of reinforcement learning path planning in complex environments and unpredictable path conflicts.

[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0010] A first aspect of the present invention provides a hybrid path planning method for avoiding multi-vehicle path conflicts, comprising:

[0011] S101: Create a two-dimensional planar regional map, divide the two-dimensional planar regional map into M×N equally spaced grid cells, and calculate a comprehensive risk assessment value for each grid cell based on the passability attribute field and other unmanned vehicle trajectories; where M is the number of horizontal grid cells and N is the number of vertical grid cells;

[0012] S102: For any unmanned vehicle, calculate the current planned action and predict the paths of other unmanned vehicles in the next several time steps at each moment; determine whether the current planned action of the unmanned vehicle conflicts with the paths of other unmanned vehicles. If a conflict occurs, calculate the conflict penalty term;

[0013] S103: Selecting a greedy strategy, a DQN strategy, or a fusion strategy based on the comprehensive risk assessment value of the grid cell where the unmanned vehicle is currently located to obtain the optimal action; if the corresponding strategy is a fusion strategy, a conflict penalty term is included;

[0014] S104: Execute the optimal action and update the current unmanned vehicle status information and comprehensive risk assessment value;

[0015] S105: Repeat steps S102 to S104 until all unmanned vehicles reach their target points or the mission is completed.

[0016] A second aspect of the present invention provides a hybrid path planning system for multi-vehicle path conflict avoidance, comprising:

[0017] The environmental modeling and risk assessment module is configured to: create a two-dimensional planar regional map, divide the two-dimensional planar regional map into M×N equally spaced grid cells, and calculate a comprehensive risk assessment value for each grid cell based on the passability attribute field and other unmanned vehicle trajectories; where M is the number of horizontal grid cells and N is the number of vertical grid cells;

[0018] The path conflict prediction and penalty calculation module is configured to: for any unmanned vehicle, calculate the current planned action and predict the path trajectories of other unmanned vehicles in the future several time steps at each moment; determine whether the current planned action of the unmanned vehicle conflicts with the path trajectories of other unmanned vehicles, and if so, calculate the conflict penalty term;

[0019] An adaptive strategy selection module is configured to select a greedy strategy, a DQN strategy, or a fusion strategy based on the comprehensive risk assessment value of the grid cell where the unmanned vehicle is currently located to obtain the optimal action; if the corresponding strategy is a fusion strategy, a conflict penalty term is included;

[0020] An action execution and status update module, which is configured to: execute the optimal action and update the current unmanned vehicle status information and comprehensive risk assessment value;

[0021] The task cycle control module is configured to: plan a path for each unmanned vehicle until all unmanned vehicles reach their target point or the mission is completed.

[0022] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a hybrid path planning method for multi-vehicle path conflict avoidance as described in the first aspect of the present invention.

[0023] A fourth aspect of the present invention provides an electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps of a hybrid path planning method for avoiding multi-vehicle path conflicts as described in the first aspect of the present invention are implemented.

[0024] The above technical solution has the following advantages or beneficial effects:

[0025] (1) This paper proposes a risk scoring mechanism based on multi-factor fusion. It integrates multiple environmental features such as obstacle distribution, dynamic target prediction, historical high collision frequency, and other unmanned vehicle trajectories into a unified model, and calculates the risk score of each grid point through linear weighting. The scoring system supports real-time updates and has good dynamic performance.

[0026] (2) The present invention constructs a path conflict scoring mechanism, which uses the path broadcast or predicted trajectory of other vehicles to determine whether the current candidate action may overlap or conflict with other unmanned vehicles in the future, and assigns the action a corresponding conflict penalty value accordingly. The scoring suppression mechanism effectively constrains the selection probability of potential conflicting actions, thereby improving the safety and path stability of the multi-vehicle system without the need for global coordination or central control.

[0027] (3) This invention overcomes the drawbacks of traditional greedy and reinforcement learning strategies that are isolated or randomly mixed. It designs a risk-aware adaptive strategy fusion mechanism system that automatically adjusts the weight ratio between the greedy strategy and the deep network output according to the risk score value of the current position, achieving a dynamic path control behavior of "low-risk rapid advancement and high-risk robust evaluation". At the same time, the path conflict penalty term is integrated to construct a unified fusion scoring function to select the best from multiple candidate actions. This mechanism not only improves the overall path efficiency, but also enhances the robustness and interpretability of path selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0029] Figure 1 A flow chart of a hybrid path planning method for avoiding multi-vehicle path conflicts provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0030] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0031] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments of the present invention. The terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0032] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0033] All data in this embodiment is obtained in compliance with laws and regulations and based on the consent of the user, and is used legally.

[0034] Example 1

[0035] See also Figure 1 , Figure 1 A flow chart of a hybrid path planning method for avoiding multi-vehicle path conflicts provided by an embodiment of the present invention.

[0036] This embodiment provides a hybrid path planning method for avoiding multi-vehicle path conflicts, including the following steps:

[0037] S101: Create a two-dimensional planar regional map, divide the two-dimensional planar regional map into M×N equally spaced grid cells, and calculate a comprehensive risk assessment value for each grid cell based on the passability attribute field and other unmanned vehicle trajectories; where M is the number of horizontal grid cells and N is the number of vertical grid cells;

[0038] S102: For any unmanned vehicle, calculate the current planned action and predict the paths of other unmanned vehicles in the next several time steps at each moment; determine whether the current planned action of the unmanned vehicle conflicts with the paths of other unmanned vehicles. If a conflict occurs, calculate the conflict penalty term;

[0039] S103: Selecting a greedy strategy, a DQN strategy, or a fusion strategy based on the comprehensive risk assessment value of the grid cell where the unmanned vehicle is currently located to obtain the optimal action; if the corresponding strategy is a fusion strategy, a conflict penalty term is included;

[0040] S104: Execute the optimal action and update the current unmanned vehicle status information and comprehensive risk assessment value;

[0041] S105: Repeat steps S102 to S104 until all unmanned vehicles reach their target points or the mission is completed.

[0042] S101 specifically includes:

[0043] S101-1: A two-dimensional image of the unmanned vehicle's surroundings is captured using an onboard camera and mapped into physical space to construct a two-dimensional planar map covering the entire navigation area. Specifically, a world coordinate system is established based on the unmanned vehicle's initial positioning information, and the image pixel coordinates are aligned with the physical coordinates to complete the coordinate system alignment and construct a two-dimensional planar map covering the entire navigation area.

[0044] The two-dimensional plane area map is grid-divided according to the preset map resolution parameter, and the two-dimensional plane area map is divided into M×N equally spaced grid units. In this embodiment, each grid unit is set to represent a real range of r=0.1m×0.1m (i.e., each grid is 10cm). For a physical space area of ​​X×Y=5m×5m, the two-dimensional plane area map will be divided into horizontal grid cells, Vertical grid cells are formed to form a two-dimensional grid map consisting of 50×50 equally spaced grid cells.

[0045] Each grid cell sets the passability attribute field according to the sensor observation results. The passability attribute field includes the static obstacle identifier obstacle_static, the dynamic obstacle identifier obstacle_dynamic, and the historical pass heat counter visit_count.

[0046] Static obstacles refer to fixed inaccessible areas, such as walls, columns, and shelves. Static obstacle identification is obtained by identifying fixed inaccessible areas from the depth map.

[0047] Dynamic obstacle identification is based on the results of target detection and trajectory prediction to determine the area occupied by dynamic targets;

[0048] The historical heat counter records the frequency of visits to the location during the task, which is used to reflect historical risks or congestion trends.

[0049] S101-2: Initialize the risk score value R0(x,y) for each equally spaced grid cell for subsequent risk modeling and path assessment.

[0050] Taking into account the distribution of static obstacles, dynamic target activities, historical collision heat, and other unmanned vehicle trajectory coverage, a multi-factor weighted model is constructed to initialize the risk score value R0(x,y):

[0051] R0(x,y)=α1R obs (x,y)+α2R dyn (x,y)+α3R hist (x,y)+α4R veh (x,y) (1)

[0052] Where R obs (x, y) indicates whether the grid is a static obstacle. If it is an obstacle grid, the value is 1, and if it is not, the value is 0; R dyn (x, y) represents the probability that the current grid will be occupied by a dynamic obstacle in the future time period; R hisr (x,y) represents the heat value of the grid that has experienced collision, sudden stop, or frequent passing in the historical mission, reflecting the potential risk density of the area, expressed as R hist (x,y); R veh (x, y) represents the coverage probability of other unmanned vehicles planning to pass through this grid in the future time step; α1~α4 are weighting coefficients, satisfying α1+α2+α4+α4=1.

[0053] Among them, R dyn The specific calculation method of (x, y) is as follows: use the target detection and trajectory prediction algorithm to generate the path trajectory of the dynamic target in the next 5 to 10 time steps H; traverse the grid points covered in the trajectory and record the frequency n of its occurrence in the prediction sequenced (x,y); calculate the probability of dynamic obstacle occupancy through normalization

[0054] R hist The specific calculation method of (x, y) is as follows: during the path execution process, the number of visits to each grid by the vehicle is counted in real time, and the maximum number of passes V set by experience is used. max Perform normalization to obtain the heat value

[0055] R veh The specific calculation method of (x, y) is as follows: the system obtains the future trajectory prediction results of all other unmanned vehicles to form a trajectory set Π others (t), where each trajectory contains the path trajectory within the next 5 to 10 time steps H, and the system counts the frequency n of grid points appearing in all trajectories veh (x,y), and normalized to get

[0056] Each factor is weighted according to its impact on path safety. After normalization, each score is weighted and combined to form a comprehensive risk assessment value R(x,y) for each equally spaced grid cell, R(x,y)∈[0,1], where 0 indicates complete safety and 1 indicates a high-risk prohibited area.

[0057] For example, in indoor warehousing scenarios, more attention needs to be paid to the impact of static obstacles on path safety. Therefore, the weights of each item can be set to: α1 = 0.5, α2 = 0.3, α3 = 0.1, and α4 = 0.1.

[0058] The other unmanned vehicles referred to in the present invention are unmanned vehicles that are in motion or in the task execution stage in all current driving directions, and do not include unmanned vehicles that are stationary or have not yet been assigned a path task.

[0059] This paper introduces a multi-source environmental risk assessment model based on the construction of a two-dimensional grid map, providing a real-time risk score for each grid cell. By integrating multiple environmental features, such as obstacle distribution, dynamic target prediction, historical high collision frequency, and other unmanned vehicle trajectories, the risk score R(x,y) for each grid point is calculated using linear weighting. This scoring system supports real-time updates and possesses excellent dynamic adaptability, improving the foresight and robustness of path planning in complex and dynamic environments.

[0060] S102 specifically includes:

[0061] S102-1: Calculate the current unmanned vehicle planned action s′: For the current unmanned vehicle in state s (t)Each candidate action a under A(s), A(s) is a candidate action set; based on the current position state s = (x, y) and the action direction offset (Δx a ,Δy a ), calculate the next state s′=s after execution by adding the position coordinates (t+1) =(x+Δx a ,y+Δy a ), that is, planned action.

[0062] Among them, the state s (t) Indicates the current perception state of the unmanned vehicle in the tth control cycle, including the current position grid coordinates (x t ,y t ), the comprehensive risk score R(x t ,y t ), the current candidate action set A(s); the candidate action set A(s) represents the unmanned vehicle in the current state s (t) The basic movement directions that can be taken are defined as 4 basic actions, which represent up, down, left and right respectively;

[0063] S102-2: Construct a set of trajectories of other unmanned vehicles within the set prediction time window H, denoted as Π others (t), used for conflict detection.

[0064] The process of constructing the trajectory set of other unmanned vehicles is as follows: Through V2X communication, each unmanned vehicle, including the current one, broadcasts its planned path in real time. The current unmanned vehicle receives, extracts, and integrates the paths into a standardized trajectory sequence. Ultimately, the trajectories of all other vehicles are combined into a trajectory set. The prediction time window consists of multiple time steps. For example, if each time step is set to one minute, a one-hour prediction time window contains 60 time steps.

[0065] According to the trajectory set π of other unmanned vehicles others (t), predict the path trajectories of other unmanned vehicles in the next several time steps at each moment. In this embodiment, the path trajectories of other unmanned vehicles in the next 5 to 10 time steps are predicted at each moment.

[0066] S102-3: Determine whether the planned action of the current unmanned vehicle conflicts with the path trajectory of other unmanned vehicles. If a conflict occurs, determine the conflict level.

[0067] The conflict level is determined as follows: if the predicted path points of s′ and other unmanned vehicles at the same time step completely overlap, it is considered a hard conflict; if the predicted path points of s′ and other unmanned vehicles are four adjacent positions (that is, the Manhattan distance is 1), it is considered a soft conflict; otherwise, there is no conflict.

[0068] S102-4: Assign a conflict penalty term C(s,a) to the action based on the conflict level:

[0069]

[0070] In step S102, the present invention establishes a path conflict scoring mechanism. This mechanism uses the path broadcasts or predicted trajectories of other vehicles to determine whether the current candidate action is likely to cause path overlap or adjacency conflicts with other vehicles in the future. Based on this, the action is assigned a corresponding conflict penalty value C(s, a). This scoring suppression mechanism effectively constrains the selection probability of potentially conflicting actions, improving the safety and path stability of the multi-vehicle system without the need for global coordination or central control.

[0071] S103 specifically includes:

[0072] For each candidate action a∈A(s) of the autonomous vehicle, the safety and effectiveness of its execution at the current location are evaluated, and one of the following three strategies is selected based on the comprehensive risk assessment value R(x,y) of the current grid:

[0073] Strategy 1: If the comprehensive risk assessment value R(x,y) of the current location is lower than the preset risk threshold T low = 0.2, the autonomous vehicle takes the greedy strategy to select the action. At this time, for each candidate action a, the state s′ after execution is calculated to the target point g(x g ,y g ) and select the action that minimizes the distance as the optimal action a * :

[0074]

[0075] Among them, x′ and y′ represent the position coordinates of the next grid point reached after executing candidate action a, that is, the position of the next state.

[0076] The advantage of choosing strategy 1 at this time is that the current grid point comprehensive risk score R(x,y) is lower than the set low risk threshold T low When the system determines that the current environment is relatively safe, with no obvious obstacles, conflicts, or dynamic risks, a greedy strategy can be used to improve path advancement efficiency. The greedy strategy evaluates the Euclidean distance between each candidate action and the target point after execution and selects the action that minimizes this distance as the current control instruction. This strategy offers low computational overhead and intuitive pathing. Compared to deep or fusion strategies, the greedy strategy offers advantages in low-risk areas, such as lightweight computation, rapid response, and good path continuity. It is suitable for rapid path advancement in large, open areas.

[0077] Strategy 2: If the risk score R(x,y) of the current location is higher than the preset risk threshold T high= 0.8, the reinforcement learning strategy inputs the current state s through the trained DQN network and outputs the action value function Q DQN (s,a), select the action with the largest Q value as the optimal action a * :

[0078] a * =arg max a∈A(s) Q DQN (s,a) (3)

[0079] The DQN network is a trained state-action value function estimation model. During the training phase, the state transition samples are iteratively optimized based on the Q-learning principle. The system first takes the current position state S as the network input. The network structure can be a multi-layer perceptron or a convolutional neural network, and outputs the action value Q corresponding to each candidate action a∈A(s). DQN During training, the system samples state transition sequences in a simulated environment, builds an experience pool, and optimizes the loss using the expected return of the future optimal strategy as the target value:

[0080] Q(s,a)←r+γ·max a′ Q(s′,a′)(4)

[0081] Where r is the reward function, defined as follows:

[0082]

[0083] The advantage of choosing strategy 2 at this time is that the comprehensive risk score R(x,y) at the current grid location is higher than the preset high risk threshold T high =0.8, the system determines that the current area has a high path risk, such as a concentration of dynamic obstacles, a high probability of collision with other vehicles, etc. At this time, the system adopts a reinforcement learning strategy, calling the trained DQN network to evaluate the action value of the current state s, and selects the action a with the largest Q value. * This strategy possesses global path optimization capabilities and implements a risk-averse-first path selection strategy based on historical experience and long-term reward trade-offs. Compared to the greedy strategy, the reinforcement learning strategy exhibits greater robustness and risk aversion in high-risk areas, effectively avoiding local optimal paths and potential collisions.

[0084] Strategy 3: If the risk score R(x,y) of the current location is in the medium range T low <R(x,y)<T high When , a weighted fusion method of greedy and reinforcement learning strategies is used to score actions, that is, for each candidate action a∈A(s), the greedy score, DQN score and conflict penalty term are comprehensively considered to construct a unified fusion score function Qtotal (s,a):

[0085] Q total (s,a)=λ g Q greedy (s,a)+(1-λ g )Q DQN (s,a)-λ c C(s,a) (5)

[0086] Where λ g =1-R(x,y), which represents the risk perception greed factor, which decreases as the risk increases; Q greedy (s,a) is calculated by the greedy strategy; Q DQN (s,a) is output by the DQN network; C(s,a) is the penalty term for actions that cause path conflicts in the multi-vehicle system; λ c is the weight coefficient of path conflict penalty, λ c Usually the value range is λ c ∈[2,10], which can be set according to the safety requirements and traffic density of the mission scenario. For example, in a scenario with an open environment and low collision probability, λ can be set c 2 to 4. In scenarios where moderate avoidance efficiency is prioritized, λ can be set c 4 to 6, which can be set when there is a high risk of conflict when multiple vehicles are running at the same time. c 7 to 10.

[0087] Among them, Q greedy (s,a) reflects the Euclidean distance from the target point after the candidate action a is executed. λ c The value range is usually c ∈[2,10];

[0088] Finally, the action with the highest fusion score is selected as the optimal action a * :

[0089] a * =arg max a∈A(s) Q total (s,a) (6)

[0090] The specific process of selecting the action with the highest fusion score as the final control instruction is as follows: For example, in a certain control cycle, the vehicle is currently at position (20, 25), and the comprehensive risk score of this position is R = 0.6, which is in the preset medium risk range. The system calls the fusion strategy module to calculate the greed score Q of each candidate action separately. greedy , Deep Q Score Q DQN And the path conflict penalty C(s,a), and according to the fusion function Q total(s, a) is used to calculate the total score. Finally, the fusion scores of all actions are compared, and the action with the highest score is selected as the current control instruction. In this example, the system selects "right" as the optimal action, achieving a dynamic balance between target path efficiency and risk aversion.

[0091] In the action decision stage of strategy three, a weighted fusion of reinforcement learning strategy function and greedy heuristic function is adopted to dynamically adjust the strategy selection according to the risk level of the current position. At the same time, a path conflict prediction mechanism is introduced to evaluate the degree of overlap between candidate actions and the future paths of other unmanned vehicles and assign corresponding penalty items, thereby ensuring path efficiency while improving the safety of the planning process and the coordination of the multi-vehicle system.

[0092] S104 specifically includes:

[0093] Execute the optimal action a * ; and update the current unmanned vehicle status information and comprehensive risk assessment value.

[0094] The state information of the unmanned vehicle includes: the current position grid coordinate s = (x, y), and the optimal action a performed by the unmanned vehicle in the current cycle * ; After the action is executed, the unmanned vehicle updates its current position to the next state s′=s+a * , serving as the input state for the next cycle. Simultaneously, based on sensor data and multi-vehicle path trajectory information, the comprehensive risk assessment value for each location in the current grid map is updated, providing environmental input for the next action selection. A cycle refers to the complete cycle of perception, decision-making, execution, and updating that the unmanned vehicle completes during path planning and control.

[0095] This method dynamically adjusts action selection based on varying environmental risk levels by introducing a risk-scoring-based strategy switching mechanism and a path conflict penalty. This method combines the efficiency of greedy search with the global optimization capabilities of reinforcement learning, offering excellent environmental adaptability and safety control capabilities, and has broad application prospects.

[0096] Example 2

[0097] This embodiment provides a hybrid path planning system for avoiding multi-vehicle path conflicts.

[0098] A hybrid path planning system for multi-vehicle path conflict avoidance, comprising: an environment modeling and risk assessment module, a path conflict prediction and penalty calculation module, an adaptive strategy selection module, and an action execution and state update module;

[0099] The environmental modeling and risk assessment module is configured to: create a two-dimensional planar regional map, divide the two-dimensional planar regional map into M×N equally spaced grid cells, and calculate a comprehensive risk assessment value for each grid cell based on the passability attribute field and other unmanned vehicle trajectories; where M is the number of horizontal grid cells and N is the number of vertical grid cells;

[0100] The path conflict prediction and penalty calculation module is configured to: for any unmanned vehicle, calculate the current planned action and predict the path trajectories of other unmanned vehicles in the future several time steps at each moment; determine whether the current planned action of the unmanned vehicle conflicts with the path trajectories of other unmanned vehicles, and if so, calculate the conflict penalty term;

[0101] An adaptive strategy selection module is configured to select a greedy strategy, a DQN strategy, or a fusion strategy based on the comprehensive risk assessment value of the grid cell where the unmanned vehicle is currently located to obtain the optimal action; if the corresponding strategy is a fusion strategy, a conflict penalty term is included;

[0102] An action execution and status update module, which is configured to: execute the optimal action and update the current unmanned vehicle status information and comprehensive risk assessment value;

[0103] The task cycle control module is configured to: plan a path for each unmanned vehicle until all unmanned vehicles reach their target point or the mission is completed.

[0104] It should be noted that the aforementioned environment modeling and risk assessment module, path conflict prediction and penalty calculation module, adaptive strategy selection module, and action execution and state update module correspond to steps S101 to S105 in Example 1. The examples and application scenarios implemented by these modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the aforementioned modules, as part of the system, can be executed in a computer system, such as a set of computer-executable instructions.

[0105] The description of each embodiment in the above embodiments has different emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0106] The proposed system can be implemented in other ways. For example, the system embodiment described above is merely illustrative. For example, the above module division is only a logical function division. In actual implementation, other division methods may be used. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not implemented.

[0107] Example 3

[0108] This embodiment also provides an electronic device, comprising: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in the above embodiment one.

[0109] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), off-the-shelf field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0110] The memory may include a read-only memory and a random access memory, and provides instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0111] During implementation, each step of the above method may be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software.

[0112] Those skilled in the art will appreciate that the units and algorithm steps of the various examples described in conjunction with this embodiment can be implemented using electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0113] Example 4

[0114] This embodiment further provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first embodiment is performed.

[0115] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A hybrid path planning method for multi-vehicle path conflict avoidance, characterized by: The following steps are involved: S101: Establish a two-dimensional planar regional map, divide the two-dimensional planar regional map into M×N equally spaced grid cells, and calculate a comprehensive risk assessment value for each grid cell based on the passability attribute field and other unmanned vehicle trajectories; S102: For any unmanned vehicle, calculate the current planned action and predict the paths of other unmanned vehicles in the next several time steps at each moment; determine whether the current planned action of the unmanned vehicle conflicts with the paths of other unmanned vehicles. If a conflict occurs, calculate the conflict penalty term; S103: Selecting a greedy strategy, a DQN strategy, or a fusion strategy based on the comprehensive risk assessment value of the grid cell where the unmanned vehicle is currently located to obtain the optimal action; if the corresponding strategy is a fusion strategy, a conflict penalty term is included; S104: Execute the optimal action and update the current unmanned vehicle status information and comprehensive risk assessment value; S105: Repeat steps S102 to S104 until all unmanned vehicles reach their target points or the mission is completed.

2. The hybrid path planning method for multi-vehicle path conflict avoidance according to claim 1, characterized in that: The two-dimensional plane area map is established by dividing the two-dimensional plane area map into M×N equally spaced grid units, specifically: The vehicle's onboard camera collects a two-dimensional image of the vehicle's surroundings and maps the image into the actual physical space to construct a two-dimensional planar area map covering the entire navigation area. The two-dimensional plane area map is grid-divided according to the preset map resolution parameters, and the two-dimensional grid map is divided into M×N equally spaced grid units; Each grid cell is assigned a passability attribute field based on sensor observations.

3. The hybrid path planning method for multi-vehicle path conflict avoidance according to claim 1, characterized in that: The comprehensive risk assessment value is calculated for each grid cell based on the passability attribute field and other unmanned vehicle trajectories. The specific steps are as follows: Construct a multi-factor weighted model with the initial risk score R0(x,y): R0(x,y)=α1R obs (x,y)+α2R dyn (x,y)+α3R hist (x,y)+α4R veh (x,y); Among them, R obs (x, y) indicates whether the grid is a static obstacle. If it is an obstacle grid, the value is 1, and if it is not, the value is 0; R dyn (x, y) represents the probability that the current grid will be occupied by a dynamic obstacle in the future time period; R hist (x,y) represents the heat value of the grid that has experienced collision, sudden stop, or frequent passing in the historical mission, reflecting the potential risk density of the area, expressed as R hist (x,y); R veh (x, y) represents the coverage probability of other unmanned vehicles planning to pass through this grid in the future time step; α1~α4 are weighting coefficients, satisfying α1+α2+α4+α4=1; Each factor is weighted according to its impact on path safety. After normalization, each score is weighted and combined to form a comprehensive risk assessment value R(x,y) for each equally spaced grid cell, R(x,y)∈[0,1], where 0 indicates complete safety and 1 indicates a high-risk prohibited area.

4. The hybrid path planning method for multi-vehicle path conflict avoidance according to claim 1, characterized in that: The step S102 is specifically as follows: Calculate the current planned action s′ of the unmanned vehicle; Construct the trajectory set Π of other unmanned vehicles within the set prediction time window H others (t), predicting the path trajectories of other autonomous vehicles in the next several time steps at each moment; Determine whether the planned action of the current unmanned vehicle conflicts with the path trajectory of other unmanned vehicles. If a conflict occurs, determine the conflict level. If s′ completely overlaps with the predicted path points of other unmanned vehicles at the same time step, it is determined to be a hard conflict. If s′ and the predicted path points of other unmanned vehicles are in four adjacent positions, it is determined to be a soft conflict. Otherwise, it is considered to be no conflict. A conflict penalty term C(s,a) is assigned to the action based on the conflict level: when it is judged to be a hard conflict, C(s,a) is 1.0; when it is judged to be a soft conflict, C(s,a) is 0.5; when it is judged to be no conflict, C(s,a) is 0.

5. The hybrid path planning method for multi-vehicle path conflict avoidance according to claim 1, characterized in that: The greedy strategy is selected based on the comprehensive risk assessment value of the current position of the unmanned vehicle to obtain the optimal action, specifically: If the comprehensive risk assessment value R(x,y) of the current location is lower than the preset risk threshold T low = 0.2, the autonomous vehicle adopts a greedy strategy for action selection; for each action a, the state s′ after execution is calculated to the target point g(x g ,y g ) and select the action that minimizes the distance as the optimal action a * : Among them, x′ and y′ represent the position coordinates of the next grid point reached after executing candidate action a, that is, the position of the next state; A(s) is the candidate action set.

6. The hybrid path planning method for multi-vehicle path conflict avoidance according to claim 1, characterized in that: The DQN strategy is selected based on the comprehensive risk assessment value of the current position of the unmanned vehicle to obtain the optimal action, specifically: If the risk score R(x,y) of the current location is higher than the preset risk threshold T high = 0.8, the reinforcement learning strategy inputs the current state s through the trained DQN network and outputs the action value function Q DQN (s,a), select the action with the largest Q value as the optimal action a * : a * =argmax a∈A(s) Q DQN (s,a)。 7. The hybrid path planning method for multi-vehicle path conflict avoidance according to claim 1, characterized in that: The fusion strategy is selected based on the comprehensive risk assessment value of the current position of the unmanned vehicle to obtain the optimal action, specifically: If the risk score R(x,y) of the current location is in the medium range T low <R(x,y)<T high When , a weighted fusion method of greedy and reinforcement learning strategies is used to score actions, while considering the path conflict penalty term to construct a unified fusion scoring function: Q total (s,a)=λ g Q greedy (s,a)+(1-λ g )Q DQN (s,a)-λ c C(s,a); Finally, the action with the highest fusion score is selected as the optimal action a * : Among them, λ g =1-R(x,y), which represents the risk perception greed factor, which decreases as the risk increases; Q greedy (s,a) is calculated by the greedy strategy; Q DQN (s,a) is output by the DQN network; C(s,a) is the penalty term for actions that cause path conflicts in the multi-vehicle system; λ c is the weight coefficient of path conflict penalty; a * =argmax a∈A(s) Q total (s,a)。 8. A hybrid path planning system for multi-vehicle path conflict avoidance, characterized in that: include: Environmental modeling and risk assessment module, path conflict prediction and penalty calculation module, adaptive strategy selection module, and action execution and state update module; The environmental modeling and risk assessment module is configured to: create a two-dimensional planar regional map, divide the two-dimensional planar regional map into M×N equally spaced grid cells, and calculate a comprehensive risk assessment value for each grid cell based on the passability attribute field and other unmanned vehicle trajectories; where M is the number of horizontal grid cells and N is the number of vertical grid cells; The path conflict prediction and penalty calculation module is configured to: for any unmanned vehicle, calculate the current planned action and predict the path trajectories of other unmanned vehicles in the future several time steps at each moment; determine whether the current planned action of the unmanned vehicle conflicts with the path trajectories of other unmanned vehicles, and if so, calculate the conflict penalty term; An adaptive strategy selection module is configured to select a greedy strategy, a DQN strategy, or a fusion strategy based on the comprehensive risk assessment value of the grid cell where the unmanned vehicle is currently located to obtain the optimal action; if the corresponding strategy is a fusion strategy, a conflict penalty term is included; An action execution and status update module, which is configured to: execute the optimal action and update the current unmanned vehicle status information and comprehensive risk assessment value; The task cycle control module is configured to: plan a path for each unmanned vehicle until all unmanned vehicles reach their target point or the mission is completed.

9. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the method according to any one of claims 1 to 7 are implemented.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Multi-transportation equipment cooperative motion control method based on multi-stage mixed learning and storage medium

    CN121325624A

  • Equipment track real-time re-planning method and system for complex scene

    CN121900488A