Multi-vehicle homogeneous cooperative behavior decision-making method and device
By adopting a region-driven multi-vehicle cooperative behavior decision-making method, vehicle information and environmental perception information are acquired, energy intensity distribution is determined, and conflict areas are resolved. This solves the multi-vehicle cooperative decision-making problem in complex scenarios in existing technologies and achieves efficient, generalized, and safe traffic behavior decision-making.
Patent Information
- Application Number
- CN202410907142.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2026-01-09
AI Technical Summary
Existing multi-vehicle collaborative decision-making methods struggle to achieve efficient, generalized, and safe behavioral decisions in complex scenarios. These methods also suffer from safety issues and inefficient strategies in real-world traffic situations, failing to cope with dynamically changing traffic conditions.
A region-driven behavioral decision-making mechanism is adopted. By acquiring vehicle information and environmental perception information, the energy intensity distribution in the area in front of the vehicle is determined, the first and second feasible areas are calculated, and the conflict areas are resolved through a cooperative strategy to achieve multi-vehicle cooperative behavioral decision-making.
It achieves homogeneous, generalized, efficient, and safe collaborative behavior decision-making for multiple vehicles in complex scenarios, enabling it to cope with dynamically changing traffic scenarios and improve traffic efficiency and safety.
Smart Images

Figure CN121291483A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving decision-making technology, and in particular to a method and apparatus for multi-vehicle homogeneous cooperative behavior decision-making. Background Technology
[0002] Existing multi-vehicle cooperative decision-making can be broadly categorized into three types. The first type is vehicle-to-infrastructure (V2I) cooperation based on roadside infrastructure, which uses technologies such as traffic light control to indirectly control the behavior of traffic participants and achieve macro-level coordination. However, this method can only simply coordinate traffic flow and cannot control the specific behavior of intelligent agents, still leading to numerous safety issues and inefficient strategies in real-world scenarios. The second type is equidistant platooning, widely used in commercial vehicles, where several intelligent connected vehicles travel at a constant speed in a single-line formation. This method can reduce vehicle energy consumption in simple, interference-free scenarios; however, its drawback is its limited functionality, making it difficult to apply to other scenarios. The third type is fixed platooning, which extends equidistant platooning by first defining a grid formation and then designing several formations such as triangles, rhombuses, and double-row staggered formations, selecting the appropriate formation for passage in various scenarios. This method effectively achieves cooperative passage in simple interaction scenarios; however, its generalization ability is poor, requiring the definition of specific formations for specific scenarios, and it cannot cope with the uncontrollable behavior of surrounding vehicles.
[0003] Most existing multi-vehicle cooperative decision-making methods adopt a task-driven development approach, which sets the feasible set as several tasks. However, real-world traffic scenarios are dynamically changing and involve numerous factors, making it difficult to cover real road requirements with a limited number of tasks. Therefore, there is currently no efficient, generalizable, and safe method for multi-vehicle cooperative behavior decision-making in complex scenarios. Summary of the Invention
[0004] In view of this, this application provides a method and apparatus for multi-vehicle homogeneous collaborative behavior decision-making. The method adopts a region-driven behavior decision-making mechanism to achieve efficient, generalized, and safe group collaborative behavior decision-making tasks.
[0005] In a first aspect, embodiments of this application provide a multi-vehicle homogeneous cooperative behavior decision-making method, including:
[0006] Acquire information and environmental perception information of multiple vehicles within a target area; the vehicles include controllable and uncontrollable vehicles.
[0007] Based on information about multiple vehicles within the target area, determine the energy intensity distribution in the area in front of any target controllable vehicle;
[0008] Based on the environmental perception information and the energy intensity distribution in the area in front of any target controllable vehicle, a first feasible area for the target controllable vehicle is determined;
[0009] Based on the incentive value of the first feasible region of any target controllable vehicle, determine the second feasible region of any target controllable vehicle;
[0010] According to the preset coordination strategy, the second feasible area of controllable vehicles with conflicts is resolved.
[0011] In one possible implementation, the vehicle information includes: the vehicle's position and speed;
[0012] Based on information from multiple vehicles within the target area, determine the energy intensity distribution in the area in front of any controllable target vehicle, including:
[0013] Determine the first static potential energy and dynamic potential energy generated by all uncontrollable vehicles within the target area at the target point in the area in front of any target controllable vehicle.
[0014] Based on the first static potential energy field and dynamic potential energy generated by all uncontrollable vehicles at the target point, the energy intensity of the uncontrollable vehicles at the target point is determined.
[0015] Determine the second static potential energy generated at the target point by all controllable vehicles within the target area;
[0016] Based on the time delay distribution probability function and the second static potential energy generated by all controllable vehicles at the target point, the energy intensity of the controllable vehicles at the target point is determined.
[0017] The energy intensity of the target point is determined based on the energy intensity of the uncontrollable vehicles and the energy intensity of the controllable vehicles at the target point.
[0018] In one possible implementation, determining the first static potential energy and dynamic potential energy generated by all uncontrollable vehicles within the target area at a target point in the area in front of any target controllable vehicle includes:
[0019] For a target point with two-dimensional coordinates (x, y) in front of any controllable target vehicle, the first static potential energy generated at that target point by any uncontrollable vehicle within the target area. for:
[0020]
[0021]
[0022] Where G is a preset parameter, M a Let k1 be the mass of the uncontrollable vehicle, k1 be the first parameter, and r be the distance between the uncontrollable vehicle and the target point; (x a ,y a () represents the two-dimensional coordinates of the centroid of the uncontrollable vehicle;
[0023] The dynamic potential energy generated by the uncontrollable vehicle at the target point for:
[0024]
[0025] Where k2 is the second parameter, k3 is the third parameter, and v r Let θ be the relative speed between the uncontrollable vehicle and any of the target controllable vehicles. a Let be the relative angle between the center of mass of any uncontrollable vehicle and the target point.
[0026] In one possible implementation, determining the second static potential energy generated at the target point by all controllable vehicles within the target area includes:
[0027] For a target point with two-dimensional coordinates (x, y) in front of any controllable vehicle, the second static potential energy generated by any controllable vehicle within the target area at that target point. for:
[0028]
[0029] Where G is a preset parameter, M b For the mass of the controllable vehicle, k1 is the first parameter, (x b ,y b Let S be the two-dimensional coordinates of the centroid of the controllable vehicle, and let S be the coordinates of the centroid of the vehicle. b ,y b S is a circular region with a preset radius, where (m,n) are the two-dimensional coordinates of any point within the circular region S; p(m,n) is the probability value.
[0030] p(m,n)=f(m,n,x b ,y b )
[0031] Where f is a two-dimensional normal distribution.
[0032] In one possible implementation, the energy intensity of the controllable vehicles at the target point is determined based on the time delay distribution probability function and the second static potential energy generated by all controllable vehicles at the target point; including:
[0033] Treating the time delay as a Markov process, determine the probability function p of the time delay distribution. delay (t), where t is time;
[0034] The energy intensity of any controllable vehicle at the target point is obtained by performing a time-diffusion-weighted calculation on the second static potential energy generated by the vehicle at the target point.
[0035]
[0036] Among them, v b Let θ be the speed of any controllable vehicle. b Let be the relative angle between the center of mass of any controllable vehicle and the target point;
[0037] The sum of the energy intensities of all controllable vehicles at the target point is taken as the energy intensity of the controllable vehicles at the target point.
[0038] In one possible implementation, the environmental perception information includes lane line type, traffic light phase, and traffic identifier;
[0039] The area in front of any target controllable vehicle includes six sub-areas of the same size: a first straight lane sub-area, a second straight lane sub-area, a first left lane sub-area, a second left lane sub-area, a first right lane sub-area, and a second right lane sub-area.
[0040] Based on the environmental perception information and the energy intensity distribution in the area in front of any target controllable vehicle, a first feasible area for the target controllable vehicle is determined, including:
[0041] Based on the environmental perception information, the passage status of the straight lane, left lane and right lane of any target controllable vehicle is determined respectively, thereby determining the sub-area of the first state, where the first state is passable;
[0042] Based on the energy intensity distribution in the area in front of any of the target controllable vehicles, calculate the average energy intensity of the sub-region in the first state;
[0043] When the average energy intensity of a sub-region in the first state is greater than a preset threshold, the first state of the sub-region is updated to the second state; the second state is impassable.
[0044] All sub-regions of the first state are combined to form the first feasible region of the target controllable vehicle.
[0045] In one possible implementation, determining a second feasible region for the target controllable vehicle based on an excitation value of a first feasible region of the target controllable vehicle includes:
[0046] When the first feasible region is the first straight lane sub-region, the first left lane sub-region, or the first right lane sub-region, the grid length parameter E is determined to be the length of the sub-region; the average energy intensity R is determined to be the average energy intensity of the sub-region.
[0047] When the first feasible region is the second straight lane sub-region, the second left lane sub-region, or the second right lane sub-region, the grid length parameter E is determined to be twice the length of the sub-region; the average energy intensity R is determined to be the average of the energy intensity of the sub-region and its adjacent sub-regions along the lane line direction.
[0048] When the first feasible area is a left lane sub-area, a first right lane sub-area, a second left lane sub-area, or a second right lane sub-area, the lane change penalty item C is determined to be a preset parameter greater than 0; when the first feasible area is a straight lane sub-area or a second straight lane sub-area, the lane change penalty item C is determined to be 0.
[0049] Calculate the incentive function reward for the first feasible region:
[0050] reward = w e Ew r Rw c C
[0051] Among them, w e As the first weight, w r As the second weight, w c It is the third weight;
[0052] The sub-region corresponding to the highest excitation value in the first feasible region of any target controllable vehicle is determined as the second feasible region.
[0053] In one possible implementation, the second feasible region of the conflicting controllable vehicles is resolved according to a preset cooperative strategy, including:
[0054] When it is detected that the second feasible areas of two controllable vehicles overlap, the two controllable vehicles include the first controllable vehicle and the second controllable vehicle, and the following processing is performed:
[0055] If the second feasible area of the first controllable vehicle is located in the straight lane of the first controllable vehicle, then the second feasible area of the second controllable vehicle is set to impassable.
[0056] If the second feasible area of the first controllable vehicle is located in the left lane of the first controllable vehicle and the second feasible area of the second controllable vehicle is located in the right lane of the second controllable vehicle, then the second feasible area of the second controllable vehicle is set to impassable.
[0057] If the second feasible area of the second controllable vehicle is located in the straight lane of the second controllable vehicle, then the second feasible area of the first controllable vehicle is set to impassable.
[0058] If the second feasible area of the second controllable vehicle is located in the left lane of the first controllable vehicle and the second feasible area of the first controllable vehicle is located in the right lane of the second controllable vehicle, then the second feasible area of the first controllable vehicle is set to impassable.
[0059] Secondly, embodiments of this application provide a multi-vehicle homogeneous collaborative behavior decision-making device, comprising:
[0060] The acquisition unit is used to acquire information and environmental perception information of multiple vehicles within a target area; the vehicles include controllable vehicles and uncontrollable vehicles.
[0061] The first determining unit is used to determine the energy intensity distribution in the area in front of any target controllable vehicle based on information about multiple vehicles within the target area.
[0062] The second determining unit is used to determine the first feasible area of the target controllable vehicle based on the environmental perception information and the energy intensity distribution in the area in front of the target controllable vehicle.
[0063] The third determining unit is used to determine the second feasible region of any target controllable vehicle based on the excitation value of the first feasible region of any target controllable vehicle;
[0064] The collaborative behavior decision-making unit is used to resolve the second feasible area of controllable vehicles with conflicts according to a preset collaborative strategy.
[0065] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor, wherein the memory stores an executable program, and the processor executes the executable program to implement the steps of the method of embodiments of this application.
[0066] This application enables homogeneous, generalized, efficient, and safe collaborative behavior decision-making for multiple vehicles in complex scenarios. Attached Figure Description
[0067] Figure 1 This is a flowchart of the multi-vehicle homogeneous collaborative behavior decision-making method according to an embodiment of this application;
[0068] Figure 2 This is a schematic diagram of six sub-regions in front of any target controllable vehicle in an embodiment of this application;
[0069] Figure 3 This is a flowchart of step 102 of the multi-vehicle homogeneous cooperative behavior decision-making method according to an embodiment of this application;
[0070] Figure 4 This is a schematic diagram of the time delay distribution probability function according to an embodiment of this application;
[0071] Figure 5This is a flowchart of step 103 of the multi-vehicle homogeneous cooperative behavior decision-making method according to an embodiment of this application;
[0072] Figure 6 This is a schematic diagram illustrating conflict resolution in an embodiment of this application;
[0073] Figure 7 This is a schematic diagram of the extended feasible area according to an embodiment of this application;
[0074] Figure 8 This is a functional structure diagram of the multi-vehicle homogeneous collaborative behavior decision-making device according to an embodiment of this application;
[0075] Figure 9 This is a functional structure diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0076] Various embodiments and features of this application are described herein with reference to the accompanying drawings.
[0077] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered as limiting, but merely as an example of embodiments. Other modifications within the scope and spirit of this application will be apparent to those skilled in the art.
[0078] The accompanying drawings, which are included in and form part of this specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.
[0079] These and other features of this application will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.
[0080] It should also be understood that although this application has been described with reference to some specific examples, those skilled in the art can certainly implement many other equivalent forms of this application.
[0081] The above and other aspects, features and advantages of this application will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.
[0082] Specific embodiments of this application are described thereafter with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are merely examples of this application, which can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that could obscure the application. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but merely serve as the basis and representative basis for the claims to teach those skilled in the art to use this application in a variety of substantially any suitable detailed structures.
[0083] This specification may use the phrases “in one embodiment,” “in another embodiment,” “in yet another embodiment,” or “in other embodiments,” all of which may refer to one or more of the same or different embodiments according to this application.
[0084] First, a brief introduction to the technical terms used in the embodiments of this application will be given.
[0085] Scene understanding and evaluation for intelligent connected vehicles (ICVs) is one of the key technologies for autonomous driving. After receiving real-time and historical information about the surrounding traffic environment, in order to make decisions and control autonomous driving commands, it is necessary to further refine the raw physical information to achieve semantic understanding and evaluation of complex traffic scenarios.
[0086] With the advancement of vehicle-road-cloud integration, multi-vehicle collaboration is considered the next generation of technology that can revolutionize intelligent mobility. Numerous studies have shown that collaborative decision-making can significantly improve traffic efficiency and safety compared to single-vehicle decision-making.
[0087] Most existing multi-vehicle cooperative decision-making algorithms adopt a task-driven development approach, which sets the feasible set as several tasks. However, real-world traffic scenarios are dynamically changing and involve numerous factors, making it difficult to cover real road requirements with a limited number of tasks. Therefore, there is a lack of an efficient, generalizable, and safe multi-vehicle cooperative behavior decision-making algorithm for handling complex scenarios.
[0088] To address the aforementioned technical challenges, this application employs a region-driven design approach, decomposing behavioral decisions into combinations of relative regions. Simultaneously, it designs a collaborative driving risk field to achieve multi-vehicle collaborative decision-making. Finally, it performs conflict detection and correction, realizing homogeneous, generalized, efficient, and safe collaborative behavioral decision-making for multiple vehicles in complex scenarios, thus overcoming the constraints of strong rules such as fixed distances and speeds.
[0089] After introducing the application scenarios and design concepts of the embodiments of this application, the technical solutions provided by the embodiments of this application will be described below.
[0090] This application provides a multi-vehicle homogeneous cooperative behavior decision-making method, which can be applied to electronic devices such as computer terminals and executed by the processor of the electronic device.
[0091] Figure 1 A flowchart illustrating the multi-vehicle homogeneous collaborative behavior decision-making method provided in this application embodiment. Figure 1 As shown, the multi-vehicle homogeneous cooperative behavior decision-making method of this application embodiment specifically includes steps 101-105:
[0092] Step 101: Obtain information on multiple vehicles and environmental perception information within the target area; the vehicles include controllable vehicles and uncontrollable vehicles;
[0093] For example, the controllable vehicle is an autonomous vehicle, and the uncontrollable vehicle is a human-driven vehicle. This embodiment involves collaborative decision-making regarding the driving behaviors of multiple controllable vehicles.
[0094] In this embodiment, vehicle information includes: vehicle position, speed, and mass. Environmental perception information includes lane line type, traffic light phase, and traffic identifiers.
[0095] Step 102: Based on information about multiple vehicles within the target area, determine the energy intensity distribution in the area in front of any target controllable vehicle;
[0096] like Figure 2 As shown, the area in front of any target controllable vehicle includes six sub-areas of the same size: the first straight lane sub-area 1, the second straight lane sub-area 4, the first left lane sub-area 0, the second left lane sub-area 3, the first right lane sub-area 2, and the second right lane sub-area 5.
[0097] In this sub-region, the length l is twice the speed of any target controllable vehicle. For example, if the speed is 10m / s, the length l is 20m. The length increases proportionally with the speed. The width w is the width of one lane. 0, 1, and 2 represent slow driving areas, and 3, 4, and 5 represent fast driving areas.
[0098] For example, energy intensity distribution refers to the energy intensity value of each point in the area ahead. This energy intensity value is based on the energy value generated by the vehicles around that point. The larger the value, the higher the risk of static collision and the risk of conflict between behavior strategy for any target controllable vehicle.
[0099] Step 103: Based on the environmental perception information and the energy intensity distribution in the area in front of any target controllable vehicle, determine the first feasible area of the target controllable vehicle;
[0100] For example, environmental perception information is used to determine whether the left lane, straight lane and right lane of any target controllable vehicle are passable. For passable lanes, energy intensity distribution is used to further determine passable sub-regions as the first feasible region, which includes one or more sub-regions.
[0101] Step 104: Based on the excitation value of the first feasible region of any target controllable vehicle, determine the second feasible region of any target controllable vehicle;
[0102] For example, if the number of subregions of the first feasible region is greater than 1, then a unique traversable subregion needs to be determined from them as the second feasible region.
[0103] Step 105: According to the preset coordination strategy, resolve the second feasible area of the controllable vehicles with conflict.
[0104] For example, when the second feasible areas of two controllable vehicles of multiple cooperative controllable vehicles overlap, it is necessary to coordinate and resolve such conflicts in accordance with a preset cooperative strategy.
[0105] This embodiment adopts a region-driven design concept, decomposes behavioral decisions into combinations of relative regions, and designs a collaborative driving energy distribution field and collaborative strategy to achieve efficient, generalized, and safe multi-vehicle collaborative decision-making. It can be applied to complex road scenarios with dynamic changes and numerous elements in traffic scenarios.
[0106] like Figure 3 As shown, step 102, which determines the energy intensity distribution in the area in front of any target controllable vehicle based on information from multiple vehicles within the target area, includes:
[0107] Step A1: Determine the first static potential energy and dynamic potential energy generated by all uncontrollable vehicles within the target area at the target point in front of any target controllable vehicle.
[0108] Step A2: Based on the first static potential energy field and dynamic potential energy generated by all uncontrollable vehicles at the target point, determine the energy intensity of the uncontrollable vehicles at the target point;
[0109] Step A3: Determine the second static potential energy generated at the target point by all controllable vehicles within the target area;
[0110] Step A4: Based on the time delay distribution probability function and the second static potential energy generated by all controllable vehicles at the target point, determine the energy intensity of the controllable vehicles at the target point;
[0111] Step A5: Determine the energy intensity of the target point based on the energy intensity of the uncontrollable vehicles and the energy intensity of the controllable vehicles at the target point.
[0112] In step A1 above, for uncontrollable vehicles, the risk arising from the conflict between static collision risk and behavioral strategy is quantified as continuous energy in a local area; specifically including:
[0113] For a target point with two-dimensional coordinates (x, y) in front of any controllable target vehicle, the first static potential energy generated at that target point by any uncontrollable vehicle within the target area. for:
[0114]
[0115] Where G is a preset parameter, M aThe mass of the uncontrollable vehicle is given by k1, which is the first parameter and is usually set to 2; r is the distance between the uncontrollable vehicle and the target point; (x a ,y a () represents the two-dimensional coordinates of the centroid of the uncontrollable vehicle;
[0116] Dynamic potential energy generated by any uncontrollable vehicle at the target point for:
[0117]
[0118] Where k2 is the second parameter, k3 is the third parameter, and v r Let θ be the relative speed between the uncontrollable vehicle and any of the target controllable vehicles. a Let be the relative angle between the center of mass of any uncontrollable vehicle and the target point.
[0119] In step A3 above, for controllable vehicles, risk energy modeling is performed on their static collision risk, perception error, and control delay; the static potential energy field is based on the static potential energy model of uncontrollable vehicles, and spatial weighting of the perception risk is considered within the error region, including:
[0120] For a target point with two-dimensional coordinates (x, y) in front of any controllable vehicle, the second static potential energy generated by any controllable vehicle within the target area at that target point. for:
[0121]
[0122] Where G is a preset parameter, M b For the mass of the controllable vehicle, k1 is the first parameter, (x b ,y b Let S be the two-dimensional coordinates of the centroid of the controllable vehicle, and let S be the coordinates of the centroid of the vehicle. b ,y b Let S be a circular region with a preset radius, where (m,n) are the two-dimensional coordinates of any point within the circular region S; p(m,n) is the probability value.
[0123] p(m,n)=f(m,n,x b ,y b )
[0124] Where f is a two-dimensional normal distribution.
[0125] In step A4 above, the energy intensity of the controllable vehicles at the target point is determined based on the time delay distribution probability function and the second static potential energy generated by all controllable vehicles at the target point; including:
[0126] Treating the time delay as a Markov process, determine the probability function p of the time delay distribution.delay (t), where t is time;
[0127] The energy intensity of any controllable vehicle at the target point is obtained by performing a time-diffusion-weighted calculation on the second static potential energy generated by the vehicle at the target point.
[0128]
[0129] Among them, v b Let θ be the speed of any controllable vehicle. b Let be the relative angle between the center of mass of any controllable vehicle and the target point;
[0130] The sum of the energy intensities of all controllable vehicles at the target point is taken as the energy intensity of the controllable vehicles at the target point.
[0131] In the above steps, the time delay is considered as a Markov process, meaning that given the initial time delay, the probability of the delay unit decreasing or increasing at the next time step is the same. Based on this assumption, statistical simulation of the time delay yields the final time delay distribution as follows: Figure 4 As shown, the time delay distribution probability function p is obtained from this. delay (t). Based on the time-delay distribution, the spatially weighted energy distribution is further diffuse-weighted in time. When a time delay occurs, the vehicle will continue to travel at the speed of the previous moment until a new command is sent, thus the time scale can be converted into a spatial scale.
[0132] In this embodiment, for any target point with two-dimensional coordinates (x, y) in front of any controllable vehicle, different calculation methods are used to calculate the energy intensity generated by the uncontrollable vehicle at that point and the energy intensity generated by the controllable vehicle at that point, resulting in more objective and accurate calculation results.
[0133] like Figure 5 As shown, step 103, which determines the first feasible area of any target controllable vehicle based on the environmental perception information and the energy intensity distribution in the area in front of the target controllable vehicle, includes:
[0134] Step B1: Based on the environmental perception information, determine the passage status of the straight lane, left lane and right lane of any target controllable vehicle, thereby determining the sub-area of the first state, where the first state is passable;
[0135] For example, when the left lane is impassable, it is also impassable in sub-areas 0 and 3.
[0136] Step B2: Based on the energy intensity distribution in the area in front of any of the target controllable vehicles, calculate the average energy intensity R of the sub-region in the first state:
[0137]
[0138] Where region represents a sub-region, E(x,y) represents the energy intensity value at (x,y), and S region This represents the area of the sub-region.
[0139] Step B3: When the average energy intensity of the sub-region in the first state is greater than a preset threshold, update the first state of the sub-region to the second state; the second state is impassable.
[0140] Step B4: Combine all the sub-regions of the first state into the first feasible region of the target controllable vehicle.
[0141] In some embodiments, determining a second feasible region for any target controllable vehicle based on an incentive value for a first feasible region of that target controllable vehicle includes:
[0142] When the first feasible region is the first straight lane sub-region, the first left lane sub-region, or the first right lane sub-region, the grid length parameter E is determined to be the length of the sub-region; the average energy intensity R is determined to be the average energy intensity of the sub-region.
[0143] When the first feasible region is the second straight lane sub-region, the second left lane sub-region, or the second right lane sub-region, the grid length parameter E is determined to be twice the length of the sub-region; the average energy intensity R is determined to be the average of the energy intensity of the sub-region and its adjacent sub-regions along the lane line direction.
[0144] When the first feasible area is a left lane sub-area, a first right lane sub-area, a second left lane sub-area, or a second right lane sub-area, the lane change penalty item C is determined to be a preset parameter greater than 0; when the first feasible area is a straight lane sub-area or a second straight lane sub-area, the lane change penalty item C is determined to be 0.
[0145] Calculate the incentive function reward for the first feasible region:
[0146] reward = w e Ew r Rw c C
[0147] Among them, w e As the first weight, w r As the second weight, w c It is the third weight;
[0148] The sub-region corresponding to the highest excitation value in the first feasible region of any target controllable vehicle is determined as the second feasible region.
[0149] In some embodiments, the second feasible area of the conflicting controllable vehicles is resolved according to a preset coordination strategy, including:
[0150] When it is detected that the second feasible areas of two controllable vehicles overlap, the two controllable vehicles include the first controllable vehicle and the second controllable vehicle, and the following processing is performed:
[0151] If the second feasible area of the first controllable vehicle is located in the straight lane of the first controllable vehicle, then the second feasible area of the second controllable vehicle is set to impassable.
[0152] If the second feasible area of the first controllable vehicle is located in the left lane of the first controllable vehicle and the second feasible area of the second controllable vehicle is located in the right lane of the second controllable vehicle, then the second feasible area of the second controllable vehicle is set to impassable.
[0153] If the second feasible area of the second controllable vehicle is located in the straight lane of the second controllable vehicle, then the second feasible area of the first controllable vehicle is set to impassable.
[0154] If the second feasible area of the second controllable vehicle is located in the left lane of the first controllable vehicle and the second feasible area of the first controllable vehicle is located in the right lane of the second controllable vehicle, then the second feasible area of the first controllable vehicle is set to impassable.
[0155] For example, Figure 6 The circled area in the upper middle diagram represents the second feasible zone for two controllable vehicles, where a collision occurred. Following the strategy of going straight > changing lanes left > changing lanes right, as follows... Figure 6 As shown in the figure below, the second feasible area for vehicles changing lanes to the left is modified from sub-area 3 to sub-area 1.
[0156] It should be noted that if the two controllable vehicles in conflict are controllable vehicle A and controllable vehicle B, and the overlapping second feasible region is assigned to controllable vehicle A, then for controllable vehicle B, a new second feasible region needs to be found, for example, the sub-region with the second largest excitation function can be used as the second feasible region.
[0157] In addition, the method also includes:
[0158] When a controllable vehicle has no second feasible area, meaning all areas ahead are closed to traffic, and if the current lane is five lanes, it can extend laterally into the area ahead. For example... Figure 7 As shown, a sub-region is extended to the left of sub-region 0 and to the right of sub-region 2. The same method is used to determine whether these two extended sub-regions are passable.
[0159] The specific implementation process of this application will be described below using a specific application scenario.
[0160] By iterating through the initial situation of 1-4 controllable vehicles and 1-2 uncontrollable vehicles, 19 test conditions are obtained, which can be divided into two categories: the convoy formation is disrupted by the fast vehicles behind, and the convoy formation overtakes the slow vehicles in front.
[0161] The platform selected for testing was SUMO, and the comparison methods used in the tests were: single-vehicle intelligence based on single-vehicle finite state machine, single-vehicle intelligence based on single-vehicle minimum action, IDM collaborative method, and fixed queue collaborative method.
[0162] The comparison results are shown in the table below. Compared with single-vehicle intelligence, IDM collaboration, and fixed queues, the method in this embodiment significantly reduces the total time (16%-19%) and greatly enhances generalization (variance decreases by 58%). In comparison, although this method is slightly slower than the baseline optimal method in some scenarios, it is always less than 1 second slower than the optimal benchmark model. All benchmark methods exhibit extremely slow behavior in certain scenarios, resulting in the overall efficiency of this method being far superior. This method enables the lead vehicle to pass through the intersection close to the tail vehicle, avoiding "falling behind"; no collisions occur in any scenario, maintaining efficiency while ensuring safety (with a margin of safety). It should be noted that this method is not programmed for specific scenarios to ensure generalization, which is the core reason for the smaller travel time between different scenarios, verifying the feasibility of changing from task-driven to region-driven approaches.
[0163] The comparative data are shown in Table 1:
[0164] Table 1
[0165]
[0166] Based on the same inventive concept, embodiments of this application provide a multi-vehicle homogeneous collaborative behavior decision-making device, such as... Figure 8 As shown, the multi-vehicle homogeneous collaborative behavior decision-making device includes:
[0167] The acquisition unit 201 is used to acquire information and environmental perception information of multiple vehicles within the target area; the vehicles include controllable vehicles and uncontrollable vehicles.
[0168] The first determining unit 202 is used to determine the energy intensity distribution in the area in front of any target controllable vehicle based on information about multiple vehicles within the target area.
[0169] The second determining unit 203 is used to determine the first feasible area of the target controllable vehicle based on the environmental perception information and the energy intensity distribution of the area in front of the target controllable vehicle.
[0170] The third determining unit 204 is used to determine the second feasible region of any target controllable vehicle based on the excitation value of the first feasible region of any target controllable vehicle;
[0171] The collaborative behavior decision-making unit 205 is used to resolve the second feasible area of controllable vehicles with conflicts according to a preset collaborative strategy.
[0172] It should be noted that the principle of the multi-vehicle homogeneous collaborative behavior decision-making device 200 provided in this application embodiment to solve the technical problem is similar to the method provided in this application embodiment. Therefore, the implementation of the multi-vehicle homogeneous collaborative behavior decision-making device 200 provided in this application embodiment can refer to the implementation of the method provided in this application embodiment, and the repeated parts will not be described again.
[0173] Based on the same inventive concept, embodiments of this application also provide an electronic device, such as... Figure 9 As shown, it includes: a memory and a processor, wherein the memory stores an executable program, and the processor executes the executable program to implement the steps of the multi-vehicle homogeneous cooperative behavior decision-making method as described above.
[0174] The aforementioned processor can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0175] The aforementioned memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0176] This application embodiment also provides a storage medium carrying one or more computer programs, which, when executed by a processor, implement the steps of the multi-vehicle homogeneous cooperative behavior decision-making method described above.
[0177] The storage medium in this embodiment may be included in an electronic device / system; or it may exist independently and not be assembled into an electronic device / system. The storage medium carries one or more programs, which, when executed, implement the multi-vehicle homogeneous cooperative behavior decision-making method according to the embodiments of this application.
[0178] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0179] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.
Claims
1. A multi-vehicle homogeneous cooperative behavior decision-making method, characterized in that, include: Acquire information and environmental perception information of multiple vehicles within a target area; the vehicles include controllable and uncontrollable vehicles. Based on information about multiple vehicles within the target area, determine the energy intensity distribution in the area in front of any target controllable vehicle; Based on the environmental perception information and the energy intensity distribution in the area in front of any target controllable vehicle, a first feasible area for the target controllable vehicle is determined; Based on the incentive value of the first feasible region of any target controllable vehicle, determine the second feasible region of any target controllable vehicle; According to the preset coordination strategy, the second feasible area of controllable vehicles with conflicts is resolved.
2. The multi-vehicle homogeneous cooperative behavior decision-making method according to claim 1, characterized in that, Vehicle information includes: vehicle location and speed; Based on information from multiple vehicles within the target area, determine the energy intensity distribution in the area in front of any controllable target vehicle, including: Determine the first static potential energy and dynamic potential energy generated by all uncontrollable vehicles within the target area at the target point in the area in front of any target controllable vehicle. Based on the first static potential energy field and dynamic potential energy generated by all uncontrollable vehicles at the target point, the energy intensity of the uncontrollable vehicles at the target point is determined. Determine the second static potential energy generated at the target point by all controllable vehicles within the target area; Based on the time delay distribution probability function and the second static potential energy generated by all controllable vehicles at the target point, the energy intensity of the controllable vehicles at the target point is determined. The energy intensity of the target point is determined based on the energy intensity of the uncontrollable vehicles and the energy intensity of the controllable vehicles at the target point.
3. The multi-vehicle homogeneous cooperative behavior decision-making method according to claim 2, characterized in that, Determine the first static potential energy and dynamic potential energy generated by all uncontrollable vehicles within the target area at a target point in the area in front of any target controllable vehicle, including: For a target point with two-dimensional coordinates (x, y) in front of any controllable target vehicle, the first static potential energy generated at that target point by any uncontrollable vehicle within the target area. for: Where G is a preset parameter, M a Let k1 be the mass of the uncontrollable vehicle, k1 be the first parameter, and r be the distance between the uncontrollable vehicle and the target point; (x a ,y a () represents the two-dimensional coordinates of the centroid of the uncontrollable vehicle; The dynamic potential energy generated by the uncontrollable vehicle at the target point for: Where k2 is the second parameter, k3 is the third parameter, and v r Let θ be the relative speed between the uncontrollable vehicle and any of the target controllable vehicles. a Let be the relative angle between the center of mass of any uncontrollable vehicle and the target point.
4. The multi-vehicle homogeneous cooperative behavior decision-making method according to claim 2, characterized in that, Determine the second static potential energy generated at the target point by all controllable vehicles within the target area, including: For a target point with two-dimensional coordinates (x, y) in front of any controllable vehicle, the second static potential energy generated by any controllable vehicle within the target area at that target point. for: Where G is a preset parameter, M b For the mass of the controllable vehicle, k1 is the first parameter, (x b ,y b Let S be the two-dimensional coordinates of the centroid of the controllable vehicle, and let S be the coordinates of the centroid of the vehicle. b ,y b S is a circular region with a preset radius, where (m,n) are the two-dimensional coordinates of any point within the circular region S; p(m,n) is the probability value. p(m,n)=f(m,n,x b ,y b ) Where f is a two-dimensional normal distribution.
5. The multi-vehicle homogeneous cooperative behavior decision-making method according to claim 4, characterized in that, Based on the time delay distribution probability function and the second static potential energy generated by all controllable vehicles at the target point, the energy intensity of the controllable vehicles at the target point is determined; including: Treating the time delay as a Markov process, determine the probability function p of the time delay distribution. delay (t), where t is time; The energy intensity of any controllable vehicle at the target point is obtained by performing a time-diffusion-weighted calculation on the second static potential energy generated by the vehicle at the target point. Among them, v b Let θ be the speed of any controllable vehicle. b Let be the relative angle between the center of mass of any controllable vehicle and the target point; The sum of the energy intensities of all controllable vehicles at the target point is taken as the energy intensity of the controllable vehicles at the target point.
6. The multi-vehicle homogeneous cooperative behavior decision-making method according to claim 2, characterized in that, The environmental perception information includes lane line type, traffic light phase, and traffic identifier; The area in front of any target controllable vehicle includes six sub-areas of the same size: a first straight lane sub-area, a second straight lane sub-area, a first left lane sub-area, a second left lane sub-area, a first right lane sub-area, and a second right lane sub-area. Based on the environmental perception information and the energy intensity distribution in the area in front of any target controllable vehicle, a first feasible area for the target controllable vehicle is determined, including: Based on the environmental perception information, the passage status of the straight lane, left lane and right lane of any target controllable vehicle is determined respectively, thereby determining the sub-area of the first state, where the first state is passable; Based on the energy intensity distribution in the area in front of any of the target controllable vehicles, calculate the average energy intensity of the sub-region in the first state; When the average energy intensity of a sub-region in the first state is greater than a preset threshold, the first state of the sub-region is updated to the second state; the second state is impassable. All sub-regions of the first state are combined to form the first feasible region of the target controllable vehicle.
7. The multi-vehicle homogeneous cooperative behavior decision-making method according to claim 6, characterized in that, Based on the incentive value of the first feasible region of any target controllable vehicle, determine the second feasible region of the target controllable vehicle, including: When the first feasible region is the first straight lane sub-region, the first left lane sub-region, or the first right lane sub-region, the grid length parameter E is determined to be the length of the sub-region; the average energy intensity R is determined to be the average energy intensity of the sub-region. When the first feasible region is the second straight lane sub-region, the second left lane sub-region, or the second right lane sub-region, the grid length parameter E is determined to be twice the length of the sub-region; the average energy intensity R is determined to be the average of the energy intensity of the sub-region and its adjacent sub-regions along the lane line direction. When the first feasible area is a left lane sub-area, a first right lane sub-area, a second left lane sub-area, or a second right lane sub-area, the lane change penalty item C is determined to be a preset parameter greater than 0; when the first feasible area is a straight lane sub-area or a second straight lane sub-area, the lane change penalty item C is determined to be 0. Calculate the incentive function reward for the first feasible region: reward=w e E-w r R-w c C Among them, w e As the first weight, w r As the second weight, w c It is the third weight; The sub-region corresponding to the highest excitation value in the first feasible region of any target controllable vehicle is determined as the second feasible region.
8. The multi-vehicle homogeneous cooperative behavior decision-making method according to claim 2, characterized in that, According to the preset coordination strategy, the second feasible area of controllable vehicles with conflicts is resolved, including: When it is detected that the second feasible areas of two controllable vehicles overlap, the two controllable vehicles include the first controllable vehicle and the second controllable vehicle, and the following processing is performed: If the second feasible area of the first controllable vehicle is located in the straight lane of the first controllable vehicle, then the second feasible area of the second controllable vehicle is set to impassable. If the second feasible area of the first controllable vehicle is located in the left lane of the first controllable vehicle and the second feasible area of the second controllable vehicle is located in the right lane of the second controllable vehicle, then the second feasible area of the second controllable vehicle is set to impassable. If the second feasible area of the second controllable vehicle is located in the straight lane of the second controllable vehicle, then the second feasible area of the first controllable vehicle is set to impassable. If the second feasible area of the second controllable vehicle is located in the left lane of the first controllable vehicle and the second feasible area of the first controllable vehicle is located in the right lane of the second controllable vehicle, then the second feasible area of the first controllable vehicle is set to impassable.
9. A multi-vehicle homogeneous collaborative behavior decision-making device, characterized in that, include: The acquisition unit is used to acquire information and environmental perception information of multiple vehicles within a target area; the vehicles include controllable vehicles and uncontrollable vehicles. The first determining unit is used to determine the energy intensity distribution in the area in front of any target controllable vehicle based on information about multiple vehicles within the target area. The second determining unit is used to determine the first feasible area of the target controllable vehicle based on the environmental perception information and the energy intensity distribution in the area in front of the target controllable vehicle. The third determining unit is used to determine the second feasible region of any target controllable vehicle based on the excitation value of the first feasible region of any target controllable vehicle; The collaborative behavior decision-making unit is used to resolve the second feasible area of controllable vehicles with conflicts according to a preset collaborative strategy.
10. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores an executable program, and the processor executes the executable program to implement the steps of the method as claimed in any one of claims 1 to 8.
Citation Information
Patent Citations
Method for multiple vehicles to cooperatively and quickly pass through road bottle neck
CN103956066A
Autonomous decision-making method for conflict risk-avoidance of cooperative fleet at frequently occurring bottleneck sections of expressway
CN110853335A
Vehicle anthropomorphic decision control method and device, vehicle and storage medium
CN115923833A
Systems and methods for using an attention buffer to improve resource allocation management
WO2018035533A2
Method and apparatus for coordinating multiple cooperative vehicle trajectories on shared road networks
WO2022040748A1