Spacetime game-based collaborative decision-making method and device, equipment and medium

By constructing a multidimensional constraint matrix and generating the target path policy magnitude, the problem of traffic congestion caused by vehicles on static paths is solved, real-time path updates are achieved, and traffic safety and energy efficiency are improved.

CN120748235BActive Publication Date: 2025-11-18PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511188487.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-18
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

In traffic scenarios with high traffic volume, existing technologies lack real-time traffic analysis, causing vehicles to travel along static paths, which can easily lead to traffic congestion, affecting driving safety and travel efficiency. Furthermore, existing collaborative decision-making technologies cannot effectively integrate multi-source heterogeneous data, resulting in conflicts between global and local objectives.

Method used

By acquiring the set of environmental parameters uploaded by vehicles, a multidimensional constraint matrix is ​​constructed to generate the target path strategy magnitude. Under the guidance of the multidimensional constraint matrix, vehicles update their path strategies in real time and generate the path for the next moment by combining the real-time environmental status, so as to avoid or mitigate traffic congestion and save energy.

Benefits of technology

It effectively avoids or mitigates traffic congestion, improves vehicle driving safety and travel efficiency, and maximizes energy conservation, thus solving the congestion problem caused by static paths in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748235B_ABST
    Figure CN120748235B_ABST
Patent Text Reader

Abstract

The application discloses a spatio-temporal game-based cooperative decision method and device, equipment and medium. The road topological information, vehicle driving state information and road emergency information uploaded by a target vehicle are converted to generate a path distance threshold and a dynamic obstacle probability distribution, and a multi-dimensional constraint matrix is constructed in combination with the path distance threshold and the dynamic obstacle probability distribution; taking maximizing the driving area covered and minimizing the total distance of the avoidance driving of the plurality of vehicles in the driving area as an optimization target, the target path strategy amplitude is simulated and calculated in combination with the multi-dimensional constraint matrix and preset posterior information; and the target path strategy amplitude and the multi-dimensional constraint matrix are sent to the target vehicle, so that the target vehicle generates a path strategy at the next moment based on the current path strategy, the real-time collected dynamic environment state and the target path strategy amplitude under the indication of the multi-dimensional constraint matrix. In this way, the driving safety and travel efficiency of the vehicle are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a collaborative decision-making method, apparatus, device and medium based on spatiotemporal game theory. Background Technology

[0002] Traffic congestion is common in busy urban areas, where a large number of vehicles pass through, which can easily lead to traffic jams or accidents.

[0003] For example, taking vehicle planning technology as an example, related technologies plan routes for each vehicle. However, this route planning is generally a static route planning between the starting point and the destination, lacking real-time traffic analysis. As a result, the planned vehicle routes are static routes. Therefore, each vehicle travels according to its own static route, which can easily cause multiple vehicles to converge on the same road segment at the same time, causing traffic congestion on that road segment. Or, due to the existence of unknown sudden accidents, traffic congestion may also occur, thereby affecting vehicle driving safety and travel efficiency. Summary of the Invention

[0004] This application provides a collaborative decision-making method, apparatus, device, and medium based on spatiotemporal game theory. It can construct a multidimensional constraint matrix by combining the set of environmental parameters uploaded by the vehicle and generate the target path strategy magnitude. Under the guidance of the multidimensional constraint matrix, the vehicle updates its future path strategy in real time based on the target path strategy magnitude to avoid traffic congestion, maximize energy saving, and improve vehicle driving safety and travel efficiency.

[0005] Firstly, this application provides a collaborative decision-making method based on spatiotemporal game theory, including:

[0006] Obtain the set of environmental parameters uploaded by the target vehicle, wherein the set of environmental parameters includes at least road topology information, vehicle driving status information, and road emergency information;

[0007] The road topology information, vehicle driving status information, and road emergency information are transformed into path distance thresholds and dynamic obstacle probability distributions. A multidimensional constraint matrix is ​​constructed by combining the path distance thresholds and the dynamic obstacle probability distributions. The multidimensional constraint matrix is ​​used to limit the drivable area at different times.

[0008] The optimization objective is to maximize the coverage of the driving area and minimize the total avoidance distance of multiple vehicles within the driving area. The target path strategy magnitude is simulated and calculated by combining the multidimensional constraint matrix and the preset posterior information.

[0009] The target path strategy magnitude and the multidimensional constraint matrix are sent to the target vehicle, so that the target vehicle generates a path strategy for the next moment based on the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude under the guidance of the multidimensional constraint matrix.

[0010] Secondly, this application provides a collaborative decision-making device based on spatiotemporal game theory, comprising:

[0011] The acquisition unit is used to acquire a set of environmental parameters uploaded by the target vehicle, wherein the set of environmental parameters includes at least road topology information, vehicle driving status information and road emergency information.

[0012] The generation unit is used to convert the road topology information, vehicle driving status information and road emergency information into path distance thresholds and dynamic obstacle probability distributions, and to construct a multi-dimensional constraint matrix by combining the path distance thresholds and the dynamic obstacle probability distributions. The multi-dimensional constraint matrix is ​​used to limit the drivable area corresponding to different times.

[0013] The calculation unit is used to simulate and calculate the target path strategy magnitude by taking the maximization of the covered driving area and the minimization of the total avoidance driving distance of multiple vehicles within the driving area as the optimization objective, and combining the multidimensional constraint matrix and preset posterior information.

[0014] The sending unit is used to send the target path strategy magnitude and the multidimensional constraint matrix to the target vehicle, so that the target vehicle generates a path strategy for the next moment based on the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude under the guidance of the multidimensional constraint matrix.

[0015] In some embodiments, the generating unit is further configured to:

[0016] Determine the shortest path distance based on the road topology information;

[0017] Determine the target distance between the vehicle location of the target vehicle and the event location contained in the road emergency information, and construct a distance influence parameter based on the ratio between the target distance and the influence radius parameter contained in the road emergency information. Generate a path distance threshold based on the shortest path distance and the distance influence parameter.

[0018] Obtain the current reference driving state information of reference vehicles around the target vehicle, calculate the mean of the predicted position of the reference vehicle at the target prediction time based on the reference driving state information and the vehicle driving state information, and calculate the covariance at the target prediction time based on the preset error coefficient and the preset diffusion coefficient.

[0019] The predicted location mean and covariance are quantified using a normal distribution method to generate a dynamic obstacle probability distribution.

[0020] In some embodiments, the computing unit is further configured to:

[0021] The target feasible region at the current moment is determined based on the multidimensional constraint matrix.

[0022] Obtain a global utility function, which is used to characterize the optimization objective as maximizing the covered driving area and minimizing the total avoidance distance of multiple vehicles within the driving area.

[0023] The policy response relationship of the target vehicle relative to the path policy magnitude issued by the local node is obtained, and a local utility function is constructed based on the preset posterior information and the policy response relationship. The local utility function is used to characterize the relationship between the path policy magnitude issued by the local node and the simulated theoretical path of the target vehicle. The simulated theoretical path is obtained by the local node through simulating the response of the target vehicle to the issued path policy.

[0024] By combining the global utility function and the local utility function, the target path strategy magnitude that conforms to the target feasible region is simulated and calculated.

[0025] In some embodiments, the spatiotemporal game-based collaborative decision-making device further includes a strategy update unit, used for:

[0026] Obtain the environmental status parameters uploaded by the target vehicle;

[0027] When the environmental state parameter is greater than or equal to the preset environmental state threshold, the target path strategy magnitude is updated to obtain the updated optimized strategy magnitude.

[0028] The optimization strategy magnitude is sent to the target vehicle, so that the target vehicle generates a path strategy for the next moment based on the current path strategy, the real-time collected dynamic environment state, and the optimization strategy magnitude under the guidance of the multidimensional constraint matrix.

[0029] In some implementations, the policy update unit is further configured to:

[0030] For the target path policy magnitude, construct the corresponding global utility expectation gradient and the corresponding policy magnitude change regularization term;

[0031] By combining the expected global utility gradient with the regularization term of the policy magnitude change, a policy magnitude update function is constructed.

[0032] With the optimization objective of maximizing the value of the policy magnitude update function, the policy magnitude of the target path is updated according to the policy magnitude update function to obtain the updated optimized policy magnitude.

[0033] In some implementations, the policy verification unit is further configured to:

[0034] The initial path strategy of the target vehicle at the current time is obtained, and multiple candidate path strategies corresponding to multiple consecutive time intervals are simulated and generated by combining the initial path strategy and the optimization strategy magnitude.

[0035] Based on the policy error between each candidate path policy and the initial path policy, the average policy error is determined;

[0036] The strategy update unit is further configured to send the optimized strategy magnitude to the target vehicle when the average strategy error is less than or equal to a preset strategy error threshold.

[0037] In some implementations, the policy update unit is further configured to:

[0038] When the average strategy error is greater than a preset strategy error threshold, the difference between the average strategy error and the preset strategy error threshold is determined.

[0039] Based on the gap value, adjust the learning rate parameter in the regularization term for the change in policy magnitude in the policy magnitude update function to obtain the updated target policy magnitude update function.

[0040] The optimization objective is to maximize the value of the target policy magnitude update function. The optimized policy magnitude is updated according to the target policy magnitude update function to obtain the updated target optimized policy magnitude.

[0041] The target optimization strategy magnitude is sent to the target vehicle.

[0042] Furthermore, this application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned collaborative decision-making method based on spatiotemporal game theory.

[0043] Furthermore, embodiments of this application also provide a computer-readable storage medium storing multiple instructions adapted for loading by a processor to execute the aforementioned collaborative decision-making method based on spatiotemporal game theory.

[0044] This application embodiment obtains a set of environmental parameters uploaded by the target vehicle. The set of environmental parameters includes at least road topology information, vehicle driving status information, and road emergency information. The road topology information, vehicle driving status information, and road emergency information are transformed into path distance thresholds and dynamic obstacle probability distributions. A multi-dimensional constraint matrix is ​​constructed by combining the path distance thresholds and dynamic obstacle probability distributions. The multi-dimensional constraint matrix is ​​used to limit the drivable area corresponding to different times. The optimization objective is to maximize the covered drivable area and minimize the total avoidance distance of multiple vehicles within the drivable area. The target path strategy magnitude is simulated and calculated by combining the multi-dimensional constraint matrix and preset posterior information. The target path strategy magnitude and the multi-dimensional constraint matrix are sent to the target vehicle so that the target vehicle can generate a path strategy for the next time moment based on the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude under the guidance of the multi-dimensional constraint matrix.

[0045] As can be seen from the above, this application receives road topology information, vehicle driving status information, and road emergency information uploaded by the target vehicle, and transforms the above information into path distance thresholds and dynamic obstacle probability distributions. Then, it combines the path distance thresholds and dynamic obstacle probability distributions to generate a multi-dimensional constraint matrix. This multi-dimensional constraint matrix is ​​used to limit the drivable area at different times. Thus, constraints can be generated by combining real-time environmental parameters. Furthermore, combining the multi-dimensional constraint matrix and preset posterior information, the optimization objective is to maximize the covered drivable area while minimizing the total avoidance distance of multiple vehicles within the drivable area. The target path strategy magnitude is simulated and calculated. Finally, the target path strategy magnitude and the multi-dimensional constraint matrix are... The multidimensional constraint matrix is ​​sent to the target vehicle, enabling the target vehicle to generate a path strategy for the next moment within the drivable area defined by the multidimensional constraint matrix, combining the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude. This is to cope with traffic congestion. In this way, compared with the related technology where vehicles only travel according to static paths, resulting in traffic congestion, this application can construct a multidimensional constraint matrix by combining the environmental parameter set uploaded by the vehicle and generate a target path strategy magnitude. This allows the vehicle to update its future path strategy in real time under the guidance of the multidimensional constraint matrix and the target path strategy magnitude, so as to avoid or mitigate traffic congestion while maximizing energy saving of multiple vehicles, improving vehicle driving safety and travel efficiency. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 A schematic diagram of a collaborative decision-making system based on spatiotemporal game theory provided in an embodiment of this application;

[0048] Figure 2 A flowchart illustrating the steps of the collaborative decision-making method based on spatiotemporal game theory provided in this application embodiment;

[0049] Figure 3 An example diagram of a collaborative decision-making scenario architecture based on spatiotemporal game theory provided in this application embodiment;

[0050] Figure 4 A schematic diagram of the structure of the collaborative decision-making device based on spatiotemporal game theory provided in the embodiments of this application;

[0051] Figure 5 This is a schematic diagram of the terminal structure provided in the embodiments of this application;

[0052] Figure 6 This is a schematic diagram of the server structure provided in an embodiment of this application. Detailed Implementation

[0053] To enable those skilled in the art to better understand the solutions of this application, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0054] It is understood that in the specific implementation of this application, relevant data such as road topology information, vehicle driving status information, road emergency information, target path strategy range, current path strategy, or path strategy at the next moment are involved. When the above embodiments of this application are applied to specific products or technologies, permission or consent from the target is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards.

[0055] Furthermore, when this application embodiment needs to obtain relevant data, it will obtain separate permission or separate consent for relevant data such as road topology information, vehicle driving status information, road emergency information, target path strategy range, current path strategy, or path strategy at the next moment through pop-up windows or redirection to a confirmation page. Only after clearly obtaining separate permission or separate consent for relevant data such as road topology information, vehicle driving status information, road emergency information, target path strategy range, current path strategy, or path strategy at the next moment will it obtain the necessary data for this application embodiment to operate normally.

[0056] It should be noted that while some processes described in the specification, claims, and accompanying drawings contain multiple steps that appear in a specific order, it should be clearly understood that these steps may not be performed in the order they appear herein, or may be performed in parallel. The step numbers are merely used to distinguish different steps and do not represent any particular order of execution. Furthermore, descriptions such as "first," "second," or "objective" in this document are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0057] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0058] This application provides a collaborative decision-making method, apparatus, device, and medium based on spatiotemporal game theory. Specifically, the collaborative decision-making method based on spatiotemporal game theory in this application can be implemented in a computer device, which can be a server or a user terminal device. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The user terminal device can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, smart home appliance, vehicle terminal, smart voice interaction device, aircraft, drone, etc., but is not limited to these.

[0059] For ease of understanding, this application will describe the implementation process of the collaborative decision-making method based on spatiotemporal game theory through several embodiments, as follows:

[0060] Traffic congestion is common in busy urban areas, where a large number of vehicles pass through, which can easily lead to traffic jams or accidents.

[0061] For example, taking vehicle planning technology as an example, related technologies plan routes for each vehicle. However, this route planning is generally a static route planning between the origin and destination, lacking real-time traffic analysis. This results in static vehicle routes, where each vehicle travels along its own static route, easily causing multiple vehicles to converge on the same road segment, leading to traffic congestion. Unexpected emergencies can also cause traffic congestion, thus affecting vehicle safety and travel efficiency. It should be noted that the inventors have found that some existing collaborative decision-making technologies are not applicable to dynamic environmental data and generally rely on black-box models for decision-making, resulting in a lack of transparency and reliable decision-making basis. Furthermore, they struggle to integrate multi-source heterogeneous data (such as road topology information, emergency event information, and vehicle driving status information), leading to conflicts between global and local objectives. For example, the global objective might be for multiple vehicles to collaboratively adjust their route strategies to avoid congestion with minimal global route strategy adjustments, while the local objective can be understood as a single vehicle conserving energy (reducing the magnitude of route strategy adjustments) centered on itself. Therefore, a conflict arises between global and local objectives.

[0062] To address the above issues, this application provides a collaborative decision-making method based on spatiotemporal game theory. This method primarily combines a set of environmental parameters uploaded by the vehicle to construct a multidimensional constraint matrix and generate a target path strategy magnitude. This allows the vehicle to update its future path strategy in real time, guided by the multidimensional constraint matrix and the target path strategy magnitude. Specifically, the method involves obtaining a set of environmental parameters uploaded by the target vehicle, including at least road topology information, vehicle driving status information, and road emergency information. This information is then converted into a path distance threshold and a dynamic obstacle probability distribution. A multidimensional constraint matrix is ​​constructed using this matrix to define the drivable area at different times. The optimization objective is to maximize the covered drivable area while minimizing the total avoidance distance for multiple vehicles within that area. The target path strategy magnitude is simulated and calculated using the multidimensional constraint matrix and pre-set posterior information. Finally, the target path strategy magnitude and the multidimensional constraint matrix are sent to the target vehicle, enabling it to generate its next-time path strategy based on its current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude, guided by the multidimensional constraint matrix. This approach aims to maximize energy savings across multiple vehicles while avoiding or mitigating traffic congestion, thereby improving driving safety and travel efficiency. Please refer to the specific implementation examples below for details.

[0063] It should be noted that this collaborative decision-making method based on spatiotemporal game theory can be executed jointly by the terminal and the server.

[0064] For example, taking a collaborative decision-making method based on spatiotemporal game theory, jointly executed by the terminal and the server, as an example, see [link to relevant documentation]. Figure 1 This is a schematic diagram of an information push system provided in an embodiment of this application. The system includes a terminal 110 and a server 120.

[0065] Each terminal 110 can have a target application installed, and the corresponding application business can be run through the target application. The target application can be called a client. Taking the terminal 110 as a target vehicle as an example, the target vehicle can collect a set of environmental parameters. The set of environmental parameters includes the road topology information of the road segment where the target vehicle is located, the vehicle driving status information, and the road emergency information. The client on the terminal 110 can upload the set of environmental parameters to the server 120.

[0066] Server 120 can be a single service node, a distributed system composed of multiple service nodes, or a service node within a distributed system. For example, server 120 can be a cloud service, an edge node, etc., without limitation.

[0067] Specifically, the server 120 executes the steps of a collaborative decision-making method based on spatiotemporal game theory. Specifically, it acquires a set of environmental parameters uploaded by the target vehicle, including at least road topology information, vehicle driving status information, and road emergency information. It then transforms the road topology information, vehicle driving status information, and road emergency information into path distance thresholds and dynamic obstacle probability distributions. A multidimensional constraint matrix is ​​constructed by combining the path distance thresholds and dynamic obstacle probability distributions, which is used to limit the drivable area at different times. The optimization objective is to maximize the covered drivable area while minimizing the total avoidance distance of multiple vehicles within the drivable area. The target path strategy magnitude is simulated and calculated by combining the multidimensional constraint matrix and preset posterior information. Furthermore, the server 120 sends the target path strategy magnitude and the multidimensional constraint matrix to the target vehicle, i.e., the server 120 sends the target path strategy magnitude and the multidimensional constraint matrix to the terminal 110. Subsequently, under the guidance of the multidimensional constraint matrix, terminal 110 (target vehicle) can generate the path strategy for the next moment based on the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude. This allows the target vehicle to generate the path strategy for the next moment within the drivable area defined by the multidimensional constraint matrix, by combining the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude.

[0068] Therefore, by receiving road topology information, vehicle driving status information, and road emergency information uploaded by the target vehicle, and transforming this information into path distance thresholds and dynamic obstacle probability distributions, a multi-dimensional constraint matrix is ​​generated by combining the path distance thresholds and dynamic obstacle probability distributions. This multi-dimensional constraint matrix is ​​used to limit the drivable area at different times. Constraint conditions can then be generated by combining real-time environmental parameters. Furthermore, by combining the multi-dimensional constraint matrix and preset posterior information, the optimization objective is to maximize the covered drivable area while minimizing the total avoidance distance of multiple vehicles within that area. The target path strategy magnitude is simulated and calculated. Finally, the target path strategy magnitude and the multi-dimensional constraint matrix are sent... The data is sent to the target vehicle so that the target vehicle, within the drivable area defined by the multidimensional constraint matrix, can generate a path strategy for the next moment by combining the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude, in order to cope with traffic congestion. Thus, compared with the related technology where vehicles only travel according to static paths, resulting in traffic congestion, this application can construct a multidimensional constraint matrix by combining the set of environmental parameters uploaded by the vehicle and generate a target path strategy magnitude. This allows the vehicle to update its future path strategy in real time under the guidance of the multidimensional constraint matrix and the target path strategy magnitude, thereby avoiding or mitigating traffic congestion while maximizing energy savings for multiple vehicles, improving vehicle driving safety and travel efficiency.

[0069] For ease of understanding, the steps of the spatiotemporal game-based collaborative decision-making method will be described in detail below. It should be noted that the order of the following embodiments is not intended to limit the preferred order of the embodiments.

[0070] See Figure 2 , Figure 2 The flowchart of the collaborative decision-making method based on spatiotemporal game theory provided in this application embodiment is shown below. In this application embodiment, the collaborative decision-making method based on spatiotemporal game theory can be executed by a computer device, such as by a server. The specific process is as follows.

[0071] 101. Obtain the set of environmental parameters uploaded by the target vehicle.

[0072] Traffic congestion is common in high-traffic areas, such as busy urban areas where heavy traffic can easily lead to congestion or accidents. Current route planning typically uses static routes, setting fixed paths from start to finish for all vehicles. When all vehicles follow these static paths, the convergence of traffic on the same road at the same time can cause congestion. Furthermore, when faced with congested roads or unforeseen emergencies, vehicles adhering to these fixed static paths can further exacerbate congestion, impacting driving safety and efficiency.

[0073] Furthermore, existing collaborative decision-making technologies generally employ black-box models (such as game tree search and Monte Carlo tree search) for decision-making. This decision-making process lacks transparency, making it impossible to obtain reliable decision-making basis. In addition, it is difficult to integrate multi-source heterogeneous data, leading to conflicts between global and local decisions. Thus, if only the optimal local decision (planning for a single vehicle) is considered, it is not conducive to the optimization of the global strategy (planning for multiple vehicles). Consequently, it is impossible to alleviate the current traffic congestion. This situation will not only affect vehicle driving safety and efficiency but also further increase the energy consumption of multiple vehicles.

[0074] To address the above issues, this application embodiment combines road topology information, vehicle driving status information, and road emergency information from the environmental parameter set uploaded by the target vehicle to generate a path distance threshold and a dynamic obstacle probability distribution. The path distance threshold is used to limit the maximum safe distance for the target vehicle to travel, and the dynamic obstacle probability distribution is used to describe the uncertainty range and probability of the location of unknown obstacles (such as surrounding vehicles, other emergencies, etc.). Logically, it can be used to describe where the obstacle is "most likely to appear" and "the range that may deviate from that location" in the future, thereby quantifying the uncertainty of obstacles in the future. Furthermore, a multidimensional constraint matrix is ​​constructed by combining path distance thresholds and dynamic obstacle probability distributions. This multidimensional constraint matrix is ​​used to define the drivable area at different times, such as the drivable area at future times. Then, the optimization objective is to maximize the area of ​​the drivable area covered by collaborative decision-making and minimize the total avoidance distance of multiple vehicles within the drivable area. Combining this multidimensional constraint matrix and preset posterior information, the target path strategy magnitude is simulated and calculated. The target path strategy magnitude and the multidimensional constraint matrix are then sent to the target vehicle, so that the target vehicle can generate the path strategy for the next time moment based on the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude under the guidance of the multidimensional constraint matrix. In this way, the target vehicle can generate the path strategy for the next time moment by combining the guidance of the multidimensional constraint matrix and the actual road conditions, so that it can drive according to the path strategy for the next time moment, thereby alleviating or avoiding traffic congestion, improving vehicle driving safety and efficiency, and reducing the energy consumption of multiple vehicles.

[0075] Therefore, in this embodiment of the application, in order to generate rules (i.e., a multidimensional constraint matrix) for constraining the path decision of vehicles, the server can receive a set of environmental parameters uploaded by the target vehicle. This set of environmental parameters includes at least road topology information, vehicle driving status information, and road emergency information. Based on this data, complete environmental information can be obtained and used as the basis for spatiotemporal constraint coding. The road topology information, vehicle driving status information, and road emergency information are then transformed into corresponding path distance thresholds and dynamic obstacle probability distributions, thereby generating a multidimensional constraint matrix for constraining the path decision of vehicles to guide the generation of path strategy magnitude in the subsequent game decision-making process.

[0076] The target vehicle can be any vehicle equipped with an onboard terminal to interact with the server, or it can interact with the server through communication terminals such as mobile phones and tablets to upload the acquired set of environmental parameters to the server.

[0077] Road topology information refers to the road structure information of the road where the target vehicle is located, describing the physical structure of the road. This includes features such as intersections, lanes, and road segment lengths. Based on road topology information, possible vehicle routes can be planned, along with the costs (such as distance and time). Without road topology information, route planning would lack a basis and would not reflect real-world traffic conditions; for example, vehicles might plan illegal routes such as "going through walls" or "driving against traffic." It should be noted that... Represents road topology information. , Represents a set of nodes (path intersections / lane points). This represents the set of edges of a road (i.e., the set of road segments or paths). This represents the weight of road segment length, where the road segment length weight is... This means assigning each road segment a calculable and comparable length or cost attribute for subsequent spatiotemporal constraint coding.

[0078] The vehicle driving status information can be understood as the driving status of the target vehicle, that is, the vehicle's behavioral state, which may include the target vehicle's position, speed, and direction of travel, etc. This represents vehicle driving status information, and the various pieces of information included in the vehicle driving status information can be represented as follows: , Indicates the target vehicle exist Location at any moment Indicates the target vehicle is in Speed ​​at any time Indicates the target vehicle is in The driving direction at any given time, i.e., the heading angle. Based on the vehicle's driving status information, the target vehicle's driving intention can be determined. Furthermore, by combining the driving status information of all vehicles, the driving status information of each vehicle can serve as the basis for predicting the obstacle trajectory of other vehicles.

[0079] The road emergency information can be any emergency on the road, such as accidents, construction, or sudden weather changes. By combining this information, temporary constraints on the driving area can be constructed, such as no-entry zones and speed-limited zones. Without this road emergency information, the constructed constraint matrix will be incomplete, leading to defects in the calculated path strategy and potentially causing delays in vehicle path adjustment, thus possibly resulting in vehicles entering prohibited or dangerous areas. Road emergency information can be described by event location, time, radius of influence, and duration. For ease of subsequent coding, [the following is used as an example]. This indicates information about road emergencies. Specifically, road emergency information can be represented as follows: , Indicates a road emergency Location, Indicates the time of occurrence of the road emergency. Indicates the radius of influence of a road emergency. Indicates the duration of a road emergency.

[0080] Through the above methods, the server can receive any set of environmental parameters uploaded by the target vehicle. This set of environmental parameters includes at least road topology information, vehicle driving status information, and road emergency information. This allows the server to use this multi-source heterogeneous data as the basis for environmental perception, to obtain complete environmental information and vehicle driving status. For example, it can obtain the target vehicle's position, speed, and heading angle at a specific moment, as well as the location (space), time, radius of influence, and duration of road emergencies. This information serves as the basis for subsequent spatiotemporal constraint coding. Subsequently, the road topology information, vehicle driving status information, and road emergency information can be transformed into corresponding path distance thresholds and dynamic obstacle probability distributions, thereby generating a multi-dimensional constraint matrix to constrain the vehicle's path decision-making, guiding the subsequent game decision-making process in generating path strategy magnitudes.

[0081] 102. Transform road topology information, vehicle driving status information and road emergency information into path distance threshold and dynamic obstacle probability distribution, and combine path distance threshold and dynamic obstacle probability distribution to construct a multidimensional constraint matrix.

[0082] In this embodiment, after receiving the set of environmental parameters uploaded by the target vehicle, the server can encode the road topology information, vehicle driving status information, and road emergency information in the environmental parameter set to generate a path distance threshold and a dynamic obstacle probability distribution. The path distance threshold is used to limit the distance between the target vehicle and a specific area or object (such as the location of a road emergency) at a specific time. The dynamic obstacle probability distribution describes the uncertainty range and probability of the location of unknown obstacles (such as surrounding vehicles, other emergencies, etc.) at a specific time. Then, the path distance threshold and the dynamic obstacle probability distribution are combined to construct a multi-dimensional constraint matrix. In this way, multi-source heterogeneous data such as road topology information, vehicle driving status information, and road emergency information can be integrated to generate constraint conditions to guide the subsequent game decision-making process in generating path strategy magnitude, so that the subsequent process of generating path strategy magnitude has a reference basis and is reliable.

[0083] The path distance threshold can be understood as a distance limit that an ego vehicle should maintain between itself and a relevant area or object at a specific time. The path distance threshold is generated based on road topology information and road event information; it is a dynamically changing value that adjusts in real time as road topology, events, and other factors change. In short, the path distance threshold specifies how far an ego vehicle should stay from certain specific areas or objects at different times; the ego vehicle can refer to the aforementioned target vehicle.

[0084] The dynamic obstacle probability distribution predicts the probability that an obstacle (other vehicles around the target vehicle) will be in different states (positions, speeds, heading angles, etc.) at a future time based on the current state information (such as position, speed, and heading angle) of the obstacle. The dynamic obstacle probability distribution does not determine whether an obstacle will definitely be in a certain position in the future, but rather gives the probability of the obstacle appearing in different positions. Dynamic obstacles generally refer to other vehicles traveling around the target vehicle.

[0085] In some implementations, road topology information and road incident information can be combined for encoding to generate a path distance threshold, and a dynamic obstacle probability distribution can be calculated by combining vehicle driving status information and reference driving status information of vehicles surrounding the target vehicle. For example, step 102, "converting road topology information, vehicle driving status information, and road incident information into a path distance threshold and a dynamic obstacle probability distribution," may include:

[0086] (102.a.1) Determine the shortest path distance based on road topology information; (102.a.2) Determine the target distance between the vehicle location of the target vehicle and the event location contained in the road emergency information, and construct a distance influence parameter based on the ratio between the target distance and the influence radius parameter contained in the road emergency information, and generate a path distance threshold based on the shortest path distance and the distance influence parameter;

[0087] (102.a.3) Obtain the current reference driving state information of reference vehicles around the target vehicle, and calculate the mean of the predicted position of the reference vehicle at the target prediction time based on the reference driving state information and the vehicle driving state information, and calculate the covariance at the target prediction time based on the preset error coefficient and the preset diffusion coefficient.

[0088] (102.a.4) The mean and covariance of the predicted location are quantified by normal distribution to generate a dynamic obstacle probability distribution.

[0089] Among them, the shortest path distance refers to the shortest distance planned for the target vehicle based on the road topology. Specifically, based on the road topology information, the road segment with the smallest length weight is selected from all road segments included in the road topology. It reflects the basic influence of the road structure itself on the path distance.

[0090] The target distance refers to the distance determined by the vehicle's position in the vehicle's driving status information and the location of the event contained in the road emergency information. It reflects the distance between the target vehicle's current driving position and the road emergency (construction, traffic control location, etc.). The distance influence parameter can be generated by combining the target distance.

[0091] Among them, the distance influence parameter is used to adjust the shortest path distance of the target vehicle. For example, the distance influence parameter is superimposed on the shortest path distance to obtain the path distance threshold, which represents the shortest path threshold when the target vehicle needs to avoid the sudden road event and chooses to detour.

[0092] To facilitate understanding steps (102.a.1) and (102.a.2), the calculation process for the path distance threshold is described below:

[0093]

[0094] in, Indicates in The path distance threshold at any given time. (In the road topology) middle, This represents the set of edges of a road (i.e., the set of road segments or paths). This indicates a specific edge (road segment or path). express The weight of the edge segment length, therefore, This represents the value of the road segment with the smallest length weight selected from all road segments. It reflects the fundamental influence of the road structure itself on the path distance and is a fixed term related to road topology.

[0095] in, This represents the impact coefficient of road emergencies, used to measure the degree of influence of emergencies on the path distance threshold. The larger the value of α, the greater the impact of the emergency on the threshold.

[0096] in, The function indicates the active time interval of a road emergency, if Constantly in the event of road emergencies Active time intervals [ If the value is within ], then the indicator function value is 1. If the time is not within the active time interval, the indicator function value is 0. Therefore, the active time interval indicator function is used to determine whether a road emergency is active at the current time, that is, whether it will have an impact on the surrounding area at the current time, or whether a road emergency exists at the current time. Penalties are only imposed within this active time interval.

[0097] in, express The position of the target vehicle at any given time. Indicates a road emergency Location, This represents the radius of influence of a road emergency. exp represents an exponential function. "" represents the target distance between the target vehicle's location in the vehicle's driving status information and the event location contained in the road emergency information. Further, the target ratio between the target distance and the influence radius of the road emergency is determined (represented as ""). "), and perform indicator calculations on the target ratio, expressed as It should be noted that, A value close to 1 will significantly increase the path distance threshold, prompting vehicles to take a longer route to avoid the event, when the distance between the vehicle and the sudden event on the road is large (far from the event). A value close to 0 (small impact) has a negligible effect on the path distance threshold, indicating that vehicles do not need to take detours. Therefore, the exponential function reflects the impact of the target distance between the vehicle and the sudden road event on the path distance threshold. The closer the target distance, the smaller the value of the exponential function term, and the more significant the increase in the path distance threshold. In this way, the "relationship between distance and event" is transformed into a "smooth impact weight," realizing the quantitative adjustment of the impact of dynamic events.

[0098] In summary, "" represents a dynamic road emergency correction term, which can include multiple road emergencies. For each road emergency, a sub-correction term is constructed by combining the active time interval indicator function, exponential function, etc. of the road emergency. The sub-correction terms corresponding to multiple road emergencies are summed and multiplied by the road emergency impact coefficient to obtain the dynamic road emergency correction term.

[0099] at last, This means selecting the road segment with the smallest length weight from all road segments. Specifically, this can be understood as the basic road distance. Therefore, this basic road distance is added to the dynamic road emergency correction term to obtain the path distance threshold.

[0100] To facilitate understanding steps (102.a.3) and (102.a.4), the calculation process of the dynamic obstacle probability distribution is described below:

[0101]

[0102] in, Conditional probability, a concept related to prior information, represents the probability when surrounding reference vehicles are known. exist The reference driving status at any time is Then predict the reference vehicles in the surrounding area. exist The predicted reference driving state at any time is The probability of a target's future state. That is, predicting the likelihood of a target's future state based on prior information and the current status of reference vehicles around the target. It should be noted that predicting the reference driving state... Indicates reference vehicle exist The state vector at any given time can include the predicted position (x, y) and velocity v, where velocity v can be represented as " ".

[0103] in, Indicates surrounding reference vehicles exist The predicted reference driving state at any time is The time follows a normal distribution, which is mainly determined by two parameters: the mean of the predicted location and the mean of the predicted location. and target prediction time covariance The predicted location mean represents the reference vehicle. exist The predicted reference driving state at any time is The most likely location, i.e., the reference vehicle's position. The most likely location at time; the covariance of the target predicted time represents the predicted reference driving state of the reference vehicle. The range of uncertainty can be understood as the range of prediction error.

[0104] It should be noted that the predicted location mean and target prediction time covariance The following motion model can be derived, as follows:

[0105] ,

[0106] Indicates the target prediction time The mean of the predicted location, This represents unknown obstacles (such as reference vehicles around the target vehicle) at the initial moment. The initial position can be obtained by sensors, either detected by the target vehicle's sensors or uploaded autonomously by the reference vehicle; no limitation is made here. The integral representing velocity can be specifically understood as... arrive At any given moment, the obstacle's displacement due to its motion. Based on this, the average predicted position can be understood as the position at which the reference vehicle was initially positioned, assuming the reference vehicle started from... arrive The system moves at a constant speed at all times, and the mean value of the predicted position is calculated accordingly.

[0107] in, Indicates the target prediction time covariance, Indicates initial uncertainty. The time diffusion coefficient is . Indicates time difference, Let represent the identity matrix. Based on this, the covariance can be understood as follows: the smaller the time difference, the smaller the error; conversely, the larger the time difference, the larger the error. For example, the larger the time difference, the more likely the reference vehicle will change lanes or decelerate. Therefore, prediction based solely on the current reference vehicle's driving state information will result in a larger error due to the larger time difference. These errors can be represented by the covariance.

[0108] Furthermore, by using a normal distribution, the mean and covariance of the predicted position are quantified to generate a dynamic obstacle probability distribution for reference vehicles around the target vehicle at future times.

[0109] In some implementations, a multidimensional constraint matrix is ​​constructed by combining the path distance threshold and the probability distribution of dynamic obstacles.

[0110] The multidimensional constraint matrix is ​​used to define the drivable area at different times, i.e., the drivable area that the target vehicle can choose to move forward in the future. This multidimensional elm tree matrix can include dimensions such as time step, spatial grid area, and spatial region number, and is represented as follows: , Indicates a time step. Represents the area of ​​the spatial grid ( Rasterized map), This represents the spatial region number. Specifically, the multidimensional constraint matrix is ​​calculated as follows:

[0111]

[0112] in, Indicates the first The multidimensional constraint matrix of a spatial region. The larger the value of the multidimensional constraint matrix, the stronger the simultaneous constraint on the spatiotemporal grid (i.e., the road region corresponding to that time and space). For example, if a person enters the region from a stationary position, the drivable area at future times can be determined based on the multidimensional constraint matrix.

[0113] in," "" represents the constraint terms for dynamic obstacles, indicating the reference vehicle exist Always The probability of occupancy of the spatial grid area is used to constrain the vehicle to avoid collisions with reference vehicles. The event constraint term representing a road emergency is only valid if the target prediction time is within the active time interval (the value of the indicator function is 1). Indicates the risk weight of road emergencies. The higher the value, the greater the risk. Spatial constraints representing road emergencies express Location of the spatial grid area Indicates a road emergency The spatial constraint term can be logically defined as only if... The spatial grid areas are all located within the influence radius of the road emergency. The area of ​​the spatial grid is effective (i.e., it has an impact).

[0114] By using the above methods, multi-source heterogeneous data such as road topology information, vehicle driving status information, and road emergency information can be integrated to generate constraints, which can then guide the subsequent game decision-making process in generating path strategy magnitudes. This makes the subsequent process of generating path strategy magnitudes reliable and has a reference basis.

[0115] 103. Taking the maximization of the covered driving area and the minimization of the total avoidance distance of multiple vehicles within the driving area as the optimization objective, the target path strategy magnitude is simulated and calculated by combining the multidimensional constraint matrix and the preset posterior information.

[0116] In this embodiment, after obtaining the multidimensional constraint matrix, maximizing the coverage area of ​​the driving region is taken as the first objective, while minimizing the total avoidance distance of multiple vehicles within that driving region is taken as the second objective. The first and second objectives are parallel, and the combination of the first and second objectives forms the overall optimization objective. This optimization objective can be understood as a global objective, the purpose of which is to consider this global objective when generating the path strategy magnitude, avoiding the generation of a locally optimal path strategy magnitude for only a single vehicle. Therefore, based on this optimization objective, combined with the multidimensional constraint matrix and preset posterior information, the target path strategy magnitude can be simulated and calculated. This target path strategy magnitude considers both the coverage area of ​​the driving region and the total avoidance distance of multiple vehicles within that coverage area. This avoids providing an optimal path strategy magnitude only for the target vehicle, which could affect the path strategies of other vehicles. Thus, subsequent target vehicles use this target path strategy magnitude as a basis, effectively preventing traffic congestion caused by a single target vehicle deciding on the optimal path strategy magnitude, thereby improving the overall driving efficiency and safety of vehicles.

[0117] The target path policy magnitude can be the adjustment range of the path policy, which can be understood as the adjustment coefficient of the path policy. This target path policy magnitude can have direction, such as driving xx distance to the left, reducing speed by xx, or driving xx distance to the right, etc., depending on the actual scenario; no specific limitation is made here. It should be noted that this target path policy magnitude can be calculated using a cooperative game model based on posterior information to obtain a global utility function, and then using this global utility function to generate a path policy magnitude that maximizes the coverage area and minimizes energy consumption.

[0118] In some implementations, the target feasible region corresponding to the current moment can be determined based on a multidimensional constraint matrix, and a global utility function and a local utility function for the target vehicle can be constructed. The global and local utility functions are then combined to simulate and calculate the target path policy magnitude that conforms to the optimal target feasible region. For example, step 103 may include: determining the target feasible region corresponding to the current moment based on the multidimensional constraint matrix; obtaining the global utility function; obtaining the policy response relationship of the target vehicle relative to the path policy magnitude issued by the local node, and constructing a local utility function based on preset posterior information and the policy response relationship; and combining the global and local utility functions to simulate and calculate the target path policy magnitude that conforms to the target feasible region.

[0119] The feasible region of the target can be the corresponding road area that is allowed to travel by the multi-dimensional constraint matrix. For example, it can be the xx spatial grid of the xx segment of road x, or it can be the xx spatial grid of xx time.

[0120] The global utility function characterizes the optimization objective as maximizing the covered driving area while minimizing the total avoidance distance of multiple vehicles within that area. The global utility function is essentially a quantified integration of the "coverage area" (global task value) and the "total avoidance distance" (energy reduction). (The larger the coverage and the smaller the avoidance distance, the higher the value of the global utility function). Furthermore, the calculation of maximizing the global expected utility based on the global utility function is expressed as follows:

[0121]

[0122] in, This represents an optimization problem involving the magnitude of the global policy (global path planning). Indicates the path strategy magnitude. , This indicates that the drivable area at different times is defined by a multidimensional constraint matrix, such as the target drivable area at a specific time t in the future. Through this constraint, it is ensured that the magnitude of the generated target path strategy does not violate environmental constraints such as dynamic events and obstacles, reflecting the boundary restriction role of the constraint matrix on the global strategy. Represents the global expected utility. This represents the optimal response distribution of the server to the target vehicle's response to the issued path policy magnitude. The optimization objective is achieved by maximizing the global expected utility.

[0123] The local utility function characterizes the relationship between the magnitude of the path policy issued by the local node and the simulated theoretical path of the target vehicle. The simulated theoretical path is obtained by the local node through simulating the path policy issued in response to the target vehicle. Let represent the local utility function. Further, the calculation of maximizing the local expected utility based on the local utility function is expressed as follows:

[0124]

[0125] in, This represents the theoretical path of the simulation. Local utility function. The construction relies on two core elements: pre-defined posterior information and policy response relationships. The pre-defined posterior information is represented as follows: , The environment cognition can be updated based on the observed signals, while the policy response relationship represents the simulation of the target vehicle obtaining the optimal path while maximizing the local utility function. That is, the simulation of the policy adjustment actions and results that the server will make to the target vehicle in response to the issued path policy magnitude, thus obtaining the simulated theoretical path. Representing local expected utility, maximizing this local expected utility can simulate the path strategy issued to the target vehicle, thereby simulating the theoretical path obtained after the target vehicle makes the optimal response. In turn, the range of optional path strategies corresponding to the simulated theoretical path can be determined, and the range of optional path strategies conforms to the target feasible region determined by the multidimensional constraint matrix.

[0126] Among them, with The likelihood function represents prior information or prior probability, and is expressed as follows: Substituting prior information and the likelihood function into Bayes' formula, i.e., posterior beliefs... In actual calculations, numerical methods (such as Markov chain Monte Carlo or variational inference algorithms) can be used, especially when θ is a high-dimensional continuous variable.

[0127] Combining the formulas for maximizing global utility expectation and maximizing local utility expectation, the target path strategy magnitude within the target feasible region is simulated and calculated. Specifically, the server constructs a global utility function with the optimization objective of maximizing the covered driving area and minimizing the total avoidance distance of multiple vehicles within that area. Based on this global utility function, a formula for maximizing global utility expectation is then constructed. Simultaneously, the optimal response of the target vehicle to the issued path strategy is simulated, using the optimal path obtained when the target vehicle makes its optimal response. This is used to construct a local utility function, and a formula for maximizing local utility expectation is then constructed based on the global utility function. Finally, combining these two formulas, with the optimization objective of maximizing the covered driving area and minimizing the total avoidance distance of multiple vehicles within that area, the target path strategy magnitude is simulated within the target feasible region to determine the optimal simulated theoretical path for the target vehicle. Indicates the target path strategy magnitude.

[0128] By using the above method, the optimization objective can be based on maximizing the coverage area of ​​the driving area and minimizing the total avoidance distance of multiple vehicles within that area. Combining a multi-dimensional constraint matrix and pre-set posterior information, the target path strategy magnitude can be simulated and calculated. In this way, considering both the coverage area of ​​the driving area and the total avoidance distance of multiple vehicles within that area, it avoids providing the optimal path strategy magnitude for the target vehicle and affecting the path strategies of other vehicles. As a result, subsequent target vehicles use this target path strategy magnitude as a basis, effectively preventing traffic congestion caused by a single target vehicle deciding on the optimal path strategy magnitude, and improving the overall driving efficiency and safety of vehicles.

[0129] 104. Send the target path strategy magnitude and multidimensional constraint matrix to the target vehicle.

[0130] In this embodiment, after calculating the target path strategy magnitude, the target path strategy magnitude and the multidimensional constraint matrix can be sent to the target vehicle. This allows the target vehicle to generate a path strategy for the next moment based on the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude, under the guidance of the multidimensional constraint matrix. In this way, when the target vehicle travels according to the path strategy for the next moment, it can avoid traffic congestion caused by the individual target vehicle deciding on the optimal path strategy magnitude, thereby improving the overall driving efficiency and safety of the vehicle and saving energy consumption to a certain extent.

[0131] It should be noted that the target vehicle, through a Stackelberg game with incomplete information, uses its original current path strategy. Based on this, combined with the collected dynamic environmental conditions Dynamically adjust path strategy Specifically, it is expressed as follows:

[0132]

[0133] In some implementations, environmental status parameters uploaded by the target vehicle can be acquired in real time to determine whether the previously issued target path policy magnitude needs to be updated. If the environment changes significantly, the target path policy magnitude needs to be updated, and the optimized policy magnitude is reissued to the target vehicle. For example, after step 104, the following may also be included:

[0134] (A.1) Obtain the environmental status parameters uploaded by the target vehicle;

[0135] (A.2) When the environmental state parameter is greater than or equal to the preset environmental state threshold, the target path strategy magnitude is updated to obtain the updated optimized strategy magnitude.

[0136] (A.3) The optimization strategy magnitude is sent to the target vehicle so that the target vehicle can generate the next path strategy based on the current path strategy, the real-time collected dynamic environment state and the optimization strategy magnitude under the guidance of the multi-dimensional constraint matrix.

[0137] The preset environmental state threshold can be a judgment value used to measure the degree of change in the environmental state of the target vehicle during driving. By comparing it with the preset environmental state threshold, the degree of change in the environmental state of the target vehicle during driving can be judged, and a decision can be made on whether to update the path strategy magnitude based on the judgment result.

[0138] It should be noted that the environmental information such as road conditions and traffic flow encountered by the vehicle during driving is changing in real time. Therefore, in order to enable the vehicle to adapt to the real-time road conditions and improve the vehicle's driving efficiency, the target vehicle can be prompted to upload real-time environmental status parameters. These environmental status parameters include, but are not limited to, the positions of surrounding vehicles, current road emergencies, road topology information, and the distance the target vehicle has traveled. Through these environmental status parameters, the current road conditions around the target vehicle can be known. Furthermore, the environmental status parameters are compared with a preset environmental status threshold. For example, assuming that the path strategy amplitude needs to be re-optimized every 3 meters the vehicle travels, the preset status parameter threshold is set to 3. Taking the distance the target vehicle has traveled as an example, the specific distance can be the distance traveled after the last target path strategy amplitude is issued. The distance is compared with the preset status parameter threshold (3). If the distance is less than the preset status parameter threshold, no update is performed. If the distance is greater than or equal to the preset status parameter threshold, the target path strategy amplitude needs to be updated to obtain the updated optimized strategy amplitude. Finally, the optimized strategy magnitude is sent to the target vehicle so that the target vehicle can generate the next path strategy based on the current path strategy, the real-time collected dynamic environment state, and the optimized strategy magnitude under the guidance of the multi-dimensional constraint matrix.

[0139] In some implementations, when deciding to update the target path policy magnitude, a global utility expectation gradient and a policy magnitude change regularization term can be constructed. Then, a policy magnitude update function is constructed based on these two terms to calculate the updated optimized policy magnitude. For example, step (A.2), "updating the target path policy magnitude to obtain the updated optimized policy magnitude," can include: constructing a corresponding global utility expectation gradient and a corresponding policy magnitude change regularization term for the target path policy magnitude; combining the global utility expectation gradient and the policy magnitude change regularization term to construct a policy magnitude update function; with maximizing the value of the policy magnitude update function as the optimization objective, updating the target path policy magnitude according to the policy magnitude update function to obtain the updated optimized policy magnitude.

[0140] Specifically, for the magnitude of the target path policy, a corresponding global utility expectation gradient is constructed. This can be achieved by obtaining multiple historical time-series sequences of environmental state parameters corresponding to the magnitude of the target path policy, and then constructing the expected gradient of the global utility function based on these sequences. For example, the global utility expectation gradient can be expressed as... The global utility expectation gradient is used to measure the change in the obtained global utility expectation. The gradient can guide the correct acquisition of the maximum value of the global utility expectation.

[0141] Specifically, the regularization term for policy magnitude change constructed for the target path policy magnitude can be expressed as follows: , This represents the learning rate parameter. This indicates the magnitude of the target path strategy previously issued.

[0142] By combining the expected gradient of global utility and the regularization term for policy magnitude change, a policy magnitude update function is constructed. This function represents the latest target path policy magnitude, i.e., the updated optimized policy magnitude, obtained when the difference between the expected gradient of global utility and the regularization term for policy magnitude change is reached. The policy magnitude update function is expressed as follows:

[0143]

[0144] The optimization objective is to maximize the value of the policy magnitude update function. The target path policy magnitude is updated according to the policy magnitude update function to obtain the updated optimized policy magnitude. Specifically, using the formula of the policy magnitude update function, the latest target path policy magnitude is obtained by reversing the process when the target difference between the expected gradient of global utility and the regularization term of policy magnitude change reaches its maximum value; that is, the updated optimized policy magnitude. This optimized policy magnitude can then be sent to the target vehicle, enabling it to generate the next time-step path policy based on the current path policy, the real-time acquired dynamic environmental state, and the optimized policy magnitude, under the guidance of the multi-dimensional constraint matrix.

[0145] It should be noted that for the policy magnitude update function, the learning rate parameter... The policy magnitude update function can be simplified, and the simplified target policy magnitude update function is expressed as: in, This represents the projection of the feasible set, and the quadratic term ensures that the adjustment range is not too large. The updated optimized policy magnitude is calculated using this simplified target policy magnitude update function.

[0146] In some implementations, after obtaining the optimized strategy magnitude, the robustness of the optimized strategy magnitude can be verified before deciding whether to activate the optimized strategy magnitude, i.e., whether to send the optimized strategy magnitude to the target vehicle. For example, before step (A.3), the following steps may be included: obtaining the target vehicle's initial path strategy at the current moment; combining the initial path strategy and the optimized strategy magnitude to simulate and generate multiple candidate path strategies corresponding to multiple consecutive time intervals; and determining the average strategy error based on the strategy error between each candidate path strategy and the initial path strategy. Then, the step "sending the optimized strategy magnitude to the target vehicle" includes: sending the optimized strategy magnitude to the target vehicle when the average strategy error is less than or equal to a preset strategy error threshold.

[0147] Specifically, it can obtain the target vehicle's initial path strategy at the current moment, in order to... This represents the initial path strategy of the target vehicle at the current moment. Combining the initial path strategy and the magnitude of the optimized strategy, multiple candidate path strategies are simulated and generated for consecutive time intervals, resulting in a candidate path strategy set. This represents the set of candidate path strategies, and the set of candidate path strategies includes... The multiple candidate path strategies corresponding to each time step are represented as follows: .

[0148] Robustness verification is performed on multiple candidate path strategies in the candidate path strategy set. The first step involves obtaining the policy error between each candidate path strategy and the initial path strategy. Indicates the current initial path strategy, This represents the policy error for each policy step. After determining the policy error between each candidate path policy and the initial path policy, the policy errors over multiple time intervals (T time steps) can be summed to obtain the total policy error. This is to enable performance evaluation through measurement, which can be specifically calculated using a cost function, as shown below:

[0149]

[0150] The second step is to calculate the robustness index based on the sum of policy errors. Specifically, the expected cost (i.e., the average value) can be used as the robustness index. The specific calculation process is as follows:

[0151]

[0152] in, This represents the robustness metric, i.e., the average policy error. This indicates the amount of policy error.

[0153] The third step is to compare the average error of the strategy with a preset error threshold. This preset error threshold can be set based on prior experience. This represents the preset policy error threshold. If the average policy error is less than or equal to the preset policy error threshold, that is, if... If the current optimization strategy range meets the requirements, then the optimization strategy range will be activated and sent to the target vehicle.

[0154] Furthermore, in some implementations, if the average policy error is greater than a preset policy error threshold, it indicates that the currently obtained optimized policy magnitude is inappropriate and a new optimized policy magnitude needs to be recalculated. For example, it may also include: when the average policy error is greater than the preset policy error threshold, determining the difference between the average policy error and the preset policy error threshold; adjusting the learning rate parameter in the regularization term for policy magnitude change in the policy magnitude update function according to the difference, to obtain an updated target policy magnitude update function; with maximizing the value of the target policy magnitude update function as the optimization objective, updating the optimized policy magnitude according to the target policy magnitude update function, to obtain the updated target optimized policy magnitude; and sending the target optimized policy magnitude to the target vehicle.

[0155] It should be noted that if the average policy error exceeds the preset policy error threshold, it indicates that the current optimized policy magnitude is inappropriate, leading to a significant error compared to expectations. This results in a failure of robustness verification, requiring a recalculation of the optimized policy magnitude. Specifically, the learning rate of the policy magnitude update function can be adjusted. For example, the learning rate parameter in the policy magnitude update function can be reduced to obtain a smaller calculated optimized policy magnitude, thus obtaining a new policy magnitude update function. This results in a smaller recalculated optimized policy magnitude, reducing the policy error between each candidate path policy and the initial path policy during the robustness verification phase. As can be seen, there is a positive correlation between the learning rate parameter and the optimized policy magnitude. Reducing the learning rate parameter also reduces the recalculated optimized policy magnitude. This prevents excessive adjustment of the path policy when simulating candidate path policies based on the recalculated optimized policy magnitude, thus avoiding large errors. Therefore, when the average policy error exceeds the preset policy error threshold, the learning rate parameter in the policy magnitude update function can be reduced.

[0156] Furthermore, to achieve precise adjustment of the learning rate parameter in the policy magnitude update function, the difference between the average policy error and the preset policy error threshold can be obtained. This difference reflects the degree of difference between the average policy error and the preset policy error threshold if the current optimized policy magnitude is used. The larger the difference, the larger the error of the average policy error. Therefore, this difference can be used to precisely adjust the learning rate parameter in the policy magnitude update function, providing a basis for the adjustment process and improving the accuracy of the learning rate parameter. This makes the learning rate parameter more consistent with the robustness verification process, resulting in the updated target policy magnitude update function.

[0157] Finally, with maximizing the value of the target policy magnitude update function as the optimization objective, the optimized policy magnitude is updated according to the target policy magnitude update function to obtain the updated target optimized policy magnitude. Specifically, the new optimized policy magnitude is recalculated using the updated target policy magnitude update function. This updated target optimized policy magnitude is then sent to the target vehicle, enabling it to generate the next target path policy based on the current path policy, the real-time collected dynamic environment state, and the target optimized policy magnitude, under the guidance of the multi-dimensional constraint matrix. This ensures that the generated target optimized policy magnitude better matches the actual driving needs of the target vehicle, improving vehicle driving efficiency and safety.

[0158] In this way, the target path strategy magnitude and multidimensional constraint matrix can be sent to the target vehicle. Under the guidance of the multidimensional constraint matrix, the target vehicle generates the path strategy for the next moment based on the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude. In this way, when the target vehicle travels according to the path strategy for the next moment, it can avoid traffic congestion caused by the optimal path strategy magnitude of a single target vehicle, thereby improving the overall driving efficiency and safety of the vehicle and saving energy consumption to a certain extent.

[0159] Figure 3 This is an example diagram of a collaborative decision-making scenario architecture based on spatiotemporal game theory provided in the embodiments of this application, combined with... Figure 3 As shown, the collaborative decision-making scenario architecture based on spatiotemporal game theory comprises three parts. The first part is the viewport constraint encoding module, which integrates multi-dimensional information (i.e., multi-source heterogeneous data). In terms of temporal constraints, it performs real-time constraints and real-time multi-agent (multiple vehicles) driving prediction. In terms of spatial constraints, it combines event location and obstacle location to calculate spatial constraints, thereby generating spatiotemporal constraint conditions, i.e., a multi-dimensional constraint matrix. The second part is the hierarchical game decision-making module, which uses an incomplete information spatiotemporal game model to make game decisions. The upper-level global optimization involves the server generating the target path strategy magnitude, as described in step 103. The lower-level local adjustment involves sending the target path strategy magnitude to the target vehicle, enabling the target vehicle to generate its next-moment path strategy based on the current path strategy, the real-time collected dynamic environment state, and the target path strategy magnitude, under the guidance of the multi-dimensional constraint matrix. In other words, the target vehicle combines the sent target path strategy magnitude, the multi-dimensional constraint matrix, and the actual dynamic environment state to generate the actual path strategy to cope with traffic congestion. The third part is the implementation update of path strategy recovery, which can adjust the multi-agent behavior environment state and dynamic Bayesian network to update the path strategy amplitude in real time, so that the target optimization strategy amplitude is more in line with the actual driving needs of the target vehicle, thereby improving vehicle driving efficiency and safety.

[0160] As can be seen from the above embodiments, this application receives road topology information, vehicle driving status information, and road emergency information uploaded by the target vehicle, and transforms the above information into path distance thresholds and dynamic obstacle probability distributions. Then, it combines the path distance thresholds and dynamic obstacle probability distributions to generate a multi-dimensional constraint matrix. This multi-dimensional constraint matrix is ​​used to limit the drivable area at different times. Thus, constraints can be generated by combining real-time environmental parameters. Furthermore, by combining the multi-dimensional constraint matrix and preset posterior information, the optimization objective is to maximize the covered drivable area and minimize the total avoidance distance of multiple vehicles within the drivable area. The target path strategy magnitude is simulated and calculated. Finally, the target path strategy magnitude and the multi-dimensional constraint matrix are combined... A multidimensional constraint matrix is ​​sent to the target vehicle, enabling the target vehicle to generate a path strategy for the next moment within the drivable area defined by the multidimensional constraint matrix, combining the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude. This is to cope with traffic congestion. In contrast to related technologies where vehicles only travel according to static paths, leading to traffic congestion, this application can construct a multidimensional constraint matrix by combining the environmental parameter set uploaded by the vehicle and generate a target path strategy magnitude. Under the guidance of the multidimensional constraint matrix, the vehicle can update its future path strategy in real time by combining the target path strategy magnitude, thereby avoiding or mitigating traffic congestion while maximizing energy savings for multiple vehicles, improving vehicle driving safety and travel efficiency.

[0161] To facilitate better implementation of the spatiotemporal game-based collaborative decision-making method provided in this application, this application also provides a collaborative decision-making device based on the aforementioned spatiotemporal game-based method. The meanings of the terms used are the same as in the spatiotemporal game-based collaborative decision-making method described above, and specific implementation details can be found in the descriptions within the method embodiments.

[0162] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of a collaborative decision-making device based on spatiotemporal game theory provided in an embodiment of this application. The collaborative decision-making device based on spatiotemporal game theory is integrated into the computer equipment of this application, such as a server. The collaborative decision-making device based on spatiotemporal game theory may include an acquisition unit 401, a generation unit 402, a calculation unit 403, and a sending unit 404.

[0163] The acquisition unit 401 is used to acquire the set of environmental parameters uploaded by the target vehicle. The set of environmental parameters includes at least road topology information, vehicle driving status information and road emergency information.

[0164] The generation unit 402 is used to convert road topology information, vehicle driving status information and road emergency information into path distance threshold and dynamic obstacle probability distribution, and to construct a multi-dimensional constraint matrix by combining the path distance threshold and dynamic obstacle probability distribution. The multi-dimensional constraint matrix is ​​used to limit the drivable area corresponding to different times.

[0165] The calculation unit 403 is used to simulate and calculate the magnitude of the target path strategy by taking the optimization objective of maximizing the coverage of the driving area and minimizing the total avoidance distance of multiple vehicles within the driving area as the optimization objective, and combining the multi-dimensional constraint matrix and preset posterior information.

[0166] The sending unit 404 is used to send the target path strategy magnitude and the multidimensional constraint matrix to the target vehicle, so that the target vehicle can generate the path strategy for the next moment based on the current path strategy, the real-time collected dynamic environment state and the target path strategy magnitude under the guidance of the multidimensional constraint matrix.

[0167] In some embodiments, the generating unit 402 is further configured to:

[0168] Determine the shortest path distance based on road topology information;

[0169] Determine the target distance between the vehicle location of the target vehicle and the event location contained in the road emergency information, and construct a distance influence parameter based on the ratio between the target distance and the influence radius parameter contained in the road emergency information. Generate a path distance threshold based on the shortest path distance and the distance influence parameter.

[0170] Obtain the current reference driving status information of reference vehicles around the target vehicle, and calculate the mean of the predicted position of the reference vehicle at the target prediction time based on the reference driving status information and the vehicle driving status information, and calculate the covariance at the target prediction time based on the preset error coefficient and the preset diffusion coefficient.

[0171] By using a normal distribution method, the mean and covariance of the predicted location are quantified to generate a dynamic obstacle probability distribution.

[0172] In some embodiments, the computing unit 403 is further configured to:

[0173] Determine the target feasible region at the current moment based on the multidimensional constraint matrix;

[0174] Obtain the global utility function, which is used to characterize the optimization objective of maximizing the covered driving area and minimizing the total avoidance distance of multiple vehicles within the driving area.

[0175] The policy response relationship of the target vehicle relative to the path policy magnitude issued by the local node is obtained, and a local utility function is constructed based on the preset posterior information and the policy response relationship. The local utility function is used to characterize the relationship between the path policy magnitude issued by the local node and the simulated theoretical path of the target vehicle. The simulated theoretical path is obtained by the local node through simulating the path policy issued by the target vehicle.

[0176] By combining global and local utility functions, the target path strategy magnitude that conforms to the target feasible region is simulated and calculated.

[0177] In some implementations, the spatiotemporal game-based collaborative decision-making device further includes a strategy update unit, used for:

[0178] Obtain the environmental status parameters uploaded by the target vehicle;

[0179] When the environmental state parameter is greater than or equal to the preset environmental state threshold, the target path strategy magnitude is updated to obtain the updated optimized strategy magnitude.

[0180] The optimized strategy magnitude is sent to the target vehicle so that the target vehicle can generate the next path strategy based on the current path strategy, the real-time collected dynamic environment state, and the optimized strategy magnitude under the guidance of the multi-dimensional constraint matrix.

[0181] In some implementations, the policy update unit is further configured to:

[0182] Construct the corresponding global utility expectation gradient for the policy magnitude of the target path, and construct the corresponding policy magnitude change regularization term;

[0183] By combining the expected gradient of global utility with the regularization term of policy magnitude change, a policy magnitude update function is constructed.

[0184] The optimization objective is to maximize the value of the policy magnitude update function. The policy magnitude of the target path is updated according to the policy magnitude update function to obtain the updated optimized policy magnitude.

[0185] In some implementations, the policy verification unit is further configured to:

[0186] Obtain the initial path strategy of the target vehicle at the current time, and combine the initial path strategy and the optimization strategy magnitude to simulate and generate multiple candidate path strategies corresponding to multiple consecutive time intervals.

[0187] The average policy error is determined based on the policy error between each candidate path policy and the initial path policy.

[0188] The strategy update unit is also used to send the optimized strategy magnitude to the target vehicle when the average strategy error is less than or equal to the preset strategy error threshold.

[0189] In some implementations, the policy update unit is further configured to:

[0190] When the average strategy error is greater than the preset strategy error threshold, the difference between the average strategy error and the preset strategy error threshold is determined.

[0191] Based on the gap value, adjust the learning rate parameter in the regularization term for policy magnitude change in the policy magnitude update function to obtain the updated target policy magnitude update function.

[0192] The optimization objective is to maximize the value of the target policy magnitude update function. The optimized policy magnitude is then updated according to the target policy magnitude update function to obtain the updated target optimized policy magnitude.

[0193] Send the target optimization strategy magnitude to the target vehicle.

[0194] As described above, this application receives road topology information, vehicle driving status information, and road emergency information uploaded by the target vehicle, and transforms this information into path distance thresholds and dynamic obstacle probability distributions. These are then combined to generate a multidimensional constraint matrix, which is used to define the drivable area at different times. Real-time environmental parameters are then used to generate constraints. Furthermore, by combining the multidimensional constraint matrix and preset posterior information, the optimization objective is to maximize the covered drivable area while minimizing the total avoidance distance of multiple vehicles within that area. The target path strategy magnitude is simulated and calculated. Finally, the target path strategy magnitude and the multidimensional constraints are... The matrix is ​​sent to the target vehicle, enabling the target vehicle to generate a path strategy for the next moment within the drivable area defined by the multidimensional constraint matrix, combining the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude. This is to cope with traffic congestion. In contrast to related technologies where vehicles only travel according to static paths, leading to traffic congestion, this application can construct a multidimensional constraint matrix by combining the environmental parameter set uploaded by the vehicle and generate the target path strategy magnitude. Under the guidance of the multidimensional constraint matrix, the vehicle can update its future path strategy in real time by combining the target path strategy magnitude, thereby avoiding or mitigating traffic congestion while maximizing energy savings for multiple vehicles, improving vehicle driving safety and travel efficiency.

[0195] The specific implementation of each of the above units can be found in the previous embodiments, and will not be repeated here.

[0196] Figure 5To implement the partial structural block diagram of the terminal 110 in this application embodiment, the terminal 110 includes: a radio frequency (RF) circuit 510, a memory 515, an input unit 530, a display unit 540, a sensor 550, an audio circuit 560, a wireless fidelity (WiFi) module 570, a processor 580, and a power supply 590, etc. Those skilled in the art will understand that the terminal 110 structure shown in the figures does not constitute a limitation on a mobile phone or computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0197] The RF circuit 510 can be used to receive and transmit signals during information transmission or calls. In particular, it receives downlink information from the base station and processes it with the processor 580; in addition, it transmits uplink data to the base station.

[0198] The memory 515 can be used to store software programs and modules. The processor 580 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the memory 515.

[0199] The input unit 530 can be used to receive input numeric or character information, and to generate key signal inputs related to the terminal's settings and function control. Specifically, the input unit 530 may include a touch panel 531 and other input devices 532.

[0200] The display unit 540 can be used to display input or provided information, as well as various menus of the terminal. The display unit 540 may include a display panel 541.

[0201] Audio circuit 560, speaker 561, and microphone 562 provide an audio interface.

[0202] In this embodiment, the processor 580 included in the terminal 110 can execute the spatiotemporal game-based collaborative decision-making method of the previous embodiment.

[0203] The terminal 110 in this application embodiment includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, and aircraft. This invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0204] Figure 6This is a partial structural block diagram of a server 120 implementing an embodiment of this application. The server 120 can vary significantly due to different configurations or performance characteristics, and may include one or more central processing units (CPUs) 622 (e.g., one or more processors) and memory 632, and one or more storage media 620 (e.g., one or more mass storage devices) for storing application programs 642 or data 644. The memory 632 and storage media 620 may be temporary or persistent storage. The program stored in the storage media 620 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the server 120. Furthermore, the CPU 622 may be configured to communicate with the storage media 620 and execute the series of instruction operations in the storage media 620 on the server 120.

[0205] Server 120 may also include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0206] The central processing unit 622 in server 120 can be used to execute the spatiotemporal game-based collaborative decision-making method of this application embodiment.

[0207] This application also provides a computer-readable storage medium for storing program code, which is used to execute the spatiotemporal game-based collaborative decision-making method of the foregoing embodiments.

[0208] This application also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, causing the computer device to perform the aforementioned collaborative decision-making method based on spatiotemporal game theory.

[0209] Furthermore, the terms “comprising” and “including”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, apparatus, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are expressly listed, but may include other steps or units that are not expressly listed or that are inherent to such process, method, product or device.

[0210] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0211] It should be understood that in the description of the embodiments of this application, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.

[0212] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.

[0213] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0214] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0215] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0216] It should also be understood that the various implementation methods provided in this application can be combined arbitrarily to achieve different technical effects.

[0217] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0218] The above is a detailed description of the embodiments of this application. However, this application is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A collaborative decision-making method based on spatiotemporal game theory, characterized in that, include: Obtain the set of environmental parameters uploaded by the target vehicle, wherein the set of environmental parameters includes at least road topology information, vehicle driving status information, and road emergency information; The road topology information, vehicle driving status information, and road emergency information are transformed into path distance thresholds and dynamic obstacle probability distributions. A multidimensional constraint matrix is ​​constructed by combining the path distance thresholds and the dynamic obstacle probability distributions. The multidimensional constraint matrix is ​​used to limit the drivable area at different times. The optimization objective is to maximize the coverage of the driving area and minimize the total avoidance distance of multiple vehicles within the driving area. The target path strategy magnitude is simulated and calculated by combining the multidimensional constraint matrix and the preset posterior information. The target path strategy magnitude and the multidimensional constraint matrix are sent to the target vehicle, so that the target vehicle generates a path strategy for the next moment based on the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude under the guidance of the multidimensional constraint matrix.

2. The method according to claim 1, characterized in that, The process of converting the road topology information, vehicle driving status information, and road emergency information into path distance thresholds and dynamic obstacle probability distributions includes: Determine the shortest path distance based on the road topology information; Determine the target distance between the vehicle location of the target vehicle and the event location contained in the road emergency information, and construct a distance influence parameter based on the ratio between the target distance and the influence radius parameter contained in the road emergency information. Generate a path distance threshold based on the shortest path distance and the distance influence parameter. Obtain the current reference driving state information of reference vehicles around the target vehicle, calculate the mean of the predicted position of the reference vehicle at the target prediction time based on the reference driving state information and the vehicle driving state information, and calculate the covariance at the target prediction time based on the preset error coefficient and the preset diffusion coefficient. The predicted location mean and covariance are quantified using a normal distribution method to generate a dynamic obstacle probability distribution.

3. The method according to claim 1, characterized in that, The optimization objective is to maximize the coverage area of ​​the driving area and minimize the total avoidance distance of multiple vehicles within the driving area. Combining the multi-dimensional constraint matrix and preset posterior information, the target path strategy magnitude is simulated and calculated, including: The target feasible region at the current moment is determined based on the multidimensional constraint matrix. Obtain a global utility function, which is used to characterize the optimization objective as maximizing the covered driving area and minimizing the total avoidance distance of multiple vehicles within the driving area. The policy response relationship of the target vehicle relative to the path policy magnitude issued by the local node is obtained, and a local utility function is constructed based on the preset posterior information and the policy response relationship. The local utility function is used to characterize the relationship between the path policy magnitude issued by the local node and the simulated theoretical path of the target vehicle. The simulated theoretical path is obtained by the local node through simulating the response of the target vehicle to the issued path policy. By combining the global utility function and the local utility function, the target path strategy magnitude that conforms to the target feasible region is simulated and calculated.

4. The method according to any one of claims 1 to 3, characterized in that, After sending the target path strategy magnitude and the multidimensional constraint matrix to the target vehicle, the method further includes: Obtain the environmental status parameters uploaded by the target vehicle; When the environmental state parameter is greater than or equal to the preset environmental state threshold, the target path strategy magnitude is updated to obtain the updated optimized strategy magnitude. The optimization strategy magnitude is sent to the target vehicle, so that the target vehicle generates a path strategy for the next moment based on the current path strategy, the real-time collected dynamic environment state, and the optimization strategy magnitude under the guidance of the multidimensional constraint matrix.

5. The method according to claim 4, characterized in that, The step of updating the target path policy magnitude to obtain the updated optimized policy magnitude includes: For the target path policy magnitude, construct the corresponding global utility expectation gradient and the corresponding policy magnitude change regularization term; By combining the expected global utility gradient with the regularization term of the policy magnitude change, a policy magnitude update function is constructed. With the optimization objective of maximizing the value of the policy magnitude update function, the policy magnitude of the target path is updated according to the policy magnitude update function to obtain the updated optimized policy magnitude.

6. The method according to claim 5, characterized in that, Before sending the optimized strategy magnitude to the target vehicle, the method further includes: The initial path strategy of the target vehicle at the current time is obtained, and multiple candidate path strategies corresponding to multiple consecutive time intervals are simulated and generated by combining the initial path strategy and the optimization strategy magnitude. Based on the policy error between each candidate path policy and the initial path policy, the average policy error is determined; The step of sending the optimized strategy magnitude to the target vehicle includes: When the average strategy error is less than or equal to a preset strategy error threshold, the optimized strategy magnitude is sent to the target vehicle.

7. The method according to claim 6, characterized in that, The method further includes: When the average strategy error is greater than a preset strategy error threshold, the difference between the average strategy error and the preset strategy error threshold is determined. Based on the gap value, adjust the learning rate parameter in the regularization term for the change in policy magnitude in the policy magnitude update function to obtain the updated target policy magnitude update function. The optimization objective is to maximize the value of the target policy magnitude update function. The optimized policy magnitude is updated according to the target policy magnitude update function to obtain the updated target optimized policy magnitude. The target optimization strategy magnitude is sent to the target vehicle.

8. A collaborative decision-making device based on spatiotemporal game theory, characterized in that, include: The acquisition unit is used to acquire a set of environmental parameters uploaded by the target vehicle, wherein the set of environmental parameters includes at least road topology information, vehicle driving status information and road emergency information. The generation unit is used to convert the road topology information, vehicle driving status information and road emergency information into path distance thresholds and dynamic obstacle probability distributions, and to construct a multi-dimensional constraint matrix by combining the path distance thresholds and the dynamic obstacle probability distributions. The multi-dimensional constraint matrix is ​​used to limit the drivable area corresponding to different times. The calculation unit is used to simulate and calculate the target path strategy magnitude by taking the maximization of the covered driving area and the minimization of the total avoidance driving distance of multiple vehicles within the driving area as the optimization objective, and combining the multidimensional constraint matrix and preset posterior information. The sending unit is used to send the target path strategy magnitude and the multidimensional constraint matrix to the target vehicle, so that the target vehicle generates a path strategy for the next moment based on the current path strategy, the real-time collected dynamic environmental state, and the target path strategy magnitude under the guidance of the multidimensional constraint matrix.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the collaborative decision-making method based on spatiotemporal game theory as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to execute the spatiotemporal game-based collaborative decision-making method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • System and method for collaboratively navigating, investigating and monitoring unmanned aerial vehicle and intelligent vehicle

    CN104699102A

  • Intelligent driving decision-making method, decision-making device and vehicle

    CN115503756A