Vehicle path planning method, device, equipment and medium

Through the global blockchain system architecture and collaborative training, combined with the data and path planning tasks of multiple vehicles, the problem of failing to consider future traffic conditions in dynamic path planning is solved, the global nature of path planning is achieved, and vehicle travel efficiency is improved.

CN120467379BActive Publication Date: 2025-09-23PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510970314.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-23
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

Existing technologies fail to consider future traffic conditions in dynamic path planning, resulting in a lack of globality in path planning and affecting vehicle travel efficiency.

Method used

By adopting a global blockchain system architecture and collaborative training of multiple regional blockchain systems, combined with data from multiple vehicles and path planning tasks, a global loss function is constructed. The maximization of the global cumulative reward value is used as the optimization goal to guide model training and output a decision path that takes into account future traffic conditions.

Benefits of technology

The overall nature of route planning is achieved, traffic congestion caused by multiple vehicles occupying the same lane at the same time is avoided, and vehicle travel efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120467379B_ABST
    Figure CN120467379B_ABST
Patent Text Reader

Abstract

This application discloses a vehicle path planning method, apparatus, device, and medium. The method is applied to any first-region node in a first-region blockchain system; the method obtains path planning tasks for multiple vehicles in the first region from the first-region blockchain corresponding to the first-region blockchain system; a first vehicle status dataset corresponding to the multiple vehicles is constructed by combining the path planning tasks, vehicle status information, and road condition information of each vehicle; the first vehicle status dataset is input into a trained first-region model to obtain a decision path combination, which includes a decision planning path corresponding to each vehicle; the decision planning path corresponding to each vehicle included in the decision path combination is uploaded to the first-region blockchain, so that each vehicle reads the corresponding decision planning path from the first-region blockchain. In this way, path planning for multiple vehicles can be performed simultaneously by combining multiple path planning tasks, which is global and improves vehicle travel efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a vehicle path planning method, apparatus, device, and medium. Background Art

[0002] Path planning is a core technology in autonomous driving. Path planning includes static path planning, which involves calculating the optimal path based on preset constraints (such as shortest distance, minimum time, etc.) within known environmental information to plan the vehicle's driving path. However, static path planning assumes that known environmental information remains unchanged. However, environmental information generally changes in real time and contains uncertainties such as road congestion, increased traffic flow, temporary road closures, and other emergencies. This can prevent the path obtained from static path planning from being applied as intended, hindering vehicle travel.

[0003] The relevant technology uses dynamic path planning to continuously adjust the optimal path based on real-time feedback data from sensors, communication networks, etc. when environmental information changes in real time or there is uncertainty.

[0004] However, when performing dynamic path planning, related technologies mainly rely on current environmental observation data and fail to consider the impact of future traffic conditions. For example, due to dynamic traffic flow changes, the originally planned optimal path may become congested in the future, that is, it is not the optimal path in the future. The path planning lacks globality, which affects vehicle travel efficiency. Summary of the Invention

[0005] The embodiments of the present application provide a vehicle path planning method, apparatus, device, and medium, which can combine multiple path planning tasks and multiple vehicle data to simultaneously plan paths for multiple vehicles, thereby taking into account future traffic conditions, making path conversations global, and improving vehicle travel efficiency.

[0006] In a first aspect, the present application provides a vehicle path planning method, which is applied to any first regional node in a first regional blockchain system, wherein the first regional node and any second regional node in a second regional blockchain system corresponding to a second region form a global blockchain system, including:

[0007] Obtaining, from the first regional blockchain corresponding to the first regional blockchain system, path planning tasks for multiple vehicles within the first regional area, wherein each path planning task is initiated by each vehicle and uploaded to the first regional blockchain after consensus verification by the first regional blockchain system;

[0008] Determining the straight-line length of each vehicle's path based on the starting point information and the end point information carried by each path planning task, and determining the current vehicle state information of each vehicle and surrounding road condition information, and constructing a first vehicle state data set corresponding to the plurality of vehicles by combining the straight-line length of each vehicle's path, the vehicle state information, and the road condition information;

[0009] Inputting the first vehicle state data set into the trained first regional model to obtain a decision path combination, wherein the decision path combination includes a decision planning path corresponding to each vehicle;

[0010] wherein, after the first regional model outputs a first predicted decision path combination based on a first sample vehicle state data set, the first predicted decision path combination is combined with the first sample vehicle state data set, the first predicted decision path combination, and the second sample vehicle state data set and the second predicted decision path combination corresponding to the second region to construct a first regional loss, and the first regional loss and the second regional loss corresponding to the second region are combined to construct a global loss, and the maximization of the global cumulative reward value is used as the optimization goal to guide the collaborative training of the first regional model and the second regional model corresponding to the second region, where the global cumulative reward value is negatively correlated with the global loss;

[0011] The second sample vehicle state data set, the second predicted decision path combination, and the second regional loss are shared by the second regional node to the global blockchain corresponding to the global blockchain system, and then obtained by the local node from the global blockchain;

[0012] The decision planning path corresponding to each vehicle included in the decision path combination is uploaded to the first regional blockchain, so that each vehicle reads the corresponding decision planning path from the first regional blockchain.

[0013] In a second aspect, the present application provides a vehicle path planning device, which is applied to any first regional node in a first regional blockchain system, wherein the first regional node and any second regional node in a second regional blockchain system corresponding to a second region form a global blockchain system, including:

[0014] an acquiring unit, configured to acquire, from a first regional blockchain corresponding to the first regional blockchain system, path planning tasks for a plurality of vehicles within the first regional area, wherein each path planning task is initiated by each vehicle and uploaded to the first regional blockchain after consensus verification by the first regional blockchain system;

[0015] a determining unit, configured to determine a straight-line length of a path for each vehicle based on the starting point information and the ending point information carried by each path planning task, and to determine current vehicle status information of each vehicle and surrounding road condition information, and to construct a first vehicle status dataset corresponding to the plurality of vehicles by combining the straight-line length of the path for each vehicle, the vehicle status information, and the road condition information;

[0016] An input unit, configured to input the first vehicle state data set into the trained first regional model to obtain a decision path combination, wherein the decision path combination includes a decision planning path corresponding to each vehicle;

[0017] wherein, after the first regional model outputs a first predicted decision path combination based on a first sample vehicle state data set, the first predicted decision path combination is combined with the first sample vehicle state data set, the first predicted decision path combination, and the second sample vehicle state data set and the second predicted decision path combination corresponding to the second region to construct a first regional loss, and the first regional loss and the second regional loss corresponding to the second region are combined to construct a global loss, and the maximization of the global cumulative reward value is used as the optimization goal to guide the collaborative training of the first regional model and the second regional model corresponding to the second region, where the global cumulative reward value is negatively correlated with the global loss;

[0018] The second sample vehicle state data set, the second predicted decision path combination, and the second regional loss are shared by the second regional node to the global blockchain corresponding to the global blockchain system, and then obtained by the local node from the global blockchain;

[0019] The sending unit is used to upload the decision planning path corresponding to each vehicle included in the decision path combination to the first regional blockchain, so that each vehicle reads the corresponding decision planning path from the first regional blockchain.

[0020] In some embodiments, the sending unit is further configured to:

[0021] Calculating a path hash value for each decision-planning path corresponding to each vehicle included in the decision-planning path combination to obtain a path hash value combination, wherein the hash value combination includes a path hash value corresponding to each decision-planning path;

[0022] Sending the path hash value combination to other first-region nodes in the first-region blockchain system for consensus verification to obtain a consensus verification result;

[0023] When it is determined according to the consensus verification result that the number of target first-area nodes that have reached consensus is greater than a preset node number threshold, the path hash value combination and the decision planning path corresponding to each vehicle included in the decision path combination are uploaded to the first-area blockchain.

[0024] In some embodiments, the sending unit is further configured to:

[0025] Constructing a path hash tree according to the path hash value corresponding to each decision-making planning path included in the path hash value combination, wherein the path hash tree includes a path root hash;

[0026] Determine the generation timestamp information, the receiver identifier, and the sender identifier associated with each decision-making plan path, and generate a target block based on the path hash tree, the path root hash, the decision-making plan path corresponding to each vehicle included in the decision-making path combination, and the generation timestamp information, the receiver identifier, and the sender identifier associated with each decision-making plan path;

[0027] Sending the block header corresponding to the target block to other first-region nodes in the first-region blockchain system for consensus verification to obtain a consensus verification result, wherein the block header includes the generation timestamp information, the receiver identifier, the sender identifier, and the path root hash;

[0028] The sending unit is further applied to add the target block to the first regional blockchain.

[0029] In some embodiments, the first area model includes the first decision sub-model and the first evaluation sub-model, and the vehicle path planning device further includes a training unit for:

[0030] Obtaining a first sample vehicle state data set, and inputting the first sample vehicle state data set into the first decision sub-model to obtain a first predicted decision path combination;

[0031] Obtaining a second sample vehicle state data set and a second predicted decision path combination corresponding to the second area, and inputting the first sample vehicle state data set, the first predicted decision path combination, the second sample vehicle state data set, and the second predicted decision path combination into the first evaluation sub-model to obtain a first evaluation score;

[0032] Obtaining a first evaluation expected value, and constructing a first regional loss according to a difference between the first evaluation expected value and the first evaluation score;

[0033] Obtaining a second area loss corresponding to the second area, and determining a global cumulative reward value based on the first area loss and the second area loss using a preset global loss reward function;

[0034] Taking maximizing the global cumulative reward value as an optimization goal, training the first decision sub-model and the first evaluation sub-model to obtain a trained first region model;

[0035] The global cumulative reward value is maximized when the global loss corresponding to the first area loss and the second area loss is minimized.

[0036] In some embodiments, the vehicle path planning device further includes an area loss acquisition unit configured to:

[0037] Uploading the first sample vehicle state dataset and the first predicted decision path combination to the global blockchain, so that the second regional node obtains the first sample vehicle state dataset and the first predicted decision path combination from the global blockchain, and calculates a second evaluation score based on the first sample vehicle state dataset, the first predicted decision path combination, the second sample vehicle state dataset, and the second predicted decision path combination using a second regional model corresponding to the second region, and constructs a second regional loss based on the second evaluation score and a preset second evaluation expected value;

[0038] The training unit is further used to obtain the second area loss corresponding to the second area from the global blockchain, wherein the second area loss is uploaded to the global blockchain by the second area node.

[0039] In some embodiments, the training unit is further configured to:

[0040] Determining the number of vehicles and the road occupancy rate corresponding to the first area according to the first sample vehicle status data set, and calculating a current global cumulative reward value according to the number of vehicles, the road occupancy rate, and the first sample vehicle status data set;

[0041] Taking the current global cumulative reward value as the initial value, adjust the model parameters of the first decision sub-model and the first evaluation sub-model, and perform iterative training in an optimization direction in which the global cumulative reward value is positively increased until the maximum value of the global cumulative reward value is achieved when the global loss is minimized, thereby obtaining a trained first region model.

[0042] In some embodiments, the training unit is further configured to:

[0043] Obtaining from the global blockchain the number of vehicles in the second area and the second area road occupancy rate corresponding to the second area, the second area vehicle number and the second area road occupancy rate being determined by the second area node based on the second vehicle status dataset and uploaded to the global blockchain;

[0044] Calculate the current global cumulative reward value according to the number of vehicles in the second area, the road occupancy rate in the second area, and the first sample vehicle status data set;

[0045] Taking the current global cumulative reward value as the initial value, adjust the model parameters of the first decision sub-model and the first evaluation sub-model, and perform iterative training in an optimization direction in which the global cumulative reward value is positively increased until the maximum value of the global cumulative reward value is achieved when the global loss is minimized, thereby obtaining a trained first region model.

[0046] In addition, an embodiment of the present application also provides a computer device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the above-mentioned vehicle path planning method is implemented.

[0047] In addition, an embodiment of the present application also provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for loading by a processor to execute the above-mentioned vehicle path planning method.

[0048] The embodiment of the present application is applied to any first area node in the first area blockchain system, and the first area node and any second area node in the second area blockchain system corresponding to the second area form a global blockchain system; by obtaining the path planning tasks of multiple vehicles in the first area from the first area blockchain corresponding to the first area blockchain system, wherein each path planning task is initiated by each vehicle and uploaded to the first area blockchain after consensus verification by the first area blockchain system; the straight-line length of the path of each vehicle is determined according to the starting point information and end point information carried by each path planning task, and the current vehicle status information of each vehicle and the surrounding road condition information are determined, and a first vehicle status data set corresponding to multiple vehicles is constructed in combination with the straight-line length of the path, vehicle status information and road condition information of each vehicle; the first vehicle status data set is input into the trained first area model to obtain a decision path combination, and the decision path combination includes the decision planning path corresponding to each vehicle; wherein, the first area model After outputting the first predicted decision path combination based on the first sample vehicle state data set, the model combines the first sample vehicle state data set, the first predicted decision path combination, the second sample vehicle state data set corresponding to the second area, and the second predicted decision path combination to construct the first area loss, and combines the first area loss and the second area loss corresponding to the second area to construct the global loss, and uses the maximization of the global cumulative reward value as the optimization goal to guide the first area model and the second area model corresponding to the second area to perform collaborative training, and the global cumulative reward value is negatively correlated with the global loss; wherein, after the second sample vehicle state data set, the second predicted decision path combination and the second area loss are shared by the second area node to the global blockchain corresponding to the global blockchain system, the local node obtains them from the global blockchain; the decision planning path corresponding to each vehicle contained in the decision path combination is uploaded to the first area blockchain, so that each vehicle reads the corresponding decision planning path from the first area blockchain.

[0049] As can be seen from the above, the present application includes a two-layer blockchain system architecture, namely the regional blockchain system layer and the global blockchain system layer. The regional blockchain system is used to manage the traffic path planning of the corresponding area, and the global blockchain system is used to manage and coordinate the traffic path planning of all areas. Taking the first regional blockchain as an example, first, the vehicles in the first area can submit path planning tasks to the first regional blockchain system to be recorded on the first regional blockchain. Then, the first regional node can obtain the path planning tasks of multiple vehicles from the first regional blockchain, determine the straight-line length of the path of each vehicle according to each path planning task, and construct a first vehicle status data set for multiple vehicles in combination with the straight-line length of the path of each vehicle, vehicle status information, surrounding path information, etc., and then input the first vehicle status data set containing the status data of multiple vehicles into the trained first regional model to obtain a decision path combination, each decision path combination containing the decision planning path of each vehicle; it should be noted that since the first regional model is a combination of at least one The training process is collaboratively trained with models in other regions (such as the second regional model). During the collaborative training process, the first sample vehicle state dataset and the first predicted decision path combination in the first region are combined with the sample vehicle state datasets and corresponding predicted decision path combinations in the other regions to construct the first regional loss. That is, each regional model combines local sample data and sample data from other regions to construct a local regional loss. The first regional loss and the losses of the other regions are combined to construct a global loss. A global cumulative reward value is constructed as a supervisory guide for the training process, and maximizing the global cumulative reward value is the optimization goal of the training process. This allows the first regional model to be collaboratively trained with the regional models in the other regions. Therefore, the trained first regional model can consider multiple vehicle scenarios when outputting a decision path combination during path planning, avoiding traffic congestion caused by multiple vehicles occupying the same lane at the same time. Finally, the decision path combination is uploaded to the blockchain for access by all vehicles in the first region. In this way, multiple path planning tasks and multiple vehicle data are combined to simultaneously plan paths for multiple vehicles, thus considering future traffic conditions, making path planning global and improving vehicle travel efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0051] Figure 1 A schematic diagram of a scenario of a vehicle path planning system provided in an embodiment of the present application;

[0052] Figure 2 A schematic diagram of the steps of the vehicle path planning method provided in an embodiment of the present application;

[0053] Figure 3 An example diagram of the dynamic path planning architecture of the federated blockchain system provided in an embodiment of the present application;

[0054] Figure 4 This is a diagram of the collaborative training architecture of regional models provided in the embodiment of the present application;

[0055] Figure 5 A schematic diagram of the structure of a vehicle path planning device provided in an embodiment of the present application;

[0056] Figure 6 A schematic diagram of the structure of a terminal provided in an embodiment of the present application;

[0057] Figure 7 A schematic diagram of the structure of the server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of this application.

[0059] It can be understood that in the specific implementation of this application, it involves path planning tasks, starting point information and end point information, vehicle status information, surrounding road conditions information, vehicle status data sets, decision path combinations, global cumulative reward values ​​and other related data. When the above embodiments of this application are applied to specific products or technologies, it is necessary to obtain the object's permission or consent, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards.

[0060] In addition, when the embodiment of the present application needs to obtain relevant data, it will obtain separate permission or separate consent for path planning tasks, starting point information and end point information, vehicle status information, surrounding road conditions information, vehicle status data set, decision path combination, global cumulative reward value and other related data through pop-up windows or jumping to a confirmation page. After clearly obtaining separate permission or separate consent for path planning tasks, starting point information and end point information, vehicle status information, surrounding road conditions information, vehicle status data set, decision path combination, global cumulative reward value and other related data, the necessary data for the normal operation of the embodiment of the present application will be obtained.

[0061] It should be noted that some processes described in the specification, claims, and figures above include multiple steps that appear in a specific order. However, it should be understood that these steps may be executed in a different order than the order in which they appear herein or in parallel. The step numbers are used solely to distinguish between the different steps and do not themselves represent any order of execution. Furthermore, terms such as "first," "second," or "target" are used herein to distinguish similar objects and are not necessarily used to describe a specific order or precedence.

[0062] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0063] The embodiments of the present application provide a vehicle path planning method, apparatus, device and medium. Specifically, the vehicle path planning method of the embodiments of the present application can be implemented by a computer device, which can be a server or a user terminal device. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The user terminal device can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, smart home appliance, vehicle-mounted terminal, intelligent voice interaction device, aircraft, drone, etc., but is not limited to this.

[0064] A vehicle path planning method provided by an embodiment of the present application is applied to any first area node in a first area blockchain system, and the first area node and any second area node in a second area blockchain system corresponding to the second area form a global blockchain system. Specifically, the path planning tasks of multiple vehicles in the first area are obtained from the first area blockchain corresponding to the first area blockchain system, wherein each path planning task is initiated by each vehicle and uploaded to the first area blockchain after consensus verification by the first area blockchain system; the straight-line length of the path of each vehicle is determined according to the starting point information and end point information carried by each path planning task, and the current vehicle status information of each vehicle and the surrounding road condition information are determined, and a first vehicle status data set corresponding to multiple vehicles is constructed by combining the straight-line length of the path, vehicle status information and road condition information of each vehicle; the first vehicle status data set is input into the trained first area model to obtain a decision path combination, and the decision path combination includes the decision planning path corresponding to each vehicle; wherein, after the first area model outputs a first predicted decision path combination based on the first sample vehicle status data set, it combines the first sample vehicle status data set with the first sample vehicle status data set. The vehicle state dataset, the first predicted decision path combination, the second sample vehicle state dataset corresponding to the second region, and the second predicted decision path combination are combined to construct a first region loss. The first region loss and the second region loss corresponding to the second region are combined to construct a global loss. The global cumulative reward value is maximized as the optimization goal to guide the first region model and the second region model corresponding to the second region to be collaboratively trained. The global cumulative reward value is negatively correlated with the global loss. The second sample vehicle state dataset, the second predicted decision path combination, and the second region loss are shared by the second region node to the global blockchain corresponding to the global blockchain system, and then the local node obtains them from the global blockchain. The decision planning path corresponding to each vehicle included in the decision path combination is uploaded to the first region blockchain, so that each vehicle reads the corresponding decision planning path from the first region blockchain. Please refer to the following specific embodiments for details.

[0065] It should be noted that the vehicle path planning method can be executed jointly by the terminal and the server.

[0066] For example, taking the vehicle path planning method executed by the terminal and the server as an example, see Figure 1 , is a scenario diagram of the information push system provided in an embodiment of the present application, the system includes a terminal 110 and a server 120.

[0067] Among them, taking the terminal 110 as a vehicle (or a vehicle-mounted terminal) as an example, the target application can be installed on the vehicle, and the corresponding application business can be run through the target application, such as requesting a path planning task. The terminal 110 can collect vehicle status information and surrounding road conditions information, and upload the path planning task, vehicle status information and surrounding road conditions information to the server 120.

[0068] Among them, the server 120 can be a single service node, a distributed system composed of multiple service nodes, or a service node in a distributed system. Further, taking the blockchain technology scenario as an example, the distributed system includes multiple service nodes. The multiple service nodes in the distributed system can be divided into service nodes corresponding to different regions according to geographical location. The multiple service nodes corresponding to each region together constitute a regional blockchain system, and each service node serves as one of the blockchain nodes in the corresponding regional blockchain system. For example, assuming that the multiple service nodes in the distributed system are divided into service nodes belonging to the first region, service nodes belonging to the second region, service nodes belonging to the third region, ... service nodes belonging to the nth region according to geographical location, it should be noted that the multiple service nodes belonging to the first region together constitute the first regional blockchain system, the multiple service nodes belonging to the second region together constitute the second regional blockchain system, the multiple service nodes belonging to the third region together constitute the third regional blockchain system, and the multiple service nodes belonging to the nth region together constitute the nth regional blockchain system. Furthermore, it should be noted that the present application includes a two-layer blockchain system architecture. The first-layer blockchain system includes the regional blockchain systems of the above-mentioned regions, and the second-layer blockchain system is a global blockchain system. The global blockchain system includes at least one service node in the regional blockchain system corresponding to each region, that is, the global regional blockchain system is composed of any service node in each regional blockchain system.

[0069] For example, taking server 120 as the first regional node in the first regional blockchain system, the first regional node can be understood as the first regional blockchain node, then the first regional node and any second regional node in the second regional blockchain system corresponding to the second region constitute a global blockchain system.

[0070] The server 120 executes the steps of the vehicle path planning method. Specifically, the server 120 can obtain the path planning tasks of multiple vehicles in the first area from the first area blockchain corresponding to the first area blockchain system, wherein each path planning task is initiated by each vehicle and uploaded to the first area blockchain after consensus verification by the first area blockchain system; determine the straight line length of each vehicle's path according to the starting point information and end point information carried by each path planning task, and determine the current vehicle status information of each vehicle and the surrounding road condition information, and construct a first vehicle status data set corresponding to multiple vehicles in combination with the straight line length of each vehicle's path, vehicle status information and road condition information; input the first vehicle status data set into the trained first area model to obtain a decision path combination, which includes the decision planning path corresponding to each vehicle; wherein the first area model is based on the first sample vehicle After the vehicle state dataset outputs the first predicted decision path combination, the first regional loss is constructed by combining the first sample vehicle state dataset, the first predicted decision path combination, the second sample vehicle state dataset corresponding to the second region, and the second predicted decision path combination. The global loss is constructed by combining the first regional loss and the second regional loss corresponding to the second region. The global cumulative reward value is maximized as the optimization goal to guide the first regional model and the second regional model corresponding to the second region to be collaboratively trained. The global cumulative reward value is negatively correlated with the global loss. After the second sample vehicle state dataset, the second predicted decision path combination, and the second regional loss are shared by the second regional node to the global blockchain corresponding to the global blockchain system, the local node obtains them from the global blockchain. The decision planning path corresponding to each vehicle included in the decision path combination is uploaded to the first regional blockchain. Thereafter, each vehicle reads the corresponding decision planning path from the first regional blockchain, which is not limited here.

[0071] Therefore, the present application includes a two-layer blockchain system architecture, namely a regional blockchain system layer and a global blockchain system layer. The regional blockchain system is used to manage the traffic path planning of the corresponding area, and the global blockchain system is used to manage and coordinate the traffic path planning of all areas. Taking the first regional blockchain as an example, first, vehicles in the first area can submit path planning tasks to the first regional blockchain system to be recorded on the first regional blockchain. Then, the first regional node can obtain the path planning tasks of multiple vehicles from the first regional blockchain, determine the straight-line length of each vehicle's path according to each path planning task, and construct a first vehicle status data set for multiple vehicles in combination with the straight-line length of each vehicle's path, vehicle status information, surrounding path information, etc. Then, the first vehicle status data set containing the status data of multiple vehicles is input into the trained first regional model to obtain a decision path combination, each decision path combination containing the decision planning path of each vehicle; it should be noted that since the first regional model is combined with at least one The first regional model is trained collaboratively with the models of other regions (such as the second regional model). During the collaborative training process, the first sample vehicle state dataset and the first predicted decision path combination of the first region are combined with the sample vehicle state datasets and the corresponding predicted decision path combinations of the other regions to construct the first regional loss. That is, each regional model combines local sample data and sample data from other regions to construct a local regional loss, and the first regional loss is combined with the losses of the other regions to construct a global loss. A global cumulative reward value is constructed as a supervisory guide for the training process, and maximizing the global cumulative reward value is used as the optimization goal of the training process. This allows the first regional model to be collaboratively trained with the regional models of the other regions. Therefore, the trained first regional model can consider multiple vehicle scenarios when outputting decision path combinations during path planning, avoiding traffic congestion caused by multiple vehicles occupying the same lane at the same time. Finally, the decision path combinations are uploaded to the blockchain for access by all vehicles in the first region. In this way, multiple path planning tasks and multiple vehicle data are combined to simultaneously plan paths for multiple vehicles, thereby considering future traffic conditions, making path planning global and improving vehicle travel efficiency.

[0072] For ease of understanding, each step of the vehicle path planning method will be described in detail below. It should be noted that the order of the following embodiments is not intended to limit the preferred order of the embodiments.

[0073] See also Figure 2 , Figure 2 This is a flowchart of the steps of the vehicle path planning method provided in an embodiment of the present application. In the embodiment of the present application, the vehicle path planning method can be executed by a computer device, such as a server or a terminal. The specific process is as follows:

[0074] 101. Obtain path planning tasks for multiple vehicles in a first area from a first area blockchain corresponding to a first area blockchain system.

[0075] Path planning is a core technology in autonomous driving. Path planning includes static path planning, which involves calculating the optimal path based on preset constraints (such as shortest distance, minimum time, etc.) within known environmental information to plan the vehicle's driving path. However, static path planning assumes that known environmental information remains unchanged. However, environmental information generally changes in real time and contains uncertainties such as road congestion, increased traffic flow, temporary road closures, and other emergencies. This can prevent the path obtained from static path planning from being applied as intended, hindering vehicle travel.

[0076] Related technologies use dynamic path planning to continuously adjust the optimal path based on real-time feedback from sensors, communication networks, and other sources, even when environmental information is changing or uncertain. However, these dynamic path planning processes primarily rely on current environmental observation data and fail to consider the impact of future traffic conditions. For example, dynamic traffic flow changes could cause the originally planned optimal path to become congested in the future, making it suboptimal for the future. This lack of comprehensiveness in path planning can impact vehicle travel efficiency.

[0077] In order to address the above problems, an embodiment of the present application obtains vehicle status information, road condition information, destination information, etc. of multiple vehicles in a geographical location area through blockchain to construct a vehicle status data set of multiple vehicles in the geographical location area, and outputs a decision path combination based on the vehicle status data set through a regional model corresponding to the local area (such as a first regional model corresponding to the first area). The decision path combination includes the decision planning path corresponding to each vehicle, and thus, the decision planning path corresponding to each vehicle is returned to each vehicle through blockchain. In this way, when planning the path, the decision path combination outputted can be considered from the perspective of multiple vehicles, avoiding the situation where multiple vehicles occupy the same lane at the same time, resulting in traffic congestion in the future scenario, thereby considering the future traffic status, making the path planning global and improving the vehicle travel efficiency.

[0078] When executing the vehicle path planning method, the embodiment of the present application can use blockchain technology to transmit relevant data. For example, each vehicle can upload its own path planning tasks, vehicle status information, road condition information, etc. to the server through the blockchain, and the server can return the path planned for each vehicle to each vehicle through the blockchain. The server acts as a blockchain node in the blockchain technology.

[0079] It should be noted that in order to avoid executing the vehicle's path planning task through cross-regional service nodes, for example, the path planning of the vehicle in area A is sent to the service node in area B for path planning. This will not only increase the transmission time and communication overhead of the path planning task, occupy more network bandwidth resources, thereby causing network congestion, but also because cross-regional processing requires passing through more routing links, it puts greater pressure on network bandwidth resources, reducing the efficiency of vehicle path planning. In this regard, the embodiment of the present application adopts a regional processing method for vehicle path planning tasks. For example, the service node in area A processes the path planning tasks of vehicles located in area A, and the service node in area B processes the path planning tasks of vehicles located in area B. In this way, the phenomenon of executing the vehicle's path planning task through cross-regional service nodes can be avoided, reducing network communication overhead and pressure on network bandwidth resources, thereby improving the efficiency of vehicle path planning. Specifically, the blockchain system in the embodiment of the present application can be constructed according to geographical location areas. For example, multiple service nodes in a geographical location area jointly construct a blockchain system for the area, which can be called a regional blockchain system. In this way, each geographical location area corresponds to a regional blockchain system. Assuming that there are n areas divided according to geographical location, it includes a first regional blockchain system corresponding to the first area, a second regional blockchain system corresponding to the second area, a third regional blockchain system corresponding to the third area, ... an nth blockchain system corresponding to the nth area. The regional blockchain system corresponding to each area is used to manage the traffic path planning of the corresponding area. On this basis, in order to enable the sharing of traffic data or related information between all areas, any service node in each area can be combined to form a global blockchain system. Therefore, the blockchain system architecture of the embodiment of the present application includes a two-layer blockchain system. The first layer contains regional blockchain systems for each local area, that is, regional blockchain systems for multiple areas. The second layer contains a global blockchain system. The global blockchain system can be located in the cloud and can be understood as a cloud-based global blockchain system.

[0080] For example, taking the first-region blockchain system corresponding to the first region as an example, assuming that the server of the embodiment of the present application is one of the first-region nodes (first blockchain node) constituting the first-region blockchain system, assuming that the other regions are second regions, that is, the second region generally refers to other regions except the first region. Similarly, the second-region blockchain system is composed of the second-region nodes in the region where it is located, and the first-region node and the second-region nodes in other second-region blockchain systems together constitute the global blockchain system. It should be noted that the above-mentioned other second regions can be multiple other regions. For details, please refer to the description of the "first-region blockchain system", "second-region blockchain system", "nth-region blockchain system" and "global blockchain system" in the previous embodiment, and they are not listed here one by one.

[0081] The above regional blockchain system and the global blockchain system as a whole can be called a federal blockchain system. The embodiment of the present application can perform dynamic path planning based on the list blockchain system. Figure 3 The example diagram of the dynamic path planning architecture of the federated blockchain system provided in this application embodiment is combined with Figure 3 As shown, dynamic path planning based on a federated blockchain system can improve the privacy and security of vehicle data (such as path planning tasks, decision-making paths, or vehicle status data), and ensure the dynamic real-time and global efficiency of decision-making paths during the dynamic path planning process. During the dynamic path planning process, distributed planning can be performed based on regional blockchain systems in multiple regions. Specifically, since the federated blockchain system of the embodiment of the present application includes regional blockchain systems divided by regions and a global blockchain system located above the regional blockchain systems, it belongs to a secure collaborative architecture. Therefore, for the path planning tasks of multiple vehicles in each region, a blockchain sharding method can be adopted (the path planning tasks are assigned to the regional nodes of each corresponding regional blockchain system according to the geographical location information of the starting point and the end point, namely the "sharding strategy" in the figure; and the corresponding consensus mechanism), combined with a federated learning method, so that the regional models corresponding to each regional blockchain system can have personalized and corresponding weight distribution strategies. In addition, distributed planning also involves "collaborative decision-making distributed planning", which mainly uses reinforcement learning to collaboratively train multiple agents (regional models corresponding to multiple regions), such as using the "MADPG algorithm" to adjust the model parameters of each regional model to optimize the loss. According to the above structure, multi-vehicle path planning is performed to improve vehicle travel efficiency.

[0082] In the embodiments of the present application, a region refers to a local area divided by geographic location. For example, it can be a geographical region divided by province, prefecture-level city, region, or street division. The specific division granularity can be determined according to actual circumstances and is not limited here. The path planning of vehicles in each region can be processed by one or more regional nodes (i.e., blockchain nodes) in the corresponding regional blockchain system. Therefore, vehicles traveling in each region can upload their path planning tasks, vehicle status information, and surrounding road condition information to the regional blockchain. For example, taking the first region as an example, each vehicle in the first region can upload its path planning task to the first regional blockchain. In this way, the first regional blockchain corresponding to the first regional blockchain system can record the path planning tasks of each vehicle in the first region. Thereafter, the first regional node can obtain the path planning tasks of multiple vehicles in the first region from the first regional blockchain corresponding to the first regional blockchain system, so as to subsequently perform path planning based on the path planning tasks of multiple vehicles, the vehicle status information of each vehicle, and the surrounding road condition information. In this way, multiple path planning tasks and multiple vehicle data are combined to simultaneously plan paths for multiple vehicles, thereby considering future traffic conditions, making path planning global and improving vehicle travel efficiency.

[0083] For example, the car owner enters the destination information on the target application (such as a related navigation application or map application) on the vehicle terminal, and can also set the starting point information. After confirmation, a path planning request can be generated. The path planning request is the path planning task. The path planning task includes but is not limited to the starting point information and end point information of the current vehicle, that is, the starting address information and the destination address information. After each vehicle submits the path planning task, it will be uploaded to the regional blockchain corresponding to the local area. For example, taking the first area as an example, the path planning task submitted by each vehicle in the first area will be uploaded to the first regional blockchain corresponding to the first regional blockchain system.

[0084] In some implementations, clustering algorithms (such as DBSCAN) can be used to adjust region boundaries in real time based on traffic density (e.g., narrowing the region when vehicles are dense in urban areas during the morning rush hour and expanding it when vehicles are sparse in suburban areas). This allows blockchain node coverage to align with traffic load and reduce cross-regional data exchange latency. For example, the vehicle flow density within a first region is determined based on the starting and ending information of each task plan. The vehicle flow density reflects the number of vehicles required to travel within the first region. When the vehicle flow density exceeds a preset vehicle flow density threshold, the first region boundaries are adjusted in real time using a clustering algorithm, so that vehicle planning tasks originally belonging to the first region are assigned to the second region blockchain system for path planning.

[0085] Each path planning task is initiated by each vehicle and uploaded to the first-region blockchain after consensus verification by the first-region blockchain system. Specifically, each vehicle can send its own path planning task to any first-region node within the first-region blockchain system. During the transmission, to ensure the security of the path planning task, the vehicle terminal can sign the path planning task using the vehicle private key in the public-private key pair to obtain first vehicle signature information, and then package the path planning task and the first vehicle signature information together and send them to any first-region node. After receiving the path planning task and the first vehicle signature information, the first-region node can decrypt the first vehicle signature information using the vehicle public key in the public-private key pair of the current vehicle to obtain first decrypted information, and compare the first decrypted information with the path planning task of the current vehicle. If the two are consistent, the verification is passed. At this time, the first-region node can send the path planning task of the current vehicle, the first vehicle signature information, and the vehicle public key to other first-region nodes in the first region for consensus verification. If the consensus verification is passed, the path planning task of the current vehicle is uploaded to the first-region blockchain for execution.

[0086] In addition, each vehicle can also directly upload its own path planning tasks to the first area blockchain, which is not limited here.

[0087] Through the above method, the path planning tasks of multiple vehicles in the first area can be obtained from the first area blockchain corresponding to the first area blockchain system, so that path planning can be performed subsequently based on the path planning tasks of multiple vehicles and the vehicle status information of each vehicle and the surrounding road condition information. In this way, it is possible to combine multiple path planning tasks and multiple vehicle data to simultaneously perform path planning for multiple vehicles, thereby taking into account future traffic conditions, making path planning global and improving vehicle travel efficiency.

[0088] 102. Determine the straight-line length of each vehicle's path based on the starting point information and end point information carried by each path planning task, and determine the current vehicle status information of each vehicle and the surrounding road condition information. Combine the straight-line length of each vehicle's path, vehicle status information, and road condition information to construct a first vehicle status data set corresponding to multiple vehicles.

[0089] In an embodiment of the present application, after obtaining the path planning task corresponding to each vehicle, various vehicle-related data can also be obtained to construct a vehicle status data set in this area. For example, the straight-line length of the path of each vehicle can be constructed based on the starting point information and end point information carried by the path planning task of each vehicle, and the vehicle status information and road condition information of each vehicle can be obtained. Then, the straight-line length of the path of each vehicle, the vehicle status information and road condition information represent the vehicle status data of the corresponding vehicle. Combined with the vehicle status data of each vehicle, a vehicle status data set corresponding to multiple vehicles in the area is constructed. For example, a first vehicle status data set corresponding to multiple vehicles in the first area is constructed, and a second vehicle status data set corresponding to multiple vehicles in the second area is constructed for the second area. They are not listed one by one here. In this way, it is possible to subsequently realize the combination of multiple vehicle data to simultaneously perform path planning for multiple vehicles in the area, avoid the phenomenon that a single vehicle path planning leads to future road traffic congestion, and thus consider future traffic conditions, so that path planning has globality and improves vehicle travel efficiency.

[0090] Among them, the starting point information and the end point information can be at least the address information carried by the corresponding vehicle when initiating the path planning task (request), indicating the driving target of the corresponding vehicle. The starting point information and the end point information can be specifically in the form of an actual address, for example, it can be "No. xx, xx Road, xx Street, xx District, xx City, xx Province", or it can be in the form of longitude and latitude. In addition, it can also be in the form of world coordinates, which is not limited here.

[0091] The straight-line length of the path can be a straight-line length determined based on the starting point information and the end point information of the corresponding vehicle, and can represent the straight-line distance between the starting point and the end point. It can be understood that the longer the straight-line length of the path, the farther the end point is from the starting point. Specifically, after obtaining the path planning task for each vehicle, the straight-line length of the path can be directly calculated based on the starting point information and the end point information carried by the path planning task for each vehicle. For example, taking the starting point information and the end point information as world coordinates, the straight-line distance between the two coordinates can be directly calculated based on the starting point world coordinates and the end point world coordinates to obtain the straight-line length of the path.

[0092] Among them, the vehicle status information can be a set of parameters representing the current state of the corresponding vehicle. For example, the vehicle status information can include the vehicle's current direction state parameters, the coordinate parameters of the next intersection on the road where the vehicle is currently located, predicted time parameters, etc. The predicted time parameters can be the sending time of the path planning task of the corresponding vehicle, and / or the predicted path planning travel time determined based on the straight length of the path and the limited speed of the corresponding actual road. These parameter sets together represent the state of the vehicle. The road condition information around the vehicle can be information representing the road conditions around the vehicle's current position. The road condition information specifically reflects the road congestion around the vehicle's current position, pedestrian trajectories, and the trajectories of the vehicle and surrounding vehicles. That is, the road condition information can include road congestion parameters, pedestrian trajectory parameters, and vehicle trajectory parameters, which are not limited here.

[0093] Specifically, the vehicle status information and path information can be uploaded by the corresponding vehicle to the blockchain so that the first-region node (i.e., server) can obtain it. For example, each vehicle can send its own vehicle status information and path information to any first-region node within the first-region blockchain system. During the transmission, to ensure the security of the path planning task, the vehicle terminal can use the vehicle private key to sign its own vehicle status information and path information to obtain a second vehicle signature information, and then package the vehicle status information, path information, second vehicle signature information, and vehicle public key together and send it to any first-region node. After receiving the vehicle status information, path information, and second vehicle signature information, the first-region node can decrypt the second vehicle signature information using the vehicle public key in the current vehicle's public-private key pair to obtain a second decrypted information, and compare the second decrypted information with the vehicle status information and path information of the current vehicle. If the two are consistent, the verification is passed. At this time, the first-region node can send the vehicle status information, path information, second vehicle signature information, and vehicle public key of the current vehicle to other first-region nodes in the first region for consensus verification. If the consensus verification is passed, the vehicle status information and path information of the current vehicle are uploaded to the first-region blockchain for execution. In addition, the vehicle terminal may also send the vehicle status information and path information together with the aforementioned "path planning task" to any first-region node of the first-region blockchain system, or directly upload the vehicle status information and path information together with the aforementioned "path planning task" to the first-region blockchain. This is not limited here.

[0094] Therefore, the first regional node (server) of the embodiment of the present application can obtain the path planning task of each vehicle from the first regional blockchain, and calculate the straight-line length of the path corresponding to each vehicle based on the starting point information and end point information carried by the path planning task of each vehicle. Then, the current vehicle status information of each vehicle and the surrounding road condition information are determined. Specifically, the vehicle status information and road condition information corresponding to each vehicle can be directly obtained from the first regional blockchain, or the vehicle status information and road condition information sent by each vehicle can be obtained by directly receiving. This is not limited here. At this point, the first regional node obtains the corresponding straight-line length of the path, vehicle status information, and road condition information for each vehicle requesting path planning in the first region. The straight-line length of the path, vehicle status information, and road condition information corresponding to each vehicle together represent the status of the corresponding vehicle, which is called vehicle status data. Furthermore, combined with the straight-line length of the path, vehicle status information and road condition information of each vehicle in the first area, a first vehicle status data set corresponding to multiple vehicles in the first area is constructed. The first vehicle status data set represents a collection of vehicle status data of multiple vehicles requesting path planning in the first area, which includes vehicle status data of multiple vehicles. The first vehicle status data set can reflect the status of multiple vehicles and the overall road conditions in the first area. Subsequently, path planning can be performed for multiple vehicles in the first area based on the first vehicle status data set, thereby avoiding the phenomenon of future road traffic congestion caused by individual vehicle path planning, making path planning global and improving vehicle travel efficiency.

[0095] Through the above method, the vehicle status data of each vehicle in the first area can be combined to construct a vehicle status data set corresponding to multiple vehicles in the first area, so as to facilitate the subsequent combination of multiple vehicle data to simultaneously plan paths for multiple vehicles in the area, avoiding the phenomenon of future road traffic congestion caused by individual vehicle path planning, and taking into account future traffic conditions, so that path planning has globality and improves vehicle travel efficiency.

[0096] 103. Input the first vehicle state data set into the trained first regional model to obtain a decision path combination.

[0097] In an embodiment of the present application, after obtaining a first vehicle status data set of multiple vehicles in a first area, the first vehicle status data set can be input into a trained first area model, so that the trained first area model can perform path planning for multiple vehicles based on the first vehicle status data set, thereby outputting a path decision combination, which includes a decision-making planning path corresponding to each vehicle. In this way, it is possible to combine multiple vehicle data to simultaneously perform path planning for multiple vehicles in the area, avoiding the phenomenon of future road traffic congestion caused by separate vehicle path planning, thereby taking into account future traffic conditions, avoiding traffic road congestion, or slowing down traffic road congestion as much as possible, so that path planning is global, and the travel efficiency of multiple vehicles in the first area can be improved subsequently.

[0098] The decision path combination may be in the form of a set or array, and the decision path combination includes the decision planning path corresponding to each vehicle. For example, in the form of an array, the multiple decision planning paths corresponding to multiple vehicles are arranged in sequence in the array. The multiple decision planning paths corresponding to the multiple vehicles may be arranged in the order of the transmission time of the path planning tasks of each vehicle. That is, the multiple decision planning paths included in the decision path combination are arranged in the order of the transmission time of the multiple path planning tasks, and each path planning task corresponds to a decision planning path.

[0099] In the embodiment of the present application, due to the adoption of a regional blockchain system divided into regions, each regional blockchain system is used to perform path planning for multiple vehicles in the region, and each regional blockchain system mainly performs path planning for multiple vehicles based on the vehicle status data sets corresponding to the multiple vehicles through the corresponding regional model. Based on this, in order to avoid traffic congestion caused by the separate training of each regional model when vehicles travel across regions, for example, assuming that a regional model of region A and a regional model of region B are included, if the regional model of region A and the regional model of region B are trained independently, the regional model of region A will only consider the local optimal path planning in region A when performing path planning, and will not consider the situation that when the vehicles in region A cross the region to enter region B, it may cause congestion on the traffic roads in region B, which is not conducive to the smooth traffic in the entire region and causes troubles for the travel of other vehicles. Similarly, the path planning of region B may also bring congestion to the road traffic in region A. , resulting in the phenomenon that each region becomes a data island, making the path planning lack of globality; in this regard, the embodiment of the present application adopts a method of collaborative training between regional models of multiple regions, so that multiple regional models corresponding to multiple regions are trained together, and through global loss aggregation, cross-regional data sharing and a unified parameter system, the regional model parameters of each region are shared, so that the regional models of each region can have unified path planning performance, so that the model can not only optimize local path planning, but also coordinate with other regions to avoid global congestion, and ultimately achieve the goal of "secure collaboration and global optimization between regional blockchain systems in multiple regions under a dynamic environment".

[0100] After the first regional model outputs a first predicted decision path combination based on the first sample vehicle state dataset, it then combines the first sample vehicle state dataset, the first predicted decision path combination, the second sample vehicle state dataset corresponding to the second region, and the second predicted decision path combination to construct a first regional loss. The first regional loss and the second regional loss corresponding to the second region are then combined to construct a global loss. The first regional model and the second regional model corresponding to the second region are then collaboratively trained with maximizing the global cumulative reward value as the optimization objective. The global cumulative reward value and the global loss are negatively correlated. That is, when the global loss is minimized, the global cumulative reward value is maximized. In this way, the collaborative training of regional models across multiple regions is completed. Through global loss aggregation, cross-regional data sharing, and a unified parameter system, regional model parameters across regions are shared, resulting in unified path planning performance for regional models across regions. This allows the model to optimize local path planning while collaborating with other regions to avoid global congestion, thereby improving vehicle travel efficiency after subsequent vehicle path planning.

[0101] It should be noted that after the second regional node shares the second sample vehicle state dataset, the second predicted decision path combination, and the second regional loss with the global blockchain corresponding to the global blockchain system, the local node obtains them from the global blockchain. For an introduction to the first and second sample vehicle state datasets, please refer to the description of the "First Vehicle State Dataset" above. For an introduction to the first and second predicted decision path combinations, please refer to the description of the "Decision Path Combination" above, and we will not repeat them here.

[0102] Among them, the first area loss can be the current reward loss of the first area model in the training stage. The first area loss is constructed based on the difference between the first evaluation expected value and the evaluation score after evaluating the first sample vehicle state data set, the first predicted decision path combination, and the sample vehicle state data sets and predicted decision path combinations of all other areas.

[0103] Among them, the first evaluation expectation value can be determined based on the immediate reward value and future reward value of the path planning of multiple vehicles currently in the first area. The future reward value is evaluated in combination with the first sample vehicle state data set, the first predicted decision path combination and the sample vehicle state data sets and predicted decision path combinations of all other areas.

[0104] For example, assuming a first region model corresponding to a first region and a second region model corresponding to a second region, where the second region model corresponding to the second region generally refers to the region models of all regions except the first region, the first region model is then trained collaboratively with the second region model corresponding to the second region to maximize the global cumulative reward value as the optimization goal. The global cumulative reward value is negatively correlated with the global loss; as the global cumulative reward value increases, the global loss decreases.

[0105] The global loss is determined by combining the regional losses corresponding to each region. Each regional loss is constructed based on the difference between the expected evaluation value of the region and the evaluation score output by the evaluation sub-model for that region. The evaluation sub-model for each region combines the first sample vehicle state dataset and the first predicted decision path combination of the first region, and the second sample vehicle state dataset and the second predicted decision path combination of the second region, to output the corresponding evaluation score. Exemplarily, the first regional loss is determined by combining the first sample vehicle state dataset and the first predicted decision path combination of the first region, and the second sample vehicle state dataset and the second predicted decision path combination of the second region, to obtain the corresponding first evaluation score, and then constructing it based on the difference between the first expected evaluation value and the first evaluation score. Similarly, the second regional loss is determined by combining the first sample vehicle state dataset and the first predicted decision path combination of the first region, and the second sample vehicle state dataset and the second predicted decision path combination of the second region, to obtain the corresponding second evaluation score, and then constructing it based on the difference between the second expected evaluation value and the second evaluation score. Furthermore, the global loss is constructed by combining the first regional loss and the other second regional losses.

[0106] For ease of understanding, the following introduces the training process in combination with the structure of the regional model. For example, taking the first regional model as an example, the first regional model includes a first decision sub-model and a first evaluation sub-model. The first decision sub-model is an "Actor network" model, and the first evaluation sub-model is a "Critic network" model. After the first decision sub-model outputs the first predicted decision path combination based on the first sample vehicle state data set, the first evaluation sub-model outputs the first evaluation score based on the first sample vehicle state data set, the first predicted decision path combination, and the second sample vehicle state data set and the second predicted decision path combination corresponding to each second region, and constructs the first regional loss based on the difference between the first evaluation expected value and the first evaluation score, and combines the first regional loss with the second regional loss of each second region to construct the full regional loss. The global loss can be understood as the expected loss of the global cumulative reward value, which is used to guide the collaborative training of the first decision sub-model and the first evaluation sub-model, as well as the second decision sub-model and the second evaluation sub-model corresponding to each second region with the maximization of the global cumulative reward value as the optimization goal. During the collaborative training, the model parameters of the first decision sub-model and the first evaluation sub-model are adjusted, and the second regional nodes of each second region will also adjust the model parameters of the corresponding second decision sub-model and the second evaluation sub-model until the global cumulative reward value is maximized, the global loss is minimized, and the collaborative training ends, so that the regional nodes in each region obtain the trained decision sub-model and evaluation sub-model. For example, the first regional node obtains the trained first regional model.

[0107] Furthermore, taking the first region model as an example, the training process of the first region model is introduced as follows:

[0108] In some embodiments, the first region model includes a first decision sub-model and a first evaluation sub-model, and the training process of the first region model is as follows:

[0109] (A.1) Obtaining a first sample vehicle state data set, and inputting the first sample vehicle state data set into a first decision sub-model to obtain a first predicted decision path combination;

[0110] (A.2) Obtaining a second sample vehicle state dataset and a second predicted decision path combination corresponding to the second region, and inputting the first sample vehicle state dataset, the first predicted decision path combination, the second sample vehicle state dataset, and the second predicted decision path combination into the first evaluation sub-model to obtain a first evaluation score;

[0111] (A.3) obtaining a first evaluation expected value, and constructing a first regional loss according to a difference between the first evaluation expected value and the first evaluation score;

[0112] (A.4) Obtaining a second region loss corresponding to the second region, and determining a global cumulative reward value based on the first region loss and the second region loss using a preset global loss-reward function;

[0113] (A.5) Taking maximizing the global cumulative reward value as the optimization goal, the first decision sub-model and the first evaluation sub-model are trained to obtain a trained first region model; wherein the global cumulative reward value is maximized when the global loss corresponding to the first region loss and the second region loss is minimized.

[0114] The first sample vehicle state data set includes a plurality of first sample vehicle state data corresponding to a plurality of sample vehicles in the first area, and each first sample vehicle state data includes a sample path straight length of the corresponding sample vehicle. , sample direction state parameters ( , sample coordinate parameters of the next intersection , sample prediction time parameter , sample traffic information (Sample road congestion parameters , sample pedestrian trajectory parameters And the sample vehicle trajectory parameters ),Right now The first sample vehicle status dataset can be expressed as , is 1.

[0115] Among them, the second sample vehicle status data set includes multiple second sample vehicle status data corresponding to multiple sample vehicles in the second area, and each second sample vehicle status data includes the sample path straight-line length, sample direction state parameters, sample coordinate parameters of the next intersection, sample prediction time parameters, and sample road condition information (such as sample road congestion parameters, sample pedestrian trajectory parameters, and sample vehicle trajectory parameters) of the corresponding sample vehicle.

[0116] The first predicted decision path combination includes the predicted decision planning path corresponding to each sample vehicle in the first area, and the second predicted decision path combination includes the predicted decision planning path corresponding to each sample vehicle in the second area.

[0117] For example, assume that there are q regional models corresponding to q regional blockchain systems, and each regional model Responsible for processing the path planning task of the vehicles in the corresponding area. Assuming that there are n vehicles in the area, the regional model The path planning task is represented as , the sample vehicle state data of each sample vehicle in the area is expressed as , regional model The predicted decision planning path output for each sample vehicle state data is expressed as Since the input area model is the sample vehicle status dataset corresponding to multiple sample vehicles in the area, then the regional model The output is a set of strategies , that is, a combination of predicted decision paths containing multiple sample vehicles’ predicted decision planning paths, by finding a set of strategies , to maximize the expectation of the global cumulative reward of all regional models. The expectation of the global cumulative reward is expressed as ,in, It's in time Instant rewards received, (0,1) is the discount factor, "And the immediate rewards later" ” refer to the same meaning.

[0118] For example, taking the training of the first region model as an example, the first region model may be composed of a first decision sub-model (first Actor network) and a first evaluation sub-model (first Critic network).

[0119] The first step is to take the first sample vehicle status dataset Input to the first decision sub-model (first Actor network), so that the first decision sub-model (first Actor network) outputs the first predicted decision path combination .

[0120] The second step is to obtain a second sample vehicle state data set and a second predicted decision path combination corresponding to the other second area. The second sample vehicle state data set is represented as , the second prediction decision path combination is expressed as , the first sample vehicle state data set, the first predicted decision path combination, the second sample vehicle state data set and the second predicted decision path combination are input into the first evaluation sub-model (the first critic network) to obtain the first evaluation score, which is expressed as ,in, Indicates the current sample The sample vehicle status data set of all regions under is specifically expressed as , Indicates the current sample The set of prediction decision path combinations for all regions under , specifically expressed as .

[0121] The third step is to obtain the first evaluation expected value, which is based on the current sample The instant reward value of the first area with the target first review score Specifically, the calculation process of the first evaluation expected value is as follows:

[0122]

[0123] in, represents the first evaluation expected value of the first region under the current sample, Represents a regional model In the current sample The instant reward value of the first area, Indicates that in the current sample Next target first evaluation score, Indicates the Regional model.

[0124] Among them, the target first evaluation score is the target first evaluation sub-model (target first critic network) based on the current sample Under the "and" " is calculated, which is actually the same as " "and" It should be noted that the parameters of the target first critic network are regularly synchronized from the first critic network through "soft update" or "hard update". By delaying the update of parameters, a reliable optimization target is provided for the current first critic network, which not only maintains the correlation of parameters but also avoids the instability caused by real-time updates.

[0125] The immediate reward value is determined by weighted summation of the local reward value and the global reward value. For example, in the current sample The instant reward value of the first area The calculation process is as follows:

[0126]

[0127] Represents a regional model In the current sample The instantaneous reward value of the first region in the next time step T, represents the local reward value, represents the global reward value, and Represent the weights of local reward value and global reward value respectively.

[0128] The calculation process of the local reward value is:

[0129]

[0130] The global reward value can be calculated based on the congestion threshold of the current road segment. , the number of vehicles on the road , average waiting time on the road , road occupancy rate , vehicle interaction waiting time , the absolute value of the deviation between the predicted multi-agent traffic behavior (such as road congestion, pedestrian trajectory, vehicle trajectory) and the actual state , the difference between the vehicle's traveled distance and the remaining distance " to determine the global reward value. The calculation process of the global reward value is:

[0131]

[0132] Congestion judgment threshold Indicates the congestion trigger factor, which is used to identify whether the current road section has reached the congestion threshold. The congestion judgment threshold is expressed as follows:

[0133]

[0134] represents the time lost by the vehicle in decision sharing; and Respectively represent the vehicle's travel distance from the starting point and the travel distance from the end point; It represents the average waiting time on the road and is directly related to traffic efficiency; Indicates the number of vehicles on the road at the current moment, reflecting the congestion level of the road section; It represents the road occupancy rate, that is, the proportion of the road section occupied by vehicles, which measures the intensity of road resource utilization; Indicates the judgment threshold; Indicates the weight coefficient of each indicator, which is used to adjust the influence of different factors on the global reward. For example, , to strengthen the priority of congestion avoidance, Represents the travel distance of the static path planning algorithm.

[0135] In addition, the global reward value can also be calculated based on the congestion threshold of the current road segment. , the number of vehicles on the road , average waiting time on the road , road occupancy rate , vehicle interaction waiting time , the absolute value of the deviation between the predicted multi-agent traffic behavior (such as road congestion, pedestrian trajectory, vehicle trajectory) and the actual state , the difference between the vehicle's traveled distance and the remaining distance ", and energy consumption value To determine the global reward value.

[0136]

[0137] Energy consumption It can represent the vehicle energy consumption under the decision-making planning path, so that the regional model can give priority to routes with lower energy consumption (such as avoiding steep slopes and sections with dense traffic lights) when planning vehicle paths, thereby achieving the dual goals of "efficient passage and low-carbon travel".

[0138] Combining the above parameters, we can calculate the instant reward for participating in the first evaluation expectation. , which can be specifically expressed as Then, the first evaluation expected value is calculated. Thereafter, the first regional loss is constructed based on the difference between the first evaluation expected value and the first evaluation score. The specific calculation process is as follows:

[0139]

[0140] Indicates that for Regional loss of regional models, e.g. represents the first region loss corresponding to the first region model, represents the sample batch size, Indicates the corresponding sample.

[0141] The fourth step is to obtain the second area loss corresponding to the second area. It should be noted that the construction process of the second area loss of other areas can refer to the calculation process of the "first area loss" mentioned above. It belongs to the second area node calculation of other second areas and will not be repeated here.

[0142] Furthermore, a global cumulative reward value is determined based on the first region loss and the second region loss through a preset global loss reward function; the global cumulative reward value is the expected value of the global cumulative reward associated with the global loss function, and the global cumulative reward value is negatively correlated with the global loss value of the global loss function. Global Loss Function It is expressed as follows:

[0143]

[0144] Indicates the model parameters The global loss value under represents the expectation of the global cumulative reward value, represents the expected value of maximizing the global cumulative reward, represents the model parameters, is the model parameter space, Indicates selecting the optimal model parameters in the model parameter space to maximize the expected value of the global cumulative reward, Represents the inverse function. From the global loss function, it can be obtained that the global cumulative reward value is negatively correlated with the global loss value of the global loss function.

[0145] The fifth step is to maximize the global cumulative reward value as the optimization goal, and train the first decision sub-model and the first evaluation sub-model to adjust the model parameters of the first decision sub-model and the first evaluation sub-model respectively. Update the model parameters of the first evaluation sub-model An update is performed to obtain a trained first region model; wherein the global cumulative reward value is maximized when the global loss corresponding to the first region loss and the second region loss is minimized.

[0146] According to the first to fifth steps above, the training of the first region model is completed. It should be noted that the first region model is trained in collaboration with other second region models. Figure 4 The collaborative training architecture diagram of each regional model provided in the embodiment of this application is combined with Figure 4 As shown, it contains N regional models. represents the first decision sub-model in the first regional model, Q1 represents the first evaluation sub-model in the first regional model, Represents the first decision sub-model in the Nth regional model, QN represents the Nth evaluation sub-model in the first regional model. During the collaborative training of these N regional models, the first regional node adjusts the model parameters of the first decision sub-model and the first evaluation sub-model, and the second regional nodes of each second region also adjust the model parameters of the corresponding second decision sub-model and the second evaluation sub-model until the global cumulative reward value is maximized and the global loss is minimized. The collaborative training ends, so that the regional nodes in each region obtain the trained decision sub-model and evaluation sub-model. For example, the first regional node obtains the trained first regional model.

[0147] In an embodiment of the present application, the second regional model of other second regions can be combined with the second sample vehicle state data set, the second predicted decision path combination, and the first sample vehicle state data set and the first predicted decision path combination to calculate the second evaluation score, and construct the second regional loss based on the difference between the second evaluation expected value and the second evaluation score, thereby sharing the second regional loss with each regional node, such as sharing the second regional loss with the first regional node through the global blockchain, so that the first regional node combines the first regional loss and the second regional losses of other second regions to construct a global loss. It should be noted that the introduction of "second evaluation expected value", "second evaluation score", and "second regional loss" can refer to the description of "first evaluation expected value", "first evaluation score", and "first regional loss", which will not be repeated here.

[0148] In some embodiments, a local node (a first-region node) may upload the first sample vehicle state dataset and the first predicted decision path combination to a global blockchain for sharing with other second-region nodes in the second region, and obtain the second-region loss shared by the second-region nodes from the global blockchain. For example, prior to step (A.4), the process may include: uploading the first sample vehicle state dataset and the first predicted decision path combination to the global blockchain, allowing the second-region node to obtain the first sample vehicle state dataset and the first predicted decision path combination from the global blockchain, and calculating a second evaluation score based on the first sample vehicle state dataset, the first predicted decision path combination, the second sample vehicle state dataset, and the second predicted decision path combination using a second-region model corresponding to the second region, and constructing a second-region loss based on the second evaluation score and a preset second evaluation expected value. Furthermore, the step (A.4) of "obtaining the second-region loss corresponding to the second region" may include: obtaining the second-region loss corresponding to the second region from the global blockchain, wherein the second-region loss is uploaded to the global blockchain by the second-region node.

[0149] Specifically, regional nodes in different regions can achieve cross-regional data sharing and collaborative computing through a global blockchain. For example, a node in the first region uploads its local first sample vehicle status dataset (containing status information and road condition data for vehicles in the first region) and a first predicted decision path combination (the vehicle planned path output by the first regional model) to the global blockchain. This ensures that nodes in the second region can access this data from the global blockchain, breaking down information silos between regions and enabling data sharing. Before uploading the first sample vehicle status dataset to the global blockchain, Gaussian noise can be added to the dataset to generate a Gaussian-noised first sample vehicle status dataset. The Gaussian-noised first sample vehicle status dataset and the first predicted decision path combination can then be uploaded to the global blockchain to enhance information privacy. Furthermore, Gaussian noise can be added to sensitive fields (such as vehicle ID and precise location information) to ensure that the data meets collaborative training requirements while preventing the inference of specific vehicle information, preventing privacy leaks and improving the security of the first sample vehicle dataset.

[0150] The second regional node then performs calculations based on its own second regional model, combining two data sets: a first sample vehicle state dataset (which may contain Gaussian noise) and a first predicted decision path combination obtained from the global blockchain; and a local second sample vehicle state dataset (vehicle and road condition data within the second region) and a second predicted decision path combination (the planned path output by the second regional model). Using this data, the second regional model calculates a second evaluation score, which is used to measure the effectiveness of cross-regional path planning collaboration. For example, this score measures whether congestion at the intersection of two regions is avoided or whether multi-vehicle interaction is smooth. Furthermore, the second regional node compares this calculated second evaluation score with a preset second evaluation expected value. The difference between the two constitutes the second regional loss. The smaller the loss, the closer the second regional model's decision-making in cross-regional collaboration is to the optimal goal. Furthermore, the second regional node uploads this second regional loss to the global blockchain for access by other regional nodes, including the first regional node.

[0151] Finally, after the first regional node obtains this second regional loss from the global blockchain, it can combine it with its own first regional loss to construct a global loss, providing a unified optimization basis for the collaborative training of multi-region models, ensuring that each regional model takes into account global traffic efficiency during training. In this way, data sharing and loss transmission are achieved through the global blockchain, allowing the training of each regional model to not only rely on local data but also integrate collaborative information from other regions, ultimately achieving global optimization of cross-regional path planning.

[0152] In some embodiments, the first decision sub-model and the first evaluation sub-model are trained with maximizing the global cumulative reward value as the optimization objective, thereby obtaining a trained first regional model. For example, step (A.5) may include: determining the number of vehicles and road occupancy corresponding to the first region based on the first sample vehicle status dataset, and calculating the current global cumulative reward value based on the number of vehicles, road occupancy, and the first sample vehicle status dataset; adjusting the model parameters of the first decision sub-model and the first evaluation sub-model using the current global cumulative reward value as the initial value, and iteratively training in an optimization direction that optimizes the global cumulative reward value to a positive growth, until the global cumulative reward value is maximized while minimizing the global loss, thereby obtaining the trained first regional model.

[0153] First, the first regional node extracts and determines the number of vehicles in the first region (the total number of vehicles in the current region, which can be determined based on the number of first vehicle status data) and the road occupancy rate (the proportion of roads in the region occupied by vehicles) from the first sample vehicle status data set corresponding to the local first region. For example, the road occupancy rate can be calculated as follows:

[0154] First, calculate the real-time position and size of vehicles within a single lane. Specifically, use on-board sensors (such as GPS) or roadside units (RSUs) to collect the coordinates (such as longitude and latitude, relative position within the lane) and body length and width of all vehicles in the area, and calculate the proportion of space occupied by a single vehicle within the lane.

[0155] Second, determine the total length and effective width of the lane. Specifically, based on the preset road map data, the physical dimensions of the target lane (such as a length of 1,000 meters and an effective width of 3.5 meters) can be determined as a calculation basis.

[0156] Third, calculate road occupancy: Road occupancy = (Total space occupied by all vehicles in a single lane) ÷ (Total available space in that lane) × 100%. For example, if a lane is 1000 meters long and there are 10 vehicles, all 5 meters long, the total occupied space is 5 × 10 = 50 meters. The road occupancy = 50 ÷ 1000 100% = 5%. This metric directly reflects lane congestion and serves as a key parameter in the global reward (RG) to quantify road resource utilization efficiency. If the road occupancy is too high (e.g., exceeding 80%), a reward-penalty mechanism is triggered, guiding the agent to avoid that lane when planning routes to avoid congestion.

[0157] Then, these data are combined with the first sample vehicle status data set (including detailed information such as vehicle status and road conditions) to calculate the current global cumulative reward value. The global cumulative reward value is calculated as follows: ,in, It's in time Instant rewards received, (0,1) is the discount factor, "And the instant reward mentioned above" ” refer to the same meaning. The specific calculation process can be combined with the description above and is not repeated here. This global cumulative reward value comprehensively reflects the operating efficiency of traffic flow in the first area (such as whether the vehicle density is reasonable and whether the road resource utilization is efficient). It is the core indicator for measuring the contribution of the current regional decision to the global traffic.

[0158] Furthermore, using the currently calculated global cumulative reward as an initial benchmark, we begin adjusting the parameters of the first decision submodel (responsible for planning vehicle routes) and the first evaluation submodel (responsible for evaluating the value of routing decisions) in the first regional model. The core goal of training is to ensure a continuous positive growth in the global cumulative reward. This means that after each round of training, the global cumulative reward generated by the new decisions must be higher than that of the previous round. This guides the model to continuously optimize routing strategies and iterate towards improving global traffic efficiency.

[0159] Finally, iterative training continues until the global loss (composed of the first and second region losses) reaches a minimum. Since the global cumulative reward is negatively correlated with the global loss (the smaller the global loss, the larger the global cumulative reward), the global cumulative reward simultaneously reaches its maximum value at this point. The parameters of the first decision sub-model and the first evaluation sub-model tend to stabilize, and the trained first region model can output an optimal path planning result that balances local traffic efficiency with global coordinated optimization. Thus, through parameter iteration guided by the global cumulative reward, the training of the first region model is based on both local traffic conditions and global traffic optimization, ultimately achieving efficient path planning in a dynamic environment.

[0160] In some embodiments, for cross-regional situations, it is necessary to calculate the cumulative reward in combination with the number of vehicles and road occupancy rates in other regions for training. For example, step (A.5) may include: obtaining the number of vehicles in the second region and the road occupancy rate in the second region corresponding to the second region from the global blockchain, the number of vehicles in the second region and the road occupancy rate in the second region being determined by the second region node based on the second vehicle status dataset and uploaded to the global blockchain; calculating the current global cumulative reward value based on the number of vehicles in the second region, the road occupancy rate in the second region, and the first sample vehicle status dataset; using the current global cumulative reward value as the initial value, adjusting the model parameters of the first decision sub-model and the first evaluation sub-model, and iteratively training in the optimization direction of a positive growth of the global cumulative reward value until the maximum global cumulative reward value is achieved when the global loss is minimized, thereby obtaining the trained first region model.

[0161] Specifically, during the training of the first regional model, each regional node implements cross-regional data interaction and collaborative optimization through the global blockchain. For example, the first regional node first obtains key traffic data uploaded by the second regional node from the global blockchain, such as the number of second regional vehicles (the total number of vehicles in the current region) and the second regional road occupancy rate (the proportion of road sections occupied by vehicles). This data is analyzed by the second regional node based on the local second vehicle status dataset. After being shared through the global blockchain, it becomes an important basis for the first regional model to evaluate the global traffic status.

[0162] The first regional node then calculates the current global cumulative reward value by combining two pieces of information: the number of vehicles and road occupancy in the second region, obtained from the global blockchain, and the local first sample vehicle status dataset (including vehicle status, road conditions, and other data from the first region). The calculation of the global cumulative reward value comprehensively considers factors such as the balance of cross-regional traffic flow (such as whether there is a risk of congestion transmission between the two regions) and the efficiency of road resource utilization (such as whether the road occupancy rate is reasonable). The higher the value, the more significant the improvement in global traffic efficiency achieved by the current collaborative decision-making between the two regions.

[0163] Next, using the currently calculated global cumulative reward as the initial benchmark, iterative training begins for the first decision-making sub-model (responsible for path planning decisions) and the first evaluation sub-model (responsible for evaluating the value of decisions) within the first regional model. During training, the model output is continuously optimized by adjusting the parameters of these two sub-models (such as path weights in the decision logic and value estimation coefficients in the evaluation logic). The core focus of optimization is to ensure a continuous positive growth in the global cumulative reward. Specifically, after each round of parameter adjustments, the global cumulative reward generated by the new decisions must be higher than the previous round, ensuring that the model iterates towards improving global efficiency.

[0164] Finally, iterative training continues until the global loss reaches a minimum. Since the global cumulative reward is negatively correlated with the global loss (the smaller the global loss, the larger the global cumulative reward), the global cumulative reward simultaneously reaches its maximum value at this point. The parameters of the first decision-making sub-model and the first evaluation sub-model stabilize, and the trained first regional model is able to output an optimal path planning strategy that balances local and global efficiency in a dynamic traffic environment. In this way, through cross-regional data sharing and global reward-driven training, the first regional model's training breaks through the limitations of focusing solely on local traffic and actively adapts to the traffic conditions of the second region, ultimately achieving the core goal of "maximizing global cumulative reward."

[0165] Through the above method, the first vehicle state data set can be input into the trained first area model, so that the trained first area model can perform path planning for multiple vehicles based on the first vehicle state data set, thereby outputting a path decision combination, which includes the decision planning path corresponding to each vehicle. In this way, it is possible to combine multiple vehicle data to simultaneously perform path planning for multiple vehicles in the area, avoiding the phenomenon of future road traffic congestion caused by separate vehicle path planning, thereby taking into account future traffic conditions, avoiding traffic road congestion, or slowing down traffic road congestion as much as possible, making path planning global, and subsequently improving the travel efficiency of multiple vehicles in the first area.

[0166] 104. Upload the decision planning path corresponding to each vehicle included in the decision path combination to the first regional blockchain, so that each vehicle reads the corresponding decision planning path from the first regional blockchain.

[0167] In an embodiment of the present application, after obtaining a path decision combination including a decision-making planning path corresponding to each vehicle, the first area node can upload the decision-making planning path corresponding to each vehicle contained in the decision path combination to the first area blockchain, so that each vehicle reads the corresponding decision-making planning path from the first area blockchain. In this way, each vehicle can travel according to its own decision-making planning path, avoiding traffic congestion, or alleviating traffic congestion as much as possible, thereby improving the travel efficiency of multiple vehicles in the first area.

[0168] Among them, each decision-making planning path can be a path that includes the starting point information and the end point information carried by the path planning task of the corresponding vehicle, which is planned in combination with dynamic traffic road condition information.

[0169] In some embodiments, after obtaining a path decision combination including a decision-making planned path corresponding to each vehicle, the first regional node may sign each decision-making planned path using its own first regional node private key to obtain a hash value for each path, and then send each path hash value and the path hash value combination to other first regional nodes in the first regional blockchain system for consensus verification. After consensus verification is passed, the path hash value combination and the decision-making planned path corresponding to each vehicle included in the decision path combination are uploaded to the first regional blockchain. For example, step 104 may include:

[0170] (104.1) Calculate a path hash value for each decision-planning path corresponding to each vehicle included in the decision-planning path combination to obtain a path hash value combination, where the hash value combination includes the path hash value corresponding to each decision-planning path;

[0171] (104.2) Sending the path hash value combination to other first-region nodes in the first-region blockchain system for consensus verification, and obtaining a consensus verification result;

[0172] (104.3) When it is determined based on the consensus verification result that the number of target first-region nodes that have reached consensus is greater than a preset node number threshold, the path hash value combination and the decision planning path corresponding to each vehicle included in the decision path combination are uploaded to the first-region blockchain.

[0173] Specifically, after obtaining a path decision combination containing the decision-making planning path corresponding to each vehicle, the decision path combination is uploaded to the blockchain. During the process of uploading the decision path combination to the blockchain in the first region, a hash verification and consensus mechanism is required to ensure the reliability and immutability of the decision path. For example, first, for the decision path combination output by the first region model (including the planned paths of all vehicles in the first region), each path hash value is calculated separately for the decision-making planning path of each vehicle. Specifically, each path hash value is signed and generated using a specific encryption algorithm (such as SHA-256) combined with the private key of the first region node in the company key pair of the first region node. Each decision-making planning path corresponds to a unique path hash value. These path hash values ​​are integrated to form a path hash value combination, which serves as the "digital fingerprint" of the path data.

[0174] The first-region node then sends the path hash value combination and decision path combination to other first-region nodes within the first-region blockchain system (such as surrounding roadside units and collaborative decision nodes) to initiate consensus verification. Other first-region nodes recalculate the hash value of each vehicle's decision-planned path using the same hash algorithm and compare it with the received path hash value combination. Alternatively, they decrypt the data using the first-region node's public key and compare it with the decision path combination to verify the integrity and integrity of the path data. Ultimately, they provide feedback on the consensus verification results.

[0175] Finally, the number of target first-region nodes that have reached consensus is counted. If this number exceeds a preset node threshold (e.g., more than 1 / 3 of the total number of first-region nodes in the region), the path data is considered verified. At this point, the first-region node uploads the path hash value combination and the corresponding decision-making and planned paths of all vehicles to the first-region blockchain for on-chain storage. This ensures the integrity of the path data through the uniqueness of the hash value, and prevents data tampering through multi-node consensus verification. This ensures that all vehicles in the region can obtain trusted planned paths from the blockchain, providing a reliable data foundation for multi-vehicle coordinated driving.

[0176] In some embodiments, the path hash value combination and the decision path combination can be packaged to generate a target block, and the block header corresponding to the target block can be broadcasted to other first-region nodes in the first-region blockchain system for consensus verification. When the consensus verification passes, the target region is added to the first-region blockchain, so that each vehicle can read the corresponding decision-planning path from the first-region blockchain. For example, step (104.2) can include: constructing a path hash tree based on the path hash value corresponding to each decision-planning path included in the path hash value combination, the path hash tree including a path root hash; determining the generation timestamp information, the receiver identifier, and the sender identifier associated with each decision-planning path; and generating the target block based on the path hash tree, the path root hash, the decision-planning path corresponding to each vehicle included in the decision path combination, and the generation timestamp information, the receiver identifier, and the sender identifier associated with each decision-planning path; and sending the block header corresponding to the target block to other first-region nodes in the first-region blockchain system for consensus verification, thereby obtaining a consensus verification result, the block header including the generation timestamp information, the receiver identifier, the sender identifier, and the path root hash.

[0177] Furthermore, when it is determined based on the consensus verification result that the number of target first-region nodes that have reached consensus is greater than a preset node number threshold, step (104.3) may include: adding the target block to the first-region blockchain.

[0178] Specifically, after obtaining the path hash value combination, a path hash tree (also known as a Merkle tree) is first constructed based on the path hash values ​​corresponding to each decision-making path in the path hash value combination. The path hash tree integrates all path hash values ​​through layer-by-layer hashing operations, ultimately generating a unique path root hash. This path root hash is the root hash of the Merkle tree and serves as the "top-level digital fingerprint" of the entire path hash value combination, used to quickly verify the integrity of batch path data.

[0179] Next, the metadata associated with each decision-making path is extracted, such as the timestamp (record of the time when the path was planned), the receiver identifier (such as the ID of the vehicle receiving the path), and the sender identifier (such as the ID of the first regional node that generated the path). This metadata is integrated with the path hash tree, the path root hash, and the decision-making paths of all vehicles in the decision-making path combination to generate a target block containing complete path data and verification information.

[0180] Furthermore, the first-region node sends the target block's block header to other first-region nodes within the first-region blockchain system for consensus verification. The block header includes a generation timestamp, a recipient identifier, a sender identifier, and a path root hash. Other nodes verify the header information (for example, verifying that the path root hash matches the locally calculated result and that the timestamp is valid) and then provide feedback on the consensus verification results.

[0181] Finally, the number of nodes in the target first region that have reached consensus is counted. If this number exceeds a preset node threshold (e.g., the consensus rule of "1 / 3 of the nodes in the first region agree"), the target block is officially added to the first region's blockchain, completing the on-chain storage of the entire block. This optimizes batch data verification efficiency through the path hash tree and block header, while preserving complete path data and metadata, ensuring that the on-chain blocks are traceable and tamper-proof, providing reliable historical path data support for collaborative decision-making among multiple agents within the region.

[0182] Through the above method, the decision planning path corresponding to each vehicle included in the decision path combination can be uploaded to the first area blockchain, so that each vehicle can read the corresponding decision planning path from the first area blockchain. In this way, each vehicle can travel according to its own decision planning path, avoiding traffic congestion or alleviating traffic congestion as much as possible, thereby improving the travel efficiency of multiple vehicles in the first area.

[0183] As can be seen from the above, the vehicle path planning method of the embodiment of the present application is applied to any first-area node in the first-area blockchain system, and the first-area node and any second-area node in the second-area blockchain system corresponding to the second area constitute a global blockchain system; by obtaining path planning tasks of multiple vehicles in the first area from the first-area blockchain corresponding to the first-area blockchain system, wherein each path planning task is initiated by each vehicle and uploaded to the first-area blockchain after consensus verification by the first-area blockchain system; determining the straight-line length of each vehicle's path based on the starting point information and end point information carried by each path planning task, and determining the current vehicle status information of each vehicle and the surrounding road condition information, and constructing a first vehicle status data set corresponding to multiple vehicles in combination with the straight-line length of each vehicle's path, vehicle status information, and road condition information; inputting the first vehicle status data set into the trained first-area model to obtain a decision path combination, which includes the decision planning path corresponding to each vehicle; Among them, after the first regional model outputs the first predicted decision path combination based on the first sample vehicle state data set, the first predicted decision path combination is combined with the first sample vehicle state data set, the first predicted decision path combination, the second sample vehicle state data set corresponding to the second region, and the second predicted decision path combination to construct the first regional loss, and the first regional loss and the second regional loss corresponding to the second region are combined to construct the global loss, and the maximization of the global cumulative reward value is used as the optimization goal to guide the first regional model and the second regional model corresponding to the second region to be collaboratively trained, and the global cumulative reward value is negatively correlated with the global loss; among them, after the second sample vehicle state data set, the second predicted decision path combination and the second regional loss are shared by the second regional node to the global blockchain corresponding to the global blockchain system, the local node obtains them from the global blockchain; the decision planning path corresponding to each vehicle contained in the decision path combination is uploaded to the first regional blockchain, so that each vehicle reads the corresponding decision planning path from the first regional blockchain.

[0184] Based on this, the present application may include a two-layer blockchain system architecture, namely a regional blockchain system layer and a global blockchain system layer. The regional blockchain system is used to manage the traffic path planning of the corresponding area, and the global blockchain system is used to manage and coordinate the traffic path planning of all areas. Taking the first regional blockchain as an example, first, vehicles in the first area can submit path planning tasks to the first regional blockchain system to be recorded on the first regional blockchain. Then, the first regional node can obtain the path planning tasks of multiple vehicles from the first regional blockchain, determine the straight-line length of the path of each vehicle according to each path planning task, and construct a first vehicle status data set for multiple vehicles in combination with the straight-line length of the path of each vehicle, vehicle status information, surrounding path information, etc., and then input the first vehicle status data set containing the status data of multiple vehicles into the trained first regional model to obtain a decision path combination, each decision path combination containing the decision planning path of each vehicle; it should be noted that since the first regional model is a combination of at least one The training process is collaboratively trained with models in other regions (such as the second regional model). During the collaborative training process, the first sample vehicle state dataset and the first predicted decision path combination in the first region are combined with the sample vehicle state datasets and corresponding predicted decision path combinations in the other regions to construct the first regional loss. That is, each regional model combines local sample data and sample data from other regions to construct a local regional loss. The first regional loss and the losses of the other regions are combined to construct a global loss. A global cumulative reward value is constructed as a supervisory guide for the training process, and maximizing the global cumulative reward value is the optimization goal of the training process. This allows the first regional model to be collaboratively trained with the regional models in the other regions. Therefore, the trained first regional model can consider multiple vehicle scenarios when outputting a decision path combination during path planning, avoiding traffic congestion caused by multiple vehicles occupying the same lane at the same time. Finally, the decision path combination is uploaded to the blockchain for access by all vehicles in the first region. In this way, multiple path planning tasks and multiple vehicle data are combined to simultaneously plan paths for multiple vehicles, thus considering future traffic conditions, making path planning global and improving vehicle travel efficiency.

[0185] The specific implementation of the above steps can be found in the previous embodiments and will not be repeated here.

[0186] To facilitate better implementation of the vehicle path planning method provided in the embodiment of the present application, the embodiment of the present application also provides a vehicle path planning device based on the above-mentioned method. The meanings of the terms herein are the same as those in the above-mentioned vehicle path planning method, and the specific implementation details can be referred to the description in the method embodiment.

[0187] See also Figure 5 , Figure 5This is a structural diagram of a vehicle path planning device provided in an embodiment of the present application. The vehicle path planning device is integrated into the computer device of the present application. Specifically, it is applied to any first regional node in the first regional blockchain system. The first regional node and any second regional node in the second regional blockchain system corresponding to the second region constitute a global blockchain system, wherein the vehicle path planning device may include an acquisition unit 401, a determination unit 402, an input unit 403, and a sending unit 404.

[0188] An acquisition unit 401 is configured to acquire, from a first regional blockchain corresponding to a first regional blockchain system, routing tasks for a plurality of vehicles within a first regional area, wherein each routing task is initiated by each vehicle and uploaded to the first regional blockchain after consensus verification by the first regional blockchain system;

[0189] a determining unit 402 for determining the straight-line length of each vehicle's path based on the start and end point information carried by each path planning task, and determining the current vehicle state information of each vehicle and surrounding road condition information, and constructing a first vehicle state dataset corresponding to the plurality of vehicles by combining the straight-line length of each vehicle's path, the vehicle state information, and the surrounding road condition information;

[0190] An input unit 403 is configured to input the first vehicle state data set into the trained first regional model to obtain a decision path combination, where the decision path combination includes a decision planning path corresponding to each vehicle;

[0191] Among them, after the first regional model outputs the first predicted decision path combination based on the first sample vehicle state data set, the first regional loss is constructed by combining the first sample vehicle state data set, the first predicted decision path combination, and the second sample vehicle state data set and the second predicted decision path combination corresponding to the second region. The first regional loss and the second regional loss corresponding to the second region are combined to construct a global loss. The maximization of the global cumulative reward value is used as the optimization goal to guide the collaborative training of the first regional model and the second regional model corresponding to the second region. The global cumulative reward value is negatively correlated with the global loss.

[0192] The second sample vehicle state data set, the second predicted decision path combination, and the second regional loss are shared by the second regional node to the global blockchain corresponding to the global blockchain system, and then obtained by the local node from the global blockchain;

[0193] The sending unit 404 is used to upload the decision planning path corresponding to each vehicle included in the decision path combination to the first regional blockchain, so that each vehicle reads the corresponding decision planning path from the first regional blockchain.

[0194] In some implementations, the sending unit 404 is further configured to:

[0195] For each decision-making planning path corresponding to each vehicle included in the decision-making path combination, a path hash value is calculated to obtain a path hash value combination, where the hash value combination includes the path hash value corresponding to each decision-making planning path;

[0196] Sending the path hash value combination to other first-region nodes in the first-region blockchain system for consensus verification to obtain a consensus verification result;

[0197] When it is determined based on the consensus verification result that the number of target first-area nodes that have reached consensus is greater than the preset node number threshold, the path hash value combination and the decision planning path corresponding to each vehicle included in the decision path combination are uploaded to the first-area blockchain.

[0198] In some implementations, the sending unit 404 is further configured to:

[0199] Construct a path hash tree based on the path hash value corresponding to each decision planning path included in the path hash value combination, where the path hash tree includes the path root hash;

[0200] Determine the generation timestamp information, receiver identifier, and sender identifier associated with each decision-making plan path, and generate a target block based on the path hash tree, the path root hash, the decision-making plan path corresponding to each vehicle included in the decision-making path combination, and the generation timestamp information, receiver identifier, and sender identifier associated with each decision-making plan path;

[0201] Sending the block header corresponding to the target block to other first-region nodes in the first-region blockchain system for consensus verification to obtain a consensus verification result, where the block header includes generation timestamp information, receiver identifier, sender identifier, and path root hash;

[0202] The sending unit is further applied to add the target block to the first area blockchain.

[0203] In some embodiments, the first area model includes a first decision sub-model and a first evaluation sub-model, and the vehicle path planning apparatus further includes a training unit for:

[0204] Obtaining a first sample vehicle state data set, and inputting the first sample vehicle state data set into a first decision sub-model to obtain a first predicted decision path combination;

[0205] Obtaining a second sample vehicle state data set and a second predicted decision path combination corresponding to the second area, and inputting the first sample vehicle state data set, the first predicted decision path combination, the second sample vehicle state data set, and the second predicted decision path combination into the first evaluation sub-model to obtain a first evaluation score;

[0206] Obtaining a first evaluation expected value, and constructing a first regional loss according to a difference between the first evaluation expected value and the first evaluation score;

[0207] Obtaining a second region loss corresponding to the second region, and determining a global cumulative reward value based on the first region loss and the second region loss using a preset global loss-reward function;

[0208] Taking maximizing the global cumulative reward value as the optimization goal, the first decision sub-model and the first evaluation sub-model are trained to obtain the trained first region model;

[0209] The global cumulative reward value is maximized when the global loss corresponding to the first area loss and the second area loss is minimized.

[0210] In some embodiments, the vehicle path planning device further includes an area loss acquisition unit configured to:

[0211] Uploading the first sample vehicle state data set and the first predicted decision path combination to the global blockchain, so that the second regional node obtains the first sample vehicle state data set and the first predicted decision path combination from the global blockchain, and calculates a second evaluation score based on the first sample vehicle state data set, the first predicted decision path combination, the second sample vehicle state data set, and the second predicted decision path combination using a second regional model corresponding to the second region, and constructs a second regional loss based on the second evaluation score and a preset second evaluation expected value;

[0212] The training unit is further used to obtain the second area loss corresponding to the second area from the global blockchain, wherein the second area loss is uploaded to the global blockchain by the second area node.

[0213] In some embodiments, the training unit is further configured to:

[0214] Determining the number of vehicles and the road occupancy rate corresponding to the first area according to the first sample vehicle status data set, and calculating the current global cumulative reward value according to the number of vehicles, the road occupancy rate, and the first sample vehicle status data set;

[0215] Taking the current global cumulative reward value as the initial value, adjust the model parameters of the first decision sub-model and the first evaluation sub-model, and perform iterative training in the optimization direction of positive growth of the global cumulative reward value until the maximum global cumulative reward value is achieved when the global loss is minimized, and obtain the trained first regional model.

[0216] In some embodiments, the training unit is further configured to:

[0217] Obtaining from the global blockchain the number of vehicles in the second area and the road occupancy rate of the second area corresponding to the second area, wherein the number of vehicles in the second area and the road occupancy rate of the second area are determined by the second area node based on the second vehicle status dataset and uploaded to the global blockchain;

[0218] Calculate the current global cumulative reward value based on the number of vehicles in the second area, the road occupancy rate in the second area, and the first sample vehicle state data set;

[0219] Taking the current global cumulative reward value as the initial value, adjust the model parameters of the first decision sub-model and the first evaluation sub-model, and perform iterative training in the optimization direction of positive growth of the global cumulative reward value until the maximum global cumulative reward value is achieved when the global loss is minimized, and obtain the trained first regional model.

[0220] As can be seen from the above, the embodiment of the present application includes a two-layer blockchain system architecture, namely a regional blockchain system layer and a global blockchain system layer. The regional blockchain system is used to manage the traffic path planning of the corresponding area, and the global blockchain system is used to manage and coordinate the traffic path planning of all areas. Taking the first regional blockchain as an example, first, vehicles in the first area can submit path planning tasks to the first regional blockchain system to be recorded on the first regional blockchain. Then, the first regional node can obtain the path planning tasks of multiple vehicles from the first regional blockchain, determine the straight-line length of the path of each vehicle according to each path planning task, and construct a first vehicle status data set for multiple vehicles in combination with the straight-line length of the path of each vehicle, vehicle status information, surrounding path information, etc., and then input the first vehicle status data set containing the status data of multiple vehicles into the trained first regional model to obtain a decision path combination, each decision path combination containing the decision planning path of each vehicle; it should be noted that since the first regional model is combined to The first regional model is trained collaboratively with at least one model from another region (e.g., the second regional model). During the collaborative training process, the first sample vehicle state dataset and the first predicted decision path combination from the first region are combined with the sample vehicle state datasets and corresponding predicted decision path combinations from the other regions to construct the first regional loss. That is, each regional model combines local sample data from other regions to construct a local regional loss, and the first regional loss is combined with the losses from the other regions to construct a global loss. A global cumulative reward value is constructed as a supervisory guide for the training process, and maximizing the global cumulative reward value is the optimization goal of the training process. This allows the first regional model to be collaboratively trained with the regional models from the other regions. Therefore, the trained first regional model can consider multiple vehicle scenarios when outputting a decision path combination during path planning, avoiding traffic congestion caused by multiple vehicles occupying the same lane at the same time. Finally, the decision path combination is uploaded to the blockchain for access by all vehicles in the first region. In this way, multiple path planning tasks and multiple vehicle data are combined to simultaneously plan paths for multiple vehicles, thereby considering future traffic conditions, making path planning global and improving vehicle travel efficiency.

[0221] The specific implementation of each of the above units can be found in the previous embodiments and will not be described again here.

[0222] Figure 6The following is a block diagram of a portion of the terminal 110 for implementing an embodiment of the present disclosure. The terminal 110 includes components such as a radio frequency (RF) circuit 510, a memory 515, an input unit 530, a display unit 540, a sensor 550, an audio circuit 560, a wireless fidelity (WiFi) module 570, a processor 580, and a power supply 590. Those skilled in the art will appreciate that the structure of the terminal 110 shown in the figure does not limit the structure of a mobile phone or a computer, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0223] The RF circuit 510 may be used for receiving and sending signals during information transmission or calls. In particular, after receiving downlink information from the base station, it is sent to the processor 580 for processing. In addition, the uplink data is sent to the base station.

[0224] The memory 515 may be used to store software programs and modules. The processor 580 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the memory 515 .

[0225] The input unit 530 may be configured to receive input digital or character information and generate key signal input related to terminal settings and function control. Specifically, the input unit 530 may include a touch panel 531 and other input devices 532 .

[0226] The display unit 540 may be configured to display input information or provided information and various menus of the terminal. The display unit 540 may include a display panel 541 .

[0227] The audio circuit 560 , the speaker 561 , and the microphone 562 may provide an audio interface.

[0228] In this embodiment, the processor 580 included in the terminal 110 can execute the vehicle path planning method of the previous embodiment.

[0229] The terminal 110 of the embodiment of the present disclosure includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. The embodiment of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc.

[0230] Figure 7This is a block diagram of the structure of a portion of the server 120 for implementing an embodiment of the present disclosure. The server 120 may vary greatly due to different configurations or performance, and may include one or more central processing units (CPUs) 622 (for example, one or more processors) and memories 632, and one or more storage media 620 (for example, one or more mass storage devices) for storing application programs 642 or data 644. Among them, the memories 632 and the storage media 620 may be temporary storage or permanent storage. The program stored in the storage medium 620 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 120. Furthermore, the central processing unit 622 may be configured to communicate with the storage medium 620 to execute a series of instruction operations in the storage medium 620 on the server 120.

[0231] The server 120 may also include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input and output interfaces 658, and / or one or more operating systems 641, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0232] The central processor 622 in the server 120 can be used to execute the vehicle path planning method of the embodiment of the present disclosure.

[0233] The embodiments of the present disclosure further provide a computer-readable storage medium, which is used to store program code, and the program code is used to execute the vehicle path planning method of each of the aforementioned embodiments.

[0234] The present disclosure also provides a computer program product, which includes a computer program. A processor of a computer device reads and executes the computer program, so that the computer device implements the above-mentioned vehicle path planning method.

[0235] In addition, the terms "comprises" and "comprising" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, apparatus, product or apparatus that comprises a series of steps or elements is not necessarily limited to those steps or elements expressly listed but may include other steps or elements not expressly listed or inherent to such process, method, product or apparatus.

[0236] It should be understood that in the present disclosure, "at least one (item)" refers to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0237] It should be understood that in the description of the embodiments of the present disclosure, the meaning of multiple (or multiple items) is more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.

[0238] In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0239] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0240] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0241] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present disclosure. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0242] It should also be understood that the various implementations provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.

[0243] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or portion of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal. It can be implemented in whole or in part using software, hardware (such as processing circuits or memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the functionality of the module or unit.

[0244] The above is a specific description of the implementation methods of the present disclosure, but the present disclosure is not limited to the above implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without violating the spirit of the present disclosure. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present disclosure.

Claims

1. A vehicle path planning method, characterized in that: Applied to any first-region node in a first-region blockchain system, the first-region node and any second-region node in a second-region blockchain system corresponding to a second region form a global blockchain system, including: Obtaining, from the first regional blockchain corresponding to the first regional blockchain system, path planning tasks for multiple vehicles within the first regional area, wherein each path planning task is initiated by each vehicle and uploaded to the first regional blockchain after consensus verification by the first regional blockchain system; Determining the straight-line length of each vehicle's path based on the starting point information and the end point information carried by each path planning task, and determining the current vehicle state information of each vehicle and surrounding road condition information, and constructing a first vehicle state data set corresponding to the plurality of vehicles by combining the straight-line length of each vehicle's path, the vehicle state information, and the road condition information; Inputting the first vehicle state data set into the trained first regional model to obtain a decision path combination, wherein the decision path combination includes a decision planning path corresponding to each vehicle; wherein, after the first regional model outputs a first predicted decision path combination based on a first sample vehicle state data set, the first predicted decision path combination is combined with the first sample vehicle state data set, the first predicted decision path combination, and the second sample vehicle state data set and the second predicted decision path combination corresponding to the second region to construct a first regional loss, and the first regional loss and the second regional loss corresponding to the second region are combined to construct a global loss, and the maximization of the global cumulative reward value is used as the optimization goal to guide the collaborative training of the first regional model and the second regional model corresponding to the second region, where the global cumulative reward value is negatively correlated with the global loss; The second sample vehicle state data set, the second predicted decision path combination, and the second regional loss are shared by the second regional node to the global blockchain corresponding to the global blockchain system, and then obtained by the local node from the global blockchain; The decision planning path corresponding to each vehicle included in the decision path combination is uploaded to the first regional blockchain, so that each vehicle reads the corresponding decision planning path from the first regional blockchain.

2. The vehicle path planning method according to claim 1, characterized in that: The uploading of the decision planning path corresponding to each vehicle included in the decision path combination to the first regional blockchain includes: Calculating a path hash value for each decision-planning path corresponding to each vehicle included in the decision-planning path combination to obtain a path hash value combination, wherein the hash value combination includes a path hash value corresponding to each decision-planning path; Sending the path hash value combination to other first-region nodes in the first-region blockchain system for consensus verification to obtain a consensus verification result; When it is determined according to the consensus verification result that the number of target first-area nodes that have reached consensus is greater than a preset node number threshold, the path hash value combination and the decision planning path corresponding to each vehicle included in the decision path combination are uploaded to the first-area blockchain.

3. The vehicle path planning method according to claim 2, characterized in that: The sending of the path hash value combination to other first-region nodes in the first-region blockchain system for consensus verification to obtain a consensus verification result includes: Constructing a path hash tree according to the path hash value corresponding to each decision-making planning path included in the path hash value combination, wherein the path hash tree includes a path root hash; Determine the generation timestamp information, the receiver identifier, and the sender identifier associated with each decision-making plan path, and generate a target block based on the path hash tree, the path root hash, the decision-making plan path corresponding to each vehicle included in the decision-making path combination, and the generation timestamp information, the receiver identifier, and the sender identifier associated with each decision-making plan path; Sending the block header corresponding to the target block to other first-region nodes in the first-region blockchain system for consensus verification to obtain a consensus verification result, wherein the block header includes the generation timestamp information, the receiver identifier, the sender identifier, and the path root hash; Then, uploading the path hash value combination and the decision planning path corresponding to each vehicle included in the decision path combination to the first regional blockchain includes: Adding the target block to the first regional blockchain.

4. The vehicle path planning method according to claim 1, characterized in that: The first regional model includes a first decision sub-model and a first evaluation sub-model. Before inputting the vehicle state data set into the trained first regional model to obtain a decision path combination, wherein the decision path combination includes a decision planning path corresponding to each vehicle, the training process of the first regional model is as follows: Obtaining a first sample vehicle state data set, and inputting the first sample vehicle state data set into the first decision sub-model to obtain a first predicted decision path combination; Obtaining a second sample vehicle state data set and a second predicted decision path combination corresponding to the second area, and inputting the first sample vehicle state data set, the first predicted decision path combination, the second sample vehicle state data set, and the second predicted decision path combination into the first evaluation sub-model to obtain a first evaluation score; Obtaining a first evaluation expected value, and constructing a first regional loss according to a difference between the first evaluation expected value and the first evaluation score; Obtaining a second area loss corresponding to the second area, and determining a global cumulative reward value based on the first area loss and the second area loss using a preset global loss reward function; Taking maximizing the global cumulative reward value as an optimization goal, training the first decision sub-model and the first evaluation sub-model to obtain a trained first region model; The global cumulative reward value is maximized when the global loss corresponding to the first area loss and the second area loss is minimized.

5. The vehicle path planning method according to claim 4, characterized in that: Before obtaining the second area loss corresponding to the second area, the method further includes: Uploading the first sample vehicle state dataset and the first predicted decision path combination to the global blockchain, so that the second regional node obtains the first sample vehicle state dataset and the first predicted decision path combination from the global blockchain, and calculates a second evaluation score based on the first sample vehicle state dataset, the first predicted decision path combination, the second sample vehicle state dataset, and the second predicted decision path combination using a second regional model corresponding to the second region, and constructs a second regional loss based on the second evaluation score and a preset second evaluation expected value; Then, obtaining the second area loss corresponding to the second area includes: Obtain a second regional loss corresponding to the second regional from the global blockchain, wherein the second regional loss is uploaded to the global blockchain by the second regional node.

6. The vehicle path planning method according to claim 4, characterized in that: The first decision sub-model and the first evaluation sub-model are trained with maximizing the global cumulative reward value as the optimization goal to obtain the trained first region model, including: Determining the number of vehicles and the road occupancy rate corresponding to the first area according to the first sample vehicle status data set, and calculating a current global cumulative reward value according to the number of vehicles, the road occupancy rate, and the first sample vehicle status data set; Taking the current global cumulative reward value as the initial value, adjust the model parameters of the first decision sub-model and the first evaluation sub-model, and perform iterative training in an optimization direction in which the global cumulative reward value is positively increased until the maximum value of the global cumulative reward value is achieved when the global loss is minimized, thereby obtaining a trained first region model.

7. The vehicle path planning method according to claim 4, characterized in that: The first decision sub-model and the first evaluation sub-model are trained with maximizing the global cumulative reward value as the optimization goal to obtain the trained first region model, including: Obtaining from the global blockchain the number of vehicles in the second area and the second area road occupancy rate corresponding to the second area, the second area vehicle number and the second area road occupancy rate being determined by the second area node based on the second vehicle status dataset and uploaded to the global blockchain; Calculate the current global cumulative reward value according to the number of vehicles in the second area, the road occupancy rate in the second area, and the first sample vehicle status data set; Taking the current global cumulative reward value as the initial value, adjust the model parameters of the first decision sub-model and the first evaluation sub-model, and perform iterative training in an optimization direction in which the global cumulative reward value is positively increased until the maximum value of the global cumulative reward value is achieved when the global loss is minimized, thereby obtaining a trained first region model.

8. A vehicle path planning device, characterized in that: Applied to any first-region node in a first-region blockchain system, the first-region node and any second-region node in a second-region blockchain system corresponding to a second region form a global blockchain system, including: an acquiring unit, configured to acquire, from a first regional blockchain corresponding to the first regional blockchain system, path planning tasks for a plurality of vehicles within the first regional area, wherein each path planning task is initiated by each vehicle and uploaded to the first regional blockchain after consensus verification by the first regional blockchain system; a determining unit, configured to determine a straight-line length of a path for each vehicle based on the starting point information and the ending point information carried by each path planning task, and to determine current vehicle status information of each vehicle and surrounding road condition information, and to construct a first vehicle status dataset corresponding to the plurality of vehicles by combining the straight-line length of the path for each vehicle, the vehicle status information, and the road condition information; An input unit, configured to input the first vehicle state data set into the trained first regional model to obtain a decision path combination, wherein the decision path combination includes a decision planning path corresponding to each vehicle; wherein, after the first regional model outputs a first predicted decision path combination based on a first sample vehicle state data set, the first predicted decision path combination is combined with the first sample vehicle state data set, the first predicted decision path combination, and the second sample vehicle state data set and the second predicted decision path combination corresponding to the second region to construct a first regional loss, and the first regional loss and the second regional loss corresponding to the second region are combined to construct a global loss, and the maximization of the global cumulative reward value is used as the optimization goal to guide the collaborative training of the first regional model and the second regional model corresponding to the second region, where the global cumulative reward value is negatively correlated with the global loss; The second sample vehicle state data set, the second predicted decision path combination, and the second regional loss are shared by the second regional node to the global blockchain corresponding to the global blockchain system, and then obtained by the local node from the global blockchain; The sending unit is used to upload the decision planning path corresponding to each vehicle included in the decision path combination to the first regional blockchain, so that each vehicle reads the corresponding decision planning path from the first regional blockchain.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the vehicle path planning method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, which are suitable for being loaded by a processor to execute the vehicle path planning method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Vehicle-road cooperation multi-vehicle path planning and road right decision-making method and system and roadbed unit

    CN117651848A

  • Global path planning method and device for an unmanned vehicle

    US20220196414A1