Path planning method and device, electronic equipment and storage medium

By acquiring vehicle and passenger status information and using a large-scale multimodal model for vehicle-passenger matching and route planning, the problem of low scheduling efficiency in existing autonomous on-demand mobility systems is solved, achieving more efficient and safer route planning and scheduling.

CN120970673APending Publication Date: 2025-11-18HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511153188.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing autonomous on-demand mobility systems do not fully consider the complexity of urban road networks in vehicle dispatching, resulting in low dispatching efficiency and a lack of real-time route replanning mechanisms, which affects vehicle travel time estimation and user experience.

Method used

By acquiring vehicle and passenger status information, a large-scale multimodal model is used to match vehicles and passengers. Real-time road conditions and traffic rules are combined to perform route planning, generate the optimal driving route, and risk assessment and adjustment are carried out through model predictive control algorithms.

Benefits of technology

It improves the scheduling efficiency of the autonomous on-demand travel system, ensures the safety and real-time performance of route planning, and enhances user experience and system operating efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120970673A_ABST
    Figure CN120970673A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a path planning method and device, electronic equipment and a storage medium, and belongs to the technical field of intelligent traffic. The method comprises the steps that vehicle state information collected by a plurality of vehicles and passenger state information of a plurality of passengers waiting for the vehicles in the coverage range of the collaborative driving platform are acquired, the vehicle state information comprises the current position of each vehicle, and the passenger state information comprises the starting position of each passenger and the destination of each passenger; according to the vehicle state information and the passenger state information, the multiple vehicles and the multiple passengers are paired, multiple pairing results are obtained, and each pairing result comprises one vehicle and one passenger; and for each pairing result, performing path planning according to the current position of the vehicle and the starting position of the passenger in the pairing result to obtain a first target path of the vehicle driving from the current position to the starting position of the passenger in the pairing result. According to the embodiment of the invention, the path planning of the autonomous on-demand travel system for the intelligent connected automobile is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent transportation technology, and in particular to a path planning method, apparatus, electronic device and storage medium. Background Technology

[0002] With the acceleration of global urbanization, urban traffic congestion, energy consumption, and environmental pollution are becoming increasingly serious problems, urgently requiring efficient and sustainable mobility solutions. Autonomous Mobility-on-Demand (AMoD) systems, relying on connected and autonomous vehicles (CAVs), represent a shared and automated future transportation mode, demonstrating enormous potential to significantly alleviate traffic pressure, improve travel efficiency, and enhance user experience. By dynamically responding to passenger demand and optimizing vehicle scheduling, AMoD systems are expected to reduce empty mileage, lower emissions, and reshape urban transportation networks.

[0003] However, existing AMoD systems mostly use the straight-line distance between vehicles and passenger requests as the core optimization objective, without fully considering the impact of the complexity of urban road networks (such as road topology, intersection signals, one-way traffic restrictions, etc.) on scheduling efficiency, resulting in low scheduling efficiency for cooperative driving path planning. Summary of the Invention

[0004] The main objective of this application is to propose a path planning method, apparatus, electronic device, and storage medium, which aims to solve the technical problem of low scheduling efficiency of existing autonomous on-demand mobility systems for intelligent connected vehicles, thereby optimizing the path planning of autonomous on-demand mobility systems for intelligent connected vehicles and improving scheduling efficiency.

[0005] To achieve the above objectives, a first aspect of this application proposes a path planning method applied to a cooperative driving platform, the method comprising:

[0006] The system acquires vehicle status information collected from multiple vehicles within the coverage area of ​​the collaborative driving platform, as well as passenger status information from multiple passengers waiting to use the vehicle. The vehicle status information includes the current location of each vehicle, and the passenger status information includes the starting location and destination of each passenger.

[0007] Based on the vehicle status information and the passenger status information, multiple vehicles are paired with multiple passengers to obtain multiple pairing results, each pairing result including one vehicle and one passenger;

[0008] For each pairing result, path planning is performed based on the current position of the vehicle and the starting position of the passenger in the pairing result to obtain the first target path for the vehicle to travel from the current position to the starting position of the passenger in the pairing result.

[0009] In some embodiments, the vehicle status information also includes environmental image information collected for each vehicle, passenger status of each vehicle, and driving status of each vehicle.

[0010] The step of matching multiple vehicles with multiple passengers based on the vehicle status information and the passenger status information to obtain multiple matching results includes:

[0011] The environmental image information collected from each vehicle is fused to obtain a bird's-eye view of the coverage area of ​​the collaborative driving platform;

[0012] Assign an identifier to each of the vehicles and each of the passengers;

[0013] Based on the vehicle status information and the passenger status information, mark the passenger-carrying status of each vehicle, the driving status of each vehicle, the identification of each vehicle, the starting position of each passenger, and the identification of each passenger on the bird's-eye view.

[0014] A large-scale multimodal model is used to pair the identifiers corresponding to multiple vehicles with the identifiers corresponding to multiple passengers based on the bird's-eye view, resulting in multiple pairing results.

[0015] In some embodiments, the process of pairing multiple vehicle-related identifiers with multiple passenger-related identifiers based on the bird's-eye view using a large multimodal model to obtain multiple pairing results includes:

[0016] By using a large-scale multimodal model to infer the distance between each vehicle and each passenger, the driving direction of each vehicle, and the road layout in the bird's-eye view, matching results are obtained for the identifiers corresponding to multiple vehicles and the identifiers corresponding to multiple passengers.

[0017] Based on the matching results of the identifiers corresponding to the multiple vehicles and the identifiers corresponding to the multiple passengers, a pairing result is obtained for each passenger and the corresponding vehicle.

[0018] In some embodiments, the passenger status information also includes the waiting time for each passenger;

[0019] The method involves using a large-scale multimodal model to infer the distance between each vehicle and each passenger, the driving direction of each vehicle, and the road layout in the bird's-eye view, to obtain matching results between the identifiers corresponding to multiple vehicles and the identifiers corresponding to multiple passengers, including:

[0020] The distance matrix between each vehicle and each passenger is obtained by using the large multimodal model to determine the relative positions of each vehicle and each passenger.

[0021] The matching degree between the driving direction of each vehicle and the starting position of each passenger is obtained through the large multimodal model.

[0022] The road layout is analyzed using the large-scale multimodal model to obtain the estimated time for each vehicle to reach the starting position of each passenger.

[0023] The waiting time of each passenger is evaluated using the large-scale multimodal model to obtain a matching priority order;

[0024] The matching result is obtained based on the distance matrix, the matching degree, the estimated time, and the priority order.

[0025] In some embodiments, for each pairing result, path planning is performed based on the current position of the vehicle and the starting position of the passenger in the pairing result to obtain a first target path for the vehicle to travel from its current position to the starting position of the passenger in the pairing result, including:

[0026] For each pairing result, path planning is performed based on the current position of the vehicle and the starting position of the passenger in the pairing result, and the first driving trajectory corresponding to each pairing result is obtained through the trajectory generation function;

[0027] Based on the multiple first driving trajectories corresponding to the multiple pairing results, a first global trajectory is obtained;

[0028] A risk assessment is performed on the first global trajectory using a large-scale multimodal model to obtain a first risk assessment result, which is used to indicate the risk of a vehicle collision in the first global trajectory.

[0029] Based on the first risk assessment result, the first global trajectory is adjusted by a model predictive control algorithm to obtain the first target path corresponding to each pairing result.

[0030] In some embodiments, the step of adjusting the first global trajectory based on the first risk assessment result using a model predictive control algorithm to obtain a first target path corresponding to each pairing result includes:

[0031] The first risk assessment result is input into the model predictive control algorithm to generate a local reference trajectory;

[0032] Based on the local reference trajectory, each of the first driving trajectories in the first global trajectory is dynamically adjusted to obtain the first target path corresponding to each pairing result.

[0033] In some embodiments, after performing path planning for each pairing result based on the current position of the vehicle and the starting position of the passenger in the pairing result to obtain a first target path for the vehicle to travel from its current position to the starting position of the passenger in the pairing result, the method further includes:

[0034] For each pairing result, a route is planned based on the passenger's starting position and destination in the pairing result, and a second driving trajectory corresponding to each pairing result is obtained through a trajectory generation function;

[0035] Based on the multiple second driving trajectories corresponding to the multiple pairing results, a second global trajectory is obtained;

[0036] The risk assessment of the second global trajectory is performed using a large-scale multimodal model to obtain a second risk assessment result, which is used to indicate the risk of vehicle collision in the second global trajectory.

[0037] Based on the second risk assessment result, the second global trajectory is adjusted by a model predictive control algorithm to obtain a second target path for the vehicle in each pairing result, from the passenger's starting position to the passenger's destination.

[0038] To achieve the above objectives, a second aspect of this application provides a path planning device applied to a cooperative driving platform, the device comprising:

[0039] The acquisition module is used to acquire vehicle status information collected by multiple vehicles within the coverage area of ​​the collaborative driving platform and passenger status information of multiple passengers waiting to use the vehicle. The vehicle status information includes the current location of each vehicle, and the passenger status information includes the starting location and destination of each passenger.

[0040] The matching module is used to match multiple vehicles with multiple passengers based on the vehicle status information and the passenger status information, and obtain multiple matching results, each of the matching results including one vehicle and one passenger;

[0041] The planning module is used to perform path planning for each pairing result based on the current position of the vehicle and the starting position of the passenger in the pairing result, so as to obtain a first target path for the vehicle to travel from the current position to the starting position of the passenger in the pairing result.

[0042] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0043] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0044] The path planning method, apparatus, electronic device, and storage medium proposed in this application involve a cooperative driving platform first acquiring vehicle status information of multiple vehicles within its coverage area and passenger status information of multiple passengers waiting to use the vehicle. The vehicle status information includes the current location of each vehicle, and the passenger status information includes the starting location and destination of each passenger. This information is uploaded to the platform in real time via in-vehicle equipment and passenger mobile devices. Based on the acquired vehicle and passenger status information, the platform pairs multiple vehicles with multiple passengers. The pairing process considers factors such as the distance between the current location of the vehicle and the starting location of the passenger, and the vehicle's driving direction. An algorithm is used to obtain the optimal vehicle-passenger pairing result, and each pairing result contains a correspondence between one vehicle and one passenger. For each pairing result, the platform performs path planning based on the current location of the paired vehicle and the starting location of the passenger. The planning process considers factors such as real-time road conditions and traffic rules to generate the optimal driving path from the current location of the vehicle to the starting location of the passenger, i.e., the first target path. This application effectively solves the technical problem of low scheduling efficiency of existing autonomous on-demand mobility systems for intelligent connected vehicles by dynamically matching real-time vehicle status with passenger demand information, constructing a global bird's-eye view by combining multimodal environmental perception data, and generating safe paths using a collaborative optimization algorithm. This optimizes the path planning of autonomous on-demand mobility systems for intelligent connected vehicles and improves scheduling efficiency. Attached Figure Description

[0045] Figure 1 This is a flowchart illustrating the path planning method provided in an embodiment of this application;

[0046] Figure 2 This is a schematic diagram of the cooperative driving visualization framework provided in the embodiments of this application;

[0047] Figure 3This is a schematic diagram of the path planning device provided in the embodiments of this application;

[0048] Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0050] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0052] With the acceleration of global urbanization and the continuous rise in urban population density, problems such as traffic congestion, excessive energy consumption, and increased carbon emissions are becoming increasingly prominent, making the demand for efficient and sustainable urban transportation solutions more urgent. Autonomous On-Demand Mobility (AMoD) systems, based on connected vehicles (CAVs), provide shared and on-demand mobility services and are considered a transformative technological direction for addressing these challenges. They not only promise to significantly reduce traffic congestion and carbon emissions but also improve the overall operational efficiency of urban transportation networks, reshaping future urban travel patterns and transportation infrastructure.

[0053] However, the deployment and optimization of current AMoD systems in real-world scenarios still face several key technical challenges. Current CAV scheduling strategies are often oversimplified, relying primarily on the Euclidean distance between vehicles and passenger requests for allocation, failing to fully consider the significant impact of complex road network topologies (such as one-way streets, intersections, and congested sections) on actual travel time and scheduling efficiency. This simplification often results in suboptimal scheduling outcomes in real-world road network environments, leading to increased vehicle empty mileage, longer passenger waiting times, and decreased overall system efficiency. Existing systems, when planning CAV routes, often assume vehicles can accurately reach their destinations based on static historical data, ignoring the significant impact of real-time traffic flow dynamics (such as sudden congestion and accidents) on travel time. The lack of an effective real-time route replanning mechanism makes it difficult for vehicles to adapt to dynamic environments, resulting in inaccurate travel time estimates and large fluctuations in completion times, impacting service reliability and user experience. Cooperative motion planning (such as multi-vehicle cooperative obstacle avoidance and platooning) is a crucial downstream task of AMoD systems after vehicle scheduling, essential for achieving efficient and safe shared autonomous driving. However, existing research and technology platforms generally suffer from a significant deficiency: they fail to effectively integrate macroscopic vehicle scheduling (including destination allocation) with microscopic cooperative motion planning. Scheduling decisions fail to provide sufficient environmental context and cooperative constraints for subsequent motion planning, and the results of motion planning are difficult to provide timely feedback to influence scheduling strategies, creating "information silos." Existing mainstream AMoD simulation platforms, such as those for CAV scheduling (MATSim, CAVSim, AIMSUN, etc.), primarily focus on macroscopic traffic modeling and lack the detailed vehicle-level modeling required for motion planning, thus failing to meet the needs of AMoD systems. Platforms for motion planning and control (such as SUMO, CARLA, LG SVL, Apollo, CarSim, etc.) are insufficient in terms of graphics capabilities, support for complex sensor simulations, motion planning functions, and large-scale scheduling scenario simulation tools.

[0054] Based on this, embodiments of this application provide a route planning method, apparatus, electronic device, and storage medium, aiming to optimize CAV route planning and improve scheduling efficiency by incorporating road network details into the scheduling algorithm; to achieve real-time CAV route planning and improve completion efficiency by making adaptive decisions based on real-time traffic information; and to build a simulation platform that supports multimodal environments, supports the integration of LMMs, and promotes the development of AMoD systems.

[0055] The path planning method, apparatus, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the path planning method in the embodiments of this application is described.

[0056] The path planning method provided in this application relates to the field of intelligent transportation technology. The path planning method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the path planning method, but is not limited to the above forms.

[0057] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0058] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0059] Figure 1 This is an optional flowchart of the path planning method provided in the embodiments of this application. Figure 1 The method described above can be applied to a collaborative driving platform and may include, but is not limited to, steps S100 to S300.

[0060] Step S100: Obtain vehicle status information and passenger status information of multiple vehicles within the coverage area of ​​the cooperative driving platform. The vehicle status information includes the current location of each vehicle, and the passenger status information includes the starting location and destination of each passenger.

[0061] In this embodiment, the cooperative driving platform is an integrated platform for Autonomous On-Demand Mobility (AMoD). This platform achieves deep integration at both the micro and macro levels, focusing not only on the specific driving behavior of individual intelligent connected vehicles (CAVs), including detailed operations such as real-time path adjustment, obstacle avoidance, and throttle / steering control, ensuring the safe and precise movement of vehicles in complex environments; but also on global resource optimization, including how to efficiently allocate CAVs to passenger requests, balance vehicle distribution within an area, and reduce overall waiting time, thereby improving the operational efficiency of the entire transportation system. Leveraging the powerful perception capabilities of CARLA (a high-fidelity autonomous driving simulation platform), the platform generates rich multimodal data outputs to support traffic system decision-making. It comprehensively reconstructs the vehicle's environment using multimodal data such as point clouds generated by LiDAR (LiDAR, accurately perceiving the 3D position of surrounding objects) and images / videos captured by cameras (capturing visual scene details). Then, relying on large-scale multimodal models (LMMs), it performs comprehensive analysis of this multimodal data (e.g., identifying traffic congestion, pedestrian behavior, road signs), providing more intelligent references for vehicle scheduling, route planning, and other operations, enhancing the system's understanding of complex scenarios. The platform features adaptive bird's-eye view (BEV) generation, which visually displays vehicle distribution, passenger request locations, road networks, and other information within the entire traffic area from a top-down perspective, helping dispatchers or algorithms quickly grasp the overall status. Simultaneously, the BEV can be dynamically updated based on real-time data (e.g., new passenger requests, vehicle location changes), ensuring that the visualized content is consistent with the actual system state, providing intuitive support for scheduling decisions.

[0062] In this embodiment, the perception module of the cooperative driving platform collects vehicle status information of multiple vehicles and passenger status information of multiple waiting passengers within the platform's coverage area. Vehicle status information includes each vehicle's current location (e.g., GPS-based latitude and longitude coordinates or coordinates on a local map of the cooperative driving platform), speed, and direction of travel, and may also include vehicle type, idle / occupied status, etc. Passenger status information includes each passenger's starting location (i.e., pick-up point, such as roadside, community entrance, etc.), destination (i.e., drop-off point), and may also include waiting time, number of passengers, etc.

[0063] In a preferred embodiment, the cooperative driving platform possesses multimodal perception capabilities. It utilizes the perception capabilities of high-fidelity simulators such as CARLA to generate multimodal outputs including LiDAR point clouds, multi-view RGB images, and videos. Specifically, the platform adaptively generates a bird's-eye view (BEV) image by fusing data from vehicle surround-view cameras or directly acquiring vehicle position relationships via vehicle-to-vehicle (V2V) communication. This BEV image visually displays the entire urban road network from a top-down perspective, the real-time location and orientation of all vehicles (each vehicle is labeled with a unique ID), and the distribution of pending passenger requests (each passenger is also labeled with a unique ID), providing rich spatial contextual information for subsequent pairing decisions and route planning.

[0064] Step S200: Based on the vehicle status information and the passenger status information, multiple vehicles are paired with multiple passengers to obtain multiple pairing results. Each pairing result includes one vehicle and one passenger.

[0065] In this embodiment, the platform executes a vehicle-passenger pairing algorithm based on the acquired vehicle and passenger status information. Each pairing result includes a correspondence between one vehicle and one passenger. During the pairing process, the scheduling module of the cooperative driving platform comprehensively considers factors such as the road network distance (non-linear distance, requiring calculation based on road topology) between the vehicle's current location and the passenger's starting location, the passenger's waiting time (prioritizing matching passengers with longer waiting times), the directional relationship between the vehicle's driving direction and the passenger's starting location (e.g., prioritizing pairing when the vehicle's current driving direction is consistent with the direction to the passenger's starting location to reduce detours), and the real-time status of the road network (e.g., whether the path from the vehicle to the passenger's starting location is congested, or whether there are temporary traffic controls), to match vehicles and passengers.

[0066] Specifically, the pairing process fully utilizes the aforementioned generated BEV images and related text data (such as vehicle ID, passenger ID, distance matrix, and passenger waiting time). The platform can employ various algorithms for pairing decisions, including heuristic methods such as First-Come, First-Served (FCFS) and Distance-First (DF), or Large Multimodal Models (LMMs). The LMM-based approach inputs both the BEV images and structured text data into the LMM (such as GPT-4V). LMMs leverage their powerful contextual understanding, image recognition, and logical reasoning capabilities to comprehensively consider multiple complex factors such as vehicle-passenger proximity, driving direction consistency, road network congestion, and passenger waiting time, performing chain-of-thought (CoT) reasoning to ultimately output a globally optimal or near-optimal vehicle-passenger pairing scheme. This method can handle complex scenarios that are difficult to quantify using traditional algorithms, achieving more intelligent and human-centered scheduling. After pairing is complete, the system updates the status of each vehicle assigned a task (e.g., from "Idle" to "Hired" or "Hire") and updates the status of the corresponding passenger (e.g., from "Waiting" to "Assigned"). The pairing results and visualizations will be stored by the system for subsequent analysis and reference.

[0067] Step S300: For each pairing result, perform path planning based on the current position of the vehicle and the starting position of the passenger in the pairing result to obtain the first target path for the vehicle to travel from the current position to the starting position of the passenger in the pairing result.

[0068] In this embodiment, for each vehicle-passenger pairing result obtained in step S200, the platform initiates the path planning module. This module uses the current location of the paired vehicle as the starting point and the passenger's starting location (boarding point) as the ending point to perform path planning, aiming to generate an optimal "first target path," that is, the best route for the vehicle to travel from its current location to the passenger's boarding point. During path planning, the planning module constructs a road network map based on road topology information (such as lane distribution, intersection connections, and speed limit information), using the vehicle's current location and the passenger's starting location as nodes in the road network map; it employs path search algorithms (such as A algorithm, Jump Point Search (JPS), Iterative Deepening A (IDA), etc.), combined with real-time traffic conditions (such as road congestion coefficients obtained through the cooperative driving platform and the real-time locations of other vehicles), to search for the optimal path; the path optimization objective can be the shortest travel time, the fewest number of turns, or the highest road capacity (such as prioritizing road segments with more lanes), and can be dynamically adjusted according to actual needs.

[0069] In a preferred embodiment, the path planning process also considers dynamic environmental factors. The platform can employ various path planning algorithms, such as the A* algorithm, Jump Point Search (JPS) algorithm, and Iterative Deepening A* (IDA) algorithm, to generate a theoretically optimal initial trajectory in the global road network, without considering collision constraints with surrounding vehicles. Subsequently, to ensure driving safety and efficiency, the platform utilizes methods such as Model Predictive Control (MPC), Proximal Policy Optimization (PPO), or Finite State Machine (FSM) to perform risk assessment and local optimization on the generated initial trajectory. Taking MPC as an example, it identifies potential collision hazards or inefficient driving situations based on the vehicle dynamics model, real-time traffic flow, road conditions, and the predicted trajectories of other surrounding vehicles. By solving an optimization problem that includes multiple objectives such as trajectory tracking, energy consumption optimization, and comfort evaluation, it generates a smooth, safe, and efficient final driving strategy (i.e., a local reference trajectory) within a local spacetime. This final driving strategy constitutes the "first target path" in this step. Throughout the process, the platform achieves unified management of CAVs and passenger requests through an organized dictionary. Its control mode supports precise location changes and throttle / steering inputs, and facilitates subsequent route updates and coordination based on passenger destination information after boarding.

[0070] This embodiment, through deep integration with the cooperative driving platform, utilizes the platform's global perception data (multiple vehicle statuses, real-time road conditions) to make route planning more adaptable to the dynamic traffic environment, laying the foundation for subsequent trip planning after the vehicle carries passengers, and significantly improving the operating efficiency and coordination of the entire AMoD system.

[0071] In some embodiments, the vehicle status information may further include environmental image information collected for each vehicle, passenger status of each vehicle, and driving status of each vehicle.

[0072] Step S200 may include, but is not limited to, steps S210 to S240:

[0073] Step S210: Perform image fusion on the environmental image information collected by each vehicle to obtain a bird's-eye view of the coverage area of ​​the collaborative driving platform;

[0074] Step S220: Assign an identifier to each of the vehicles and each of the passengers;

[0075] Step S230: Based on the vehicle status information and the passenger status information, mark the passenger-carrying status of each vehicle, the driving status of each vehicle, the identification of each vehicle, the starting position of each passenger, and the identification of each passenger on the bird's-eye view.

[0076] Step S240: Using a large multimodal model, pair the identifiers corresponding to multiple vehicles with the identifiers corresponding to multiple passengers based on the bird's-eye view to obtain multiple pairing results.

[0077] In this embodiment, the collaborative driving platform acquires vehicle status information collected by multiple vehicles within its coverage area, as well as passenger status information of multiple waiting passengers. Environmental image information refers to image data about the surrounding environment collected by each vehicle through its onboard cameras (such as surround-view cameras, forward-view cameras, etc.). These images are the vehicle's most direct perception of the real world. Passenger status refers to the current operating status of each vehicle, such as "idle" (no passengers, available for tasks), "on the way" (assigned tasks, heading to pick up passengers or to the destination), "carrying passengers" (passengers picked up, heading to the destination), etc. Driving status refers to the dynamic motion information of each vehicle, such as current speed, acceleration, and direction of travel.

[0078] In this embodiment, after receiving environmental image information (such as multi-view RGB camera images and LiDAR point cloud data) uploaded by all vehicles, the perception generation module of the cooperative driving platform generates a bird's-eye view (BEV) covering the platform's service area using multi-sensor data fusion technology. Specifically, an algorithm based on the BEV (Bird's-Eye View) Transformer can be used, which extracts features and transforms perspectives from 2D images from different vehicles and perspectives through deep neural networks. Finally, these scattered and local visual information are fused and projected into a unified, global coordinate system to generate a bird's-eye view covering the entire service area of ​​the cooperative driving platform. The BEV image can intuitively present global information such as road networks, vehicle distribution, and passenger positions; the generation of the bird's-eye view supports dynamic updates and can reflect environmental changes in real time (such as the addition of new vehicles or changes in passenger positions).

[0079] In this embodiment, the scheduling module of the cooperative driving platform assigns a unique vehicle identifier to each vehicle participating in the scheduling and a unique passenger identifier to each passenger waiting to use the vehicle. The assignment of identifiers follows the principle of uniqueness, ensuring that each vehicle and each passenger can be accurately located during the subsequent pairing process; the identifier information is stored in association with the status information of the vehicle and passenger (e.g., managed through an ordered dictionary) for easy retrieval and retrieval.

[0080] In this embodiment, based on vehicle status information and passenger status information, the generated bird's-eye view is labeled with information such as the passenger-carrying status, driving status, vehicle identification, and the starting position and passenger identification of each vehicle. The labeled bird's-eye view can intuitively show the spatial distribution relationship between vehicles and passengers, as well as the real-time operating status of vehicles, providing a visual reference for pairing decisions.

[0081] In this embodiment, the cooperative driving platform invokes large-scale multimodal models (LMMs), inputting an annotated bird's-eye view and associated textual data (such as passenger waiting time and the road network distance matrix between vehicles and passengers). It then pairs vehicle identifiers with passenger identifiers, outputting multiple pairing results (each result representing a "vehicle identifier - passenger identifier" correspondence). The LMMs analyze visual information (such as vehicle position, driving direction, and passenger distribution) and textual data (such as waiting time and road topology) from the bird's-eye view, performing contextual reasoning to analyze the positional relationship between vehicles and passengers, road conditions, and other information, outputting the optimal vehicle-passenger pairing result. The pairing result must meet the "one-to-one" principle, meaning one vehicle is matched with only one passenger, and one passenger is picked up by only one vehicle.

[0082] This embodiment uses image fusion technology to convert scattered, multi-view environmental image information collected by vehicles into a global bird's-eye view. This expands the basis for scheduling decisions from single numerical data to multimodal data containing rich visual context, greatly improving the completeness and accuracy of decision information. By using the labeled BEV map as input to LMMs, the powerful image understanding, spatial reasoning, and complex decision-making capabilities of LMMs are introduced into the core scheduling link of the AMoD system, enabling pairing decisions to respond quickly to changes in the traffic environment (such as sudden congestion or new passenger requests), thus improving the intelligence level of decision-making.

[0083] In some embodiments, step S240 may include, but is not limited to, steps S241 to S242:

[0084] Step S241: Using a large-scale multimodal model, reason about the distance between each vehicle and each passenger, the driving direction of each vehicle, and the road layout in the bird's-eye view to obtain matching results between the identifiers corresponding to multiple vehicles and the identifiers corresponding to multiple passengers.

[0085] Step S242: Based on the matching results of the identifiers corresponding to the multiple vehicles and the identifiers corresponding to the multiple passengers, obtain the pairing result of each passenger and the corresponding vehicle.

[0086] In this embodiment, the collaborative driving platform calls large multimodal models (LMMs), takes the annotated bird's-eye view as input, and analyzes the distance between vehicles and passengers, vehicle driving direction, and road layout contained in the bird's-eye view.

[0087] Specifically, LMMs first identify icons for all available vehicles and all passengers on the BEV image. Based on the road network topology (such as lane distribution and intersection connections) in the bird's-eye view, they calculate the actual road network distance (non-straight-line distance) between each vehicle (including its icon) and each passenger (including their icon). LMMs not only identify vehicle positions but also assess directional convenience based on their direction of travel. For a vehicle and a passenger, the model determines whether the vehicle's current orientation makes it convenient to drive directly to the passenger. For example, if the arrow of a vehicle (e.g., A01) is pointing towards a passenger (e.g., B02), even if another vehicle (e.g., A02) is closer to passenger B02 but facing the opposite direction, LMMs may infer that A01 is the better choice. LMMs evaluate the feasibility of the path from the vehicle to the passenger's starting position by analyzing the road network structure in the BEV image, such as whether it is a one-way street, whether there are construction sections, and the timing of traffic lights at intersections. Then, based on the above reasoning results and combined with preset optimization objectives (such as minimizing passenger waiting time and maximizing vehicle utilization), LMMs generate multiple matching results between vehicle identifiers and passenger identifiers. The matching results must cover all passengers, ensuring that each passenger has at least one potentially matching vehicle.

[0088] For each vehicle-passenger pairing, the platform uses a path planning module to generate a first target path, starting from the current location of the paired vehicle and ending at the passenger's starting location. This path is achieved by employing global path planning algorithms such as A* and JPS, combined with local trajectory optimization methods such as MPC and PPO. Taking into account factors such as vehicle dynamics, real-time traffic, and obstacle avoidance, the platform guides the vehicle to the passenger's boarding point.

[0089] This embodiment achieves intelligent matching between vehicles and passengers by comprehensively considering multiple factors such as distance, driving direction, road layout, and waiting time, thereby improving the efficiency of vehicle dispatching. The application of a large-scale multimodal model enables the system to process complex scene information and make more reasonable matching decisions.

[0090] In some embodiments, the passenger status information may also include the waiting time for each passenger.

[0091] Step S241 may include, but is not limited to, the following steps:

[0092] The distance matrix between each vehicle and each passenger is obtained by using the large multimodal model to determine the relative positions of each vehicle and each passenger.

[0093] The matching degree between the driving direction of each vehicle and the starting position of each passenger is obtained through the large multimodal model.

[0094] The road layout is analyzed using the large-scale multimodal model to obtain the estimated time for each vehicle to reach the starting position of each passenger.

[0095] The waiting time of each passenger is evaluated using the large-scale multimodal model to obtain a matching priority order;

[0096] The matching result is obtained based on the distance matrix, the matching degree, the estimated time, and the priority order.

[0097] In this embodiment, the passenger status information also includes the waiting time of each passenger (i.e., the time elapsed from when the passenger initiated the ride request to the present). This information can be collected in real time through the passenger terminal (such as a mobile APP) of the collaborative driving platform and synchronized to the dispatch module.

[0098] In this embodiment, large-scale multimodal models (LMMs) calculate the road network distance between each vehicle and each passenger based on the relative positions of vehicles and passengers in the bird's-eye view, combined with the road network topology (such as lane connections and intersection distribution), and present it in matrix form (i.e., distance matrix). LMMs quantify the matching degree (value range can be set to 0-1, where 1 represents a perfect match and 0 represents a complete opposite) by analyzing the orientation relationship between the vehicle's marked driving direction (e.g., arrow direction) and the passenger's starting position in the bird's-eye view. For example, if the vehicle's driving direction is consistent with the optimal path direction to the passenger's starting position, the matching degree is 0.8-1; if a small-angle turn (≤30°) is required, the matching degree is 0.5-0.7; if a large-angle turn (>30°) or U-turn is required, the matching degree is 0-0.4. LMMs analyze the road layout in the bird's-eye view (such as road type, number of lanes, presence of congested sections, intersection traffic lights, etc.), and calculate the estimated time for each vehicle to reach each passenger's starting position by combining the vehicle's current speed and the road network distance in the distance matrix. LMMs evaluate passenger waiting times and convert them into matching priorities, assigning higher priority based on the principle that longer waiting times result in higher priority. Finally, LMMs use the distance matrix, matching degree, estimated time, and priority as input to generate matching results between vehicle identifiers and passenger identifiers.

[0099] This embodiment transforms abstract factors into calculable quantitative indicators through distance matrix, matching degree, estimated time, and priority order, making LMMs inference more operational; it integrates the estimated time affected by road layout with the matching degree of driving direction to avoid inefficient matching caused by "distance-only" and ensure that vehicles can quickly reach passenger locations; it reflects the impact of passenger waiting time through priority order, improving overall service efficiency while ensuring that passengers with long waiting times receive priority service; thus achieving more intelligent and efficient vehicle scheduling.

[0100] In some embodiments, step S300 may include, but is not limited to, steps S310 to S340:

[0101] Step S310: For each pairing result, path planning is performed based on the current position of the vehicle and the starting position of the passenger in the pairing result, and the first driving trajectory corresponding to each pairing result is obtained through the trajectory generation function;

[0102] Step S320: Obtain the first global trajectory based on the multiple first driving trajectories corresponding to the multiple pairing results;

[0103] Step S330: Perform a risk assessment on the first global trajectory using a large-scale multimodal model to obtain a first risk assessment result. The first risk assessment result is used to indicate the risk of a vehicle collision in the first global trajectory.

[0104] Step S340: Based on the first risk assessment result, the first global trajectory is adjusted using a model prediction control algorithm to obtain the first target path corresponding to each pairing result.

[0105] In this embodiment, for each pairing result, the planning module of the cooperative driving platform calls the trajectory generation function to generate the vehicle's first driving trajectory based on the vehicle's current position and the passenger's starting position. The trajectory generation function can employ path search algorithms such as A* algorithm, Jump Point Search (JPS), and Iterative Deepening A* (IDA). Based on road topology information (such as lane distribution, intersection connections, and speed limit signs), it constructs a road network map and finds the theoretically shortest or fastest path connecting the starting point and the ending point on a known static road network map, with the vehicle's current position as the starting point and the passenger's starting position as the ending point.

[0106] In this embodiment, the planning module summarizes the first driving trajectories corresponding to all pairing results to form a first global trajectory covering the service scope of the cooperative driving platform. The first global trajectory can be visualized through a bird's-eye view (BEV), with different colored lines marking the trajectories of each vehicle, intuitively displaying potential conflict points such as trajectory intersections and parallelism.

[0107] In this embodiment, the cooperative driving platform invokes large multimodal models (LMMs) to perform a risk assessment on the first global trajectory and outputs the first risk assessment result. The platform inputs the integrated first global trajectory and the latest bird's-eye view reflecting the current real-time environment into the LMMs for analysis and identification of collision risks. Against the backdrop of the BEV map, the LMMs analyze the trajectories of all vehicles in the first global trajectory, identifying which vehicle trajectories have spatial intersections (such as at the same intersection or in the same lane). For each identified spatial intersection, the LMMs further analyzes the estimated time for the vehicles involved to reach that intersection. If two or more vehicles are expected to arrive at the same point within a very close time window (e.g., less than a safety threshold), the model determines that a "spatiotemporal conflict" exists. Then, the LMMs assesses a risk level for each identified spatiotemporal conflict. The risk level assessment can be based on multiple factors, such as: the smaller the expected time difference between vehicle intersections, the higher the risk; the higher the road complexity at the intersection (e.g., an intersection without traffic lights has a higher risk than a straight road); the faster the speed of the vehicles involved, the higher the risk, etc.

[0108] In this embodiment, the planning module uses the model predictive control (MPC) algorithm to adjust the first global trajectory based on the first risk assessment result, thereby obtaining the first target path corresponding to each pairing result.

[0109] Model-based control (MPC) is a model-based optimization control method that can predict the dynamic behavior of the system over a future time domain at each control moment and determine the current optimal control input by solving an online optimization problem. In this embodiment, MPC can use "reducing collision risk" as the core constraint and "minimizing the increase in trajectory length" as the optimization objective to locally correct the trajectory in high-risk areas.

[0110] Specifically, the platform uses the initial global trajectory as a reference and high-risk points from the initial risk assessment as "hard constraints" or "high-cost penalty items" that must be avoided, constructing a large-scale cooperative trajectory optimization problem. The goal of this optimization problem is to minimize the total travel time, total energy consumption, or total deviation from the original reference trajectory for all vehicles while satisfying all safety constraints (collision avoidance). Then, the MPC (Multi-Purpose Optimization) comprehensively considers the dynamic models of all vehicles and their coupling constraints (i.e., collision avoidance constraints), coordinating the adjustment of all trajectories in the initial global trajectory. For example, by fine-tuning the vehicle speed curves, one vehicle slightly accelerates while another slightly decelerates, thus "staggering" their passage through the conflict point in time and eliminating the collision risk. Finally, based on the MPC's cooperative optimization, the originally risky initial global trajectory is adjusted into a new trajectory that balances safety and efficiency. From this new trajectory, the platform extracts the corresponding adjusted and optimized first target path for each vehicle.

[0111] This embodiment constructs a first global trajectory and uses LMMs for risk assessment, which can identify potential collision risks between multiple vehicles in advance and avoid the problem of single-vehicle optimization but global conflict. By transforming the path planning problem of multiple vehicles from an isolated individual optimization problem into a global, system-level collaborative optimization problem, it fundamentally solves the persistent problem of multi-vehicle trajectory conflict in traditional methods and lays the foundation for the large-scale, high-density deployment of AMoD systems.

[0112] In some embodiments, step S340 may include, but is not limited to, steps S341 to S342:

[0113] Step S341: Input the first risk assessment result into the model prediction and control algorithm to generate a local reference trajectory;

[0114] Step S342: Dynamically adjust each of the first driving trajectories in the first global trajectory according to the local reference trajectory to obtain the first target path corresponding to each pairing result.

[0115] In this embodiment, the planning module of the cooperative driving platform inputs the first risk assessment result (such as the location of high-risk areas, vehicle identification, risk level, etc.) into the Model Predictive Control (MPC) algorithm as the core constraint for trajectory adjustment. Based on the input first risk assessment result and the vehicle dynamics model, the MPC algorithm generates a local reference trajectory for the high-risk areas. The planning module then dynamically adjusts each first driving trajectory involving high risk in the first global trajectory according to the generated local reference trajectory, obtaining the first target path corresponding to each pairing result.

[0116] This embodiment improves the safety and efficiency of path planning by dynamically adjusting the first global trajectory; the generation of the local reference trajectory takes into account the risk assessment results, which can effectively avoid potential collision risks; the dynamic adjustment process ensures the real-time performance and adaptability of the trajectory, and can cope with complex and ever-changing traffic environments.

[0117] In some embodiments, steps S300 may be followed by, but are not limited to, steps S301 to S304:

[0118] Step S301: For each pairing result, perform path planning based on the passenger's starting position and destination in the pairing result, and obtain the second driving trajectory corresponding to each pairing result through a trajectory generation function;

[0119] Step S302: Obtain the second global trajectory based on the multiple second driving trajectories corresponding to the multiple pairing results;

[0120] Step S303: Perform a risk assessment on the second global trajectory using a large-scale multimodal model to obtain a second risk assessment result. The second risk assessment result is used to indicate the risk of a vehicle collision in the second global trajectory.

[0121] Step S304: Based on the second risk assessment result, the second global trajectory is adjusted using a model predictive control algorithm to obtain a second target path for the vehicle in each pairing result, from the passenger's starting position to the passenger's destination.

[0122] In this embodiment, for each pairing result (i.e., the established vehicle-passenger correspondence), after the vehicle arrives at the passenger's starting position along the first target path, the planning module of the cooperative driving platform calls the trajectory generation function to generate the vehicle's second driving trajectory based on the passenger's starting position and destination. The trajectory generation function can employ the same path search algorithm as the first stage (such as the A* algorithm, jump point search, etc.). The planning module aggregates the second driving trajectories corresponding to all pairing results to form a second global trajectory covering the service area of ​​the cooperative driving platform. The second global trajectory can be visualized through the bird's-eye view (BEV) adaptive generation function, using dynamic lines to mark the driving direction of each vehicle and the estimated arrival time at key nodes (such as intersections, toll stations), visually displaying potential trajectory intersections or parallel areas. The cooperative driving platform calls large multimodal models (LMMs) to perform a risk assessment on the second global trajectory and outputs the second risk assessment result. The planning module uses the model predictive control (MPC) algorithm to adjust the second global trajectory based on the second risk assessment result, obtaining the second target path (i.e., the final path the vehicle takes from the passenger's starting position to the destination) corresponding to each pairing result. The specific risk assessment and trajectory adjustment process is the same as in the first phase, and will not be repeated here.

[0123] This embodiment generates an initial path through a trajectory generation function, then uses a large-scale multimodal model for risk assessment, and finally uses model predictive control for trajectory optimization, effectively avoiding potential collisions between vehicles. It seamlessly extends the collaborative path planning framework from the "pick-up" stage to the "drop-off" stage, forming a closed-loop optimization system that covers the entire lifecycle of a vehicle from accepting a task to completing the service, ensuring that the AMoD system can maintain efficient and safe operation at any service stage.

[0124] The cooperative driving visualization framework in this embodiment is as follows: Figure 2 As shown, the collaborative driving platform mainly consists of four key modules: perception generation module, scheduling module, update module, and planning module.

[0125] In one specific implementation of this embodiment, the core objective of the perception generation module is to provide a comprehensive and accurate representation of the driving environment for connected vehicles (CAVs). The BEV image generation module generates a bird's-eye view (BEV) image by fusing multi-sensor data of the environment. The multi-sensor data includes data generated by the surround-view cameras mounted on the vehicle body, which can be converted into a BEV image via the BEV Transformer. Furthermore, this embodiment can directly obtain vehicle position relationships through vehicle-to-vehicle (V2V) communication to generate a global (the entire urban road network) BEV image, enabling the LMM to process and understand the image information. Assume the data set acquired by the sensors is D = {d1, d2, ..., d...} n}, where d i Let I represent the data from the i-th sensor. This data is then transformed into a BEV image I using a specific fusion algorithm F. BEV , that is I BEV =F(D). BEV images can provide an overall overview of road networks, vehicle locations, and passenger activity, providing crucial spatial information for subsequent decision-making from a top-down perspective.

[0126] To ensure the quality and consistency of the images used in the playback buffer, the module employs a complex transform chain. First, the generated BEV image is normalized. Let image I... BEV The pixel value in is p ij If the mean of the dataset is and the standard deviation is σ, then the normalized pixel values ​​are... It can be represented as: This normalization process helps maintain data integrity in different scenarios.

[0127] The normalized image is processed using advanced embedding techniques, such as Vision Transformers (ViT) or Swin Transformers. Taking ViT as an example, it segments the image into multiple non-overlapping image blocks P = {p1, p2, ..., p...}. m Each image patch is mapped to a fixed-length embedding vector e. i Feature extraction is performed using the Transformer architecture to obtain the final feature representation E = {e1, e2, ..., e}. m The extracted features are implicit features learned by ViT itself, indicating the uniqueness of each image. Subsequent scheduling modules can use these extracted feature representations to calculate the similarity between different images, obtaining historical data similar to the current vehicles to be scheduled, thus aiding in guiding current scheduling decisions. These features provide rich information for subsequent analysis and decision-making.

[0128] In addition to BEV images, the AMoD simulation platform also supports the generation of multimodal sensing data, including point cloud data P and multi-view RGB camera images. These data originate from the digital assets of the underlying autonomous driving simulator (CARLA and Unreal Engine 4) and are rendered using computer graphics. This data can be seamlessly integrated into the perception module of a single vehicle for testing the collaborative perception capabilities of CAVs. Target traffic element information can be extracted by applying advanced instance segmentation techniques to multi-view RGB camera images. Let the segmentation function be S, then the extracted traffic element information... Then, the point cloud and camera image are aligned through coordinate frame transformation to ensure accurate positioning of traffic elements.

[0129] Finally, the acquired location data is uploaded to the cloud map database for the generation or updating of high-definition maps. Location data refers to the vehicle's position information when collecting current multimodal perception data. Geographic location information (local coordinates based on GPS or locally maintained BEV maps) helps in the maintenance and updating of the global map.

[0130] In one specific implementation of this embodiment, the main task of the scheduling module is to optimize the allocation of available vehicles to passenger requests, thereby improving service efficiency and reducing passenger waiting time. The scheduling module further includes the following steps:

[0131] 1. Data Collection: First, collect real-time status data of vehicles and passengers, including geographical location L={l v1 ,l v2 ,…,l vn ,l p1 ,l p2 ,…,l pm} and the current other states S = {s v1 ,s v2 ,…,s vn ,s p1 ,s p2 ,…,s pm}, where l vi and s vi Let l represent the position and state of the i-th vehicle, respectively. pj and s pj Let represent the position and state of the j-th passenger, respectively.

[0132] 2. BEV Image Generation: Generate BEV images using the collected data. The system's current operational status is displayed intuitively, highlighting the locations of idle vehicles, occupied vehicles, and pending passenger requests. Specifically, data generated by the vehicle's surround-view cameras is used, and the images are converted into BEV images using a BEV Transformer. Furthermore, the method currently used in this invention directly obtains vehicle location relationships via V2V, generating a global (entire city's passable road network) BEV image for LMM to process and understand the image information. The BEV image with sched is specifically optimized for vehicle scheduling problems, showing the direction arrow for each vehicle with an ID; the BEV image without sched is a generalized BEV image, containing the sched BEV.

[0133] 3. Vehicle-Passenger Pairing Decision: Based on the generated BEV images and text data, a specified algorithm is used to make vehicle-passenger pairing decisions. The text data can be obtained through V2V and includes the vehicle ID, passenger ID, waiting time for each passenger, and the distance matrix between vehicles and passengers. This text data can be customized according to different vehicle scheduling objectives. Taking an LMM-based method as an example, let the decision function be D, then the pairing result... This decision-making process comprehensively considers multiple factors, such as proximity, driving direction, and road network conditions. The final scheduling decision and visualization results will be stored by the system for subsequent analysis and reference. The decision function refers to the macroscopic decision function, which uses LMM (Leveled Model) for Chain of Thought (CoT) reasoning to obtain the decision result (pairing of vehicle IDs with passenger IDs). In addition, heuristic decision functions such as first-come-first-served and proximity principles can be used for vehicle scheduling and pairing.

[0134] In one specific implementation of this embodiment, the update module is responsible for real-time updates of the system state, ensuring that the system can accurately reflect changes in the dynamic environment. This is achieved by updating the states of vehicles and pedestrians. This module first collects the latest state information S of all entities (vehicles, passengers) within the system. new ={s new1 ,s new2 ,…,s newq Then, based on the vehicle's current state and the passenger's destination information, the state is updated. State information refers to the vehicle's current location, the passenger's current location, and the waiting time. Let the state update function be U, then the updated state is S. updated =U(S) new E env ), where E env This indicates the latest environmental data (obtained via the simulation environment's API). After updating the status, a new BEV image is generated. To reflect these changes and ensure the relevance and informativeness of the visual context, updated states and corresponding images will be archived to support continuous analysis and optimization. (Finite state machine: After a passenger arrives at their destination, the corresponding vehicle is set to idle, and the passenger is set to arrived. After a vehicle receives a dispatch task, it is set to hired. After a passenger is picked up by a vehicle, the passenger is set to picked).

[0135] In one specific implementation of this embodiment, the planning module is responsible for developing and evaluating potential driving strategies to achieve cooperative driving behavior among autonomous vehicles. First, it collects the current state data of all CAVs and passengers waiting to be picked up, including information such as location, speed, and driving direction, denoted as D. plan ={d plan1 ,d plan2 ,…,d planr Based on the collected data, and considering factors such as vehicle dynamics, traffic flow, road conditions, and mutual avoidance, a potential driving trajectory T = {t1, t2, ..., t} is generated. g Let the trajectory generating function be G, then T = G(D). plan This platform supports algorithms such as A* algorithm, Jump Point Search (JPS), and Iterative Deepening A* (IDA) for dynamic updating of the global trajectory. Secondly, after obtaining the global trajectory, methods such as Model Predictive Control (MPC), Proximal Policy Optimization (PPO), and Finite State Machine (FSM) are used to perform a comprehensive risk assessment of the generated trajectory. Taking MPC as an example, let the risk assessment function be R, then the risk assessment result R0... result =R(T), the risk assessment function can be implemented using LMM. Subsequent driving strategies utilize MPC to perform local planning and execution on the generated global trajectory, ensuring no collisions occur during multi-vehicle driving. Through assessment, potential collision hazards or inefficient situations are identified. Based on the risk assessment results, the optimal driving strategy S, balancing safety and efficiency, is selected. opt Driving strategy refers to the local reference trajectory generated by MPC. Balancing safety and efficiency is achieved through the weights of different factors in the MPC's objective function and the safety constraints in the constraints. This balance is achieved by solving an optimization problem. Let the strategy selection function be C, then S... opt =C(R) result Specifically, this can be achieved through the weights of different factors (trajectory tracking, energy consumption optimization, comfort evaluation, etc.) in the MPC objective function and the safety constraints in the constraints. At the end of each decision-making and vehicle control step, the results are visualized and recorded for further analysis and continuous improvement of the driving strategy.

[0136] Through the coordinated operation of these four modules, the AMoD platform enables efficient, safe, and intelligent autonomous vehicle operation, promoting the development of intelligent transportation systems.

[0137] This application embodiment dynamically matches real-time vehicle status with passenger demand information, constructs a global bird's-eye view using multimodal environmental perception data, and generates safe paths using a collaborative optimization algorithm. This effectively solves the problems of insufficient utilization of multi-source data and high path planning conflict rate in traditional scheduling systems, improving connection efficiency and reducing collision risks. Furthermore, this application embodiment also develops a collaborative driving autonomous on-demand mobility platform that supports visual content context generation. It is equipped with a series of diverse development tools. These tools include a multimodal context generation function, which integrates various data such as traffic flow, road conditions, and weather information to construct a comprehensive driving context. The passenger request allocation function optimizes task allocation for intelligent connected vehicles based on multiple factors. Regarding motion planning, the platform's built-in motion planning function for intelligent connected vehicles enables real-time route generation. For example, in simulation tests, compared to traditional non-integrated motion planning systems, in complex urban traffic scenarios, the proposed AMoD platform can simultaneously process the trajectory generation of multiple intelligent connected vehicles, taking into account spatiotemporal obstacle avoidance between them. This is achieved by continuously adjusting routes based on real-time traffic conditions, such as avoiding congested sections detected by traffic sensors. A replay buffer for data analysis allows for in-depth study of past driving data. This helps in the continuous optimization and improvement of motion planning algorithms. From an economic perspective, with the help of this platform, researchers can optimize the task allocation of connected vehicles based on distance, waiting time, and passenger preferences. This optimization makes connected vehicles more efficient. For example, in a pilot project in a medium-sized city, the platform reduced the operating costs of a fleet of 100 connected vehicles. The cost reduction was mainly due to reduced empty trips and lower energy consumption due to better route planning. The platform also has the potential to increase the utilization rate of connected vehicles, thereby increasing revenue for transportation service providers. At the social level, the platform plays a crucial role. By generating real-time driving routes for connected vehicles that take into account traffic conditions and individual requests, it can improve the overall travel experience for passengers. In addition, it strongly supports cooperative motion planning and control, which can improve traffic safety. In simulation experiments, the platform strives to seamlessly integrate into the dynamic urban environment, providing a comprehensive platform for possible solutions for future shared transportation systems, thereby contributing to the creation of a more sustainable and convenient urban transportation ecosystem.

[0138] Please see Figure 3This application also provides a path planning device 400, applied to a cooperative driving platform, which can implement the above-mentioned path planning method. The path planning device 400 includes:

[0139] The acquisition module 10 is used to acquire vehicle status information collected by multiple vehicles within the coverage area of ​​the collaborative driving platform and passenger status information of multiple passengers waiting to use the vehicle. The vehicle status information includes the current location of each vehicle and the passenger status information includes the starting location and destination of each passenger.

[0140] The pairing module 20 is used to pair multiple vehicles with multiple passengers based on the vehicle status information and the passenger status information, and obtain multiple pairing results, each pairing result including one vehicle and one passenger;

[0141] The planning module 30 is used to perform path planning for each pairing result based on the current position of the vehicle and the starting position of the passenger in the pairing result, so as to obtain a first target path for the vehicle to travel from the current position to the starting position of the passenger in the pairing result.

[0142] The specific implementation of this path planning device is basically the same as the specific implementation of the path planning method described above, and will not be repeated here.

[0143] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described path planning method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0144] Please see Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0145] The processor 801 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0146] The memory 802 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 802 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 using the path planning method of the embodiments of this application.

[0147] The 803 input / output interface is used to implement information input and output.

[0148] The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0149] Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804);

[0150] The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.

[0151] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described path planning method.

[0152] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0153] The path planning method, path planning device, electronic device, and storage medium provided in this application embodiment first acquire vehicle status information of multiple vehicles within the coverage area and passenger status information of multiple passengers waiting to use the vehicle. The vehicle status information includes the current location of each vehicle, and the passenger status information includes the starting location and destination of each passenger. This information is uploaded to the platform in real time through in-vehicle devices and passenger mobile terminals. Based on the acquired vehicle status information and passenger status information, the platform pairs multiple vehicles with multiple passengers. The pairing process considers factors such as the distance between the current location of the vehicle and the starting location of the passenger, and the direction of vehicle travel. An algorithm is used to obtain the optimal vehicle-passenger pairing result. Each pairing result contains a correspondence between one vehicle and one passenger. For each pairing result, the platform performs path planning based on the current location of the paired vehicle and the starting location of the passenger. The planning process considers factors such as real-time road conditions and traffic rules to generate the optimal driving path from the current location of the vehicle to the starting location of the passenger, i.e., the first target path. This application effectively solves the technical problem of low scheduling efficiency of existing autonomous on-demand mobility systems for intelligent connected vehicles by dynamically matching real-time vehicle status with passenger demand information, constructing a global bird's-eye view by combining multimodal environmental perception data, and generating safe paths using a collaborative optimization algorithm. This optimizes the path planning of autonomous on-demand mobility systems for intelligent connected vehicles and improves scheduling efficiency.

[0154] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0155] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0157] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0158] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0159] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0160] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0161] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0162] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0163] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0164] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A path planning method, characterized in that, Applied to a collaborative driving platform, the method includes: The system acquires vehicle status information collected from multiple vehicles within the coverage area of ​​the collaborative driving platform, as well as passenger status information from multiple passengers waiting to use the vehicle. The vehicle status information includes the current location of each vehicle, and the passenger status information includes the starting location and destination of each passenger. Based on the vehicle status information and the passenger status information, multiple vehicles are paired with multiple passengers to obtain multiple pairing results, each pairing result including one vehicle and one passenger; For each pairing result, path planning is performed based on the current position of the vehicle and the starting position of the passenger in the pairing result to obtain the first target path for the vehicle to travel from the current position to the starting position of the passenger in the pairing result.

2. The method according to claim 1, characterized in that, The vehicle status information also includes environmental image information collected for each vehicle, passenger status of each vehicle, and driving status of each vehicle. The step of matching multiple vehicles with multiple passengers based on the vehicle status information and the passenger status information to obtain multiple matching results includes: The environmental image information collected from each vehicle is fused to obtain a bird's-eye view of the coverage area of ​​the collaborative driving platform; Assign an identifier to each of the vehicles and each of the passengers; Based on the vehicle status information and the passenger status information, mark the passenger-carrying status of each vehicle, the driving status of each vehicle, the identification of each vehicle, the starting position of each passenger, and the identification of each passenger on the bird's-eye view. A large-scale multimodal model is used to pair the identifiers corresponding to multiple vehicles with the identifiers corresponding to multiple passengers based on the bird's-eye view, resulting in multiple pairing results.

3. The method according to claim 2, characterized in that, The process involves using a large-scale multimodal model to pair the identifiers corresponding to multiple vehicles with the identifiers corresponding to multiple passengers based on the bird's-eye view, resulting in multiple pairing results, including: By using a large-scale multimodal model to infer the distance between each vehicle and each passenger, the driving direction of each vehicle, and the road layout in the bird's-eye view, matching results are obtained for the identifiers corresponding to multiple vehicles and the identifiers corresponding to multiple passengers. Based on the matching results of the identifiers corresponding to the multiple vehicles and the identifiers corresponding to the multiple passengers, a pairing result is obtained for each passenger and the corresponding vehicle.

4. The method according to claim 3, characterized in that, The passenger status information also includes the waiting time for each passenger; The method involves using a large-scale multimodal model to infer the distance between each vehicle and each passenger, the driving direction of each vehicle, and the road layout in the bird's-eye view, to obtain matching results between the identifiers corresponding to multiple vehicles and the identifiers corresponding to multiple passengers, including: The distance matrix between each vehicle and each passenger is obtained by using the large multimodal model to determine the relative positions of each vehicle and each passenger. The matching degree between the driving direction of each vehicle and the starting position of each passenger is obtained through the large multimodal model. The road layout is analyzed using the large-scale multimodal model to obtain the estimated time for each vehicle to reach the starting position of each passenger. The waiting time of each passenger is evaluated using the large-scale multimodal model to obtain a matching priority order; The matching result is obtained based on the distance matrix, the matching degree, the estimated time, and the priority order.

5. The method according to claim 1, characterized in that, For each pairing result, path planning is performed based on the vehicle's current position and the passenger's starting position in the pairing result to obtain a first target path for the vehicle to travel from its current position to the passenger's starting position, including: For each pairing result, path planning is performed based on the current position of the vehicle and the starting position of the passenger in the pairing result, and the first driving trajectory corresponding to each pairing result is obtained through the trajectory generation function; Based on the multiple first driving trajectories corresponding to the multiple pairing results, a first global trajectory is obtained; A risk assessment is performed on the first global trajectory using a large-scale multimodal model to obtain a first risk assessment result, which is used to indicate the risk of a vehicle collision in the first global trajectory. Based on the first risk assessment result, the first global trajectory is adjusted by a model predictive control algorithm to obtain the first target path corresponding to each pairing result.

6. The method according to claim 5, characterized in that, The step of adjusting the first global trajectory based on the first risk assessment result using a model predictive control algorithm to obtain the first target path corresponding to each pairing result includes: The first risk assessment result is input into the model predictive control algorithm to generate a local reference trajectory; Based on the local reference trajectory, each of the first driving trajectories in the first global trajectory is dynamically adjusted to obtain the first target path corresponding to each pairing result.

7. The method according to claim 1, characterized in that, After performing path planning for each pairing result based on the vehicle's current position and the passenger's starting position in the pairing result to obtain a first target path for the vehicle to travel from its current position to the passenger's starting position, the method further includes: For each pairing result, a route is planned based on the passenger's starting position and destination in the pairing result, and a second driving trajectory corresponding to each pairing result is obtained through a trajectory generation function; Based on the multiple second driving trajectories corresponding to the multiple pairing results, a second global trajectory is obtained; The risk assessment of the second global trajectory is performed using a large-scale multimodal model to obtain a second risk assessment result, which is used to indicate the risk of vehicle collision in the second global trajectory. Based on the second risk assessment result, the second global trajectory is adjusted by a model predictive control algorithm to obtain a second target path for the vehicle in each pairing result, from the passenger's starting position to the passenger's destination.

8. A path planning device, characterized in that, The device, applied to a collaborative driving platform, includes: The acquisition module is used to acquire vehicle status information collected by multiple vehicles within the coverage area of ​​the collaborative driving platform and passenger status information of multiple passengers waiting to use the vehicle. The vehicle status information includes the current location of each vehicle, and the passenger status information includes the starting location and destination of each passenger. The matching module is used to match multiple vehicles with multiple passengers based on the vehicle status information and the passenger status information, and obtain multiple matching results, each of the matching results including one vehicle and one passenger; The planning module is used to perform path planning for each pairing result based on the current position of the vehicle and the starting position of the passenger in the pairing result, so as to obtain a first target path for the vehicle to travel from the current position to the starting position of the passenger in the pairing result.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the path planning method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the path planning method according to any one of claims 1 to 7.