Intelligent automatic driving vehicle scheduling system based on multi-mode perception and dynamic environment modeling

By designing an intelligent autonomous vehicle scheduling system based on multimodal perception and dynamic environment modeling, the problem of existing systems lacking multimodal perception and dynamic environment scheduling capabilities in high-density traffic or large-scale activity areas is solved, and higher driving safety and efficiency are achieved.

CN120071616AActive Publication Date: 2025-05-30GUANGZHOU SMART BODY TECH CO LTD

Patent Information

Application Number
CN202510219116.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing autonomous vehicle systems lack the ability to schedule multimodal perception and dynamic environment in high-density traffic or large-scale activity areas, resulting in problems such as path conflict, inefficiency and traffic congestion.

Method used

An intelligent autonomous driving vehicle scheduling system based on multimodal perception and dynamic environment modeling is designed, including a situational perception module, feature extraction and fusion module, a semantic and action decision module, an adaptive cluster transfer module, a trajectory generation and optimization module and an embodied intelligent learning module. Through technical means such as multimodal Transformer model and Occupancy Network network, vehicle scheduling optimization is achieved.

Benefits of technology

Through multimodal perception and dynamic environment modeling, the system significantly improves the driving safety and efficiency of autonomous driving vehicles in complex scenarios, enhances adaptability and scheduling capabilities, and solves the problems of path conflicts and inefficiency of existing systems in high-density traffic or large-scale activity areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071616A_ABST
    Figure CN120071616A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent automatic driving vehicle scheduling system based on multi-mode perception and dynamic environment modeling, and belongs to the technical field of automatic driving. The system comprises a situation awareness module, a feature extraction and fusion module, a semantic and action decision module, a self-adaptive cluster pick-up and drop-off module, a track generation and optimization module and a self intelligent learning module. The automatic driving technology of the system relates to intelligent, VLA multi-mode architecture and the like, and an end-to-end data processing mode is realized on the basis, so that an automatic driving automobile can quickly generate a corresponding strategy according to input and make a response according to the strategy; and the driving of the automatic driving vehicle which realizes environmental understanding, intelligent decision making and accurate control in a complex scene is safer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and more particularly, to an intelligent autonomous driving vehicle scheduling system based on multi-modal perception and dynamic environment modeling. Background Art

[0002] As one of the important applications in the field of artificial intelligence, autonomous driving technology is gradually entering our lives and becoming an important development direction for future transportation. It plays an important role in improving traffic safety, optimizing traffic efficiency, and enriching travel mode options. At present, autonomous driving technology has been diversely applied in different fields such as public transportation, taxis, logistics and distribution, and urban infrastructure.

[0003] The basic process of autonomous driving is divided into three parts: perception, decision-making, and control. By integrating the data of various sensors through the perception system, driving plans are made by making decisions on the information output by the perception layer with the help of different algorithms and supporting software, and finally the control system completes the control behavior of the vehicle. Currently, there are two mainstream technical routes: one is a multi-sensor fusion scheme dominated by cameras; the other is a technical scheme dominated by lidar with other sensors as auxiliary.

[0004] The current autonomous driving vehicle system usually performs single perception and dynamic environment scheduling and pick-up / drop-off tasks, and rarely considers the coordination problem between multiple perceptions and dynamic environments. For high-density traffic or large-scale activity areas, it faces challenges such as path conflicts, low efficiency, and traffic jams. Existing systems lack the scheduling ability of multi-modal perception and dynamic environment, and lack the driving scheduling decision-making of autonomous driving vehicles with environmental understanding, intelligent decision-making, and precise control in complex scenarios. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above problems existing in the prior art and greatly improve its technical effect on the basis of the original technology; the present invention provides an intelligent autonomous driving vehicle scheduling system based on multi-modal perception and dynamic environment modeling, which system includes: a situation perception module, a feature extraction and fusion module, a semantic and action decision module, an adaptive cluster pick-up / drop-off module, a trajectory generation and optimization module, and an embodied intelligence learning module;

[0006] The situation perception module is used to perceive the internal and external situation data of the autonomous driving vehicle in real time, and the internal and external situation data includes passenger needs, traffic conditions, road conditions, weather data, and vehicle status;

[0007] The feature extraction and fusion module is used to extract features from the internal and external situation data, fuse the extracted features, and generate environmental features, and the environmental features at least include BEV space features;

[0008] The semantic and action decision-making module is used to generate intelligent behavior decisions by combining the environmental features in the feature extraction and fusion module through the VLA large model;

[0009] The adaptive cluster pick-up and drop-off module is used to convert the final trajectory and intelligent behavior decisions into vehicle scheduling;

[0010] The trajectory generation and optimization module is used to construct the Occ model of the environmental occupancy grid in real time based on the Occupancy Network network. The Occ model predicts and updates the free and occupied areas in the environment by combining the BEV space features in the environmental features, and uses self-supervised learning methods to optimize the path planning and vehicle scheduling planning through the free and occupied areas. At the same time, reinforcement learning is introduced to optimize the real-time vehicle scheduling generated by the adaptive cluster pick-up and drop-off module, and the optimization results are input into the adaptive cluster pick-up and drop-off module in real time for real-time vehicle scheduling;

[0011] The embodied intelligence learning module is used to realize real-time adjustment of the path planning and vehicle scheduling planning of the autonomous driving vehicle through the interaction between the autonomous driving vehicle and the environment, and learn according to the received real-time environmental feedback, the environmental features, and the optimization results of the trajectory generation and optimization module, and input the learning results into the adaptive cluster pick-up and drop-off module in real time for real-time vehicle scheduling.

[0012] Furthermore, the situation awareness module includes a passenger demand awareness sub-module, a traffic dynamics awareness sub-module, and a vehicle status awareness sub-module;

[0013] The passenger demand awareness sub-module is used to obtain the pick-up and drop-off requests, preferences, and expected arrival times of passengers in real time, and generate passenger demand awareness data by combining the historical behavior data of passengers;

[0014] The traffic dynamics awareness sub-module is used to obtain the traffic flow in real time through traffic sensors, GPS maps, and road cameras, and generate traffic dynamics awareness data; the traffic flow is used to analyze congestion or temporary road closure situations;

[0015] The vehicle status awareness sub-module is used to collect the power data, passenger number data, and current pick-up and drop-off task progress data of each vehicle, and generate vehicle status awareness data.

[0016] Furthermore, the feature extraction and fusion module includes a feature extraction sub-module and a fusion sub-module;

[0017] The feature extraction sub-module is used to extract features from the passenger demand awareness data, traffic dynamics awareness data, and vehicle status awareness data respectively, and generate passenger demand awareness feature data, traffic dynamics awareness feature data, and vehicle status awareness feature data;

[0018] The fusion sub-module is used to fuse the passenger demand perception feature data, traffic dynamic perception feature data, and vehicle state perception feature data by using a multi-modal Transformer model.

[0019] Furthermore, the specific implementation process of the fusion sub-module includes:

[0020] Using a multi-modal Transformer model, the corresponding feature data of the passenger demand perception feature data, traffic dynamic perception feature data, and vehicle state perception feature data are fused through the self-attention mechanism in the multi-modal Transformer model. Among them, the image features in the passenger demand perception feature data, traffic dynamic perception feature data, and vehicle state perception feature data are all based on the bird's-eye view BEV combined with the Transformer model for image feature extraction to generate BEV space features;

[0021] The formula of the self-attention mechanism of the multi-modal Transformer model is:

[0022]

[0023] Among them, Q, K, and V represent the query matrix, key matrix, and value matrix respectively, and d k represents the dimension of the key vector;

[0024] The formula for image feature extraction based on generating a bird's-eye view BEV combined with the Transformer model is:

[0025] P bev (X t ) = Transformer bev (X t ),

[0026] Among them, X t represents the image sequence features in the passenger demand perception feature data, traffic dynamic perception feature data, and vehicle state perception feature data, and P bev (X t ) represents the output bird's-eye view.

[0027] Further, the semantic and action decision-making module includes: obtaining intelligent behavior decisions including target task scheduling, control instructions, and target task scheduling execution strategies by combining the environmental features in the feature extraction and fusion module with the VLA large model; the steps are as follows: First, perform data fusion analysis and processing on the environmental features through the VLA large model to analyze the passengers' needs and the riding environment; Second, combine the current road condition information and the vehicle's own state information obtained by the traffic dynamic perception sub-module and the vehicle state perception sub-module to initially generate target task scheduling, and generate corresponding control instructions and target task scheduling execution strategies for the vehicle according to the generated target task scheduling.

[0028] Further, the embodied intelligent learning module includes: real-time perception and decision-making, environmental feedback mechanism, and reinforcement learning and self-supervised learning; the real-time perception and decision-making includes: obtaining the surrounding environment information of the autonomous vehicle in real time through the perception system, and making decisions in a timely manner based on the obtained information; the perception system integrates multiple sensors, and obtains the surrounding environment data through the sensors, and establishes an environmental model by fusing and processing the sensor data. The established model formula is:

[0029] H t =F p (S t ,v p )

[0030] Wherein, H t is the environmental state at time t, and the environmental state is the internal and external context data; F p is the environmental perception function, indicating how to convert the sensor data into the environmental state; S t is the data collected by the sensor at time t, and v p is the parameter of the perception model;

[0031] The environmental feedback mechanism includes: obtaining feedback through interaction with the environment, and adjusting the scheduling task plan according to the feedback, and then adjusting the scheduling task; First, the system generates corresponding decisions through the dynamic environmental model established by the context perception module, and interacts with the targets in the corresponding dynamic environment; Subsequently, after the interaction, after receiving the feedback from the targets in the dynamic environment, the system adjusts the scheduling task plan according to the feedback results, and decides the next decision of the vehicle; Finally, according to the re-adjusted scheduling task plan, adjust the scheduling task plan, and input the final result into the adaptive cluster pick-up and drop-off module in real time for real-time vehicle scheduling;

[0032] The reinforcement learning and self-supervised learning include: optimizing scheduling tasks through a reinforcement learning framework; optimizing the scheduling action value at the current moment based on the scheduling action value at the previous moment; through continuous updates, and learning and then preferentially selecting according to the received real-time environment feedback, the environmental features, and the optimization results of the trajectory generation and optimization module, and inputting the final result into the adaptive cluster pick-up and drop-off module in real time for real-time vehicle scheduling.

[0033] Further, the trajectory generation and optimization module includes: Occ model construction and prediction and self-supervised learning;

[0034] The Occ model construction and prediction include: based on the Occupancy Network, constructing an Occ model of the environmental occupancy grid in real time. The Occ model predicts and updates the free and occupied areas in the environment by combining the BEV space features in the environmental features. The training formula of the Occupancy Network:

[0035] P occ (s t ) = OccupancyNetwork(P bev (X t ))), where P occ (s t ) represents the state of the occupancy grid at time t, and P bev (X t ) represents the bird's-eye view output by the occupancy grid at time t;

[0036] The loss function of the Occupancy Network is:

[0037] η(Occ t ) = ||Occ t - Occ' t || 2 ,

[0038] where Occ t is the environmental state at the current moment, Occ' t is the environmental state predicted by the system, || || is the norm, and η() is the function that minimizes the loss; the Occupancy Network adaptively learns and optimizes the path planning from the environment in this way;

[0039] The self-supervised learning includes using self-supervised learning methods to optimize path planning and vehicle scheduling planning through free and occupied areas, introducing reinforcement learning to optimize vehicle scheduling at the same time, and inputting the optimization results into the adaptive cluster pick-up and drop-off module in real time for real-time vehicle scheduling.

[0040] Furthermore, the semantic and action decision-making module, the adaptive cluster pick-up and drop-off module, the trajectory generation and optimization module, and the embodied intelligence learning module are designed as an end-to-end automatic decision-making module. After the data of the surroundings of the vehicle and the passenger information collected by the sensors are input into the automatic decision-making module, the automatic decision-making module directly generates a series of driving decisions, and the series of driving decisions include: task planning, path planning, control instructions, task scheduling, and scheduling optimization.

[0041] The beneficial effects of the present invention are:

[0042] The present invention provides an intelligent autonomous driving vehicle scheduling system based on multi-modal perception and dynamic environment modeling, which has the following advantages:

[0043] 1. The present invention extracts features from internal and external context data through the self-attention mechanism of the multi-modal Transformer model. Compared with the individual processing of internal and external context data, the obtained internal and external context data will be more accurate.

[0044] 2. The present invention provides an end-to-end optimized autonomous driving pick-up and drop-off scheduling system that combines embodied intelligence, the VLA multi-modal architecture, from environmental perception to task execution. The adaptive ability and scheduling ability will be greatly enhanced, making the driving of autonomous vehicles safer.

[0045] 3. At the same time, self-supervised learning and the Occupancy Network are used to optimize the path optimization planning and vehicle scheduling optimization planning of autonomous vehicles, improving the safety of autonomous vehicles. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a schematic diagram of an intelligent autonomous driving vehicle scheduling system based on multi-modal perception and dynamic environment modeling of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0047] The following describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments given here are only for illustrating and explaining the present invention and cannot be used to limit the present invention.

[0048] It should be noted that many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention may have other embodiments and variations, and therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0049] Such as Figure 1As shown, a schematic diagram of an intelligent autonomous driving vehicle scheduling system based on multi-modal perception and dynamic environment modeling according to an embodiment of the present invention; the schematic diagram includes: a situation awareness module A1, a feature extraction and fusion module A2, a semantic and action decision-making module A3, an adaptive cluster pick-up and drop-off module A4, a trajectory generation and optimization module A5, and an embodied intelligence learning module A6;

[0050] The situation awareness module A1 is used to perceive the internal and external situation data of the autonomous driving vehicle in real time, and the internal and external situation data includes passenger demand, traffic conditions, road conditions, weather data, and vehicle status;

[0051] The feature extraction and fusion module A2 is used to extract features from the internal and external situation data, fuse the extracted features, and generate environmental features, and the environmental features at least include BEV space features;

[0052] The semantic and action decision-making module A3 is used to generate intelligent behavior decisions by combining the environmental features in the feature extraction and fusion module A2 through the VLA large model;

[0053] The adaptive cluster pick-up and drop-off module A4 is used to convert the final trajectory and intelligent behavior decisions into vehicle scheduling;

[0054] The trajectory generation and optimization module A5 is used to construct an Occ model of the environmental occupancy grid in real time based on the Occupancy Network network. The Occ model predicts and updates the free and occupied areas in the environment by combining the BEV space features in the environmental features, and uses a self-supervised learning method to optimize the path planning and vehicle scheduling planning through the free and occupied areas. At the same time, reinforcement learning is introduced to optimize the real-time vehicle scheduling generated by the adaptive cluster pick-up and drop-off module, and the optimization results are input into the adaptive cluster pick-up and drop-off module A4 in real time for real-time vehicle scheduling;

[0055] The embodied intelligence learning module A6 is used to realize real-time adjustment of the path planning and vehicle scheduling planning of the autonomous driving vehicle through the interaction between the autonomous driving vehicle and the environment, and learn according to the received real-time environmental feedback, the environmental features, and the optimization results of the trajectory generation and optimization module A5, and input the learning results into the adaptive cluster pick-up and drop-off module A4 in real time for real-time vehicle scheduling.

[0056] In the above embodiment, specifically, the situation awareness module A1 includes a passenger demand perception sub-module, a traffic dynamic perception sub-module, and a vehicle status perception sub-module;

[0057] The passenger demand perception sub-module is used to obtain the pick-up and drop-off requests, preferences, and estimated arrival times of passengers in real time, and generate passenger demand perception data by combining the historical behavior data of passengers;

[0058] The traffic dynamic perception sub-module is used to obtain traffic flow in real time through traffic sensors, GPS maps, and road cameras, and generate traffic dynamic perception data; the traffic flow is used to analyze congestion situations or temporary road closures;

[0059] The vehicle status perception sub-module is used to collect power data, passenger number data, and the progress data of the current pick-up and drop-off tasks for each vehicle, and generate vehicle status perception data.

[0060] In the above embodiment, specifically, the feature extraction and fusion module A2 includes a feature extraction sub-module and a fusion sub-module;

[0061] The feature extraction sub-module is used to perform feature extraction on passenger demand perception data, traffic dynamic perception data, and vehicle status perception data respectively, and generate passenger demand perception feature data, traffic dynamic perception feature data, and vehicle status perception feature data;

[0062] The fusion sub-module is used to use a multi-modal Transformer model to fuse passenger demand perception feature data, traffic dynamic perception feature data, and vehicle status perception feature data;

[0063] Specifically, the specific implementation process of feature extraction of passenger demand perception data is as follows:

[0064] Obtain passenger pick-up and drop-off requests, passenger preferences, and passenger historical behavior data through the passenger's mobile device. The passenger historical behavior data at least includes the passenger's historical travel patterns, and the passenger historical behavior data is stored in the InfluxDB database;

[0065] Extract features from the passenger pick-up and drop-off requests, passenger preferences, and passenger historical behavior data through a temporal convolutional network (TCN) model;

[0066] Specifically, the specific implementation process of feature extraction of traffic dynamic perception data is as follows:

[0067] The traffic dynamic perception sub-module includes a model construction unit and a model analysis unit;

[0068] The specific implementation process of the model construction unit is as follows:

[0069] Use a graph neural network (GNN) to model the traffic flow, generate a graph neural network (GNN) model. The graph neural network (GNN) model represents the traffic road network as a graph structure, where nodes represent road segments and edges represent traffic flow. The graph neural network (GNN) model propagates information through graph convolution operations. The calculation formula of the graph neural network (GNN) model is:

[0070]

[0071] Among them, represents the hidden state of node i at the m-th layer, N(i) represents the neighbor set of node i, and W m represents the weight matrix of the m-th layer, δ represents the activation function, and d i , d j respectively represent the degrees of nodes i and j;

[0072] Specifically, the specific implementation process of the feature extraction of the graph neural network GNN model is as follows:

[0073] Combining the vehicle status and road condition information reported by the vehicle, the graph neural network GNN model analyzes the congestion situation or temporary road closure situation by calling the K-Means++ mean clustering algorithm through two indicators of the average speed of the road section and the load degree of the road section, and generates traffic dynamic perception feature data;

[0074] The specific implementation process of the feature extraction of the vehicle status perception data is as follows:

[0075] Each vehicle is embedded with an edge computing unit responsible for local computing and regularly uploading the internal state data of the vehicle. The internal state of the vehicle includes in-vehicle CAN bus data, power, and the current task progress;

[0076] The dynamic Bayesian network DBN model is used to extract features from the vehicle internal state data to generate vehicle status perception feature data.

[0077] In the above embodiment, specifically, the specific implementation process of the fusion sub-module includes:

[0078] Using a multi-modal Transformer model, the passenger demand perception feature data, traffic dynamic perception feature data, and vehicle status perception feature data are fused through the self-attention mechanism in the multi-modal Transformer model. Among them, the image features included in the passenger demand perception feature data, traffic dynamic perception feature data, and vehicle status perception feature data are all based on the bird's-eye view BEV combined with the Transformer model for image feature extraction to generate BEV space features;

[0079] The formula of the self-attention mechanism of the multi-modal Transformer model is:

[0080]

[0081] Among them, Q, K, and V respectively represent the query matrix, key matrix, and value matrix, and d k represents the dimension of the key vector;

[0082] The formula for image feature extraction based on generating the bird's-eye view BEV combined with the Transformer model is:

[0083] P bev (X t ) = Transformer bev (X t )

[0084] Among them, X t represents the image sequence feature among the passenger demand perception feature data, traffic dynamic perception feature data, and vehicle state perception feature data, and P bev (X t ) represents the output bird's-eye view.

[0085] In the above embodiment, specifically, the semantic and action decision-making module A3 includes: obtaining intelligent behavior decisions including target task scheduling, control instructions, and target task scheduling execution strategies by combining the environmental features in the feature extraction and fusion module through the VLA large model; the steps are as follows: First, perform data fusion analysis and processing on the environmental features through the VLA large model to analyze the passenger's needs and riding environment; secondly, combine the current road condition information and the vehicle's own state information obtained by the traffic dynamic perception sub-module and the vehicle state perception sub-module to initially generate a target task scheduling, and generate the corresponding control instructions and target task scheduling execution strategies for the vehicle according to the generated target task scheduling.

[0086] In the above embodiment, specifically, the embodied intelligent learning module A6 includes: real-time perception and decision-making, environmental feedback mechanism, and reinforcement learning and self-supervised learning; the real-time perception and decision-making includes: obtaining the surrounding environment information of the autonomous vehicle in real time through the perception system, and making decisions in a timely manner based on the obtained information; the perception system integrates a variety of sensors, and obtains the surrounding environment data through the sensors, and establishes an environmental model by fusing and processing the sensor data. The established model formula is:

[0087] H t = F p (S t , v p )

[0088] Among them, H t is the environmental state at time t, and the environmental state is the internal and external context data; F p is the environmental perception function, indicating how to convert the sensor data into the environmental state; S t is the data collected by the sensor at time t, and v p is the parameter of the perception model;

[0089] The environmental feedback mechanism includes: obtaining feedback through interaction with the environment, adjusting the scheduling task plan according to the feedback, and then adjusting the scheduling task. First, the system generates corresponding decisions through the dynamic environment model established by the situation awareness module A1 and interacts with the targets in the corresponding dynamic environment. Subsequently, after the interaction, when receiving the feedback from the targets in the dynamic environment, the system adjusts the scheduling task plan according to the feedback results and decides the next decision of the vehicle. Finally, according to the re-adjusted scheduling task plan, the scheduling task plan is adjusted, and the final result is input into the adaptive cluster pick-up and drop-off module A4 in real time for real-time vehicle scheduling.

[0090] The reinforcement learning and self-supervised learning include: optimizing the scheduling task through the reinforcement learning framework; optimizing the scheduling action value in the current state through the scheduling action value in the previous state; through continuous updating, and learning and then preferentially selecting according to the received real-time environmental feedback, the environmental characteristics, and the optimization results of the trajectory generation and optimization module A5, and inputting the final result into the adaptive cluster pick-up and drop-off module A4 in real time for real-time vehicle scheduling.

[0091] In the above embodiment, specifically, the trajectory generation and optimization module A5 includes: Occ model construction and prediction and self-supervised learning.

[0092] The Occ model construction and prediction include: based on the Occupancy Network network, constructing an Occ model of the environmental occupancy grid in real time. The Occ model predicts and updates the free and occupied areas in the environment by combining the BEV space features in the environmental characteristics. The training formula of the Occupancy Network network:

[0093] P occ (s t )=OccupancyNetwork(P bev (X t )),

[0094] where, P occ (s t ) represents the state of the occupancy grid at time (t), and P bev (X t ) represents the bird's-eye view output by the occupancy grid at time (t).

[0095] The loss function of the Occupancy Network network is:

[0096] η(Occ t )=||Occ t -Occ' t || 2 ,

[0097] where Occ t is the environmental state at the current moment, Occ' t is the environmental state predicted by the system, || || is the norm, and η() is the function that minimizes the loss; the Occupancy Network learns and optimizes the path planning adaptively from the environment in this way;

[0098] The self-supervised learning includes using self-supervised learning methods to optimize the path planning and vehicle scheduling planning through free and occupied areas, introducing reinforcement learning to optimize the vehicle scheduling, and inputting the optimization results into the adaptive cluster pick-up and drop-off module in real time for real-time vehicle scheduling.

[0099] In the above embodiment, specifically, the semantic and action decision-making module A3, the adaptive cluster pick-up and drop-off module A4, the trajectory generation and optimization module A5, and the embodied intelligent learning module A6 are designed as an end-to-end automatic decision-making module. After the data of the vehicle surroundings and passenger information collected by the sensor are input into the automatic decision-making module, the automatic decision-making module directly generates a series of driving decisions, and the series of driving decisions include: task planning, path planning, control instructions, task scheduling, and scheduling optimization.

[0100] It should be noted that the above end-to-end automatic decision-making module is constructed through the OpenPilot open-source autonomous driving technology. Through the data collection of the sensor and the output of the module, the decision to control the autonomous vehicle is directly generated, making the decision more rapid and accurate.

[0101] In the above embodiment, specifically, the adaptive cluster pick-up and drop-off module converts the final trajectory and behavior strategy into a vehicle scheduling plan, and the mathematical calculation formula is:

[0102]

[0103] Constraint conditions:

[0104]

[0105] where w ij represents the task assignment matrix, d ij represents the distance between the vehicle and the task, min() represents the final trajectory and behavior strategy optimization function, i and j respectively represent node i and node j, and N and M respectively represent the total numbers corresponding to i and j;

[0106] When multiple preferred trajectories are selected, the similarity degree between different driving paths is compared by using the trajectory similarity measurement method, so as to select the optimal driving path. The trajectory similarity measurement method in this embodiment includes Euclidean distance, dynamic time warping, etc.

Claims

1. An intelligent autonomous driving vehicle dispatching system based on multimodal perception and dynamic environment modeling, characterized in that: The system includes a situational awareness module, a feature extraction and fusion module, a semantic and action decision module, an adaptive cluster pick-up module, a trajectory generation and optimization module, and an embodied intelligent learning module; The situational awareness module is used to perceive the internal and external situational data of the autonomous driving vehicle in real time, wherein the internal and external situational data include passenger demand, traffic conditions, road conditions, weather data and vehicle status; The feature extraction and fusion module is used to extract features from the internal and external situational data, fuse the extracted features, and generate environmental features, wherein the environmental features at least include BEV space features; The semantic and action decision module is used to generate intelligent behavior decisions by combining the environmental features in the feature extraction and fusion module through the VLA large model; The adaptive cluster pick-up module is used to convert the final trajectory and intelligent behavior decision into vehicle scheduling; The trajectory generation and optimization module is used to construct an Occupancy Network model of the environment occupancy grid in real time. The Occupancy Network predicts and updates the idle and occupied areas in the environment by combining the BEV spatial features in the environmental features, and uses a self-supervised learning method to optimize the path planning and vehicle scheduling planning through the idle and occupied areas. At the same time, reinforcement learning is introduced to optimize the real-time vehicle scheduling generated by the adaptive cluster pick-up module, and the optimization results are input into the adaptive cluster pick-up module in real time for real-time vehicle scheduling; The embodied intelligent learning module is used to achieve real-time adjustment of the path planning and vehicle scheduling planning of the autonomous driving vehicle through the interaction between the autonomous driving vehicle and the environment, and to learn based on the received real-time environmental feedback, the environmental characteristics, and the optimization results of the trajectory generation and optimization module, and input the learning results into the adaptive cluster pick-up module in real time for real-time vehicle scheduling.

2. The intelligent autonomous driving vehicle dispatching system based on multimodal perception and dynamic environment modeling according to claim 1 is characterized in that: The situational awareness module includes a passenger demand awareness submodule, a traffic dynamics awareness submodule and a vehicle status awareness submodule; The passenger demand perception submodule is used to obtain the passenger's pick-up and drop-off requests, preferences, and estimated arrival time in real time, and generate passenger demand perception data in combination with the passenger's historical behavior data; The traffic dynamics perception submodule is used to obtain traffic flow in real time through traffic sensors, GPS maps and road cameras to generate traffic dynamics perception data; the traffic flow is used to analyze congestion or temporary road closures; The vehicle status perception submodule is used to collect the power data, passenger number data and current pick-up and drop-off task progress data of each vehicle to generate vehicle status perception data.

3. The intelligent automatic driving vehicle dispatching system based on multimodal perception and dynamic environment modeling according to claim 2 is characterized in that: The feature extraction and fusion module includes a feature extraction submodule and a fusion submodule; The feature extraction submodule is used to extract features from the passenger demand perception data, the traffic dynamics perception data and the vehicle state perception data respectively, and generate passenger demand perception feature data, traffic dynamics perception feature data and vehicle state perception feature data; The fusion submodule is used to fuse the passenger demand perception feature data, the traffic dynamics perception feature data and the vehicle status perception feature data using a multimodal Transformer model.

4. The intelligent automatic driving vehicle dispatching system based on multimodal perception and dynamic environment modeling according to claim 3 is characterized in that: The specific implementation process of the fusion submodule includes: A multimodal Transformer model is used to fuse passenger demand perception feature data, traffic dynamic perception feature data, and vehicle status perception feature data through the self-attention mechanism in the multimodal Transformer model. The image features contained in the passenger demand perception feature data, traffic dynamic perception feature data, and vehicle status perception feature data are extracted based on the bird's-eye view BEV combined with the Transformer model to generate BEV spatial features. The self-attention mechanism formula of the multimodal Transformer model is: Among them, Q, K, and V represent the query matrix, key matrix, and value matrix respectively. k represents the dimension of the key vector; The formula for image feature extraction based on generating a bird's-eye view BEV combined with the Transformer model is: P bev (X t )=Transformer bev (X t ), Among them, X t represents the image sequence features in the passenger demand perception feature data, traffic dynamics perception feature data, and vehicle state perception feature data, P bev (X t ) represents a bird’s-eye view of the output.

5. The intelligent automatic driving vehicle dispatching system based on multimodal perception and dynamic environment modeling according to claim 4 is characterized in that: The semantic and action decision module includes: combining the environmental features in the feature extraction and fusion module with the VLA large model to obtain intelligent behavior decisions including target task scheduling, control instructions and target task scheduling execution strategy; the steps are: first, performing data fusion analysis and processing on the environmental features through the VLA large model to analyze the passengers' needs and the riding environment; secondly, combining the current road condition information and the vehicle's own status information obtained by the traffic dynamic perception submodule and the vehicle status perception submodule to preliminarily generate the target task scheduling, and generate the control instructions and target task scheduling execution strategy corresponding to the vehicle according to the generated target task scheduling.

6. The intelligent automatic driving vehicle dispatching system based on multimodal perception and dynamic environment modeling according to claim 5 is characterized in that: The embodied intelligent learning module includes: real-time perception and decision-making, environmental feedback mechanism, and reinforcement learning and self-supervised learning; the real-time perception and decision-making includes: obtaining the surrounding environment information of the autonomous driving vehicle in real time through the perception system, and making timely decisions based on the obtained information; the perception system integrates a variety of sensors, and obtains the surrounding environment data through the sensors, and establishes an environmental model by fusing and processing the sensor data. The established model formula is: H t =F p (S t ,v p ) Among them, H t is the environmental state at time t, and the environmental state is the internal and external situation data; F p is the environment perception function, which indicates how to convert the sensor data into the environment state; S t is the data collected by the sensor at time t, v p are the parameters of the perception model; The environmental feedback mechanism includes: obtaining feedback through interaction with the environment, and adjusting the scheduling task plan according to the feedback, and then adjusting the scheduling task; first, the system generates corresponding decisions through the dynamic environment model established by the context perception module, and interacts with the target in the corresponding dynamic environment; then, after the interaction, after receiving the feedback from the target in the dynamic environment, the system adjusts the scheduling task plan according to the feedback result and determines the next decision of the car; finally, according to the readjusted scheduling task plan, the scheduling task plan is adjusted, and the final result is input into the adaptive cluster pick-up module in real time for real-time vehicle scheduling; The reinforcement learning and self-supervised learning include: optimizing scheduling tasks through a reinforcement learning framework; optimizing the scheduling action value at the current moment through the scheduling action value at the previous moment; and continuously updating and learning and selecting based on the received real-time environmental feedback, the environmental characteristics, and the optimization results of the trajectory generation and optimization module, and inputting the final result into the adaptive cluster pick-up module in real time for real-time vehicle scheduling.

7. The intelligent automatic driving vehicle dispatching system based on multimodal perception and dynamic environment modeling according to claim 1 is characterized in that: The trajectory generation and optimization module includes: Occ model construction and prediction and self-supervised learning; The Occ model construction and prediction includes: based on the Occupancy Network, the Occ model of the environmental occupancy grid is constructed in real time. The Occ model predicts and updates the idle and occupied areas in the environment by combining the BEV spatial features in the environmental features. The training formula of the Occupancy Network is: P occ (s t )=OccupancyNetwork(P bev (X t )), Among them, P occ (s t ) represents the state of the occupied grid at time t, P bev (X t ) represents the bird’s-eye view output of the occupancy grid at time t; The loss function of the OccupancyNetwork network is: η(Occ t )=||Occ t -Occ' t || 2 , Among them, Occ t is the current environmental state, Occ' t is the environmental state predicted by the system, || || is the norm, and η() is the function that minimizes the loss; the Occupancy Network adaptively learns from the environment and optimizes path planning in this way; The self-supervised learning includes using a self-supervised learning method to optimize path planning and vehicle scheduling planning through idle and occupied areas, while introducing reinforcement learning to optimize the real-time vehicle scheduling generated by the adaptive cluster pick-up module, and inputting the optimization results into the adaptive cluster pick-up module in real time for real-time vehicle scheduling.

8. The intelligent autonomous driving vehicle dispatching system based on multimodal perception and dynamic environment modeling according to claim 1, characterized in that: The semantic and action decision module, the adaptive cluster pick-up module, the trajectory generation and optimization module and the embodied intelligent learning module are designed as an end-to-end automatic decision module, so that after the car surrounding data and passenger information data collected by the sensor are input into the automatic decision module, the automatic decision module directly generates a series of driving decisions, and the series of driving decisions include: task planning, path planning, control instructions, task scheduling and scheduling optimization.

Citation Information

Patent Citations

  • Multi-agent parameterized Q network-based unmanned aerial vehicle communication sensing task scheduling method for Internet of Vehicles

    CN117692929A

  • Urban integrated management method based on artificial intelligence

    CN118098000A

  • System and method for motion prediction in autonomous driving

    CN119013705A

  • Driving situation prediction and adaptive strategy generation system based on cloud multi-mode large model

    CN119408566A

  • Adaptive driving behavior evaluation and training system based on multi-modal large model

    CN119513809A

Cited By

  • Traffic comprehensive query optimization method and system applied to intelligent human-computer interaction

    CN120336372A

  • Causal control VLA end-to-end system fusing perception and world model

    CN120745790A

  • A causal control vla end-to-end system that fuses perception and world model

    CN120745790B

  • Man-machine collaborative situation awareness dynamic task allocation method

    CN120817086A

  • Multi-modal fusion automatic driving decision arbitration method and system based on VLA model

    CN122275953A