Autonomous Driving Vehicle Dynamic Pick-up and Drop-off Management System Based on Multimodal Large Language Model
Through a multimodal large model system, combined with multimodal data acquisition and processing, the path and pick-up points of autonomous driving vehicles are optimized in real time in complex environments, solving the problem of inaccurate selection of paths and pick-up points in the existing technology, and improving the intelligence of autonomous driving and passenger experience.
Patent Information
- Application Number
- CN202411727528.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-11-28
AI Technical Summary
When existing autonomous driving vehicles select routes and pick-up points, they cannot optimize in real time according to environmental changes and passenger needs, resulting in inaccurate selection of paths and pick-up points.
A multimodal large model system is adopted, including multimodal data acquisition, processing, personalized pick-up point recommendation, dynamic path planning and adjustment, and intelligent passenger interaction modules. By integrating voice, image, and sensor data, the best path and pick-up point are generated in real time, and combined with reinforcement learning and collaborative filtering algorithms, personalized services are realized.
It realizes real-time response of autonomous vehicles in complex environments, provides personalized paths and pick-up point selection, improves traffic efficiency and passenger experience, and ensures safe and fast passenger pick-up.
Smart Images

Figure CN119659668B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and more particularly, to a dynamic pick-up and drop-off management system for autonomous vehicles based on a multimodal large model. Background Art
[0002] With the rapid development of autonomous driving technology, the service of automatically picking up and dropping off passengers has gradually been proposed; the service of automatically picking up and dropping off passengers through autonomous driving technology has the advantages of improving traffic safety, enhancing traffic efficiency, improving the passenger experience, saving time, and promoting the intelligence of the traffic system.
[0003] However, various problems exist during the operation of autonomous vehicles. For example, the route selected by the analysis system of existing autonomous vehicles is not the optimal route and cannot be adjusted in real time; the pick-up and drop-off points selected by autonomous vehicles are not the ones required by passengers; the pick-up and drop-off points selected by autonomous vehicles cannot be adjusted in real time according to passenger needs and environmental changes.
[0004] Therefore, it is necessary to improve the automatic pick-up and drop-off system of autonomous vehicles so that the improved system can not only generate accurate pick-up and drop-off paths and pick-up and drop-off points, but also adjust and update the pick-up and drop-off paths and pick-up and drop-off points in real time according to environmental changes and passenger needs. Summary of the Invention
[0005] The object of the present invention is to overcome the above problems existing in the prior art and greatly improve its technical effect on the basis of the original technology; to this end, the present invention provides a dynamic pick-up and drop-off management system for autonomous vehicles based on a multimodal large model, which system includes:
[0006] A multimodal data acquisition module; a multimodal large model processing module; a personalized pick-up and drop-off point recommendation module; a dynamic path planning and pick-up and drop-off point adjustment module; an intelligent passenger interaction and feedback module.
[0007] The multi-modal data acquisition module is used to receive the voice, text, image and video input information of passengers, obtain the real-time environmental data collected by vehicle sensors, and extract features from the acquired data; subsequently, the original data and all the data after feature extraction are transmitted to the multi-modal large model processing module; specifically, the multi-modal data acquisition module includes: data acquisition and data feature extraction; data acquisition refers to obtaining passenger information and real-time environmental data through various sensors of the autonomous vehicle; data feature extraction refers to obtaining the feature data of the corresponding data by extracting features from the acquired data; the data feature extraction is divided into: voice input processing, image and video input processing, and sensor data processing; voice input processing uses an end-to-end ASR model; the end-to-end ASR model is a speech recognition method that attempts to directly output text from the original speech signal through a unified neural network model, omitting multiple independent steps in the traditional ASR model; that is, the end-to-end ASR model combines the processing of speech information by the acoustic model, speech model and decoder; the steps for the end-to-end ASR model to process speech information are: pre-training stage: training in an unsupervised manner enables the model to learn the features of the speech signal to generate latent representations, thereby capturing the speech patterns in the audio; fine-tuning stage: further optimizing the model using labeled data so that the model can accurately map from the speech signal to the text; decoding stage: extracting the final speech-to-text conversion process from the trained model through the Transformer decoder in the model, generating a series of possible text outputs based on the input speech features and the conversion process, and obtaining the final recognition result by selecting the most likely sequence; the speech formula for converting the speech signal into text is:
[0008]
[0009] where X represents the speech feature sequence, X = (x1, x2,..., x T ), T represents that there are T time steps in X; W represents the predicted word sequence, P(W|X) represents the conditional probability of predicting the word sequence W given the speech sequence X; x t represents the speech feature at the current time step t, ω t represents the target output at the current time step t, h t-1 represents the hidden state at the previous time step t - 1, θ refers to the model parameters, and P(ω t x t , h t-1 , θ) represents the conditional probability of predicting ω t at the current time step t.
[0010] The image and video input processing uses a convolutional neural network (CNN) and a vision transformer (ViT) for feature extraction, and the YOLOv5 model is used to perform object recognition on the information after feature extraction. The steps of the image and video input processing are as follows: The image and video information is preliminarily processed by a convolutional neural network (CNN) to extract the low-level features of the image and video information, and the low-level features are the local features of the image and video information. The low-level features extracted by the convolutional neural network (CNN) are passed to the vision transformer (ViT) module, and the vision transformer (ViT) module uses the self-attention mechanism to further capture global information, identify long-range dependencies, and context information. Combining the features extracted by the convolutional neural network (CNN) and the vision transformer (ViT) module, the YOLOv5 model divides the input image into multiple grids according to the combined features. Each grid is responsible for predicting the objects in the image and giving the category and bounding box. Object recognition is performed through a first-order residual structure combined with multi-scale feature fusion. The first-order residual structure refers to directly adding the input and output of feature extraction by introducing skip connections, so that information can bypass some convolutional layers through these skip connections and be directly transmitted to deeper layers. The multi-scale feature fusion uses a feature pyramid network.
[0011] The sensor data processing combines PointNet++ processing and the VoxelNet algorithm for processing. First, PointNet++ is used to obtain point cloud data by hierarchical data sampling of the data in the sensor. Subsequently, the VoxelNet algorithm is used to convert the point cloud data into voxel data, and the voxel data is processed by 3D CNN technology of a three-dimensional convolutional neural network to obtain low-level local spatial features. Finally, PointNet++ aggregates the features extracted from the 3D CNN, learns global features, and combines the local spatial features and global features to generate the 3D structural features of the environmental model.
[0012] The multimodal large model processing module is used to analyze and fuse the data collected by the multimodal data acquisition module to generate intelligent decisions for pick-up and drop-off point recommendation and path planning. Specifically, the multimodal large model processing module includes: natural language processing (NLP), image and video processing, environmental perception and 3D space understanding, and language emotion recognition. The natural language processing (NLP) uses the BERT large language model and the method of introducing multi-task learning (MTL) to understand and process the feature text information. The processing steps are as follows: BERT pre-training, BERT is pre-trained through a large model autonomous driving corpus, and BERT learns general language representations; fine-tuning BERT and multi-task learning (MTL) are combined, and multiple tasks including classification, propositional entity recognition, and question answering are added to the basis of BERT, and at the same time, the shared layer of BERT and the task-specific head are optimized; joint loss optimization, through the joint loss function, optimize the objectives of different tasks to enhance the generalization ability of the model; according to the requirements of downstream tasks, output corresponding results through the layers corresponding to the tasks; The formula of the BERT pre-training model is:
[0013] H t = Transformer(H t-1 ,H t-2 ,...,H0; θ)
[0014] where, H t represents the hidden state of the Transformer at the t-th time step, H t-1 represents the hidden state of the Transformer at the (t - 1)-th time step. The time step refers to the representation of each unit of time in a discretized time series; θ represents the model parameters.
[0015] Based on the image and video processing of the multimodal data acquisition module, the image and video processing adds a three-dimensional convolutional neural network (3D-CNN) to capture the dynamic features of the target. The steps are as follows: The three-dimensional convolutional neural network (3D-CNN) sets a one-dimensional feature extraction structure and a three-dimensional feature extraction structure; the target features are obtained through the one-dimensional feature extraction structure, and the spatio-temporal features of the target changing over time are extracted through the three-dimensional feature extraction structure; the extracted target features and the corresponding spatio-temporal features are combined to capture the dynamic features of the target.
[0016] The environmental perception and 3D space understanding use LiDAR point cloud processing combined with Simultaneous Localization and Mapping (SLAM) technology to generate a real-time 3D environmental map. The steps to generate a real-time 3D environmental map are as follows: Through LiDAR point cloud processing, sequentially extract the local features of the point cloud data through steps including: collecting and obtaining point cloud data, preprocessing the point cloud data, and extracting the feature extraction of the preprocessed point cloud data; Through the map construction SLAM technology, continuously accumulate the point cloud data features at different times and update the position, construct a 3D environmental map, and compare the point cloud obtained by each LiDAR scan with the data of the previous scan to gradually update the 3D environmental map.
[0017] The language emotion recognition refers to classifying the speech emotion through an LSTM model based on emotion vectors, and adjusting the pick-up and drop-off point priorities according to the classification results. The emotion recognition model formula is:
[0018] s t =σ(W s ·x t +U s ·h t-1 +b s )
[0019] Among them, s t represents the emotional state, σ represents the Sigmoid activation function, W s and U s represent the weight parameters of the model, x t represents the input speech feature vector, h t-1 is the hidden state of the previous moment, b s represents the bias term, which is used to adjust the model output.
[0020] Furthermore, through various types of processed data, fuse and analyze different types of data to lay a foundation for subsequent path planning and pick-up and drop-off point determination.
[0021] The personalized pick-up and drop-off point recommendation module is used to generate personalized pick-up and drop-off points based on the passenger's historical behavior, preferences, and current environment. Specifically, provide a customized pick-up and drop-off point service based on the passenger's historical behavior through the collaborative filtering algorithm. The process is to decompose the passenger and pick-up and drop-off point rating matrix into multiple low-dimensional matrices through matrix factorization, so as to judge which pick-up and drop-off points are more convenient for the passenger; judge the pick-up and drop-off point position through the inner product of the feature vectors of the passenger and the pick-up and drop-off point, and take the pick-up and drop-off point with the highest rating value of the passenger feature vector as the best pick-up and drop-off point; after the pick-up and drop-off point is determined, adjust the pick-up and drop-off point according to the environmental changes and passenger needs to ensure the best experience.
[0022] The dynamic path planning and pick-up / drop-off point adjustment module is used to dynamically adjust pick-up / drop-off points and path planning by combining real-time environmental data and multimodal inputs; dynamically adjust path planning and pick-up / drop-off points according to the dynamic changes in the real-time environment; the dynamic path planning and pick-up / drop-off point adjustment module includes: The real-time environmental data includes traffic flow, road closures, and weather; dynamic path planning means that the system uses advanced path planning algorithms to select the optimal driving route and also supports dynamically adjusting pick-up / drop-off points to ensure that passengers can get on and off quickly and safely; and when there are road closures, traffic congestion, and the pick-up / drop-off point location is requested by passengers, the system can adjust the pick-up / drop-off point location in real time and feedback the adjusted pick-up / drop-off point location to passengers; using the A* path planning algorithm and reinforcement learning technology to ensure that the vehicle can select the optimal path according to external environmental changes; The A* path planning algorithm is a heuristic path search algorithm used to quickly plan the optimal path in a known map and environment; the system generates the shortest path based on the A* algorithm in combination with real-time traffic data; the A* path planning algorithm formula is:
[0023] f(n) = g(n) + h(n)
[0024] where n represents a node, representing a certain state and position of the autonomous vehicle, g(n) represents the cost from the starting point to the current node, h(n) is the estimated cost from the current node to the target point, and f(n) is the total cost.
[0025] After the optimal path is planned through the A* path planning algorithm, the path needs to be optimized through reinforcement learning, and the Q-learning algorithm in deep reinforcement learning DRL is used to achieve real-time path optimization of the vehicle in a dynamic environment.
[0026] The Q-learning algorithm formula is:
[0027]
[0028] where Q(s, a) is the state-action value function, which refers to the maximum expected cumulative reward that the autonomous vehicle can obtain by taking action a in state s, r is the immediate reward, γ is a discount factor with a value between 0 and 1, s' refers to the next state that the autonomous vehicle transfers to after executing action a, and a' refers to the actions that the autonomous vehicle may take in state s'. It refers to the maximum value of the state-action value function of all possible actions a' in the next state s'.
[0029] The intelligent passenger interaction and feedback module is used to interact with passengers in real time, share the driving route and recommended pick-up and drop-off points with passengers. Passengers view the driving route and recommended pick-up and drop-off points and put forward new demands according to their needs. The system feeds back the new demands to the dynamic route planning and pick-up and drop-off point adjustment module, and the dynamic route planning and pick-up and drop-off point adjustment module adjusts the route and pick-up and drop-off points according to the new demands of passengers. Specifically, the intelligent passenger interaction and feedback module includes: the system interacts with passengers through the passengers' mobile devices, providing real-time pick-up and drop-off point suggestions and route updates; passengers view the recommended pick-up and drop-off points, the real-time position of the vehicle, and the estimated arrival time through the mobile application; the system also helps passengers quickly find the vehicle parking position in a complex environment through augmented reality (AR) navigation; combining the visual SLAM algorithm and the camera data of the passenger device to provide accurate AR navigation; the formula for AR navigation is:
[0030] T cw =K[R|t]
[0031] where T cw is the transformation matrix from camera coordinates to world coordinates, and the world coordinates are used to describe the absolute position of objects in the scene; K is the camera intrinsic matrix, and R and t are the rotation matrix and translation vector respectively; by determining the absolute positions of objects, passengers, and autonomous vehicles in the camera, mapping the position and scene information to the AR scene, and passengers navigate through the AR-rendered interface to find the pick-up and drop-off point position of the autonomous vehicle.
[0032] The beneficial effects of the present invention are:
[0033] The present invention provides an autonomous vehicle dynamic pick-up and drop-off management system based on a multimodal large model. The system includes a multimodal data acquisition module; a multimodal large model processing module; a personalized pick-up and drop-off point recommendation module; a dynamic route planning and pick-up and drop-off point adjustment module; an intelligent passenger interaction and feedback module; the system aims to provide intelligent pick-up and drop-off point selection and dynamic adjustment capabilities for autonomous vehicles; the system realizes efficient perception and real-time response to complex environments by integrating multimodal data, and generates the best route and pick-up and drop-off points in real time; the system can automatically adjust the pick-up and drop-off points and route planning according to real-time traffic, environmental changes, and passenger demands, providing a personalized passenger experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic diagram of the autonomous vehicle dynamic pick-up and drop-off management system based on the multimodal large model of the present invention.
[0035] Figure 2 is a flowchart of the autonomous vehicle dynamic pick-up and drop-off based on the multimodal large model of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0036] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments given here are only for illustrating and explaining the present invention, and cannot be used to limit the present invention.
[0037] It should be noted that many specific details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention may have other embodiments and variations, and therefore, the scope of protection of the present invention is not limited by the specific embodiments disclosed below.
[0038] As Figure 1 FIG. is a schematic diagram of an autonomous driving vehicle dynamic pick-up and drop-off management system based on a multimodal large model according to an embodiment of the present invention. The schematic diagram includes: a multimodal data acquisition module S100; a multimodal large model processing module S200; a personalized pick-up and drop-off point recommendation module S300; a dynamic path planning and pick-up and drop-off point adjustment module S400; an intelligent passenger interaction and feedback module S500.
[0039] Specifically, the multimodal data acquisition module S100 is used to receive the voice, text, image and video input information of passengers, obtain the real-time environment data collected by vehicle sensors, and extract features from the acquired data; subsequently, all the original data and the data after feature extraction are transmitted to the multimodal large model processing module S200.
[0040] Furthermore, the multimodal data acquisition module S100 includes: the acquisition of data and the extraction of data features; the acquisition of data refers to obtaining passenger information and real-time environment data through various sensors of the autonomous driving vehicle; the extraction of data features refers to extracting feature data corresponding to the acquired data by extracting features from the acquired data; the extraction of data features is divided into: voice input processing, image and video input processing, and sensor data processing; voice input processing uses an end-to-end ASR model; the end-to-end ASR model is a speech recognition method that attempts to directly output text from the original speech signal through a unified neural network model, omitting multiple independent steps in the traditional ASR model; that is, the end-to-end ASR model combines the processing of speech information by the acoustic model, the speech model, and the decoder; the steps for the end-to-end ASR model to process speech information are: pre-training stage: training in an unsupervised manner enables the model to learn the features of the speech signal to generate a latent representation, thereby capturing the speech patterns in the audio; fine-tuning stage: further optimizing the model using labeled data so that the model can accurately map from the speech signal to the text; decoding stage: extracting the final speech-to-text conversion process from the trained model through the Transformer decoder in the model, generating a series of possible text outputs based on the input speech features and the conversion process, and obtaining the final recognition result by selecting the most likely sequence; the speech formula for converting the speech signal into text is:
[0041]
[0042] Among them, X represents the speech feature sequence, X = (x1, x2, …, x T ), T represents that there are T time steps in X; W represents the predicted word sequence, and P(W|X) represents the conditional probability of predicting the word sequence W given the speech sequence X; x t represents the speech feature at the current time step t, and ω t represents the target output at the current time step t, h t-1 represents the hidden state at the previous time step t - 1, θ refers to the model parameters, and P(ω t |x t , h t-1 , θ) represents the conditional probability of predicting ω t at the current time step t.
[0043] The image and video input processing uses a convolutional neural network CNN and a vision transformer ViT for feature extraction, and uses the YOLOv5 model to perform object recognition on the information after feature extraction; the steps of image and video input processing are as follows: the convolutional neural network CNN is used to preliminarily process the image and video information to extract the low-level features of the image and video information, and the low-level features are the local features of the image and video information; the low-level features extracted by the convolutional neural network CNN are passed to the vision transformer ViT module, and the vision transformer ViT module uses the self-attention mechanism to further capture global information and identify long-distance dependencies and context information; combining the features extracted by the convolutional neural network CNN and the vision transformer ViT module, the YOLOv5 model divides the input image into multiple grids according to the combined features, and each grid is responsible for predicting the objects in the image and giving the category and bounding box; object recognition is performed through a first-order residual structure combined with multi-scale feature fusion; the first-order residual structure refers to directly adding the input and output of feature extraction by introducing skip connections, so that information can bypass some convolutional layers through these skip connections and be directly passed to deeper layers; the multi-scale feature fusion uses a feature pyramid network.
[0044] The sensor data processing combines PointNet++ processing and VoxelNet algorithm for processing. First, PointNet++ obtains point cloud data by hierarchical data sampling of the data in the sensor. Subsequently, the VoxelNet algorithm converts the point cloud data into voxel Voxel data, and processes the voxel data through the 3D convolutional neural network 3D CNN technology to obtain low-level local spatial features. Finally, PointNet++ aggregates the features extracted from the 3D CNN, learns global features, and combines the local spatial features and global features to generate the 3D structural features of the environmental model.
[0045] In the above embodiment, the multimodal large model processing module S200 is used to analyze and fuse the data collected by the multimodal data acquisition module S100, and generate intelligent decisions for pick-up and drop-off point recommendation and path planning.
[0046] Further, the multimodal large model processing module S200 includes: natural language processing NLP, image and video processing, environmental perception and 3D space understanding, and language emotion recognition. The natural language processing NLP understands and processes the feature text information through the BERT large language model and the method of introducing multi-task learning MTL. The processing steps are as follows: BERT pre-training, BERT pre-training is carried out through the large model autonomous driving corpus, and BERT learns general language representations; fine-tuning BERT and multi-task learning MTL are combined, and multiple tasks including classification, propositional entity recognition, and question answering are added to the basis of BERT, and at the same time, the shared layer and task-specific heads of BERT are optimized; joint loss optimization, through the joint loss function, optimizes the objectives of different characters and enhances the generalization ability of the model; according to the requirements of downstream tasks, corresponding results are output through the layers corresponding to the tasks; the BERT pre-training model formula is:
[0047] H t = Transformer(H t-1 , H t-2 ,..., H0; θ)
[0048] Among them, H t represents the hidden state of the Transformer at the t-th time step, H t-1 represents the hidden state of the Transformer at the (t - 1)-th time step. The time step refers to the representation of each unit time in a discretized time series; θ represents the model parameters.
[0049] Based on the image and video processing in the multimodal data acquisition module S100, a three-dimensional convolutional neural network 3D-CNN is added to capture the dynamic features of the target. The steps are as follows: The three-dimensional convolutional neural network 3D-CNN sets up a one-dimensional feature extraction structure and a three-dimensional feature extraction structure; the one-dimensional feature extraction structure is used to obtain the target features, and the three-dimensional feature extraction structure is used to extract the spatio-temporal features of the target changing over time; the extracted target features and the corresponding spatio-temporal features are combined to capture the dynamic features of the target.
[0050] The environmental perception and 3D space understanding use LiDAR point cloud processing combined with the Simultaneous Localization and Mapping SLAM technology to generate a real-time 3D environmental map. The steps to generate a real-time 3D environmental map are as follows: Through LiDAR point cloud processing, the local features of the point cloud data are extracted successively through steps including: acquiring point cloud data, preprocessing the point cloud data, and feature extraction of the preprocessed point cloud data; through the map construction SLAM technology, the point cloud data features and the position are continuously accumulated at different times, a 3D environmental map is constructed, and the point cloud obtained by each LiDAR scan is compared with the data of the previous scan to gradually update the 3D environmental map.
[0051] The language emotion recognition refers to classifying the speech emotion through an LSTM model based on emotion vectors, and adjusting the priority of pick-up and drop-off points according to the classification results. The emotion recognition model formula is:
[0052] s t =σ(W s ·x t +U s ·h t-1 +b s )
[0053] Among them, s t represents the emotional state, σ represents the Sigmoid activation function, W s and U s represent the weight parameters of the model, x t represents the input speech feature vector, h t-1 is the hidden state of the previous moment, b s represents the bias term, which is used to adjust the model output.
[0054] Through the processed various types of data, the different types of data are fused and analyzed to lay a foundation for subsequent path planning and the determination of pick-up and drop-off points.
[0055] Specifically, for example, interactions can be carried out between the multimodal data acquisition module and the multimodal large model processing module, between the multimodal large model processing module and the dynamic path planning and pick-up / drop-off point adjustment module, between the dynamic path planning and pick-up / drop-off point adjustment module and the personalized pick-up / drop-off point recommendation module, and between the dynamic path planning and pick-up / drop-off point adjustment module and the intelligent passenger interaction and feedback module.
[0056] Furthermore, in the interaction between the multimodal data acquisition module and the multimodal large model processing module, the data of the multimodal data acquisition module and the multimodal large model processing module can be fused through a spatio-temporal alignment algorithm to ensure the time synchronization and spatial consistency of multimodal information; subsequently, multimodal feature fusion is carried out; in the interaction between the multimodal large model processing module and the dynamic path planning and pick-up / drop-off point adjustment module, the path and pick-up / drop-off points can be planned through path planning based on 3D space understanding and real-time dynamic adjustment and optimization can be carried out; in the interaction between the dynamic path planning and pick-up / drop-off point adjustment module and the personalized pick-up / drop-off point recommendation module, by combining historical data and real-time planning data, the real-time updated data of the dynamic path planning and pick-up / drop-off point adjustment module and the passenger historical data of the personalized pick-up / drop-off point recommendation module are combined to generate paths and pick-up / drop-off points closer to the passenger's needs; in the interaction between the dynamic path planning and pick-up / drop-off point adjustment module and the intelligent passenger interaction and feedback module, the paths and pick-up / drop-off points updated in real time by the dynamic path planning and pick-up / drop-off point adjustment module are combined with the passenger's real-time needs to achieve real-time optimization of the paths and pick-up / drop-off points; in the above processes, after the extracted features are fused, the results can be obtained through corresponding analysis methods.
[0057] In the above embodiment, the personalized pick-up / drop-off point recommendation module S300 is used to generate personalized pick-up / drop-off points based on the passenger's historical behavior, preferences, and current environment.
[0058] Furthermore, a customized pick-up / drop-off point service is provided based on the passenger's historical behavior through a collaborative filtering algorithm; the process is to decompose the passenger-pick-up / drop-off point rating matrix into multiple low-dimensional matrices by matrix factorization to determine which pick-up / drop-off points are more convenient for the passenger; the position of the pick-up / drop-off point is judged by the rating value obtained from the inner product of the feature vectors of the passenger and the pick-up / drop-off point, and the pick-up / drop-off point with the highest rating value of the passenger feature vector is used as the best pick-up / drop-off point; after the pick-up / drop-off point is determined, the pick-up / drop-off point is adjusted according to environmental changes and passenger needs to ensure the best experience.
[0059] At the same time, the recommended best pick-up / drop-off point and other paths and pick-up / drop-off points with relatively high ratings are recommended to the user, and within a specific time, such as within 20 seconds, the user is asked to confirm the path and pick-up / drop-off point; if the user does not confirm, the system selects the best path and the best pick-up / drop-off point.
[0060] In the above embodiments, the dynamic path planning and pick-up / drop-off point adjustment module S400 is used to dynamically adjust the pick-up / drop-off points and path planning by combining real-time environmental data and multimodal inputs; and dynamically adjust the path planning and pick-up / drop-off points according to the dynamic changes of the real-time environment.
[0061] Further, the dynamic path planning and pick-up / drop-off point adjustment module S400 includes: The real-time environmental data includes traffic flow, road closures, and weather; Dynamic path planning means that the system uses advanced path planning algorithms to select the optimal driving route and also supports dynamically adjusting the pick-up / drop-off points to ensure that passengers can get on and off quickly and safely; And when there are road closures, traffic congestion, and when the pick-up / drop-off point location is requested by passengers, the system can adjust the pick-up / drop-off point location in real time and feedback the adjusted pick-up / drop-off point location to the passengers; Using the A* path planning algorithm and reinforcement learning technology to ensure that the vehicle can select the optimal path according to external environmental changes; The A* path planning algorithm is a heuristic path search algorithm used to quickly plan the optimal path in a known map and environment; The system generates the shortest path based on the A* algorithm and combines real-time traffic data; The formula of the A* path planning algorithm is:
[0062] f(n) = g(n) + h(n)
[0063] Where n represents a node, representing a certain state and position of the autonomous vehicle, g(n) represents the cost from the starting point to the current node, h(n) is the estimated cost from the current node to the target point, and f(n) is the total cost.
[0064] After planning the optimal path through the A* path planning algorithm, it is necessary to optimize the path through reinforcement learning. The Q-learning algorithm in deep reinforcement learning DRL is used to achieve real-time path optimization of the vehicle in a dynamic environment. The formula of the Q-learning algorithm is:
[0065]
[0066] Where Q(s, a) is the state-action value function, which refers to the maximum expected cumulative reward that the autonomous vehicle can obtain by taking action a in state s, r is the immediate reward, γ is a discount factor with a value between 0 and 1, s' refers to the next state that the autonomous vehicle transfers to after executing action a, a' refers to the action that the autonomous vehicle may take in state s', Refers to the maximum value of the state-action value function of all possible actions a' in the next state s'.
[0067] In the above embodiments, the intelligent passenger interaction and feedback module S500 is used to interact with passengers in real time, share the driving route and recommended pick-up and drop-off points with passengers. Passengers view the driving route and recommended pick-up and drop-off points and put forward new demands according to their needs. The system feeds back the new demands to the dynamic route planning and pick-up and drop-off point adjustment module S400, and the dynamic route planning and pick-up and drop-off point adjustment module S400 adjusts the route and pick-up and drop-off points according to the new demands of passengers.
[0068] Furthermore, the system interacts with passengers through their mobile devices, providing real-time pick-up and drop-off point suggestions and route updates; passengers view the recommended pick-up and drop-off points, the real-time location of the vehicle, and the estimated arrival time through a mobile application; the system also uses augmented reality (AR) navigation to help passengers quickly find the vehicle's parking location in complex environments; combining the visual SLAM algorithm with the camera data of the passenger device to provide accurate AR navigation; the formula for AR navigation is:
[0069] T cw = K[R|t]
[0070] where T cw is the transformation matrix from camera coordinates to world coordinates, and the world coordinates are used to describe the absolute position of objects in the scene; K is the camera intrinsic matrix, and R and t are the rotation matrix and translation vector respectively; by determining the absolute positions of objects, passengers, and the autonomous driving vehicle in the camera, mapping the position and scene information to the AR scene, and passengers navigate through the AR-rendered interface to find the pick-up and drop-off point location of the autonomous driving vehicle.
[0071] Such as Figure 2 is the flowchart of the dynamic pick-up and drop-off of an autonomous driving vehicle based on a multimodal large model according to an embodiment of the present invention; the realization of the pick-up and drop-off passenger service in this flowchart is based on Figure 1 the system and the analysis method of this system.
[0072] Specifically, first, the multi-modal data acquisition module collects voice, text, images, videos, sensor data information, etc. of the vehicle surroundings and passengers; among them, sensor data refers to LiDAR data, radar data, GPS data, etc.; after collecting the data, preprocess and extract features from the data, and then send the data after feature extraction to the multi-modal large model processing module; the multi-modal large model processing module performs natural language processing NLP, image and video processing, environmental perception and 3D space understanding, and language emotion recognition on the data after feature extraction respectively, understands the deep semantics of different types of data information, and then fuses the deep semantics of different types of data information to generate personalized recommendations for passengers, and sends the personalized recommendations to the personalized pick-up and drop-off point recommendation module; the personalized pick-up and drop-off point recommendation module has the function of real-time recommending and adjusting the pick-up and drop-off points and optimizing the pick-up and drop-off points; the personalized pick-up and drop-off point recommendation module recommends personalized pick-up and drop-off points for passengers according to the passenger's historical pick-up and drop-off point data information combined with the environmental data information analyzed by the multi-modal large model processing module, and then feeds back the recommended personalized pick-up and drop-off points to the dynamic path planning and pick-up and drop-off point adjustment module; the dynamic path planning and pick-up and drop-off point adjustment module plans the optimal path and pick-up and drop-off points in real time according to the recommended personalized pick-up and drop-off points combined with the real-time fused multi-modal data, and feeds back the planned optimal path and pick-up and drop-off points to the passengers; the passengers view the displayed path and vehicle position through the intelligent passenger interaction and feedback module, and view the recommended pick-up and drop-off points; if the passengers are not satisfied after viewing, they put forward new requirements, and the mobile devices used by the passengers will feedback the requirements to the dynamic path planning and pick-up and drop-off point adjustment module through the passenger app connected to the system, and the dynamic path planning and pick-up and drop-off point adjustment module will adjust the path and pick-up and drop-off points according to the passengers' requirements; when adjusting the path and pick-up and drop-off points, the passengers' requirements are taken as the top priority for adjustment.
[0073] At the same time, when the dynamic path planning and pick-up and drop-off point adjustment module is planning the optimal path and pick-up and drop-off points; it also performs real-time traffic and environmental perception, optimizes the generated path through real-time monitoring of traffic and external environmental perception, and feeds back the optimized path to the multi-modal large model processing module, and the multi-modal large model processing module further processes the optimized path in combination with the real-time updated multi-modal data to generate a personalized recommended path that meets the passengers.
Claims
1. An autonomous driving vehicle dynamic pick-up and drop-off management system based on a multi-modal large model, characterized in that, The system includes: A multimodal data acquisition module; a multimodal large model processing module; a personalized pick-up and drop-off point recommendation module; a dynamic route planning and pick-up and drop-off point adjustment module; an intelligent passenger interaction and feedback module; The multimodal data acquisition module is used to receive the voice, text, image and video input information of passengers, obtain the real-time environmental data collected by vehicle sensors, and extract features from the acquired data; subsequently, all the original data and the data after feature extraction are transmitted to the multimodal large model processing module; The multimodal large model processing module is used to analyze and fuse the data collected by the multimodal data acquisition module, and generate intelligent decisions for pick-up and drop-off point recommendation and route planning; The personalized pick-up and drop-off point recommendation module is used to generate personalized pick-up and drop-off points based on the historical behavior, preferences and current environment of passengers; The dynamic route planning and pick-up and drop-off point adjustment module is used to dynamically adjust the pick-up and drop-off points and route planning in combination with the real-time environmental data and multimodal inputs; dynamically adjust the route planning and pick-up and drop-off points according to the dynamic changes of the real-time environment; The intelligent passenger interaction and feedback module is used to interact with passengers in real time, share the driving route and recommended pick-up and drop-off points with passengers. Passengers view the driving route and recommended pick-up and drop-off points and put forward new requirements according to their needs. The system feeds the new requirements back to the dynamic route planning and pick-up and drop-off point adjustment module, and the dynamic route planning and pick-up and drop-off point adjustment module adjusts the route and pick-up and drop-off points according to the new requirements of passengers.
2. The dynamic pick-up and drop-off management system for autonomous driving vehicles based on a multi-modal large model according to claim 1, wherein, The multimodal data acquisition module includes: the acquisition of data and the extraction of data features; the acquisition of data refers to obtaining passenger information and real-time environmental data through various sensors of the autonomous vehicle; the extraction of data features refers to obtaining the feature data of the corresponding data by extracting features from the acquired data; the extraction of data features is divided into: voice input processing, image and video input processing, and sensor data processing; the voice input processing uses an end-to-end ASR model; the end-to-end ASR model is a speech recognition method that attempts to directly output text from the original speech signal through a unified neural network model, omitting multiple independent steps in the traditional ASR model; that is, the end-to-end ASR model combines the processing of speech information by the acoustic model, speech model and decoder; the steps for the end-to-end ASR model to process speech information are: pre-training stage: training in an unsupervised manner enables the model to learn the features of the speech signal to generate a latent representation, thereby capturing the speech patterns in the audio; fine-tuning stage: using labeled data to further optimize the model so that the model can accurately map from the speech signal to the text; decoding stage: extracting the final speech-to-text conversion process from the trained model through the Transformer decoder in the model, generating a series of possible text outputs based on the input speech features and the conversion process, and obtaining the final recognition result by selecting the most likely sequence; the speech formula for converting the speech signal into text is: Among them, X represents the speech feature sequence, X = (x1, x2,..., x T ), T represents that there are T time steps in X; W represents the predicted word sequence, and P(W|X) represents the conditional probability of predicting the word sequence W given the speech sequence X; x t represents the speech feature at the current time step t, ω t represents the target output at the current time step t, h t-1 represents the hidden state at the previous time step t - 1, θ refers to the model parameters, and P(ω t |x t , h t-1 , θ) represents the conditional probability of predicting ω t at the current time step t; The image and video input processing uses a convolutional neural network (CNN) and a vision transformer (ViT) for feature extraction, and the YOLOv5 model is used for object recognition of the information after feature extraction. The steps of the image and video input processing are as follows: The image and video information is preliminarily processed by the CNN to extract the low-level features of the image and video information, and the low-level features are the local features of the image and video information. The low-level features extracted by the CNN are passed to the ViT module. The ViT module uses the self-attention mechanism to further capture global information, identify long-range dependencies and context information. Combining the features extracted by the CNN and the ViT module, the YOLOv5 model divides the input image into multiple grids according to the combined features. Each grid is responsible for predicting the objects in the image and giving the category and bounding box. Object recognition is performed through a first-order residual structure combined with multi-scale feature fusion. The first-order residual structure means that by introducing skip connections, the input and output of feature extraction are directly added, so that information can bypass some convolutional layers through these skip connections and be directly passed to deeper layers. The multi-scale feature fusion uses a feature pyramid network. The sensor data processing combines PointNet++ processing and the VoxelNet algorithm for processing. First, PointNet++ is used to obtain point cloud data by hierarchical data sampling of the data in the sensor. Subsequently, the VoxelNet algorithm is used to convert the point cloud data into voxel data, and the voxel data is processed by 3D CNN technology of a three-dimensional convolutional neural network to obtain low-level local spatial features. Finally, PointNet++ aggregates the features extracted from the 3D CNN, learns global features, and combines the local spatial features and global features to generate the 3D structural features of the environmental model.
3. The dynamic pick-up and drop-off management system for autonomous vehicles based on a multi-modal large model according to claim 1, characterized in that, The multi-modal large model processing module includes: natural language processing (NLP), image and video processing, environmental perception and 3D space understanding, and language emotion recognition. The NLP understands and processes the feature text information through the BERT large language model and the method of introducing multi-task learning (MTL). The processing steps are as follows: BERT pre-training, BERT pre-training is performed through a large model autonomous driving corpus, and BERT learns general language representations. Fine-tuning BERT and MTL are combined, and multiple tasks including classification, propositional entity recognition, and question answering are added to the basis of BERT, and at the same time, the shared layer of BERT and the task-specific head are optimized. Joint loss optimization, through the joint loss function, optimizes the objectives of different people and enhances the generalization ability of the model. According to the requirements of downstream tasks, the corresponding results are output through the layers corresponding to the tasks. The BERT pre-training model formula is: H t = Transformer(H t-1 , H t-2 , …, H0; θ) Among them, H t represents the hidden state of the Transformer at the t-th time step, and H t-1 represents the hidden state of the Transformer at the (t - 1)-th time step. The time step refers to the representation of each unit of time in a discretized time series; θ represents the model parameters; Based on the image and video processing of the multi-modal data acquisition module, a three-dimensional convolutional neural network 3D-CNN is added to capture the dynamic features of the target. The steps are as follows: The 3D-CNN sets up a one-dimensional feature extraction structure and a three-dimensional feature extraction structure; the one-dimensional feature extraction structure is used to obtain the target features, and the three-dimensional feature extraction structure is used to extract the spatio-temporal features of the target changing over time; the extracted target features and the corresponding spatio-temporal features are combined to capture the dynamic features of the target. The environmental perception and 3D space understanding use LiDAR point cloud processing combined with the Simultaneous Localization and Mapping (SLAM) technology to generate a real-time 3D environmental map. The steps to generate a real-time 3D environmental map are as follows: Through LiDAR point cloud processing, the local features of the point cloud data are extracted successively through steps including: acquiring point cloud data, preprocessing the point cloud data, and extracting features from the preprocessed point cloud data; through the SLAM technology for map construction, the features of the point cloud data at different times are continuously accumulated and the position is updated to construct a 3D environmental map, and the point cloud obtained from each LiDAR scan is compared with the data from the previous scan to gradually update the 3D environmental map. The language emotion recognition refers to classifying speech emotions through an LSTM model based on emotion vectors, and adjusting the priority of pick-up and drop-off points according to the classification results. The formula for the emotion recognition model is: s t = σ(W s · x t + U s · h t-1 + b s ) Among them, s t represents the emotional state, σ represents the Sigmoid activation function, W s and U s represent the weight parameters of the model, x t represents the input speech feature vector, h t-1 is the hidden state at the previous moment, b s represents the bias term, which is used to adjust the model output; By processing various types of data, the different types of data are fused and analyzed to lay a foundation for subsequent path planning and the determination of pick-up and drop-off points.
4. The dynamic pick-up and drop-off management system for autonomous driving vehicles based on a multimodal large model according to claim 1, wherein The personalized pick-up and drop-off point recommendation module includes: providing a customized pick-up and drop-off point service based on the historical behavior of passengers through a collaborative filtering algorithm. The process is to decompose the passenger and pick-up / drop-off point rating matrix into multiple low-dimensional matrices through matrix factorization to determine which pick-up and drop-off points are more convenient for passengers; the position of the pick-up and drop-off point is judged by the rating value obtained from the inner product of the feature vectors of the passenger and the pick-up and drop-off point, and the pick-up and drop-off point with the highest rating value of the passenger feature vector is used as the best pick-up and drop-off point; after the pick-up and drop-off point is determined, the pick-up and drop-off point is adjusted according to environmental changes and passenger needs to ensure the best experience.
5. The dynamic pick-up and drop-off management system for autonomous vehicles based on a multi-modal large model according to claim 1, characterized in that, The dynamic path planning and pick-up / drop-off point adjustment module includes: Real-time environmental data includes traffic flow, road closures, and weather. Dynamic path planning means that the system uses an advanced path planning algorithm to select the optimal driving route and also supports dynamic adjustment of pick-up and drop-off points to ensure that passengers can get on and off quickly and safely; and when there are road closures, traffic congestion, and the pick-up / drop-off point position requested by the passenger, the system can adjust the pick-up / drop-off point position in real time and feedback the adjusted pick-up / drop-off point position to the passenger; the A* path planning algorithm and reinforcement learning technology are used to ensure that the vehicle can select the optimal path according to external environmental changes. The A* path planning algorithm is a heuristic path search algorithm used to quickly plan the optimal path in a known map and environment. The system generates the shortest path based on the A* algorithm combined with real-time traffic data. The formula for the A* path planning algorithm is: f(n) = g(n) + h(n) Among them, n represents a node, which represents a certain state and position of the autonomous vehicle. g(n) represents the cost from the starting point to the current node, h(n) is the estimated cost from the current node to the target point, and f(n) is the total cost; After planning the optimal path through the A* path planning algorithm, it is necessary to optimize the path through reinforcement learning, and use the Q-learning algorithm in the deep reinforcement learning DRL to achieve real-time path optimization of the vehicle in a dynamic environment; The formula of the Q-learning algorithm is: Among them, Q(s, a) is the state-action value function, which refers to the maximum expected cumulative reward that an autonomous vehicle can obtain by taking action a in state s. r is the immediate reward, γ is the discount factor with a value between 0 and 1, s' refers to the next state that the autonomous vehicle transfers to after executing action a, and a' refers to the actions that the autonomous vehicle may take in state s'. It refers to the maximum value of the state-action value function of all possible actions a' in the next state s'.
6. The dynamic pick-up and drop-off management system for autonomous vehicles based on a multi-modal large model according to claim 1, wherein The intelligent passenger interaction and feedback module includes: the system interacts with passengers through the passengers' mobile devices, providing real-time pick-up and drop-off point suggestions and route updates; passengers view the recommended pick-up and drop-off points, the real-time position of the vehicle, and the estimated arrival time through the mobile application; the system also helps passengers quickly find the vehicle parking position in a complex environment through augmented reality (AR) navigation; combining the visual SLAM algorithm and the camera data of the passenger device to provide accurate AR navigation; the formula for AR navigation is: T cw = K[Rt] Among them, T cw is the transformation matrix from camera coordinates to world coordinates, and the world coordinates are used to describe the absolute positions of objects in the scene; K is the camera intrinsic matrix, and R and t are the rotation matrix and the translation vector respectively; by determining the absolute positions of the objects, passengers, and the autonomous driving vehicle in the camera, the position and scene information are mapped into the AR scene, and the passengers navigate through the AR-rendered interface to find the pick-up and drop-off point positions of the autonomous driving vehicle.
Citation Information
Patent Citations
Real-time dynamic intelligent path planning method and system based on multi-sensor information fusion
CN116678394A
Path planning method, application and device based on knowledge and data combination
CN117808180A