A cloud-road collaborative learning method for traffic multimodal large language model
Through the cloud-road collaborative learning method, a large multimodal traffic language model deployed on roadside equipment and the cloud, using model compression and knowledge distillation mechanisms, solves the detection accuracy problem of the existing model in complex scenarios, and realizes efficient traffic condition recognition and target detection.
Patent Information
- Application Number
- CN202411757238.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-12-03
AI Technical Summary
In existing technologies, the target detection model based on convolutional neural networks does not perform well in complex scenarios, especially during peak hours and unusual traffic events. The recognition accuracy decreases, making it difficult to directly deploy and continuously optimize it on roadside equipment.
Using a cloud-road collaborative learning approach, a large multimodal traffic language model is pre-built and deployed in the cloud and on the roadside after an initial round of supervised training. Model compression is then deployed on roadside devices and in the cloud, while knowledge distillation is used to update and optimize parameters.
We have achieved the deployment of a large multimodal traffic language model on roadside equipment, and through continuous optimization, we have improved the model's performance, solved the problem of the model's detection accuracy in complex scenarios, and achieved efficient traffic condition recognition and target detection.
Smart Images

Figure CN119692471B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a cloud-road collaborative learning method for a large multimodal traffic language model. Background Art
[0002] Artificial Intelligence (AI) models have a variety of application scenarios in integrated vehicle-road-cloud systems, such as detecting traffic participants, identifying traffic conditions, and identifying traffic events / accidents. Currently, the models deployed in these scenarios mostly use object detection and event recognition models based on convolutional neural networks. These models are also called "small models" because of their small parameter size. We have found through practical application that, due to limitations in model structure and computing power, these "small models" do not perform well in complex scenarios. For example, the accuracy of object detection and event recognition decreases during rush hour, and they fail to recognize or make errors in unusual traffic events / accidents.
[0003] With the maturity and application of MultiModal Large Language Models (MM-LLMs) technology, if a traffic multimodal large language model suitable for traffic application scenarios can be trained based on a multimodal large language model base, then the "small model" problem mentioned above can be better solved. However, through application practice, we have found that the model parameter scale of the traffic multimodal large language model is too large, and it is unrealistic to deploy it directly on the roadside equipment. So, how to deploy the traffic multimodal large language model on the roadside equipment and continuously optimize it is the technical problem that the present invention needs to solve. Summary of the Invention
[0004] The purpose of the present invention is to address the shortcomings of existing technologies and provide a method, electronic device, and computer-readable storage medium for collaborative cloud-road learning of a large multimodal traffic language model. The present invention pre-constructs a large multimodal traffic language model and performs an initial round of supervised training on the model training dataset. After the initial round of training, the model is deployed in the cloud as a cloud-based heavy model. The model is then compressed and deployed simultaneously on the roadside and in the cloud as the corresponding roadside / cloud-based light model. After the deployment is completed, the roadside device will input the real-time first roadside multimodal data set and the first scene recognition instruction text into the roadside light model for visual target detection and traffic condition recognition processing to obtain the first target detection map and the first scene recognition text, and the first roadside multimodal data set, the first scene recognition instruction text, the first target detection map and the first scene recognition text will form the corresponding first roadside feedback and send it to the cloud; every time the cloud server receives a first roadside feedback, it will input the first roadside multimodal data set and the first scene recognition instruction text in the current feedback into the cloud heavy model for visual target detection and traffic condition recognition processing to obtain the corresponding second target detection map and the second scene recognition text to form the corresponding first cloud feedback, and the first roadside feedback and the first cloud feedback will form a roadside-cloud record and save it; the cloud server will also regularly form the corresponding first data set of all roadside-cloud records saved within the last specified time period, and update the cloud light model parameters according to the first data set according to the knowledge distillation method, and the parameter difference △θ of this update X Send to the roadside; each time the roadside equipment receives a parameter difference △θ X , based on the current parameter difference △θ X The roadside light model is updated once. In addition, the cloud server will regularly perform data enhancement processing on the model training data set based on the real scene data submitted by the roadside, and continuously fine-tune the cloud-based heavy model according to the newly added test data. Based on the cloud-road collaborative learning mechanism provided by the present invention: on the one hand, the purpose of deploying a large multimodal language model of traffic on the roadside can be achieved by means of model compression; on the other hand, the purpose of continuously optimizing the roadside deployment model can be achieved through the knowledge distillation mechanism between the heavy and light models on the cloud side; on the other hand, the purpose of continuously improving the performance of the cloud-based heavy model can be achieved by recycling real data.
[0005] To achieve the above objectives, a first aspect of an embodiment of the present invention provides a cloud-road collaborative learning method for a large multimodal traffic language model, the method comprising:
[0006] Construct a traffic multimodal large language model; and perform a first round of training on the traffic multimodal large language model based on a preset model training data set in a supervised training manner; and after the first round of training, deploy the traffic multimodal large language model on a cloud server and record it as a corresponding cloud-based heavy model; and perform model compression on the traffic multimodal large language model based on a preset pruning rule to obtain a corresponding compressed model; and deploy the compressed model simultaneously on a roadside device and the cloud server and record it as a corresponding roadside light model and a cloud-based light model; wherein the traffic multimodal large language model is used to perform visual target detection and traffic condition recognition processing based on the input roadside multimodal data set and scene recognition instruction text and output the corresponding target detection map and scene recognition text; the model parameter θ of the cloud-based heavy model M The model parameters θ of the compression model S The parameter compression relationship between them is recorded as F M-S ,θ S =F M-S (θ M );
[0007] The roadside device uses a preset scene recognition instruction template as a corresponding first scene recognition instruction text; and inputs a real-time first roadside multimodal data set and the first scene recognition instruction text into the roadside light model to perform visual target detection and traffic condition recognition processing to obtain a corresponding first target detection map and a first scene recognition text; and the first roadside multimodal data set, the first scene recognition instruction text, the first target detection map, and the first scene recognition text form a corresponding first roadside feedback and send it to the cloud server;
[0008] Each time the cloud server receives a first roadside feedback, it inputs the first roadside multimodal dataset and the first scene recognition instruction text in the current feedback into the cloud-based re-model to perform visual target detection and traffic condition recognition processing to obtain a corresponding second target detection map and a second scene recognition text to form a corresponding first cloud-based feedback; and the current first roadside feedback and the first cloud-based feedback form a corresponding roadside-cloud record and save it;
[0009] The cloud server regularly collects all the roadside-cloud records saved within the latest first preset time period into a corresponding first data set according to a preset first time frequency; and stores the current model parameters of the cloud light model as the pre-update parameters θ X ; and update the model parameters of the cloud light model according to the first data set according to the knowledge distillation method to obtain the corresponding updated parameters And based on the updated parameters θ X and the updated parameters Confirm the corresponding parameter difference And the parameter difference Δθ X Sending to the roadside equipment;
[0010] Each time the roadside equipment receives the parameter difference Δθ X , the current model parameters of the roadside light model are recorded as the parameters before update And based on the parameters before update and the parameter difference △θ X Confirm the corresponding updated parameters And based on the updated parameters The model parameters of the roadside light model are updated.
[0011] Preferably, the roadside multimodal dataset includes a roadside real-life image, a roadside video, a roadside BEV map, and a roadside point cloud; the scene recognition text includes at least a traffic condition description text, a traffic accident / incident description text, and a target detection description text; the target detection image is an image with one or more target boxes marked on the roadside real-life image; each target box carries a target type text;
[0012] The scene recognition instruction text is a formatted instruction text of the traffic multimodal large language model; the scene recognition instruction text is used to instruct the traffic multimodal large language model to perform target detection based on the currently input roadside multimodal data set, and to identify roadside traffic conditions and traffic accidents / incidents, and to generate scene recognition text based on the target detection results and the traffic condition and accident / incident identification results, and to output the generated scene recognition text and the target detection map corresponding to the target detection results as a model processing result;
[0013] The model training dataset consists of multiple multimodal data records; the multimodal data records include a training-multimodal dataset, a training-instruction text, a label-target detection map, and a label-scene recognition text; the training multimodal dataset includes a training-roadside real-life image, a training-roadside video, a training-roadside BEV map, and a training-roadside point cloud; the training-instruction text meets the formatting requirements of the scene recognition instruction text;
[0014] The roadside device is interconnected with the cloud server; the roadside device is also interconnected with multiple roadside sensors, and the multiple roadside sensors include at least a first camera, a second camera and a first laser radar; the roadside device also has a built-in high-precision map module; the first camera is used to regularly output the latest first roadside real-shot picture of the current road / intersection to the roadside device at a preset first data update frequency; the second camera is used to regularly output the latest first roadside video of the current road / intersection to the roadside device at a preset second data update frequency; the high-precision map module is used to regularly output the latest first roadside BEV map of the current road / intersection to the roadside device at a preset third data update frequency; the first laser radar is used to regularly output the latest first roadside point cloud of the current road / intersection to the roadside device at a preset fourth data update frequency;
[0015] The real-time first roadside multimodal dataset consists of the latest first roadside real-shot image, the first roadside video, the first roadside BEV map and the first roadside point cloud.
[0016] Preferably, the traffic multimodal large language model includes a visual target detection model, a multimodal encoder, an input projector, a text feature encoder, a large language model, and a detection map output module; the multimodal encoder includes a detection box encoder, an image encoder, a video encoder, a BEV encoder, and a point cloud encoder; the input projector includes a detection box-text feature space projector, an image-text feature space projector, a video-text feature space projector, a BEV-text feature space projector, and a point cloud-text feature space projector;
[0017] The first, second, third and fourth model input ends of the traffic multimodal large language model are respectively used to receive the corresponding roadside real-life pictures, the roadside videos, the roadside BEV maps and the roadside point clouds in the input roadside multimodal dataset, the fifth model input end is used to receive the input scene recognition instruction text, and the first and second model output ends are used to output the corresponding target detection pictures and the scene recognition text;
[0018] The input end of the visual object detection model is connected to the input end of the first model, and the output end is respectively connected to the input end of the detection box encoder and the first input end of the detection map output module; the second input end of the detection map output module is connected to the input end of the first model, and the output end is connected to the output end of the first model;
[0019] The output end of the detection frame encoder is connected to the input end of the detection frame-text feature space projector; the output end of the detection frame-text feature space projector is connected to the first input end of the large language model;
[0020] The input end of the image encoder is connected to the input end of the first model, and the output end is connected to the input end of the image-text feature space projector; the output end of the image-text feature space projector is connected to the second input end of the large language model;
[0021] The input end of the video encoder is connected to the input end of the second model, and the output end is connected to the input end of the video-text feature space projector; the output end of the video-text feature space projector is connected to the third input end of the large language model;
[0022] The input end of the BEV encoder is connected to the input end of the third model, and the output end is connected to the input end of the BEV-text feature space projector; the output end of the BEV-text feature space projector is connected to the fourth input end of the large language model;
[0023] The input end of the point cloud encoder is connected to the input end of the fourth model, and the output end is connected to the input end of the point cloud-text feature space projector; the output end of the point cloud-text feature space projector is connected to the fifth input end of the large language model;
[0024] The input end of the text feature encoder is connected to the input end of the fifth model, and the output end is connected to the sixth input end of the large language model;
[0025] The output end of the large language model is connected to the output end of the second model; the large language model is also connected to an external traffic knowledge base.
[0026] Preferably, the visual target detection model is implemented based on a class of conventional target detection models; the conventional target detection models include at least RCNN series models and YOLO series models; the visual target detection model is used to perform target detection processing on the roadside real-shot image input to the model to obtain one or more corresponding target detection frames to form a corresponding target detection frame sequence and send it to the detection image output module and the detection frame encoder; each of the target detection frames includes at least the coordinates of the detection frame center point, the detection frame size, the detection frame orientation and the target type;
[0027] The detection map output module is used to draw the corresponding target frame on the roadside real shot image input by the model according to the detection frame center point coordinates, the detection frame size and the detection frame orientation of each target detection frame in the target detection frame sequence; and set a corresponding display text at a specified position on each target frame, and set the display content of the corresponding display text based on the target type text corresponding to each target frame; and output the roadside real shot image after the target frame drawing and display text setting as the corresponding target detection map;
[0028] The detection frame encoder is used to perform detection frame feature encoding on the target detection frame sequence to obtain a corresponding first encoding tensor and send it to the detection frame-text feature space projector; the image encoder is used to perform image feature encoding on the roadside real-shot image input by the model to obtain a corresponding second encoding tensor and send it to the image-text feature space projector; the video encoder is used to perform video feature encoding on the roadside video input by the model to obtain a corresponding third encoding tensor and send it to the video-text feature space projector; the BEV encoder is used to perform BEV feature encoding on the roadside BEV map input by the model to obtain a corresponding fourth encoding tensor and send it to the BEV-text feature space projector; the point cloud encoder is used to perform point cloud feature encoding on the roadside point cloud input by the model to obtain a corresponding fifth encoding tensor and send it to the point cloud-text feature space projector;
[0029] The detection box-text feature space projector is used to perform feature projection processing on the first coding tensor from the detection box feature space to the text feature space to obtain a corresponding first projection feature tensor and send it to the large language model; the image-text feature space projector is used to perform feature projection processing on the second coding tensor from the image feature space to the text feature space to obtain a corresponding second projection feature tensor and send it to the large language model; the video-text feature space projector is used to perform feature projection processing on the third coding tensor from the video feature space to the text feature space to obtain a corresponding third projection feature tensor and send it to the large language model; the BEV-text feature space projector is used to perform feature projection processing on the fourth coding tensor from the BEV feature space to the text feature space to obtain a corresponding fourth projection feature tensor and send it to the large language model; the point cloud-text feature space projector is used to perform feature projection processing on the fifth coding tensor from the point cloud feature space to the text feature space to obtain a corresponding fifth projection feature tensor and send it to the large language model;
[0030] The text feature encoder is used to perform text feature encoding on the scene recognition instruction text input by the model to obtain a corresponding first text feature tensor and send it to the large language model;
[0031] The large language model is a type of large language model implemented based on the Transformer framework structure, including at least a GPT series model; the large language model has completed model pre-training; the large language model is used to combine the first, second, third, fourth and fifth projection feature tensors and the first text feature tensor of the input into a corresponding first input tensor; and perform target detection result recognition, traffic road condition recognition and traffic accident / incident recognition based on the first input tensor and the traffic knowledge base, and perform scene recognition text generation processing based on the recognition results to obtain the corresponding scene recognition text and output it.
[0032] Preferably, the first round of training of the traffic multimodal large language model based on a preset model training data set in a supervised training manner specifically includes:
[0033] Step 51: Split the model training data set into two sub-data sets according to a preset first split ratio and record them as a corresponding first training set and a first evaluation set;
[0034] Wherein, both the first training set and the first evaluation set are composed of a plurality of the multimodal data records; the ratio of the total number of records in the first training set to the first evaluation set satisfies the first segmentation ratio;
[0035] Step 52: taking the first multimodal data record of the first training set as the corresponding current training record;
[0036] Step 53: Input the training-multimodal dataset and the training-instruction text of the current training record as the corresponding roadside multimodal dataset and the scene recognition instruction text into the traffic multimodal large language model to perform visual target detection and traffic condition recognition processing, and use the output target detection image and the scene recognition text as the corresponding first predicted detection image and first predicted text;
[0037] Step 54: The first prediction detection image and the label-target detection image of the current training record form a corresponding first prediction-label pair; and the first prediction text and the label-scene recognition text of the current training record form a corresponding second prediction-label pair; the first and second prediction-label pairs are subjected to a preset model loss function; and based on a preset first model optimizer, a round of parameter modulation is performed on the traffic multimodal large language model in a direction that minimizes the model loss function.
[0038] Wherein, the first model optimizer includes at least SGD optimizer, Adam optimizer, RMSprop optimizer, and Nadam optimizer;
[0039] Step 55: Identify whether the current training record is the last multimodal data record of the first training set; if so, proceed to step 56; if not, use the next multimodal data record of the first training set as the new current training record and return to step 53;
[0040] Step 56: Perform a round of traversal on all the multimodal data records of the first evaluation set; and during this round of traversal, use the multimodal data record currently traversed as the corresponding current evaluation record; and input the training-multimodal dataset and the training-instruction text of the current evaluation record as the corresponding roadside multimodal dataset and the scene recognition instruction text into the traffic multimodal large language model for visual target detection and traffic condition recognition processing, and use the output target detection map and the scene recognition text as the corresponding second prediction detection map and second prediction text; form a corresponding third prediction-label pair from the second prediction detection map and the label-target detection map of the current evaluation record; and form a corresponding fourth prediction-label pair from the second prediction text and the label-scene recognition text of the current evaluation record;
[0041] In step 57, after this round of traversal is completed, a comprehensive evaluation is performed on the target detection performance and scene recognition text generation quality of the traffic multimodal large language model based on all the third and fourth prediction-label pairs obtained to obtain a corresponding first evaluation result; and when the first evaluation result does not meet the preset compliance requirements, return to step 52 to continue training; and when the first evaluation result meets the compliance requirements, the model training is terminated.
[0042] Preferably, the knowledge distillation method is used to update the model parameters of the cloud-based light model according to the first data set to obtain the corresponding updated parameters θ * X , specifically including:
[0043] Step 61: Introduce multiple low-rank adapters into the cloud-based light model to form a corresponding current student model; and use the cloud-based heavy model as the corresponding current teacher model;
[0044] The model parameters of the current student model are composed of the model parameters of the cloud-based light model and the model parameters of all the low-rank adapters;
[0045] Step 62: Using each of the roadside-cloud records of the first data set as a corresponding current record; and inputting the first roadside multimodal data set and the first scene recognition instruction text of the current record into the current student model for visual target detection and traffic condition recognition processing to obtain a corresponding third target detection map and third scene recognition text; and using the second target detection map and the second scene recognition text of the current record as the corresponding current label detection map and current label recognition text, forming a corresponding fifth prediction-label pair with the third target detection map and the current label detection map, and forming a corresponding sixth prediction-label pair with the third scene recognition text and the current label recognition text;
[0046] Step 63: Substitute all the fifth and sixth prediction-label pairs obtained into a preset knowledge distillation loss function to calculate a corresponding first loss value;
[0047] Step 64: Identify whether the first loss value satisfies a preset first loss value range; if so, proceed to step 65; if not, modulate the model parameters of all the low-rank adapters based on a preset second model optimizer in a direction that minimizes the knowledge distillation loss function, while keeping the model parameters of the cloud-based light model of the current student model unchanged, and return to step 62 after this round of modulation.
[0048] Wherein, the second model optimizer includes at least an SGD optimizer and an Adam optimizer;
[0049] Step 65: Update the model parameters of the cloud-based light model based on the model parameters of the current student model; and use the updated model parameters of the cloud-based light model as the corresponding updated parameters θ * X ; and remove all low-rank adapters introduced in the cloud-based light model.
[0050] Preferably, the method further comprises:
[0051] The cloud server regularly takes each of the roadside-cloud records saved within the most recent second preset time period as the corresponding current record at a preset second time frequency; and takes the first roadside real-shot image, the first roadside video, the first roadside BEV map, the first roadside point cloud, the first scene recognition instruction text, the second target detection image and the second scene recognition text of the currently recorded image as the corresponding current roadside real-shot image, the current roadside video, the current roadside BEV map, the current roadside point cloud, the current scene recognition instruction text, the current target detection image and the current scene recognition text; and based on the preset image, video and point cloud noise addition algorithm / model, noise addition processing is performed in the corresponding current roadside real-shot image, the current roadside video and the current roadside point cloud to obtain the corresponding noisy roadside real-shot image, noisy roadside video and noisy roadside point cloud. Roadside point cloud; and the noisy roadside real shot picture, the noisy roadside video, the current roadside BEV map, and the noisy roadside point cloud corresponding to the current record are used as the corresponding training-roadside real shot picture, the training-roadside video, the training-roadside BEV map, and the training-roadside point cloud to form a corresponding training multimodal data set; and the current scene recognition instruction text, the current target detection map, and the current scene recognition text corresponding to the current record are used as the corresponding training-instruction text, the label-target detection map, and the label-scene recognition text; and the training multimodal data set, the training-instruction text, the label-target detection map, and the label-scene recognition text corresponding to the current record form a corresponding multimodal data record and are stored in the model training data set;
[0052] At the end of the first round of training, the cloud server records all the current multimodal data records in the model training data set as used records; and each time a new multimodal data record is added thereafter, the newly added multimodal data record is recorded as an unused record; and when the total number of unused records in the model training data set reaches a preset threshold for the total number of new records, all the current unused records in the model training data set are extracted to form a corresponding current training data set; and the cloud-based re-model is fine-tuned based on the current training data set; and at the end of this fine-tuning, all the multimodal data records in the model training data set corresponding to the current training data set are transferred to used records.
[0053] A second aspect of an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0054] The processor is configured to be coupled to the memory, read and execute instructions in the memory, so as to implement the method steps described in the first aspect above;
[0055] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
[0056] A third aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a computer, the computer executes the instructions of the method described in the first aspect.
[0057] Embodiments of the present invention provide a cloud-road collaborative learning method, electronic device, and computer-readable storage medium for a large multimodal traffic language model. As can be seen from the foregoing invention, embodiments of the present invention pre-construct a large multimodal traffic language model and perform an initial round of supervised training on the model training dataset. After the initial round of training, the model is deployed in the cloud as a cloud-based heavy model. The model is then compressed and deployed simultaneously on the roadside and in the cloud as the corresponding roadside / cloud-based light model. After the deployment is completed, the roadside device will input the real-time first roadside multimodal data set and the first scene recognition instruction text into the roadside light model for visual target detection and traffic condition recognition processing to obtain the first target detection map and the first scene recognition text, and the first roadside multimodal data set, the first scene recognition instruction text, the first target detection map and the first scene recognition text will form the corresponding first roadside feedback and send it to the cloud; every time the cloud server receives a first roadside feedback, it will input the first roadside multimodal data set and the first scene recognition instruction text in the current feedback into the cloud heavy model for visual target detection and traffic condition recognition processing to obtain the corresponding second target detection map and the second scene recognition text to form the corresponding first cloud feedback, and the first roadside feedback and the first cloud feedback will form a roadside-cloud record and save it; the cloud server will also regularly form the corresponding first data set of all roadside-cloud records saved within the last specified time period, and update the cloud light model parameters according to the first data set according to the knowledge distillation method, and the parameter difference △θ of this update X Send to the roadside; each time the roadside equipment receives a parameter difference △θ X , based on the current parameter difference △θ X The roadside light model is updated once. In addition, the cloud server will regularly perform data enhancement processing on the model training data set based on the real scene data submitted by the roadside, and continuously fine-tune the cloud-based heavy model according to the newly added test data. Based on the cloud-road collaborative learning mechanism given in the embodiment of the present invention: on the one hand, the purpose of deploying a large multimodal language model of traffic on the roadside is achieved through model compression; on the other hand, the purpose of continuously optimizing the roadside deployment model is achieved through the knowledge distillation mechanism between the heavy and light models on the cloud side; on the other hand, the purpose of continuously improving the performance of the cloud-based heavy model is achieved through the recycling of real data. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 A schematic diagram of a cloud-road collaborative learning method for a large multimodal traffic language model provided in Example 1 of the present invention;
[0059] Figure 2 This is a module structure diagram of the traffic multimodal language model provided in Example 1 of the present invention;
[0060] Figure 3 This is a structural diagram of an electronic device provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0061] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the embodiments described herein are merely some, rather than all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0062] The first embodiment of the present invention provides a cloud-road collaborative learning method for a traffic multimodal large language model, such as Figure 1 A schematic diagram of a cloud-road collaborative learning method for a traffic multimodal large language model provided in Example 1 of the present invention includes the following main steps:
[0063] Step 1: Construct a large multimodal traffic language model; perform a first round of training on the large multimodal traffic language model based on a preset model training dataset using a supervised training method; after the first round of training, deploy the large multimodal traffic language model on a cloud server as a corresponding cloud-based heavy model; perform model compression on the large multimodal traffic language model based on preset pruning rules to obtain a corresponding compressed model; and deploy the compressed model simultaneously on a roadside device and a cloud server as corresponding roadside light models and cloud-based light models;
[0064] Specifically including: step 11, constructing a large multimodal traffic language model;
[0065] Here, the traffic multimodal large language model of the embodiment of the present invention is used to perform visual target detection and traffic condition recognition processing based on the input roadside multimodal dataset and scene recognition instruction text and output corresponding target detection map and scene recognition text;
[0066] The roadside multimodal dataset here includes roadside real-life images, roadside videos, roadside BEV maps, and roadside point clouds; the scene recognition text includes at least traffic condition description text, traffic accident / incident description text, and target detection description text; the target detection image is an image with one or more target boxes marked on the roadside real-life image; each target box is accompanied by a target type text;
[0067] The scene recognition instruction text here is a formatted instruction text of the traffic multimodal large language model; the scene recognition instruction text is used to instruct the traffic multimodal large language model to perform target detection based on the currently input roadside multimodal dataset, identify roadside traffic conditions and traffic accidents / incidents, generate scene recognition text based on the target detection results and the traffic condition and accident / incident identification results, and output the generated scene recognition text and the target detection map corresponding to the target detection results as the model processing result;
[0068] The traffic multimodal large language model of the embodiment of the present invention includes a visual target detection model, a multimodal encoder, an input projector, a text feature encoder, a large language model, and a detection map output module; wherein the multimodal encoder includes a detection box encoder, an image encoder, a video encoder, a BEV encoder, and a point cloud encoder; the input projector includes a detection box-text feature space projector, an image-text feature space projector, a video-text feature space projector, a BEV-text feature space projector, and a point cloud-text feature space projector; for details, please refer to Figure 2 The module structure diagram of the traffic multimodal large language model provided in Example 1 of the present invention is provided for understanding. It should be noted that the implementation framework of the traffic multimodal large language model is a relatively general MM-LLMs implementation framework. For details, please refer to the technical document "MM-LLMs: Recent Advances in MultiModal Large Language Models" for understanding.
[0069] like Figure 2 As shown, the first, second, third and fourth model input ends of the traffic multimodal large language model of the embodiment of the present invention are respectively used to receive the corresponding roadside real-life pictures, roadside videos, roadside BEV maps and roadside point clouds in the input roadside multimodal dataset, the fifth model input end is used to receive the input scene recognition instruction text, and the first and second model output ends are used to output the corresponding target detection pictures and scene recognition text;
[0070] like Figure 2 As shown, the connection relationship between the components of the traffic multimodal language model according to the embodiment of the present invention is as follows:
[0071] 1) The input end of the visual object detection model is connected to the input end of the first model, and the output end is respectively connected to the input end of the detection frame encoder and the first input end of the detection image output module; the second input end of the detection image output module is connected to the input end of the first model, and the output end is connected to the output end of the first model; 2) The output end of the detection frame encoder is connected to the input end of the detection frame-text feature space projector; the output end of the detection frame-text feature space projector is connected to the first input end of the large language model; 3) The input end of the image encoder is connected to the input end of the first model, and the output end is connected to the input end of the image-text feature space projector; the output end of the image-text feature space projector is connected to the second input end of the large language model; 4) The input end of the video encoder is connected to the input end of the second model, and the output end is connected to the video-text feature space projector The input end of the video-text feature space projector is connected to the third input end of the large language model; 5) the input end of the BEV encoder is connected to the third model input end, and the output end is connected to the input end of the BEV-text feature space projector; the output end of the BEV-text feature space projector is connected to the fourth input end of the large language model; 6) the input end of the point cloud encoder is connected to the fourth model input end, and the output end is connected to the input end of the point cloud-text feature space projector; the output end of the point cloud-text feature space projector is connected to the fifth input end of the large language model; 7) the input end of the text feature encoder is connected to the fifth model input end, and the output end is connected to the sixth input end of the large language model; 8) the output end of the large language model is connected to the output end of the second model; the large language model is also connected to the external traffic knowledge base;
[0072] The functions of the components of the traffic multimodal language model according to the embodiment of the present invention are briefly described as follows:
[0073] 1) Visual object detection model:
[0074] The visual target detection model of the embodiment of the present invention is implemented based on a class of conventional target detection models; such conventional target detection models include at least the RCNN series model and the YOLO series model. The visual target detection model is used to perform target detection processing on the roadside real-life image input to the model to obtain one or more corresponding target detection frames to form a corresponding target detection frame sequence, which is sent to the detection image output module and the detection frame encoder. Each target detection frame includes at least the coordinates of the detection frame center point, the detection frame size, the detection frame orientation, and the target type.
[0075] 2) Detection image output module:
[0076] The detection image output module of the embodiment of the present invention is used to draw corresponding target frames on the roadside real-shot image input by the model based on the detection frame center point coordinates, detection frame size, and detection frame orientation of each target detection frame in the target detection frame sequence; set a corresponding display text at a specified position on each target frame, and set the display content of the corresponding display text based on the target type text corresponding to each target frame; and output the roadside real-shot image after the target frame drawing and display text setting as the corresponding target detection image;
[0077] 3) Multimodal encoder:
[0078] In the multimodal encoder of the embodiment of the present invention: a) the detection frame encoder is used to perform detection frame feature encoding processing on the target detection frame sequence to obtain a corresponding first encoding tensor and send it to the detection frame-text feature space projector; b) the image encoder is used to perform image feature encoding processing on the roadside real shot image input by the model to obtain a corresponding second encoding tensor and send it to the image-text feature space projector; c) the video encoder is used to perform video feature encoding processing on the roadside video input by the model to obtain a corresponding third encoding tensor and send it to the video-text feature space projector; d) the BEV encoder is used to perform BEV feature encoding processing on the roadside BEV map input by the model to obtain a corresponding fourth encoding tensor and send it to the BEV-text feature space projector; e) the point cloud encoder is used to perform point cloud feature encoding processing on the roadside point cloud input by the model to obtain a corresponding fifth encoding tensor and send it to the point cloud-text feature space projector;
[0079] It should be noted that the five types of encoders in the multimodal encoder can be customized based on application requirements. You can also refer to the optional encoders given in the Modality Encoder section of the technical document "MM-LLMs: Recent Advances in MultiModal Large Language Models" for configuration;
[0080] 4) Input projector:
[0081] In the input projector of the embodiment of the present invention: a) the detection box-text feature space projector is used to perform feature projection processing on the first coding tensor from the detection box feature space to the text feature space to obtain the corresponding first projected feature tensor and send it to the large language model; b) the image-text feature space projector is used to perform feature projection processing on the second coding tensor from the image feature space to the text feature space to obtain the corresponding second projected feature tensor and send it to the large language model; c) the video-text feature space projector is used to perform feature projection processing on the third coding tensor from the video feature space to the text feature space to obtain the corresponding third projected feature tensor and send it to the large language model; d) the BEV-text feature space projector is used to perform feature projection processing on the fourth coding tensor from the BEV feature space to the text feature space to obtain the corresponding fourth projected feature tensor and send it to the large language model; e) the point cloud-text feature space projector is used to perform feature projection processing on the fifth coding tensor from the point cloud feature space to the text feature space to obtain the corresponding fifth projected feature tensor and send it to the large language model;
[0082] It should be noted that the five types of projectors in the multimodal encoder can be customized based on application requirements. You can also refer to the optional projectors given in the input projector section of the technical document "MM-LLMs: Recent Advances in MultiModal Large Language Models" for configuration;
[0083] 5) Text feature encoder:
[0084] The text feature encoder of the embodiment of the present invention is used to perform text feature encoding processing on the scene recognition instruction text input by the model to obtain a corresponding first text feature tensor and send it to the large language model;
[0085] Here, the text feature encoder of the embodiment of the present invention can be a simple word segmentation and word embedding encoding module, or can be configured by selecting one of the various text encoders mentioned in the technical document "MM-LLMs: Recent Advances in MultiModal Large Language Models".
[0086] 6) Large Language Models (LLMs):
[0087] The large language model of the embodiment of the present invention is a large language model implemented based on the Transformer framework structure. This type of LLMs includes at least the GPT series model. The embodiment of the present invention configures a traffic knowledge base for the large language model and maintains continuous updates to the traffic knowledge base. The large language model of the embodiment of the present invention has completed model pre-training.
[0088] The large language model of the embodiment of the present invention is used to combine the input first, second, third, fourth, and fifth projection feature tensors and the first text feature tensor into a corresponding first input tensor; and perform target detection result recognition, traffic road condition recognition, and traffic accident / incident recognition based on the first input tensor and the traffic knowledge base, and perform scene recognition text generation processing based on the recognition results to obtain corresponding scene recognition text and output it;
[0089] Step 12: Perform the first round of training on the traffic multimodal large language model based on the preset model training dataset in a supervised training manner;
[0090] The model training dataset consists of multiple multimodal data records; the multimodal data records include training-multimodal dataset, training-instruction text, label-target detection map, and label-scene recognition text; the training multimodal dataset includes training-roadside real-life images, training-roadside videos, training-roadside BEV maps, and training-roadside point clouds; the training-instruction text meets the formatting requirements of scene recognition instruction text;
[0091] Specifically comprising: step 121, dividing the model training data set into two sub-data sets according to a preset first division ratio and recording them as a corresponding first training set and a first evaluation set;
[0092] The first split ratio is a preset ratio parameter, such as 8:2; the first training set and the first evaluation set are both composed of multiple multimodal data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first split ratio;
[0093] Step 122: using the first multimodal data record of the first training set as the corresponding current training record;
[0094] Step 123: Input the training-multimodal dataset and training-command text of the current training record as the corresponding roadside multimodal dataset and scene recognition command text into the traffic multimodal large language model for visual target detection and traffic condition recognition processing, and use the output target detection image and scene recognition text as the corresponding first predicted detection image and first predicted text;
[0095] Step 124: A first prediction-label pair is formed from the first prediction detection image and the label-target detection image of the current training record; a second prediction-label pair is formed from the first prediction text and the label-scene recognition text of the current training record; the first and second prediction-label pairs are subjected to a preset model loss function; and a parameter modulation round is performed on the traffic multimodal large language model based on a preset first model optimizer in a direction that minimizes the model loss function.
[0096] Here, the model loss function of the embodiment of the present invention is a preset loss function, which can be customized based on application requirements; the first model optimizer includes at least an SGD optimizer, an Adam optimizer, an RMSprop optimizer, and a Nadam optimizer;
[0097] Step 125: Identify whether the current training record is the last multimodal data record of the first training set; if so, go to step 126; if not, use the next multimodal data record of the first training set as the new current training record and return to step 123;
[0098] Step 126: Perform a round of traversal on all multimodal data records of the first evaluation set; and during this round of traversal, use the currently traversed multimodal data record as the corresponding current evaluation record; and input the training-multimodal dataset and training-instruction text of the current evaluation record as the corresponding roadside multimodal dataset and scene recognition instruction text into the traffic multimodal large language model to perform visual target detection and traffic condition recognition processing, and use the output target detection image and scene recognition text as the corresponding second predicted detection image and second predicted text; form a corresponding third prediction-label pair from the second predicted detection image and the label-target detection image of the current evaluation record; and form a corresponding fourth prediction-label pair from the second predicted text and the label-scene recognition text of the current evaluation record;
[0099] At step 127, after this round of traversal, a comprehensive evaluation is performed on the object detection performance and the scene recognition text generation quality of the traffic multimodal large language model based on all the third and fourth prediction-label pairs obtained to obtain a corresponding first evaluation result; and if the first evaluation result does not meet the preset compliance requirement, the training is returned to step 122 to continue; and if the first evaluation result meets the compliance requirement, the model training is terminated;
[0100] Here, the evaluation rule for comprehensively evaluating the object detection performance and scene recognition text generation quality of the traffic multimodal large language model is a pre-set evaluation rule that can be customized based on application needs. The compliance requirement is actually a series of threshold relationships corresponding to this evaluation rule.
[0101] In step 13, after the first round of training, the traffic multimodal large language model is deployed on the cloud server and recorded as the corresponding cloud heavy model; and the traffic multimodal large language model is compressed based on the preset pruning rules to obtain the corresponding compressed model; and the compressed model is deployed on the roadside equipment and the cloud server at the same time and recorded as the corresponding roadside light model and cloud light model.
[0102] In this embodiment of the present invention, the roadside equipment is interconnected with a cloud server. The roadside equipment is also interconnected with multiple roadside sensors, which include at least a first camera, a second camera, and a first laser radar. The roadside equipment also includes a built-in high-precision map module. The first camera is configured to periodically output the latest first roadside real-time image of the current road / intersection to the roadside equipment at a preset first data update frequency. The second camera is configured to periodically output the latest first roadside video of the current road / intersection to the roadside equipment at a preset second data update frequency. The high-precision map module is configured to periodically output the latest first roadside BEV map of the current road / intersection to the roadside equipment at a preset third data update frequency. The first laser radar is configured to periodically output the latest first roadside point cloud of the current road / intersection to the roadside equipment at a preset fourth data update frequency. The first, second, third, and fourth data update frequencies are four customizable time frequency parameters. The real-time first roadside multimodal dataset obtained by the roadside equipment consists of the latest first roadside real-time image, first roadside video, first roadside BEV map, and first roadside point cloud.
[0103] It should be noted that the embodiment of the present invention achieves the purpose of compressing the model through model pruning: the pruning rules of the embodiment of the present invention can also be implemented based on a variety of conventional methods, such as weight pruning (removing connections with weight values close to zero), neuron pruning (removing neurons that have little impact on model output), and structured pruning (removing the entire convolution kernel or channel, which usually requires retraining the model to restore performance). After pruning is completed, the model parameters θ of the cloud-based re-model are M and the model parameters θ of the compression model S The parameter compression relationship between them is recorded as F M-S ,θ S =F M-S (θ M ).
[0104] In step 2, the roadside device uses the preset scene recognition instruction template as the corresponding first scene recognition instruction text; and inputs the real-time first roadside multimodal data set and the first scene recognition instruction text into the roadside light model for visual target detection and traffic condition recognition processing to obtain the corresponding first target detection map and first scene recognition text; and the first roadside multimodal data set, the first scene recognition instruction text, the first target detection map and the first scene recognition text form the corresponding first roadside feedback and send it to the cloud server.
[0105] Here, the scene recognition instruction template of the embodiment of the present invention is a standard instruction text whose instruction format meets the formatting requirements of the scene recognition instruction text and whose content is pre-set. When the real-time first roadside multimodal data set and the first scene recognition instruction text are input into the roadside light model for visual target detection and traffic condition recognition processing to obtain the corresponding first target detection map and first scene recognition text, the first roadside multimodal data set and the first scene recognition instruction text are actually input into the roadside light model as the roadside multimodal data set and scene recognition instruction text input by the current model, and the roadside light model performs visual target detection and traffic condition recognition processing based on the roadside multimodal data set and scene recognition instruction text input this time, and uses the target detection map and scene recognition text output by the current model as the corresponding first target detection map and first scene recognition text.
[0106] In step 3, each time the cloud server receives a first roadside feedback, it inputs the first roadside multimodal data set and the first scene recognition instruction text in the current feedback into the cloud re-model for visual target detection and traffic condition recognition processing to obtain the corresponding second target detection image and the second scene recognition text to form the corresponding first cloud feedback; and the current first roadside feedback and the first cloud feedback form a corresponding roadside-cloud record and save it.
[0107] Step 4: The cloud server regularly collects all roadside-cloud records saved within the latest first preset time period into a corresponding first data set according to the preset first time frequency; and saves the current model parameters of the cloud light model as the pre-update parameters θ X ; And according to the knowledge distillation method, the model parameters of the cloud light model are updated according to the first data set to obtain the corresponding updated parameters And based on the updated parameters θ X and updated parameters Confirm the corresponding parameter difference And the parameter difference △θ X Send to roadside equipment;
[0108] Specifically comprising: step 41, the cloud server regularly compiles all roadside-cloud records saved within a recent first preset time period into a corresponding first data set according to a preset first time frequency;
[0109] Here, the first time frequency is a preset time frequency parameter; the first preset duration is a preset time length parameter;
[0110] Step 42, and save the current model parameters of the cloud light model as the parameters before update θ X ;
[0111] Step 43, and update the model parameters of the cloud light model according to the first data set using the knowledge distillation (KD) method to obtain the corresponding updated parameters
[0112] Specifically, the method includes: step 431, introducing multiple low-rank adapters (LRAs) into the cloud-based light model to form a corresponding current student model; and using the cloud-based heavy model as the corresponding current teacher model;
[0113] Here, the model parameters of the current student model are composed of the model parameters of the cloud-based light model and the model parameters of all low-rank adapters;
[0114] Step 432: Each roadside-cloud record of the first data set is used as a corresponding current record; the first roadside multimodal data set and the first scene recognition instruction text of the current record are input into the current student model for visual object detection and traffic condition recognition processing to obtain a corresponding third object detection map and third scene recognition text; the second object detection map and the second scene recognition text of the current record are used as the corresponding current label detection map and current label recognition text, and the third object detection map and the current label detection map form a corresponding fifth prediction-label pair, and the third scene recognition text and the current label recognition text form a corresponding sixth prediction-label pair;
[0115] Step 433: Substitute all the fifth and sixth prediction-label pairs obtained into the preset knowledge distillation loss function to calculate the corresponding first loss value;
[0116] Here, the knowledge distillation loss function is a pre-set loss function that can be customized based on application requirements;
[0117] In step 434, it is determined whether the first loss value satisfies a preset first loss value range. If so, the process proceeds to step 435. If not, the model parameters of the cloud-based light model of the current student model remain unchanged, and the model parameters of all low-rank adapters are modulated in a round based on a preset second model optimizer in a direction that minimizes the knowledge distillation loss function. After this round of modulation, the process returns to step 432.
[0118] The first loss value range is a preset loss value range; the second model optimizer includes at least an SGD optimizer and an Adam optimizer;
[0119] Step 435: Update the model parameters of the cloud-based light model based on the model parameters of the current student model; and use the updated model parameters of the cloud-based light model as the corresponding updated parameters. And remove all low-rank adapters introduced in the cloud-based light model;
[0120] Step 44, and based on the updated parameters θ X and updated parameters Confirm the corresponding parameter difference And the parameter difference △θ X Sent to roadside equipment.
[0121] Step 5: Every time the roadside equipment receives a parameter difference △θ X , the current model parameters of the roadside light model are recorded as the parameters before update And based on the parameters before update and parameter difference △θ X Confirm the corresponding updated parameters And based on the updated parameters Update the model parameters of the roadside light model.
[0122] Here, through steps 2-5 above, we can see a cloud-road collaborative learning closed loop that starts from the roadside to the cloud and then returns to the roadside. Through this closed loop, the roadside lightweight model can be continuously optimized without local training.
[0123] It should also be noted that the embodiment of the present invention will also regularly perform data enhancement processing on the model training data set based on the real scene data submitted by the roadside. The specific implementation steps are as follows: the cloud server of the embodiment of the present invention regularly uses the various roadside-cloud records saved within the most recent second preset time period as the corresponding current records according to the preset second time frequency; and uses the currently recorded first roadside real-shot picture, first roadside video, first roadside BEV map, first roadside point cloud, first scene recognition instruction text, second target detection picture and second scene recognition text as the corresponding current roadside real-shot picture, current roadside video, current roadside BEV map, current roadside point cloud, current scene recognition instruction text, current target detection picture and current scene recognition text; and based on the preset image, video and point cloud noise addition algorithm / model, in the corresponding current roadside real-shot picture , add noise to the current roadside video and the current roadside point cloud to obtain the corresponding noisy roadside real-shot picture, noisy roadside video and noisy roadside point cloud; and the noisy roadside real-shot picture, noisy roadside video, current roadside BEV map, and noisy roadside point cloud corresponding to the current record are used as the corresponding training-roadside real-shot picture, training-roadside video, training-roadside BEV map, and training-roadside point cloud to form a corresponding training multimodal dataset; and the current scene recognition instruction text, current target detection map, and current scene recognition text corresponding to the current record are used as the corresponding training-instruction text, label-target detection map, and label-scene recognition text; and the training multimodal dataset, training-instruction text, label-target detection map, and label-scene recognition text corresponding to the current record form a corresponding multimodal data record and are stored in the model training dataset.
[0124] Here, the second time frequency is a preset time frequency parameter.
[0125] It should also be noted that embodiments of the present invention may also utilize an AIGC (Artificial Intelligence Generated Content) model implemented by a multimodal retrieval-augmented generation (RAG) to query a pre-set multimodal knowledge base based on real-world feedback, i.e., multimodal information in roadside-cloud records, to generate a new set of highly simulated multimodal datasets to achieve data enhancement. Embodiments of the present invention may also utilize another large content generation model from text to image / video / point cloud / BEV to construct a highly simulated multimodal dataset of difficult cases that are difficult to collect in real life to achieve data enhancement. Such difficult cases include cases of malicious / major events / accidents, various types of critical state cases (corner cases), and the like.
[0126] It should also be noted that the embodiment of the present invention will continue to fine-tune the cloud-based re-model based on the newly added test data. The specific implementation steps are as follows: at the end of the first round of training, the cloud-based server of the embodiment of the present invention will record all current multimodal data records in the model training data set as used records; and for each subsequent multimodal data record, the currently added multimodal data record will be recorded as an unused record; and when the total number of unused records in the model training data set reaches a preset total number of new records threshold, all current unused records in the model training data set will be extracted to form a corresponding current training data set; and the cloud-based re-model will be fine-tuned based on the current training data set; and at the end of this fine-tuning, all multimodal data records in the model training data set corresponding to the current training data set will be converted into used records.
[0127] It is not difficult to see from the specific implementation steps of the above method that in order to further improve the recognition ability of the AI model in the vehicle-road-cloud integrated system for complex traffic conditions, the embodiment of the present invention customizes a multimodal large language model, namely a traffic multimodal large language model, which can perform target detection, traffic road condition recognition, and traffic event / accident recognition based on multimodal data (images, videos, BEV maps, point clouds, text); and through the lightweight deployment of the traffic multimodal large language model on the roadside, the recognition ability of the roadside nodes for complex traffic conditions is improved; and through the cloud-roadside collaborative learning process with the knowledge distillation mechanism as the core, the cloud-based heavy model is used to provide continuous optimization guarantee for the lightweight model deployed on the roadside, so that the lightweight model on the roadside can be continuously optimized even if it is not actively trained / fine-tuned locally. In addition, the embodiment of the present invention also constructs a data closed loop for the cloud-based heavy model that can continuously generate diversified data sets through continuous real data recovery and addition processing from the roadside to the cloud, which in turn provides data guarantee for the continuous improvement of the model performance of the cloud-based heavy model.
[0128] Figure 3 This is a schematic diagram of the structure of an electronic device provided in the second embodiment of the present invention. The electronic device can be a terminal device or server that implements the method of the aforementioned embodiment, or it can be a terminal device or server that implements the method of the aforementioned embodiment connected to the aforementioned terminal device or server. Figure 3As shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver 303's transceiver actions. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the aforementioned embodiment method. Preferably, the electronic device involved in the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The above-mentioned communication port 306 is used for connecting and communicating between the electronic device and other peripherals.
[0129] exist Figure 3 The system bus 305 mentioned in the figure can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to realize communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include random access memory (RAM) and may also include non-volatile memory (Non-Volatile Memory), such as at least one disk storage.
[0130] The above-mentioned processors can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0131] It should be noted that an embodiment of the present invention further provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computer, it enables the computer to execute the methods and processing procedures provided in the above embodiments.
[0132] Embodiments of the present invention provide a cloud-road collaborative learning method, electronic device, and computer-readable storage medium for a large multimodal traffic language model. As can be seen from the foregoing invention, embodiments of the present invention pre-construct a large multimodal traffic language model and perform an initial round of supervised training on the model training dataset. After the initial round of training, the model is deployed in the cloud as a cloud-based heavy model. The model is then compressed and deployed simultaneously on the roadside and in the cloud as the corresponding roadside / cloud-based light model. After the deployment is completed, the roadside device will input the real-time first roadside multimodal data set and the first scene recognition instruction text into the roadside light model for visual target detection and traffic condition recognition processing to obtain the first target detection map and the first scene recognition text, and the first roadside multimodal data set, the first scene recognition instruction text, the first target detection map and the first scene recognition text will form the corresponding first roadside feedback and send it to the cloud; every time the cloud server receives a first roadside feedback, it will input the first roadside multimodal data set and the first scene recognition instruction text in the current feedback into the cloud heavy model for visual target detection and traffic condition recognition processing to obtain the corresponding second target detection map and the second scene recognition text to form the corresponding first cloud feedback, and the first roadside feedback and the first cloud feedback will form a roadside-cloud record and save it; the cloud server will also regularly form the corresponding first data set of all roadside-cloud records saved within the last specified time period, and update the cloud light model parameters according to the first data set according to the knowledge distillation method, and the parameter difference △θ of this update X Send to the roadside; each time the roadside equipment receives a parameter difference △θ X , based on the current parameter difference △θ X Update the roadside light model once. In addition, the cloud server will regularly perform data enhancement processing on the model training data set based on the real scene data submitted by the roadside, and continuously fine-tune the cloud-based heavy model based on the newly added test data. Based on the cloud-road collaborative learning mechanism given in the embodiment of the present invention: on the one hand, the purpose of deploying a large multimodal language model of traffic on the roadside is achieved through model compression; on the other hand, the purpose of continuously optimizing the roadside deployment model is achieved through the knowledge distillation mechanism between the heavy and light models on the cloud side; on the other hand, the purpose of continuously improving the performance of the cloud-based heavy model is achieved through the recycling of real data.
[0133] Professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0134] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0135] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A cloud-road collaborative learning method for a large multimodal traffic language model, characterized by: The method comprises: Construct a traffic multimodal large language model; and perform a first round of training on the traffic multimodal large language model based on a preset model training data set in a supervised training manner; and after the first round of training, deploy the traffic multimodal large language model on a cloud server and record it as a corresponding cloud-based heavy model; and perform model compression on the traffic multimodal large language model based on a preset pruning rule to obtain a corresponding compressed model; and deploy the compressed model simultaneously on a roadside device and the cloud server and record it as a corresponding roadside light model and a cloud-based light model; wherein the traffic multimodal large language model is used to perform visual target detection and traffic condition recognition processing based on the input roadside multimodal data set and scene recognition instruction text and output the corresponding target detection map and scene recognition text; the model parameter θ of the cloud-based heavy model M The model parameters θ of the compression model S The parameter compression relationship between them is recorded as F M-S ,θ S =F M-S (θ M ); The roadside device uses a preset scene recognition instruction template as a corresponding first scene recognition instruction text; and inputs a real-time first roadside multimodal data set and the first scene recognition instruction text into the roadside light model to perform visual target detection and traffic condition recognition processing to obtain a corresponding first target detection map and a first scene recognition text; and the first roadside multimodal data set, the first scene recognition instruction text, the first target detection map, and the first scene recognition text form a corresponding first roadside feedback and send it to the cloud server; Each time the cloud server receives a first roadside feedback, it inputs the first roadside multimodal dataset and the first scene recognition instruction text in the current feedback into the cloud-based re-model to perform visual target detection and traffic condition recognition processing to obtain a corresponding second target detection map and a second scene recognition text to form a corresponding first cloud-based feedback; and the current first roadside feedback and the first cloud-based feedback form a corresponding roadside-cloud record and save it; The cloud server regularly collects all the roadside-cloud records saved within the latest first preset time period into a corresponding first data set according to a preset first time frequency; and stores the current model parameters of the cloud light model as the pre-update parameters θ X ; and update the model parameters of the cloud light model according to the first data set according to the knowledge distillation method to obtain the corresponding updated parameters And based on the updated parameters θ X and the updated parameters Confirm the corresponding parameter difference And the parameter difference Δθ X Sending to the roadside equipment; Each time the roadside equipment receives the parameter difference Δθ X , the current model parameters of the roadside light model are recorded as the parameters before update And based on the parameters before update and the parameter difference △θ X Confirm the corresponding updated parameters And based on the updated parameters The model parameters of the roadside light model are updated.
2. The cloud-road collaborative learning method for a traffic multimodal large language model according to claim 1 is characterized in that: The roadside multimodal dataset includes a roadside real-life image, a roadside video, a roadside BEV map, and a roadside point cloud; the scene recognition text includes at least a traffic condition description text, a traffic accident / incident description text, and a target detection description text; the target detection image is an image with one or more target boxes marked on the roadside real-life image; each target box has a target type text; The scene recognition instruction text is a formatted instruction text of the traffic multimodal large language model; The scene recognition instruction text is used to instruct the traffic multimodal large language model to perform target detection based on the currently input roadside multimodal dataset, and to identify roadside traffic conditions and traffic accidents / incidents, and to generate scene recognition text based on the target detection results and the traffic condition and accident / incident identification results, and to output the generated scene recognition text and the target detection graph corresponding to the target detection results as a model processing result; The model training dataset consists of multiple multimodal data records; the multimodal data records include a training-multimodal dataset, a training-instruction text, a label-target detection map, and a label-scene recognition text; the training-multimodal dataset includes a training-roadside real-life image, a training-roadside video, a training-roadside BEV map, and a training-roadside point cloud; the training-instruction text meets the formatting requirements of the scene recognition instruction text; The roadside device is interconnected with the cloud server; the roadside device is also interconnected with multiple roadside sensors, the multiple roadside sensors including at least a first camera, a second camera, and a first lidar; the roadside device also has a built-in high-precision map module; the first camera is used to regularly output the latest first roadside real-time image of the current road / intersection to the roadside device at a preset first data update frequency; the second camera is used to regularly output the latest first roadside video of the current road / intersection to the roadside device at a preset second data update frequency; The high-precision map module is used to regularly output the latest first roadside BEV map of the current road / intersection to the roadside equipment at a preset third data update frequency; the first laser radar is used to regularly output the latest first roadside point cloud of the current road / intersection to the roadside equipment at a preset fourth data update frequency; The real-time first roadside multimodal dataset consists of the latest first roadside real-shot image, the first roadside video, the first roadside BEV map and the first roadside point cloud.
3. The cloud-road collaborative learning method for a traffic multimodal large language model according to claim 2 is characterized in that: The traffic multimodal large language model includes a visual target detection model, a multimodal encoder, an input projector, a text feature encoder, a large language model, and a detection map output module; the multimodal encoder includes a detection box encoder, an image encoder, a video encoder, a BEV encoder, and a point cloud encoder; the input projector includes a detection box-text feature space projector, an image-text feature space projector, a video-text feature space projector, a BEV-text feature space projector, and a point cloud-text feature space projector; The first, second, third and fourth model input ends of the traffic multimodal large language model are respectively used to receive the corresponding roadside real-life pictures, the roadside videos, the roadside BEV maps and the roadside point clouds in the input roadside multimodal dataset, the fifth model input end is used to receive the input scene recognition instruction text, and the first and second model output ends are used to output the corresponding target detection pictures and the scene recognition text; The input end of the visual object detection model is connected to the input end of the first model, and the output end is respectively connected to the input end of the detection box encoder and the first input end of the detection map output module; the second input end of the detection map output module is connected to the input end of the first model, and the output end is connected to the output end of the first model; The output end of the detection frame encoder is connected to the input end of the detection frame-text feature space projector; the output end of the detection frame-text feature space projector is connected to the first input end of the large language model; The input end of the image encoder is connected to the input end of the first model, and the output end is connected to the input end of the image-text feature space projector; the output end of the image-text feature space projector is connected to the second input end of the large language model; The input end of the video encoder is connected to the input end of the second model, and the output end is connected to the input end of the video-text feature space projector; the output end of the video-text feature space projector is connected to the third input end of the large language model; The input end of the BEV encoder is connected to the input end of the third model, and the output end is connected to the input end of the BEV-text feature space projector; the output end of the BEV-text feature space projector is connected to the fourth input end of the large language model; The input end of the point cloud encoder is connected to the input end of the fourth model, and the output end is connected to the input end of the point cloud-text feature space projector; the output end of the point cloud-text feature space projector is connected to the fifth input end of the large language model; The input end of the text feature encoder is connected to the input end of the fifth model, and the output end is connected to the sixth input end of the large language model; The output end of the large language model is connected to the output end of the second model; the large language model is also connected to an external traffic knowledge base.
4. The cloud-road collaborative learning method for a traffic multimodal large language model according to claim 3 is characterized in that: The visual target detection model is implemented based on a class of conventional target detection models; the conventional target detection models include at least RCNN series models and YOLO series models; the visual target detection model is used to perform target detection processing on the roadside real-time image input to the model to obtain one or more corresponding target detection frames to form a corresponding target detection frame sequence, which is sent to the detection image output module and the detection frame encoder; each target detection frame includes at least the coordinates of the detection frame center point, the detection frame size, the detection frame orientation, and the target type; The detection map output module is used to draw the corresponding target frame on the roadside real shot map input by the model according to the detection frame center point coordinates, the detection frame size and the detection frame orientation of each target detection frame in the target detection frame sequence; and setting a corresponding display text at a designated position on each target frame, and setting the display content of the corresponding display text based on the target type text corresponding to each target frame; and outputting the roadside real shot image after the target frame drawing and display text setting is completed as the corresponding target detection image; The detection frame encoder is used to perform detection frame feature encoding on the target detection frame sequence to obtain a corresponding first encoding tensor and send it to the detection frame-text feature space projector; the image encoder is used to perform image feature encoding on the roadside real-shot image input by the model to obtain a corresponding second encoding tensor and send it to the image-text feature space projector; the video encoder is used to perform video feature encoding on the roadside video input by the model to obtain a corresponding third encoding tensor and send it to the video-text feature space projector; the BEV encoder is used to perform BEV feature encoding on the roadside BEV map input by the model to obtain a corresponding fourth encoding tensor and send it to the BEV-text feature space projector; the point cloud encoder is used to perform point cloud feature encoding on the roadside point cloud input by the model to obtain a corresponding fifth encoding tensor and send it to the point cloud-text feature space projector; The detection box-text feature space projector is used to perform feature projection processing on the first coding tensor from the detection box feature space to the text feature space to obtain a corresponding first projection feature tensor and send it to the large language model; the image-text feature space projector is used to perform feature projection processing on the second coding tensor from the image feature space to the text feature space to obtain a corresponding second projection feature tensor and send it to the large language model; the video-text feature space projector is used to perform feature projection processing on the third coding tensor from the video feature space to the text feature space to obtain a corresponding third projection feature tensor and send it to the large language model; the BEV-text feature space projector is used to perform feature projection processing on the fourth coding tensor from the BEV feature space to the text feature space to obtain a corresponding fourth projection feature tensor and send it to the large language model; the point cloud-text feature space projector is used to perform feature projection processing on the fifth coding tensor from the point cloud feature space to the text feature space to obtain a corresponding fifth projection feature tensor and send it to the large language model; The text feature encoder is used to perform text feature encoding on the scene recognition instruction text input by the model to obtain a corresponding first text feature tensor and send it to the large language model; The large language model is a type of large language model implemented based on the Transformer framework structure, including at least a GPT series model; the large language model has completed model pre-training; the large language model is used to combine the first, second, third, fourth and fifth projection feature tensors and the first text feature tensor of the input into a corresponding first input tensor; and perform target detection result recognition, traffic road condition recognition and traffic accident / incident recognition based on the first input tensor and the traffic knowledge base, and perform scene recognition text generation processing based on the recognition results to obtain the corresponding scene recognition text and output it.
5. The cloud-road collaborative learning method for a traffic multimodal large language model according to claim 4 is characterized in that: The first round of training of the traffic multimodal large language model based on a preset model training dataset in a supervised training manner specifically includes: Step 51: Split the model training data set into two sub-data sets according to a preset first split ratio and record them as a corresponding first training set and a first evaluation set; Wherein, both the first training set and the first evaluation set are composed of a plurality of multimodal data records; the ratio of the total number of records in the first training set to the first evaluation set satisfies the first segmentation ratio; Step 52: taking the first multimodal data record of the first training set as the corresponding current training record; Step 53: Input the training-multimodal dataset and the training-instruction text of the current training record as the corresponding roadside multimodal dataset and the scene recognition instruction text into the traffic multimodal large language model to perform visual target detection and traffic condition recognition processing, and use the output target detection image and the scene recognition text as the corresponding first predicted detection image and first predicted text; Step 54: The first prediction detection image and the label-target detection image of the current training record form a corresponding first prediction-label pair; and the first prediction text and the label-scene recognition text of the current training record form a corresponding second prediction-label pair; the first and second prediction-label pairs are subjected to a preset model loss function; and based on a preset first model optimizer, a round of parameter modulation is performed on the traffic multimodal large language model in a direction that minimizes the model loss function. Wherein, the first model optimizer includes at least SGD optimizer, Adam optimizer, RMSprop optimizer, and Nadam optimizer; Step 55: Identify whether the current training record is the last multimodal data record of the first training set; if so, proceed to step 56; if not, use the next multimodal data record of the first training set as the new current training record and return to step 53; Step 56: Perform a round of traversal on all the multimodal data records of the first evaluation set; and during this round of traversal, use the multimodal data record currently traversed as the corresponding current evaluation record; and input the training-multimodal dataset and the training-instruction text of the current evaluation record as the corresponding roadside multimodal dataset and the scene recognition instruction text into the traffic multimodal large language model for visual target detection and traffic condition recognition processing, and use the output target detection map and the scene recognition text as the corresponding second prediction detection map and second prediction text; form a corresponding third prediction-label pair from the second prediction detection map and the label-target detection map of the current evaluation record; and form a corresponding fourth prediction-label pair from the second prediction text and the label-scene recognition text of the current evaluation record; In step 57, after this round of traversal is completed, a comprehensive evaluation is performed on the target detection performance and scene recognition text generation quality of the traffic multimodal large language model based on all the third and fourth prediction-label pairs obtained to obtain a corresponding first evaluation result; and when the first evaluation result does not meet the preset compliance requirements, return to step 52 to continue training; and when the first evaluation result meets the compliance requirements, the model training is terminated.
6. The cloud-road collaborative learning method for a traffic multimodal large language model according to claim 1 is characterized in that: The knowledge distillation method is used to update the model parameters of the cloud-based light model according to the first data set to obtain corresponding updated parameters Specifically include: Step 61: Introduce multiple low-rank adapters into the cloud-based light model to form a corresponding current student model; and use the cloud-based heavy model as the corresponding current teacher model; The model parameters of the current student model are composed of the model parameters of the cloud-based light model and the model parameters of all the low-rank adapters; Step 62: Using each of the roadside-cloud records of the first data set as a corresponding current record; and inputting the first roadside multimodal data set and the first scene recognition instruction text of the current record into the current student model for visual target detection and traffic condition recognition processing to obtain a corresponding third target detection map and third scene recognition text; and using the second target detection map and the second scene recognition text of the current record as the corresponding current label detection map and current label recognition text, forming a corresponding fifth prediction-label pair with the third target detection map and the current label detection map, and forming a corresponding sixth prediction-label pair with the third scene recognition text and the current label recognition text; Step 63: Substitute all the fifth and sixth prediction-label pairs obtained into a preset knowledge distillation loss function to calculate a corresponding first loss value; Step 64: Identify whether the first loss value satisfies a preset first loss value range; if so, proceed to step 65; if not, modulate the model parameters of all the low-rank adapters based on a preset second model optimizer in a direction that minimizes the knowledge distillation loss function, while keeping the model parameters of the cloud-based light model of the current student model unchanged, and return to step 62 after this round of modulation. Wherein, the second model optimizer includes at least an SGD optimizer and an Adam optimizer; Step 65: Update the model parameters of the cloud-based light model based on the model parameters of the current student model; and use the updated model parameters of the cloud-based light model as the corresponding updated parameters θ * X ; and remove all low-rank adapters introduced in the cloud-based light model.
7. The cloud-road collaborative learning method for a traffic multimodal large language model according to claim 2 is characterized in that: The method further comprises: The cloud server regularly takes each of the roadside-cloud records saved within the most recent second preset time period as the corresponding current record at a preset second time frequency; and takes the first roadside real-shot image, the first roadside video, the first roadside BEV map, the first roadside point cloud, the first scene recognition instruction text, the second target detection image and the second scene recognition text of the currently recorded image as the corresponding current roadside real-shot image, the current roadside video, the current roadside BEV map, the current roadside point cloud, the current scene recognition instruction text, the current target detection image and the current scene recognition text; and based on the preset image, video and point cloud noise addition algorithm / model, noise addition processing is performed in the corresponding current roadside real-shot image, the current roadside video and the current roadside point cloud to obtain the corresponding noisy roadside real-shot image, noisy roadside video and noisy roadside point cloud. side point cloud; and the noisy roadside real shot picture, the noisy roadside video, the current roadside BEV map, and the noisy roadside point cloud corresponding to the current record are used as the corresponding training-roadside real shot picture, the training-roadside video, the training-roadside BEV map, and the training-roadside point cloud to form a corresponding training-multimodal data set; and the current scene recognition instruction text, the current target detection map, and the current scene recognition text corresponding to the current record are used as the corresponding training-instruction text, the label-target detection map, and the label-scene recognition text; and the training-multimodal data set, the training-instruction text, the label-target detection map, and the label-scene recognition text corresponding to the current record form a corresponding multimodal data record and are stored in the model training data set; At the end of the first round of training, the cloud server records all the current multimodal data records in the model training data set as used records; and each time a new multimodal data record is added thereafter, the newly added multimodal data record is recorded as an unused record; and when the total number of unused records in the model training data set reaches a preset threshold for the total number of new records, all the current unused records in the model training data set are extracted to form a corresponding current training data set; and the cloud-based re-model is fine-tuned based on the current training data set; and at the end of this fine-tuning, all the multimodal data records in the model training data set corresponding to the current training data set are transferred to used records.
8. An electronic device, characterized in that: include: memory, processors, and transceivers; The processor is configured to be coupled to the memory, read and execute instructions in the memory, so as to implement the method according to any one of claims 1 to 7; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Target identification method, and training method and device of target identification model
CN115019060A
Cloud cooperative control method, electronic equipment and integrated large model deployment architecture
CN118034913A