End-to-end traffic road state sensing method based on multi-modal large model
Through the end-to-end traffic road state perception method based on multimodal large model, the problem of insufficient traffic state perception ability in complex environments in the existing technology is solved, automatic labeling and scene understanding text generation are realized, and real-time and accuracy of traffic management are improved.
Patent Information
- Application Number
- CN202510017064.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The prior art is limited in real-time, accuracy and environmental adaptability, and it is difficult to meet the needs of smart city traffic management, especially in complex environments such as low light.
The end-to-end traffic road state perception method based on multimodal large model is adopted. By collecting traffic video data and text data, image features and text features are extracted, the correlation between traffic elements and features is constructed, and sliding windows and pre-trained large language models are used for training to generate scene understanding text.
It realizes automatic labeling of images, improves the road network state perception ability, enhances the reasoning ability of automatic understanding of scenes, and can realize accurate object detection and text description of road network state in complex environments, reducing the cost of manual detection.
Smart Images

Figure CN119964101A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic digital data processing, and more particularly to an end-to-end traffic road state perception method based on a multi-modal large model. Background Art
[0002] In recent years, with the increasing complexity of urban traffic, the video surveillance traffic network monitoring method based on traditional deep learning target detection is limited in real-time, accuracy and environmental adaptability, and it is difficult to meet the needs of smart city traffic management.
[0003] The road traffic state perception of traditional visual models mainly relies on image processing and computer vision technology, and extracts traffic flow information by analyzing video surveillance or image data. This method extracts key features from videos or sensors by constructing a neural network structure that adapts to the characteristics of traffic data, and identifies traffic state information such as vehicle position, speed, and traffic flow. Specifically, traditional deep learning models can identify vehicles, pedestrians, and abnormal events on the road by segmenting and detecting image data. Although traditional deep learning methods can achieve a high level of traffic state recognition targets under certain conditions, their excessive reliance on data annotation limits their performance in complex environments such as low light, which greatly affects the accuracy of traffic network state description. Therefore, it is necessary to design a large end-to-end perception model with a huge number of parameters to improve the road network state perception capability. Summary of the invention
[0004] The present invention is proposed based on the above-mentioned requirements of the prior art. The technical problem to be solved by the present invention is to provide an end-to-end traffic road state perception method based on a multimodal large model to improve the road network state perception capability.
[0005] In order to solve the above problems, the present invention is implemented by adopting the following technical solutions:
[0006] Provided is an end-to-end traffic road state perception method based on a multimodal large model, the method comprising: collecting a traffic video data set and a traffic text data set; extracting each frame of at least part of the traffic video data, performing feature extraction on each frame to obtain traffic features; annotating each frame of the image based on a traffic element set, if the image has a certain traffic element, the corresponding traffic element is annotated as 1; otherwise, it is annotated as 0; cleaning, segmenting and tokenizing the traffic text data set; counting a first ratio and a second ratio, the first ratio representing the ratio of the number of times the traffic element appears in the traffic video data set to the total number of words in the traffic vocabulary, and the second ratio representing the ratio of the number of times the traffic feature appears in the traffic video data set to the total number of words in the traffic vocabulary; calculating a joint probability based on the first ratio and the second ratio; judging whether the joint probability exceeds a threshold, if it exceeds the threshold, the corresponding traffic element is associated with the corresponding traffic feature; based on The method comprises the following steps: the first step is to extract the cleaned traffic text data set using a sliding window and obtain training samples; the second step is to pre-train a large language model based on the training samples to obtain a pre-trained large language model, including: inputting the annotated phrases corresponding to the traffic elements into the large language model so that the large language model can understand the semantic information in the traffic scene; the second step is to pre-train the training samples using the large language model so that the large language model can predict the traffic state change description words; the second step is to train the perception large model based on the traffic video data set, the corresponding traffic text data and the data pairs consisting of the annotated images and the corresponding traffic elements, the perception large model includes a pre-trained image encoder, a position embedding layer, a Q-Former, a linear layer and a pre-trained large language model in sequence, and during the training process, the parameters of the pre-trained image encoder and the pre-trained large language model are frozen; the video image to be input is input into the trained perception large model to obtain the scene understanding text.
[0007] Optionally, it is characterized in that the labeled traffic video data set is trained based on the Sigmoid cross entropy loss function to obtain a trained automatic labeling model; the image is automatically labeled using the automatic labeling model; the expression of the Sigmoid cross entropy loss function is: Among them, L represents the cross entropy loss function of image annotation, m represents the number of images after the traffic video data is extracted, n represents the sequence number of the image, and y n represents a sequence of n images, P n Represents the image annotation results of sequence n.
[0008] Optionally, the joint probability is calculated based on the first ratio and the second ratio, and its expression is:
[0009] Among them, A represents the appearance of traffic elements in the traffic video dataset, B represents the appearance of traffic features in the traffic video dataset, P(A|B) represents the joint probability, P(B|A) represents the probability of event B occurring under the premise of event A occurring, P(A) represents the first ratio, and P(B) represents the second ratio.
[0010] Optionally, during pre-training of a large language model, L2 regularization is used to control the complexity of the model.
[0011] Optionally, the Lion optimizer is used to adjust the batch size of the input data for the pre-trained large language model.
[0012] Optionally, each time the batch size is adjusted, the weights of the pre-trained large language model are initialized using a learning decay rate, where the learning decay rate is calculated as:
[0013] Among them, α represents the learning decay rate, DecayRate represents the initial large language model decay rate, EpochNumber represents the number of model training iterations, and α0 represents the initial learning rate.
[0014] Optionally, during the pre-training process, the large language model is fine-tuned based on the LoRA algorithm.
[0015] Optionally, the large language model is LLaMA3, and the network structure of the LLaMA3 includes, from top to bottom, first layer normalization, residual connection, multi-head self-attention layer, second layer normalization and fully connected layer.
[0016] Compared with the prior art, the present invention provides an end-to-end traffic road state perception method based on a multimodal large model, which can realize automatic annotation of images, and by building the correlation between traffic features and traffic elements, it increases the ability to extract correlation characteristics between different samples. In addition, the output of the end-to-end perception large model of the present invention is the scene understanding text description content, which greatly enhances the reasoning ability of automatic scene understanding, can realize end-to-end perception description of traffic status, and simultaneously realize accurate target detection of urban traffic road status and text description of road network status in complex environments, which can meet the accurate perception of urban road traffic networks and reduce the cost of manual detection of road traffic conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0018] Figure 1 This is a flow chart of an end-to-end traffic road state perception method based on a multimodal large model provided in this embodiment. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] To facilitate understanding of the embodiments of the present invention, further explanation will be given below with reference to specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation on the protection scope of the present invention.
[0021] This embodiment provides an end-to-end traffic road state perception method based on a multi-modal large model, and its process is as follows: Figure 1 As shown, including:
[0022] S1 collects traffic video datasets and traffic text datasets.
[0023] Collect traffic video datasets and traffic text datasets in the field of urban road traffic.
[0024] S2 extracts at least part of each frame of the traffic video data, and performs feature extraction on each frame to obtain traffic features.
[0025] With the convolutional network as the input layer, the pixel matrix corresponding to the labeled image is convolved. During the convolution operation, usually, the number of elements in the feature matrix after convolution is reduced compared to the input matrix, and the input image matrix needs to be padded with zeros, i.e., filled. Through this zero-padding process, the size of the feature matrix can be effectively maintained. Subsequently, the set convolution kernel will gradually slide on the padded input matrix, performing a dot product operation to extract the feature information in the image. The formula for the convolution operation is as follows:
[0026]
[0027] In the formula, X i is the feature obtained after the i-th layer of convolution, X is the input feature, W i represents the convolution kernel in the i-th convolution layer, b i represents the bias of the i-th convolution layer, Represents the operation of the convolution kernel and image features, and f represents the activation function used after the convolution operation.
[0028] After padding, we need to calculate the output image size after the convolution operation. The size of the convolution kernel is F×F, and the step size of the convolution kernel is S. Assuming the padding is P, the output size of an image of size W×W×C1 after the convolution layer is Q×Q×C2, where Q is calculated as follows:
[0029] Q=(W-F+2P) / S+1
[0030] Convolution operation can reduce the number of function parameters, but the amount of video image data is huge, and feature extraction still faces the problem of high parameter calculation. In order to reduce the amount of calculation for image feature extraction, the pooling layer is introduced to reduce the spatial size of the feature. The maximum pooling is used for repeated feature sampling. The traffic monitoring feature map contains a lot of information, including edge features, local features and main features. However, not all extracted features are actually required to be annotated. Before entering the next layer of the network, measures are taken to remove these unnecessary features and retain key information. This process can not only effectively reduce the number of parameters in the network training process, but also effectively control the overfitting phenomenon and improve the generalization ability of the model.
[0031] S3 annotates each frame of the image based on the traffic element set. If the image contains a certain traffic element, the corresponding traffic element is annotated as 1; otherwise, it is annotated as 0.
[0032] The performance of large visual models depends largely on large amounts of accurate and rich video image data for training. Video data can provide spatiotemporal information in dynamic scenes, which not only helps the model recognize objects, but also understands the object's motion trajectory, behavior pattern, and scene changes. Video image annotation is a process of adding labels to objects and scenes in the video in order to train visual models. Image annotation is the basis of video annotation. Image annotation methods include object detection annotation, semantic segmentation annotation, instance segmentation annotation, key point annotation, and tracking annotation. Through these annotation methods, video data can be structured and parsed, helping multimodal models to recognize, understand, and predict different objects and behaviors in complex dynamic scenes during training. Object detection annotation is mainly used to determine the information and location of traffic elements such as cars and pedestrians in the image, select objects by rectangular boxes, and label each box.
[0033] Assume that the sequence of frame images x of a video X is X' = {x1, x2, ..., x n Y={y1,y2,…,y j ,…,y m} represents the set of traffic elements that have been annotated in the traffic element database. Such an automatic annotation task for a video can be represented as a 1-M mapping set of N pairs of images and traffic elements after the video is extracted, represented as P = {(x1, Y1), (x2, Y2), …, (x i ,Y i ),…(x n ,Y n )},Y i is an M-dimensional vector whose element value is 0 or 1, indicating whether the image contains the element. i The y in j =1, indicating traffic element y j In the image x i When Y i The y in j =0, indicating y j In the image x i There are no markings in the .
[0034] For the problem of labeling multiple traffic elements, the Sigmoid cross entropy loss function is used to train the automatic labeling model, and its expression is:
[0035]
[0036] Among them, L represents the cross entropy loss function of image annotation, m represents the number of images after the traffic video data is extracted, n represents the sequence number of the image, and y n represents a sequence of n images, P n Represents the image annotation results of sequence n.
[0037] The derivation process of the above expression is as follows: The probability output by the Sigmoid cross entropy function needs to be mapped to the interval [0,1], and this probability needs to reflect the probability of being predicted as a positive class. The predicted output is the probability when the sample label is 1:
[0038] P n =P(y=1|x)
[0039] When the image sample is 0, the probability is as follows:
[0040] 1-P n =P(y=0|x)
[0041] Constructing the maximum likelihood value combines the formula as follows:
[0042]
[0043] When the traffic element label of the picture in the actual next video is y=0, the first term in the following formula is 1, and the probability needs to be further transformed and calculated as follows:
[0044] P(y=0|x)=1-P n
[0045] When the traffic element label of the picture in the actual next video is y=1, the probability needs to be further transformed and calculated as follows:
[0046] P(y=1|x)=P n
[0047] In these two cases, the probability expression remains unchanged. The greater the probability under the overall probability expression, the better the extraction of traffic element characteristics. In order to reduce the amount of calculation of the exponential function while not changing the monotonicity of the probability calculation formula, the log function is introduced as follows:
[0048] logP(y|x)=y n logP n +(1-y n )log(1-P n )
[0049] For m traffic element labels, a Sigmoid cross entropy loss function is constructed.
[0050] The semantic annotation of the multimodal large model divides the pixels in the image into different category areas. Each traffic element is annotated as a specific category, such as road, vehicle, pedestrian, etc., and relevant text enhancement training is performed on these categories to improve the richness of the description content and reasoning ability. The specific text enhancement training process is shown in S4-S8.
[0051] S4 cleans, tokenizes, and segment traffic text datasets.
[0052] Traffic text datasets usually contain noise, errors, and data from different sources may have inconsistent formats. Data cleaning is required to ensure the quality of data training, including removing HTML tags, processing missing data, and removing duplicate samples.
[0053] The cleaned data is tokenized and tokenized. Specifically, the traffic text data is segmented into a sequence of words or subwords. After tokenization, each word can be mapped to a corresponding identifier, and the index in the vocabulary can be used to convert the text to lowercase, remove punctuation, special characters, etc.
[0054] S5 counts the first ratio and the second ratio.
[0055] The first ratio represents the ratio of the number of times traffic elements appear in the traffic video dataset to the total number of words in the traffic vocabulary, and the second ratio represents the ratio of the number of times traffic features appear in the traffic video dataset to the total number of words in the traffic vocabulary.
[0056] S6 calculates the joint probability based on the first ratio and the second ratio; determines whether the joint probability exceeds a threshold value, and if so, associates the corresponding traffic element with the corresponding traffic feature.
[0057] The expression for calculating the joint probability P(A|B) is as follows:
[0058]
[0059] Among them, A represents the appearance of traffic elements in the traffic video dataset, B represents the appearance of traffic features in the traffic video dataset, P(A|B) represents the joint probability, P(B|A) represents the probability of event B occurring under the premise of event A occurring, P(A) represents the first ratio, and P(B) represents the second ratio.
[0060] In this embodiment, the threshold is set to 0.8. When P(A|B) exceeds 0.8 during training, it is considered that the traffic element and the traffic feature are associated.
[0061] In this embodiment, the correlation between traffic elements and traffic characteristics can also be output by training traffic text data using a large language model.
[0062] Based on the associated traffic elements and traffic characteristics, S7 uses a sliding window to extract the cleaned traffic text dataset to obtain training samples.
[0063] According to the characteristics of urban traffic, the data is organized into training samples. For large language models, continuous sequences can be extracted from the text using sliding windows based on associated traffic elements and traffic characteristics as training samples, with the goal of inferring the description words that may change the traffic status next. Methods to expand the data set by performing some random transformations on the training data. For example, the text can be randomly truncated, noise added, synonyms replaced, etc. to improve the robustness and generalization ability of the model.
[0064] S8 pre-trains the large language model based on the training samples to obtain a pre-trained large language model.
[0065] The large language model is pre-trained based on a large number of annotated traffic text datasets. Preferably, the large language model is LLaMA3, whose basic structure is the Transformer structure, whose core is the self-attention mechanism to capture the dynamic text language relevance, and the multi-head attention mechanism is used to realize the feature calculation of multiple outputs.
[0066] In order to improve the stability of LLaMA3 during training, the pre-layer normalization method was introduced, the first layer normalization was moved before the multi-head self-attention layer, the second layer normalization was also moved before the fully connected layer, and the position of the residual connection was also adjusted to after the multi-head self-attention layer and the fully connected layer. The adjusted LLaMA3 network structure includes the first layer normalization, residual connection, multi-head self-attention layer, second layer normalization and fully connected layer from top to bottom.
[0067] The RMSNorm normalization function is used in layer normalization. The expression of the RMSNorm function processing the input vector is:
[0068]
[0069] Among them, RMS(a) represents root mean square layer normalization, n represents the feature dimension, and a i represents the input feature vector, Represents the normalized output feature vector or tensor.
[0070] In position coding, the rotation position coding RoPE is used to replace the original absolute position coding. Based on complex number theory, the starting point of RoPE is to achieve relative position coding through absolute position coding. Its goal is to add absolute position information to q and k through the following operations:
[0071]
[0072] Among them, q represents the query vector, m represents the position index, represents the position of the current position encoding, and k represents the key vector. Indicates the absolute position information that the large language model brings to q, It represents the absolute position information brought by the large language model to k, and f(k,m) represents the encoding function of the absolute position information.
[0073] Based on the above method, and Brings absolute position information of text position embedding to large language models.
[0074] The pre-training process includes:
[0075] The annotated phrases corresponding to the traffic elements are input into the large language model so that the large language model can understand the semantic information in the traffic scene.
[0076] The annotation phrases include entity class, attribute class, state class and relationship class.
[0077] The training samples are pre-trained using a large language model so that the large language model can predict traffic state change description words.
[0078] Preferably, during pre-training of a large language model, L2 regularization is used to control the complexity of the model.
[0079] The L2 regularization method is used to help control the complexity of the model and calculate the original loss LOSS of the model in the data. total , weight decay is used to prevent overfitting.
[0080]
[0081] Where, LOSS data represents the data loss term, λ is the L2 regularization coefficient, which is used to control the contribution of the regularization term to the total loss; ||w|| 2 is the square of the L2 norm of the weight vector w.
[0082] Preferably, the Lion optimizer is used to adjust the batch size of the pre-trained large language model input data.
[0083] Batch input helps achieve data parallelism in large language models. During training, a larger batch size can usually improve training efficiency, using more samples each time the weight is updated, thereby reducing the number of updates, which is very important for the training speed of large language models. This embodiment takes into account the difficulty of multimodal training, uses the Lion optimizer to adjust the batch size, and introduces some functions to make the program more compact, such as the linear interpolation function interp(x,y,a), and the optimized original function is (1-a)·x+a·y.
[0084] Preferably, each time the batch size is adjusted, the weights of the pre-trained large language model are initialized using the learning decay rate.
[0085] The batch size is a hyperparameter that needs to be adjusted, and is selected based on the characteristics of the perception model architecture, task, and dataset. Experiments are needed on the traffic network to find the optimal batch size. This method starts with a larger batch size and then gradually reduces it to improve the stability of the model. The learning decay rate α is used to initialize the weights of the pre-trained model, which helps the model converge quickly. The decay formula is as follows:
[0086]
[0087] Among them, α0 represents the initial learning rate, which continues to decrease as the number of iterations increases, DecayRate represents the initial large language model decay rate, EpochNumber represents the number of model training iterations, and α0 represents the initial learning rate.
[0088] S9 trains the perception model based on traffic video datasets, corresponding traffic text data, and data pairs consisting of annotated images and corresponding traffic elements.
[0089] The end-to-end perception model is mainly a combination model that integrates the perception module and the large language model. The output content includes question and answer and perception content output, etc. The data needs to be organized into question-answer pairs, source-target pairs, etc. For supervised learning tasks, the data labels need to be associated with the input samples so that the model can be trained in a supervised manner.
[0090] The perceptual large model includes a pre-trained image encoder, a position embedding layer, a Q-Former, a linear layer and a pre-trained large language model in sequence. During the training process, the parameters of the pre-trained image encoder and the pre-trained large language model are frozen.
[0091] The core of building a large multimodal perception model is video encoding understanding and large language model output. This embodiment proposes a video Q-former to capture the temporal changes of visual scenes. The model is trained on a large number of video image caption pairs and visual instruction tuning datasets to align the output of the visual encoder with the embedding space of the LLM.
[0092] The main steps of the end-to-end perception large model are as follows:
[0093] Features are extracted from input video frames through a frozen pre-trained image encoder, which includes ViT-G / 14 from EVA-CLIP and a pre-trained Q-Former; temporal information is injected into video frames using a position embedding layer; and a visual Q-former is used to capture the temporal changes in road traffic scenes, providing a dynamic understanding of road traffic conditions. The training strategy of the visual Q-Former is as follows: on the one hand, the visual Q-Former performs multi-objective training on text pairs, aligns the visual representation and the text representation, and learns the visual representation that best matches the text; on the other hand, a linear layer is used to project the output of the visual Q-Former into a vector with the same dimension as the text embedding of the large language model, so that the visual representation learned by the Q-Former can be explained by the large language model.
[0094] The visual representation model is trained on a large dataset of various types of traffic event videos and completes the video-to-text generation task. The annotated image-traffic element pairs are also added to the pre-trained dataset to enhance the understanding of the concept of static traffic scenes. The BLIP image-text matching loss coefficient l is constructed through contrastive learning. v2t :
[0095]
[0096] Among them, vi The image embedding representation of the i-th feature, t i The text embedding representation of the i-th feature, sim(v i ,t i ), represents the similarity calculation function, and τ represents the change parameter for contrastive learning
[0097] Preferably, during the pre-training process, the large language model is fine-tuned based on the LoRA algorithm.
[0098] The text output of the large language model is based on text prediction and text classification after self-attention layer normalization. The urban traffic system has a large amount of text content in professional fields. To significantly improve the representation ability of the large model in the urban traffic road perception task, it is necessary to finely tune the large model with a certain amount of road traffic text data to improve the task applicability of the pre-trained large model. This embodiment adopts a fine-grained traffic scenario fine-tuning method based on LoRA, and the basic steps are as follows:
[0099] Freeze the parameters of the large language model and use the large number of original parameters of LLaMA3 for inference without updating to ensure the basic language logic. Then create two low-rank matrices A and B, so that the size of the matrix after multiplication will be the same as the size of the weight matrix of the model being fine-tuned. There are multiple weight matrices in the large language model of this embodiment, and similar matrix pairs are created for each weight matrix.
[0100] Only train the low-rank matrices A and B, while keeping the pre-trained weight matrix W unchanged. The pre-trained weight matrix W is a d×k matrix, which represents the weights in the original model and remains frozen during the adaptation process without training and updating. The low-rank matrix B is a d×r matrix, and the low-rank matrix A is an r×k matrix, where r << min(d,k), which is the maximum rank of LoRA. A is initialized using a random Gaussian distribution, and its elements are sampled from a Gaussian distribution with a mean of 0. B is a zero matrix at initialization. Based on this reparameterization, the expression for Lora fine-tuning is:
[0101] h = Wx + ΔWx = Wx + BAx
[0102] where ΔW represents the parameter update during fine-tuning. LoRA restricts the weight update so that ΔW = BA. W0x represents the original forward propagation process, that is, the input x passes through the weight matrix W to obtain the output. In the large language model, x can be a word embedding vector, and W is the weight matrix from the word embedding layer to the hidden layer. BAx represents the correction term introduced in the fine-tuning process, and the product BA represents the correction to the original weight matrix W.
[0103] Through this reparameterization, Lora fine-tuning can adapt to the road traffic perception content text generation task by learning the low-rank matrix BA while keeping the pre-trained weights unchanged. It greatly reduces the number of parameters that need to be trained and improves the efficiency of traffic perception. At the same time, since the matrix W is pre-trained and remains unchanged during the training process, the computational and storage costs of the perception large model can be further reduced.
[0104] S10 inputs the video image to be input into the trained perception model to obtain scene understanding text.
[0105] This embodiment provides an end-to-end traffic road state perception method based on a multimodal large model, which can realize automatic annotation of images, and by building the correlation between traffic features and traffic elements, it increases the ability to extract correlation characteristics between different samples. In addition, the output of the end-to-end perception large model of this embodiment is the scene understanding text description content, which greatly enhances the reasoning ability of automatic scene understanding, and can realize end-to-end perception description of traffic status, while realizing accurate target detection of urban traffic road status and text description of road network status in complex environments, which can meet the accurate perception of urban road traffic networks and reduce the cost of manual detection of road traffic conditions.
[0106] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An end-to-end traffic road state perception method based on a multimodal large model, characterized in that: include: Collect traffic video datasets and traffic text datasets; Extracting at least a portion of each frame of the traffic video data, and performing feature extraction on each frame to obtain traffic features; Each frame of the image is annotated based on the traffic element set. If the image contains a certain traffic element, the corresponding traffic element is annotated as 1. Otherwise, it is marked as 0; Clean, tokenize and segment traffic text datasets; Counting a first ratio and a second ratio, wherein the first ratio represents the ratio of the number of times a traffic element appears in a traffic video data set to the total number of words in a traffic vocabulary, and the second ratio represents the ratio of the number of times a traffic feature appears in a traffic video data set to the total number of words in the traffic vocabulary; Calculating a joint probability based on the first ratio and the second ratio; determining whether the joint probability exceeds a threshold, and if so, associating the corresponding traffic element with the corresponding traffic feature; Based on the associated traffic elements and traffic characteristics, the cleaned traffic text dataset is extracted using a sliding window to obtain training samples; Pre-training the large language model based on the training samples to obtain the pre-trained large language model includes: inputting the labeled phrases corresponding to the traffic elements into the large language model so that the large language model can understand the semantic information in the traffic scene; pre-training the training samples using the large language model so that the large language model can predict the traffic state change description words; Based on the traffic video data set, the corresponding traffic text data and the data pairs consisting of the annotated images and the corresponding traffic elements, the perception large model is trained, wherein the perception large model sequentially includes a pre-trained image encoder, a position embedding layer, a Q-Former, a linear layer and a pre-trained large language model, and during the training process, the parameters of the pre-trained image encoder and the pre-trained large language model are frozen; The video image to be input is input into the trained perception model to obtain the scene understanding text.
2. The end-to-end traffic road state perception method based on a multimodal large model according to claim 1 is characterized in that: The labeled traffic video dataset is trained based on the Sigmoid cross entropy loss function to obtain the trained automatic labeling model; Automatically annotate images using an automatic annotation model; The expression of the Sigmoid cross entropy loss function is: Among them, L represents the cross entropy loss function of image annotation, m represents the number of images after the traffic video data is extracted, n represents the sequence number of the image, and y n represents a sequence of n images, P n Represents the image annotation results of sequence n.
3. The end-to-end traffic road state perception method based on a multimodal large model according to claim 1 is characterized in that: The joint probability is calculated based on the first ratio and the second ratio, and its expression is: Among them, A represents the appearance of traffic elements in the traffic video dataset, B represents the appearance of traffic features in the traffic video dataset, P(A|B) represents the joint probability, P(B|A) represents the probability of event B occurring under the premise of event A occurring, P(A) represents the first ratio, and P(B) represents the second ratio.
4. The end-to-end traffic road state perception method based on a multimodal large model according to claim 1 is characterized in that: During pre-training of large language models, L2 regularization is used to control the complexity of the model.
5. The end-to-end traffic road state perception method based on a multimodal large model according to claim 1 is characterized in that: The Lion optimizer is used to adjust the batch size of the input data for the pre-trained large language model.
6. The end-to-end traffic road state perception method based on a multimodal large model according to claim 5 is characterized in that: Each time the batch size is adjusted, the weights of the pre-trained large language model are initialized using the learning decay rate, which is calculated as: Among them, α represents the learning decay rate, DecayRate represents the initial large language model decay rate, EpochNumber represents the number of model training iterations, and α0 represents the initial learning rate.
7. The end-to-end traffic road state perception method based on a multimodal large model according to claim 1 is characterized in that: During the pre-training process, the large language model is fine-tuned based on the LoRA algorithm.
8. The end-to-end traffic road state perception method based on a multimodal large model according to claim 1 is characterized in that: The large language model is LLaMA3, and the network structure of LLaMA3 includes, from top to bottom, first layer normalization, residual connection, multi-head self-attention layer, second layer normalization and full connection layer.
Citation Information
Patent Citations
Method, system and equipment for analyzing traffic data in real time based on large model and medium
CN117951348A
Traffic accident identification method and device based on multi-modal large language model
CN118525275A
Multi-modal entity recognition method based on large language model
CN119167937A
Scenario identification for validation and training of machine learning based models for autonomous vehicles
US20210356968A1
Cited By
Video image data cleaning method and system for urban rail transit engineering
CN120182636A
Video data processing method and device, electronic equipment and storage medium
CN120529132A
Traffic flow prediction method, device, equipment, storage medium and program product
CN121583125A
Large model traffic accident analysis method fusing blind area completion and physical driving reasoning
CN122416740A
A Large-Model Traffic Accident Analysis Method Integrating Blind Spot Completion and Physics-Driven Reasoning
CN122416740B