Method and device for processing, analyzing and predicting multi-modal data of AI large model and storage medium
By integrating multimodal data processing, analysis and prediction methods into AI large models, the problems of poor application of AI large models in the transportation field and excessive update frequency of prediction results in the prior art are solved, and more accurate and timely traffic conditions analysis and prediction are achieved.
Patent Information
- Application Number
- CN202510372837.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-20
AI Technical Summary
The existing AI models have poor application effects in the field of transportation, and the analysis and prediction results are biased greatly. Due to the real-time changes in traffic conditions, the frequency of prediction results is too high, resulting in untimely untimely.
The AI large model multimodal data processing, analysis and prediction method is adopted to collect road image video data, road noise data, road area weather data and pedestrian image video data. Through preprocessing, correlation analysis, regression analysis and simulation experiments, a mathematical model between the data and traffic conditions is established, and a convolutional neural network and long-term memory network are combined for model training to analyze and predict traffic conditions in real time.
It improves the accuracy and timeliness of traffic conditions analysis and prediction, avoids the problems of untimely update of prediction results and excessive update frequency, and enhances the stability and application effect of analysis and prediction results.
Smart Images

Figure CN120183192A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent transportation, and in particular, relates to a method, device, and storage medium for processing, analyzing, and predicting multi-modal data of an AI large model. Background Art
[0002] The technology of processing, analyzing, and predicting AI large model data has significant application effects in the transportation field. Through the AI large model, it is possible to predict and analyze the traffic environment based on the input data. Relevant departments can then take measures such as traffic control and early warning in advance according to the prediction results, improving traffic safety and reducing the accident rate.
[0003] Currently, with the application of AI technology, AI large models have also been gradually applied in the transportation field. By creating an AI large model, it can analyze data and predict the traffic environment based on the analysis results. It generally uses the Transformer architecture model and analyzes and predicts possible future traffic conditions through simple data sources such as traffic flow and vehicle speed. However, due to the relatively simple data sources, the analysis and prediction results of the AI model for traffic conditions have a large deviation, and the application effect for traffic prediction is not good. Secondly, since traffic conditions change in real time, if the AI model is based on real-time changes in traffic conditions, it will lead to a high frequency of change in traffic prediction results. If the method of obtaining prediction results from input data is used, it may result in untimely traffic prediction. Summary of the Invention
[0004] The object of the present invention is to address the above problems and provide a method, device, and storage medium for processing, analyzing, and predicting multi-modal data of an AI large model.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A method for processing, analyzing, and predicting multi-modal data of an AI large model, the analysis and prediction method includes: Collect data; wherein, collecting data includes at least road image and video data information, road noise data information, road area weather data, and pedestrian image and video data information; Preprocess the collected data; Conduct a correlation analysis of each preprocessed data with traffic condition indicators, and perform a regression analysis to establish a mathematical model between each data and traffic condition indicators; Conduct a simulation experiment in the model. First, fix all the data, and gradually change the value of one of the data, then observe the change in the traffic condition results output by the model, calculate the relative sensitivity index of this changing data to traffic conditions, and then calculate the relative sensitivity indices of other data to traffic conditions in turn; Extract road vehicle-related features, road noise-related features, road area weather-related features, and pedestrian-related features. At the same time, extract the relative sensitivity features of road vehicles, the relative sensitivity features of road noise, the relative sensitivity features of road weather, and the relative sensitivity features of pedestrians; Select a model that combines a convolutional neural network and a long short-term memory network, and perform model training; Connect the terminals of each collected data to the data input end of the model. When the model analyzes and predicts the result of a traffic condition, the model sets the current data as anchor points, and based on the relative sensitivity indicators corresponding to each data, after at least one data changes beyond the corresponding relative sensitivity indicator, the model re-analyzes and predicts a new traffic condition result.
[0006] In the above method for multi-modal data processing, analysis, and prediction of an AI large model, the road image and video data information at least includes: the number, type, speed, driving direction, lane information, and vehicle distance of vehicles; The road noise data information at least includes: the sound of vehicle driving, the sound generated by vehicle honking, and the sound made by pedestrians; The road area weather data at least includes: temperature, humidity, precipitation, visibility, wind speed; The pedestrian image and video data information at least includes: the flow rate of pedestrians, walking speed, walking direction, and whether they cross the road illegally.
[0007] In the above method for multi-modal data processing, analysis, and prediction of an AI large model, when preprocessing the road image and video data information and the pedestrian image and video data information, perform image enhancement processing on the video data to improve clarity, and perform time synchronization and frame rate adjustment. Clean the data of the extracted vehicle information and pedestrian information to remove outliers; When preprocessing the road noise data, perform noise reduction processing on the noise data, calibrate the deviation of the data collection terminal device, and align the noise data according to the time series; When preprocessing the road area weather data, perform standardization processing on the weather data and perform one-hot encoding on the categorical data.
[0008] In the above method for multi-modal data processing, analysis, and prediction of an AI large model, the road vehicle-related features include: vehicle density at a specific road section during a certain period, the proportion of different types of vehicles, average vehicle speed, vehicle speed standard deviation, the degree of chaos of vehicle directions, and statistical features of vehicle following distances; Road noise-related features include: the mean, standard deviation, peak frequency, noise change rate of the noise, and the correlation between the noise and vehicle density and vehicle speed; The weather-related features of the road area include: dummy variables of weather types, interaction terms of visibility and vehicle speed, and relationship features between precipitation intensity and accident incidence rate; The relevant features of pedestrians include: pedestrian flow density, the number of conflict points between pedestrians and vehicles, the frequency of pedestrians crossing illegally, etc.
[0009] In the above method for processing, analyzing and predicting multi-modal data of an AI large model, when training the model, each data set is divided into a training set, a validation set and a test set according to a certain ratio. The training set is used to train the model, and the hyperparameters of the model are adjusted on the validation set to prevent overfitting. Appropriate loss functions and optimization algorithms are used for model training.
[0010] In the above method for processing, analyzing and predicting multi-modal data of an AI large model, the collected data also includes collecting text information from social platforms, extracting relevant information describing the road traffic conditions, and supplementing this information into the model.
[0011] An apparatus for processing, analyzing and predicting multi-modal data of an AI large model includes a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the computer program, it implements the above method for processing, analyzing and predicting multi-modal data of an AI large model.
[0012] A computer-readable storage medium stores a computer program, and when the computer program is run by a processor, the processor is caused to execute the above method for processing, analyzing and predicting multi-modal data of an AI large model.
[0013] Compared with the existing technology, the advantages of a method, apparatus and storage medium for processing, analyzing and predicting multi-modal data of an AI large model are as follows: By collecting road image and video data information, road noise data information, road area weather data and pedestrian image and video data information, the AI large model can have multi-modal data, improving the accuracy of traffic condition analysis and prediction. Secondly, by calculating the relative sensitivity indicators of each data, and after analyzing and predicting a traffic condition result, setting the current data as anchor points, based on the relative sensitivity indicators, new traffic condition results can be obtained, which can not only avoid the situation of untimely update of traffic condition analysis and prediction results, but also avoid the situation of too high update frequency of traffic condition analysis and prediction results due to the real-time change of traffic data, improving the stability and timeliness of analysis and prediction results, and having better application effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a flowchart of the steps of a method for processing, analyzing and predicting multi-modal data of an AI large model provided by the present invention. DETAILED DESCRIPTION
[0015] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0016] A method for processing, analyzing and predicting multi-modal data of an AI large model includes: S1: Collect data; wherein, the collected data at least includes road image and video data information, road noise data information, road area weather data, and pedestrian image and video data information. The road image and video data information at least includes: the number, type, speed, driving direction, lane information, and vehicle distance of vehicles; the road noise data information at least includes: the sound of vehicle driving, the sound generated by vehicle honking, and the sound made by pedestrians; the road area weather data at least includes: temperature, humidity, precipitation, visibility, wind speed; the pedestrian image and video data information at least includes: the flow rate, walking speed, walking direction, and whether pedestrians cross the road illegally. The collected data also includes collecting text information on social platforms, extracting relevant information describing the road traffic conditions, and supplementing this information into the model; S2: Preprocess the collected data. In this preprocessing, when preprocessing the road image and video data information and the pedestrian image and video data information, perform image enhancement processing on the video data to improve clarity, and perform time synchronization and frame rate adjustment. Clean the extracted vehicle information and pedestrian information to remove outliers; when preprocessing the road noise data, perform noise reduction processing on the noise data, calibrate the deviation of the data acquisition terminal device, and align the noise data according to the time series; when preprocessing the road area weather data, standardize the weather data and perform one-hot encoding on the categorical data; S3: Conduct a correlation analysis between each preprocessed data and traffic condition indicators respectively, and conduct a regression analysis to establish a mathematical model between each data and traffic condition indicators; S4: Conduct a simulation experiment in the model. First, fix the existing data, and gradually change the value of one of the data, and then observe the change in the traffic condition result output by the model, calculate the relative sensitivity index between this changed data and the traffic condition, and then calculate the relative sensitivity indexes between other data and the traffic condition in turn; S5: Extract road vehicle-related features, road noise-related features, road area weather-related features, and pedestrian-related features. At the same time, extract the relative sensitivity features of road vehicles, road noise, road weather, and pedestrians. Among them, the road vehicle-related features include: vehicle density on a specific section during a certain period, the proportion of different types of vehicles, average vehicle speed, vehicle speed standard deviation, the degree of chaos of vehicle directions, and statistical features of vehicle following distances; road noise-related features include: the mean value, standard deviation, peak frequency, noise change rate of noise, and the correlation between noise and vehicle density and vehicle speed; road area weather-related features include: dummy variables of weather types, interaction terms between visibility and vehicle speed, and relationship features between precipitation intensity and accident incidence rates; pedestrian-related features include: pedestrian flow density, the number of conflict points between pedestrians and vehicles, pedestrian illegal crossing frequencies, etc.; S6: Select a model that combines a convolutional neural network and a long short-term memory network, and perform model training. When training the model, divide each dataset into a training set, a validation set, and a test set according to a certain ratio. Use the training set to train the model, adjust the hyperparameters of the model on the validation set to prevent overfitting, and use an appropriate loss function and optimization algorithm to train the model; S7: Connect the terminals for collecting various data to the data input end of the model. After the model analyzes and predicts the result of a traffic condition, the model sets the current various data as anchor points, and based on the relative sensitivity indicators corresponding to each data, after at least one data changes beyond the corresponding relative sensitivity indicator, the model re-analyzes and predicts a new traffic condition result.
[0017] The operating principle of the present invention is described as follows: Obtain video data from vehicle monitoring cameras installed at key positions on the road, extract data such as the number, type, speed, driving direction, lane information, and vehicle distance of vehicles, and set up collection terminals such as noise sensors on this section of the road to collect noise level data, which can reflect traffic flow density and abnormal events; obtain local real-time weather information, including temperature, humidity, precipitation, visibility, wind speed, etc., which affect traffic conditions, and then obtain information such as the flow, walking speed, walking direction, and whether pedestrians cross the road illegally through pedestrian monitoring cameras; Then, perform image enhancement processing on the collected video data to improve clarity, and perform practice synchronization and frame rate adjustment. At the same time, clean the data information of the extracted video data to remove outliers (such as unreasonable vehicle speeds or misidentified vehicle types). For noise data, noise reduction processing can be performed on the noise data, sensor biases can be calibrated, and the noise data can be aligned according to the time series; for weather data, the weather data is standardized, and categorical data (such as sunny, rainy, etc.) is one-hot encoded (for example, if there are 3 weather types: sunny, rainy, and snowy, after one-hot encoding, "sunny" can be represented as [1, 0, 0], "rainy" as [0, 1, 0], and "snowy" as [0, 0, 1], and each category is encoded into a vector with only one bit being 1 and the rest being 0). Next, perform correlation analysis between each preprocessed data and traffic condition indicators respectively. Taking traffic flow as an example: calculate the correlation coefficient between traffic flow and traffic condition indicators (such as congestion index, accident incidence rate, etc.), and the Pearson correlation coefficient can be used. At the same time, perform correlation analysis between pedestrian flow, noise level, weather factors, etc. and traffic condition indicators respectively, and perform regression analysis to establish a mathematical model between each data and traffic condition indicators. For example, for traffic flow and vehicle speed, establish a linear regression model y = a*b + c (where y is the vehicle speed, b is the traffic flow, and a and c are model parameters), and determine the model parameters through the least squares method to quantify the influence degree of each factor on traffic conditions; In the model, conduct simulation experiments. First, fix all the data, and gradually change the value of one of the data, and then observe the change in the traffic condition results output by the model, calculate the relative sensitivity index of this changed data to traffic conditions, and then calculate the relative sensitivity indices of other data to traffic conditions in turn. Taking traffic flow as an example, the relative sensitivity e = (f1 / f2) / (g1 / g2), where f1 is the change amount of traffic condition results, f2 is the initial traffic condition result, g1 is the change amount of traffic flow, and g2 is the initial traffic flow (when calculating the relative sensitivity of traffic flow, keep other parameters such as weather, pedestrians, and noise unchanged), and then calculate the relative sensitivity indices of other parameters in turn; Next, extract vehicle-related features: including vehicle density on a specific section during a certain period, the proportion of different types of vehicles, average vehicle speed, vehicle speed standard deviation, the degree of chaos of vehicle directions, statistical features of vehicle following distances, etc.; noise-related features: the mean value, standard deviation, peak frequency, noise change rate of noise, and the correlation between noise and vehicle density and vehicle speed; weather-related features: dummy variables of weather types, interaction terms between visibility and vehicle speed, relationship features between precipitation intensity and accident incidence rate, etc.; pedestrian-related features: pedestrian flow density, the number of conflict points between pedestrians and vehicles, pedestrian illegal crossing frequency, etc.; Select a model that combines a convolutional neural network and a long short-term memory network. The convolutional neural network is used to extract features from image data, and the long short-term memory network is used to handle the long-term dependencies of sequential data. Training process: Divide the dataset into a training set, a validation set, and a test set according to a certain ratio. Use the training set to train the model, and adjust the hyperparameters of the model on the validation set, such as the learning rate, the number of layers, the number of neurons, etc., to prevent overfitting. Use a suitable loss function (such as mean squared error for predicting continuous values such as traffic flow) and an optimization algorithm (such as Adam) for training; When analyzing and predicting traffic conditions, connect each data collection terminal to the data input end of the model, and then run the program through a computer. The computer processor runs the model, and the model will predict a traffic condition result based on the data information input by each current data collection terminal. Relevant departments can issue guidance information or take appropriate traffic control measures in a timely manner according to this traffic condition. For example, based on the current vehicle, noise, weather, and pedestrian data, predict whether there will be congestion, the degree of congestion, and the duration of congestion on a specific road section in the future, and provide information to the traffic management department in advance so that they can take traffic control measures, such as adjusting the signal light duration, guiding vehicles to detour, etc. Or when the weather is bad, the vehicle density is high, there are many pedestrian violations, and the noise is abnormal, the model determines that the accident risk is high and can remind relevant departments to strengthen patrols and take preventive measures, etc. At the same time, after the model analyzes and predicts the result, the model sets the current various parameters as anchor points. Taking the traffic flow as an example, the anchor point of the current traffic flow is n1. Since the traffic flow changes in real time, and the data terminal for monitoring the traffic flow will send the real-time data n2 of the traffic flow to the model, the model calculates the difference between n2 and the set anchor point n1. When the difference exceeds the relative sensitivity index corresponding to the traffic flow, it indicates that due to the excessive change of the traffic flow, the traffic condition may change (for example, the traffic flow at the time of setting the anchor point n1 is in the night low period, and during the morning rush hour and other times, due to the significant increase in the traffic flow, the difference between the new traffic flow data n2 and the anchor point n1 exceeds the relative sensitivity index. At this time, traffic congestion may occur). At this time, the model will re-analyze and predict with the new traffic flow data to obtain a new traffic condition prediction result.
[0018] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for AI large model multimodal data processing, analysis and prediction, characterized in that: The analysis and prediction methods include: Collecting data; wherein the collected data at least includes road image video data information, road noise data information, road area weather data and pedestrian image video data information; Preprocess the collected data; Conduct correlation analysis on each pre-processed data with traffic condition indicators, and perform regression analysis to establish a mathematical model between each data and traffic condition indicators; Conduct simulation experiments in the model, first fix all the data, and gradually change the value of one of the data, then observe the changes in the traffic condition results output by the model, calculate the relative sensitivity index between the changed data and the traffic condition, and then calculate the relative sensitivity indexes of other data and the traffic condition in turn; Extract road vehicle related features, road noise related features, road area weather related features and pedestrian related features, and at the same time, extract the relative sensitivity features of road vehicles, the relative sensitivity features of road noise, the relative sensitivity features of road weather and the relative sensitivity features of pedestrians; Select a model that combines a convolutional neural network and a long short-term memory network, and perform model training; Connect each data collection terminal to the data input end of the model. When the model analyzes and predicts a traffic condition result, the model sets the current data as anchor points and based on the relative sensitivity index corresponding to each data, after at least one data change exceeds the corresponding relative sensitivity index, the model re-analyzes and predicts a new traffic condition result.
2. The method for processing, analyzing and predicting multimodal data using a large AI model according to claim 1, characterized in that: The road image video data information includes at least: the number, type, speed, driving direction, lane information and vehicle distance of vehicles; The road noise data information includes at least: the sound of vehicles running, the sound of vehicles honking, and the sound of pedestrians; The road area weather data at least includes: temperature, humidity, precipitation, visibility, and wind speed; The pedestrian image video data information at least includes: pedestrian flow, walking speed, walking direction, and whether the pedestrians cross the road illegally.
3. The method for processing, analyzing and predicting multimodal data using a large AI model according to claim 1, characterized in that: When preprocessing the road image video data information and pedestrian image video data information, image enhancement processing is performed on the video data to improve the clarity, and time synchronization and frame rate adjustment are performed, and data cleaning is performed on the extracted vehicle information and pedestrian information to remove abnormal values; When preprocessing road noise data, the noise data is denoised, the deviation of the data acquisition terminal equipment is calibrated, and the noise data is aligned in time series; When preprocessing the weather data of the road area, the weather data is standardized and the category data is one-hot encoded.
4. The method for processing, analyzing and predicting multimodal data using a large AI model according to claim 1, characterized in that: The road vehicle related characteristics include: vehicle density in a specific road section during a certain period of time, proportion of different types of vehicles, average vehicle speed, standard deviation of vehicle speed, degree of confusion of vehicle directions, and statistical characteristics of vehicle following distance; Road noise related characteristics include: noise mean, standard deviation, peak frequency, noise change rate, and the correlation between noise and vehicle density and speed; Weather-related characteristics of the road area include: dummy variables of weather type, interaction terms between visibility and vehicle speed, and relationship characteristics between precipitation intensity and accident rate; The relevant characteristics of pedestrians include: pedestrian flow density, the number of conflict points between pedestrians and vehicles, and the frequency of pedestrian illegal crossings.
5. The method for processing, analyzing and predicting multimodal data using a large AI model according to claim 1, characterized in that: When training the model, each data set is divided into a training set, a validation set and a test set according to a certain ratio. The model is trained using the training set, and the hyperparameters of the model are adjusted on the validation set to prevent overfitting. The model is trained using a suitable loss function and optimization algorithm.
6. The method for processing, analyzing and predicting multimodal data using a large AI model according to claim 1, characterized in that: The data collection also includes collecting text information from social platforms, extracting relevant information describing road traffic conditions, and adding the information to the model.
7. A device for AI large-scale multimodal data processing, analysis and prediction, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, it implements the method for processing, analyzing and predicting multimodal data of a large AI model as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the processor executes a method for processing, analyzing and predicting multimodal data of an AI large model as described in any one of claims 1 to 6.