Vehicle scheduling optimization method based on Internet of Things

By applying technical means such as Kalman filtering, interpolation, time series prediction, multi-factor regression, adaptive thresholding, support vector machines and genetic algorithms in the vehicle scheduling system, problems such as inaccurate data and difficult early warning rules setting in vehicle scheduling are solved, and efficient and accurate vehicle status monitoring and scheduling optimization are achieved.

CN120163397AActive Publication Date: 2025-06-17LUSHAN COLLEGE OF GUANGXI UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510324777.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-17
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

During the vehicle scheduling process, the data collected by the on-board sensors may have problems such as inaccuracy, incompleteness, and delay, resulting in the vehicle status information inconsistent with the actual situation. It is difficult to avoid false alarms or missed reports in the early warning rule setting, and dynamic adjustment of the scheduling strategy also faces high requirements for computing capabilities.

Method used

The Kalman filtering algorithm is used to smooth the original data, the interpolation algorithm fills the missing values, the time series prediction model is corrected in real time, combined with road conditions information and historical driving records, and dynamic evaluation is used using a multi-factor regression model, the adaptive threshold algorithm adjusts the early warning rules, the support vector machine classification algorithm filters the early warning signals, and the genetic algorithm generates the optimal scheduling scheme.

Benefits of technology

It improves the accuracy and scheduling efficiency of vehicle status monitoring, reduces false alarms and missed alarms, and realizes the rapid processing of massive data and the generation of optimal scheduling solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163397A_ABST
    Figure CN120163397A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle scheduling optimization method based on the Internet of Things, and relates to the technical field of the Internet of Things, and the method comprises the steps: obtaining the original data of a vehicle state, carrying out the smoothing of the original data through a Kalman filtering algorithm, and obtaining first processing data; according to the first processing data, filling a missing value by using an interpolation algorithm to obtain second processing data; key state parameters in the second processing data are extracted, the parameters are corrected in real time through a time sequence prediction model, and third processing data are obtained; according to the fifth processing data, performing secondary screening on the early warning signal by using a support vector machine classification algorithm to obtain sixth processing data; and extracting a vehicle scheduling demand in the sixth processing data. According to the method, the optimal scheduling scheme is quickly generated by using the Apache Spark distributed computing framework and the genetic algorithm, the accuracy of vehicle state monitoring and the scheduling efficiency are effectively improved, and comprehensive technical support is provided for motorcade management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the Internet of Things, and particularly to a vehicle scheduling optimization method based on the Internet of Things. Background Art

[0002] During the vehicle scheduling process, it is necessary to collect and analyze various state parameters of the vehicle in real time, so as to timely understand the running situation of the vehicle and adjust the scheduling strategy accordingly. However, the data collected by on-vehicle sensors may have problems such as inaccuracy, incompleteness, and delay, resulting in the vehicle state information obtained being inconsistent with the actual situation. In addition, external factors such as traffic congestion, accidents, and weather will also affect the driving state of the vehicle, making it difficult to accurately estimate the expected arrival time. When triggering the warning mechanism, how to set reasonable warning rules according to the thresholds of vehicle state parameters to avoid excessive false alarms or missed alarms is also an urgent problem to be solved. At the same time, the dynamic adjustment of the scheduling strategy needs to comprehensively consider various factors, such as the location distribution of vehicles, the cargo loading volume, the remaining fuel volume, etc. How to quickly find the optimal scheduling plan from the massive data puts forward high requirements for the computing power of the scheduling system. The solution of these technical problems requires optimization and innovation in multiple links such as data collection, transmission, storage, analysis, and decision-making to improve the accuracy and efficiency of vehicle scheduling. Summary of the Invention

[0003] The present invention provides a vehicle scheduling optimization method based on the Internet of Things, mainly including: Obtaining the original vehicle state data, using the Kalman filter algorithm to smooth the original data to obtain the first processed data; according to the first processed data, using the interpolation algorithm to fill in the missing values to obtain the second processed data; extracting the key state parameters in the second processed data, using the time series prediction model to perform real-time correction on the parameters to obtain the third processed data; according to the third processed data, combining road condition information, weather forecast data, and historical driving records, using the multi-factor regression model to dynamically evaluate the vehicle driving state to obtain the fourth processed data; extracting the vehicle location distribution, cargo loading volume, and remaining fuel volume information in the fourth processed data, using the adaptive threshold algorithm to dynamically adjust the warning rules to obtain the fifth processed data; according to the fifth processed data, using the support vector machine classification algorithm to perform secondary screening on the warning signals to obtain the sixth processed data; extracting the vehicle scheduling requirements in the sixth processed data, using the distributed computing framework to perform parallel optimization on the scheduling strategy to obtain the seventh processed data; according to the seventh processed data, using the genetic algorithm to quickly generate the optimal scheduling plan to obtain the final scheduling result.

[0004] Further, the Kalman filter algorithm is used to smooth the original data to obtain the first processed data, including: acquiring the vehicle speed, acceleration, and direction angle data collected by vehicle-mounted sensors, establishing a Kalman filter model, and setting the process noise covariance matrix and the measurement noise covariance matrix; constructing the state transition matrix, the observation matrix, and the noise matrix of the Kalman filter; inputting the original data into the Kalman filter model, and recursively estimating the optimal values of the vehicle states at each moment through prediction and update; in the prediction step, predicting the prior state estimate and the error covariance matrix at the current moment according to the state estimate value at the previous moment and the state transition matrix; in the update step, calculating the Kalman gain, and using the observation value at the current moment to correct the predicted value to obtain the posterior state estimate and the updated error covariance matrix; iteratively executing the prediction and update steps until all the original data of the vehicle states at all time points are processed; evaluating the vehicle state estimate values after Kalman filter processing, and if the error exceeds the preset threshold, adjusting the parameters of the process noise covariance matrix and the measurement noise covariance matrix of the Kalman filter.

[0005] Further, the interpolation algorithm is used to fill in the missing values to obtain the second processed data, including: constructing a data model according to the distribution characteristics and change trends of the existing data to predict the value range of the missing data; for the missing continuous numerical values, using the linear interpolation method for estimation and filling; for the missing categorical data, using the nearest neighbor algorithm for interpolation; detecting outliers in the filled data, analyzing the distribution characteristics of the outliers, setting the outlier threshold, and removing the outlier data points introduced by the interpolation process; expanding the data scale of the data after removing the outliers through data augmentation techniques; for the continuous numerical values, generating new data samples by adding random noise; for the categorical data, using the synthetic minority over-sampling technique algorithm to synthesize new minority class samples to balance the class distribution of the data set; dividing the expanded data set into a training set and a test set; applying different interpolation algorithms to fill in the missing data, comparing and evaluating the prediction performance of each algorithm on the training set and the test set, selecting the interpolation model with the best filling effect, and using it to estimate all the missing values in the original data set.

[0006] Furthermore, a time series prediction model is used to perform real-time correction on the parameters to obtain third processed data, including: obtaining key parameter values reflecting the operating state of the device; using the obtained historical parameter value sequence to establish an autoregressive integrated moving average model; adopting the established autoregressive integrated moving average model, and according to the historical state parameter value sequence of the device, performing real-time prediction and correction on the state parameters obtained with a delay at the current moment to obtain the true operating parameters of the device at the current moment; substituting the corrected current state parameters into the autoregressive integrated moving average model, and performing continuous rolling prediction on the device state parameters within a subsequent period of time in the form of a sliding time window to obtain a predicted parameter sequence for a future period of time; setting a normal value range threshold according to the predicted device operating parameter sequence, and determining whether the predicted parameters exceed the threshold; if the predicted parameters at a certain moment exceed the normal range, an early warning message is issued; obtaining the real-time operating parameters of the device, and performing real-time comparison with the parameter sequence predicted by the autoregressive integrated moving average model to calculate the deviation value between the actual parameters and the predicted parameters.

[0007] Furthermore, in combination with road condition information, weather forecast data, and historical driving records, a multi-factor regression model is used to dynamically evaluate the vehicle driving state to obtain fourth processed data, including: obtaining real-time road congestion conditions and accident information; obtaining temperature, humidity, and visibility data within a future period of time; reading the historical driving data of the vehicle, where the historical driving data includes driving speed, acceleration, and braking frequency data; inputting the obtained real-time road condition information, the weather forecast data, and the historical driving record data into a vehicle driving state evaluation model pre-trained based on a neural network algorithm; the vehicle driving state evaluation model calculates the driving state evaluation score of the current vehicle through weight assignment and comprehensive analysis of various factors; comparing the vehicle driving state score calculated by the vehicle driving state evaluation model with a preset safety threshold; if the score exceeds the safety threshold, it is determined that the current vehicle is in a dangerous driving state, and a safety warning is automatically issued to the driver; comparing and analyzing the vehicle driving state evaluation score with the data in the historical driving record database of the vehicle, and identifying whether there is an abnormal change trend in the vehicle driving state through a clustering algorithm.

[0008] Further, an adaptive threshold algorithm is used to dynamically adjust the warning rules to obtain the fifth processed data, including: dividing the vehicles into regions by using a density-based clustering algorithm according to the vehicle position distribution information to obtain the vehicle distribution density in different regions; for each region, obtaining the cargo loading amount and remaining fuel amount data of the vehicles in the region, removing outliers and redundant data to obtain a normalized feature data set; defining the judgment criteria for the vehicle state according to the vehicle historical operation data, and labeling the feature data set to obtain the label data of the vehicle state; using the label data and the feature data, training the association model between the cargo loading amount and remaining fuel amount and the vehicle state through a logistic regression algorithm to obtain the parameters and thresholds of the association model; during the vehicle operation, collecting the cargo loading amount and remaining fuel amount data of the vehicle in real time to construct time series data; filtering the time series data by using a Kalman filter algorithm to estimate the optimal values of the cargo loading amount and remaining fuel amount; inputting the filtered cargo loading amount and remaining fuel amount data into the association model to judge the vehicle state.

[0009] Further, a support vector machine classification algorithm is used to perform a secondary screening on the warning signals to obtain the sixth processed data, including: extracting the features of the warning signals to form a warning signal feature vector; using the labeled warning signal sample data, using the support vector machine algorithm, and training to obtain a support vector machine classification model by adjusting the kernel function and penalty coefficient parameters; for the new warning signals, inputting their feature vectors into the trained support vector machine classification model for binary classification prediction; the support vector machine classification model will divide the warning signals into two categories: normal signals and abnormal signals; for the warning signals predicted as abnormal signals by the support vector machine classification model, marking them as potential false alarms or missed alarms; setting a confidence threshold, and when the confidence of the abnormal signal is lower than the confidence threshold, determining that the warning signal is a false alarm or a missed alarm and filtering it out; retaining the warning signals with a confidence higher than the confidence threshold as the sixth processed data; screening out the warning signals with a higher confidence from the sixth processed data according to the classification results of the support vector machine classification model as the final effective warning signals.

[0010] Furthermore, a genetic algorithm is used to quickly generate an optimal scheduling plan and obtain the final scheduling result, including: initializing a set of scheduling plans as the initial population, where each scheduling plan contains the execution order of tasks and the allocated resource information; for each scheduling plan in the population, according to factors such as task priority, execution time, and resource utilization rate, the fitness value of the plan is calculated by weighted summation as an index to evaluate the quality of the scheduling plan; selection, crossover, and mutation operations are performed on the population according to the fitness value to generate new scheduling plans; the selection operation uses the roulette wheel selection method to select individuals from the current population according to the magnitude of the fitness value, and the larger the fitness value, the higher the probability of being selected; the crossover operation uses the two-point crossover method to randomly select two crossover points and swap the gene segments between the two selected individuals between the crossover points; the mutation operation uses the basic bit mutation method to invert certain genes of an individual with a certain probability to introduce new gene combinations; it is judged whether the maximum fitness value of the current population has not changed significantly after multiple consecutive iterations or has reached the preset maximum number of iterations; if satisfied, the scheduling plan with the highest fitness value is output as the optimal plan, the execution order of tasks and the resource allocation situation are determined, and the final scheduling result is generated.

[0011] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: The present invention discloses a vehicle scheduling optimization method based on the Internet of Things. Aiming at the problems of inaccurate, incomplete, and delayed in-vehicle sensor data, the present invention uses the Kalman filter, interpolation algorithm, and time series prediction model for data processing and correction. Combining real-time road conditions, weather, and historical records, the driving state of the vehicle is dynamically evaluated through a multi-factor regression model. To solve the problems of state threshold setting and false alarms and missed alarms, the present invention uses an adaptive threshold algorithm and a support vector machine classification algorithm to optimize the warning rules. Finally, aiming at the massive data processing and real-time scheduling requirements of multiple vehicles, multiple routes, and multiple target points, the present invention uses the Apache Spark distributed computing framework and a genetic algorithm to quickly generate an optimal scheduling plan. This method effectively improves the accuracy of vehicle state monitoring and scheduling efficiency, providing comprehensive technical support for fleet management. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a flowchart of a vehicle scheduling optimization method based on the Internet of Things according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this specification without making creative efforts shall fall within the scope of protection of this specification.

[0014] As Figure 1 , a vehicle scheduling optimization method based on the Internet of Things in this embodiment may specifically include: Step S101, obtain the original vehicle state data from in-vehicle sensors. For the problem of inaccurate data, use the Kalman filter algorithm to smooth the original data to obtain the first processed data.

[0015] Obtain the vehicle speed, acceleration, direction angle and other state data collected by the in-vehicle sensors in real time. For the problem of inaccurate data, perform data preprocessing using the Kalman filter algorithm. According to the characteristics of the vehicle state data, establish a Kalman filter model and set appropriate process noise covariance matrix Q and measurement noise covariance matrix R parameters. The covariance matrices Q and R respectively represent the uncertainty of vehicle state estimation and the uncertainty of measurement noise, and need to be reasonably set according to the vehicle motion characteristics and sensor performance. According to the set Q and R parameters, construct the state transition matrix, observation matrix and noise matrix of the Kalman filter to form a complete Kalman filter model. Input the obtained original vehicle state data into the Kalman filter model one by one at each time point, and recursively estimate the optimal value of the vehicle state at each moment through two steps: prediction and update. In the prediction step, according to the state estimation value and state transition matrix at the previous moment, predict the state prior estimate and error covariance matrix at the current moment. In the update step, calculate the Kalman gain and use the observation value at the current moment to correct the predicted value to obtain the state posterior estimate and the updated error covariance matrix. Iteratively execute the prediction and update steps until all the original vehicle state data at all time points are processed. A fixed time window, such as 5 seconds, can be set to determine the termination condition of the iteration. For the vehicle state estimation value after Kalman filter processing, use indicators such as mean square error to evaluate its smoothness and accuracy relative to the original data. If the error exceeds the preset threshold, appropriately adjust the Q and R parameters of the Kalman filter according to the evaluation results to improve the data smoothing effect. Optimization algorithms such as grid search can be used to search for the optimal combination of Q and R parameters within the value range. By continuously iterating the Kalman filter processing and parameter optimization process, finally obtain smooth and accurate vehicle state estimation data as the input basis for subsequent vehicle control and decision-making.

[0016] Specifically, the Kalman filter algorithm is an excellent data preprocessing method, especially suitable for processing vehicle sensor data. This algorithm estimates the vehicle state recursively by establishing a state space model and combining measurement values and prior estimates. In vehicle state data processing, it is first necessary to determine the state vector, which usually includes position, speed, and acceleration, etc. Taking the longitudinal motion of the vehicle as an example, the state vector can be defined as [position, speed, acceleration]. The state transition matrix describes the evolution of the vehicle state over discrete time. Assuming the sampling interval is Δt, it can be expressed as: [[1, Δt, 0.5Δt^2], [0, 1, Δt], [0, 0, 1]]. The observation matrix is determined according to the actually measurable state variables. If only the position and speed can be directly measured, the observation matrix is: [[1, 0, 0], [0, 1, 0]]. The settings of the process noise covariance matrix Q and the measurement noise covariance matrix R are crucial. Q reflects the uncertainty of the state space model, and R reflects the uncertainty of the measurement. For example, for a high-precision GPS receiver, its position measurement error may be at the meter level, and the speed measurement error is at the centimeter / second level. Based on this, the diagonal elements of the R matrix can be initially set. In practical applications, the sliding time window method can be used to process continuous data streams. Assuming the window size is 5 seconds and the sampling frequency is 10Hz, then 50 data points are processed each time. For the data within each time window, the prediction and update steps of the Kalman filter are executed. In the prediction step, the prior estimate at the current moment is calculated using the optimal estimate at the previous moment and the state transition matrix. The update step combines the observed values and corrects the prior estimate through the Kalman gain to obtain the posterior estimate. To evaluate the filtering effect, the root mean square error (RMSE) between the filtered state estimate value and the original measurement value can be calculated. If the RMSE exceeds a preset threshold, such as the position error is greater than 0.5 meters or the speed error is greater than 0.2 meters / second, then the Q and R matrices need to be adjusted. The grid search method can be used during the adjustment process. For example, search for the diagonal elements of the Q matrix within the range of [10^-6, 10^-5,..., 10^-1], and search for the diagonal elements of the R matrix within the range of [10^-2, 10^-1,..., 10^2]. Perform the Kalman filter for each set of parameter combinations, calculate the RMSE, and select the parameter combination that minimizes the RMSE. Through repeated iterative optimization processes, the finally obtained vehicle state estimation data will be smoother and more accurate. These processed data can be used for advanced driver assistance functions such as vehicle trajectory prediction and adaptive cruise control, improving driving safety and comfort. For example, in adaptive cruise control, using accurate speed and acceleration estimates, the following distance from the vehicle ahead can be adjusted more accurately, achieving a smoother acceleration and deceleration process.

[0017] Step S102, according to the first processed data, for the problem of incomplete data, use the interpolation algorithm to fill in the missing values to obtain the second processed data.

[0018] According to the distribution characteristics and changing trends of existing data, use the data mining tool SPSSModeler to construct a data model to predict the possible value ranges of missing data. For missing continuous numerical values, use the linear interpolation method for estimation and filling; for missing categorical data, use the nearest neighbor algorithm for interpolation. Export the filled data as the second data set. Perform outlier detection on the second data set, use statistical charts such as box plots to analyze the distribution characteristics of outliers, and eliminate the outlier data points introduced by the interpolation process by setting appropriate thresholds. Save the data after removing outliers as the third data set. Based on the third data set, expand the scale of training data through data augmentation techniques. For continuous numerical values, generate new data samples by adding random noise; for categorical data, use the SMOTE algorithm to synthesize new minority class samples to balance the class distribution of the data set. Randomly divide the expanded data set into a training set and a test set. Use the Scikit-learn library in Python to fill in the missing data using different interpolation algorithms such as linear interpolation and spline interpolation respectively, and compare and evaluate the prediction performance of each algorithm on the training set and the test set. By calculating evaluation indicators such as root mean square error and mean absolute error, select the interpolation model with the best filling effect, and use it to estimate all missing values in the original data set to obtain a complete data set. Finally, randomly divide the filled complete data set into a training set and a validation set according to a ratio of 8:2. Use the XGBoost algorithm to train the model on the training set and evaluate the generalization performance of the model on the validation set. Optimize the hyperparameters of XGBoost through grid search to obtain the machine learning model with the best performance for subsequent prediction tasks.

[0019] Specifically, when building a data model, the data mining tool SPSSModeler can use algorithms such as decision trees and neural networks to predict the possible value ranges of missing data. For example, for an automotive sales dataset, it may be necessary to predict the missing "sales price". By analyzing the existing data, it is found that there is a significant correlation between "vehicle model", "mileage", and "age" and the "sales price". Using these features, SPSSModeler can build a prediction model to provide a reasonable estimated range for the missing "sales price". For the missing continuous numerical values, the linear interpolation method is a simple and effective filling method. Suppose there is an environmental monitoring dataset with missing temperature data at a certain time point. The temperature values of adjacent time points can be used for linear interpolation. If the temperatures before and after the missing point are 20°C and 22°C respectively, and the time intervals are equal, then the temperature at the missing point can be estimated to be 21°C. This method assumes that the temperature changes linearly in a short period and is applicable to data with gentle changes. The k-nearest neighbor algorithm is very effective in dealing with missing categorical data. Taking medical diagnosis data as an example, if the "symptom type" information of a patient is missing, patient records with similar characteristics (such as age, gender, other symptoms) can be searched, and the symptom type that appears most frequently among these similar patients can be used to fill the missing value. This method is based on the assumption that similar patients may have similar symptoms and can maintain the internal relevance of the data. Using box plots for outlier detection is an intuitive and effective method. In a student grade dataset, box plots of each subject's grades can be drawn. Suppose the box plot of the math grade shows that the upper quartile is 90 points and the lower quartile is 60 points. Then any grade below 15 points (lower quartile minus 1.5 times the interquartile range) or above 135 points (upper quartile plus 1.5 times the interquartile range) may be regarded as an outlier. These outliers may be caused by data entry errors or improper interpolation and need to be further verified or removed. Data augmentation techniques play an important role in expanding the scale of training data. For continuous numerical values, such as stock price data, small random fluctuations can be added to the original price to generate new samples. For example, for a stock price of 100 yuan, random noise can be added within the range of ±5% to generate new price data between 95 yuan and 105 yuan. This method can simulate the small fluctuations in the market and increase the sensitivity of the model to price changes. The SMOTE algorithm is very effective in dealing with imbalanced datasets. Suppose in a credit card fraud detection dataset, fraudulent transactions only account for 1% of the total transactions. The SMOTE algorithm can be used to synthesize new fraudulent transaction samples to increase the proportion of fraudulent transactions to 10%. This balance can help the model better learn the characteristics of fraudulent transactions and improve the detection accuracy. When evaluating the performance of interpolation algorithms, the root mean square error (RMSE) and mean absolute error (MAE) are commonly used metrics. For example, in a weather forecast dataset, it may be necessary to fill in the missing rainfall data.Assume that the RMSE of linear interpolation is 2.5 mm, while the RMSE of spline interpolation is 1.8 mm. One would tend to choose spline interpolation as the final filling method. Spline interpolation can better capture the non-linear variation characteristics of rainfall, so it performs better in this case. The XGBoost algorithm performs well in dealing with structured data. In a customer churn prediction model, there may be features such as "customer age", "monthly consumption amount", "number of customer service contacts", etc. XGBoost can automatically learn the complex interaction relationships between these features, for example, discover rules like "customers aged between 30 and 40 with a monthly consumption amount less than 100 yuan are more likely to churn". By optimizing hyperparameters such as the learning rate and tree depth through grid search, a model that can accurately predict the churn risk without overfitting can be obtained.

[0020] Step S103: Extract key state parameters from the second processed data. For the data delay problem, use a time series prediction model to correct the parameters in real time to obtain the third processed data.

[0021] According to the processed data, obtain the key parameter values reflecting the operating status of the device, such as temperature, pressure, rotational speed, etc. For the possible delay problems during the data transmission process, use the obtained historical parameter value sequence to establish an ARIMA time series prediction model. Adopt the established ARIMA model, and based on the historical state parameter value sequence of the device, perform real-time prediction and correction on the state parameters obtained with delay at the current moment to obtain the true operating parameters of the device at the current moment. Substitute the corrected current state parameters into the ARIMA model, and use a sliding time window method to continuously and recursively predict the device state parameters within a subsequent period of time to obtain a predicted parameter sequence for a future period of time. According to the predicted device operating parameter sequence, set the normal value range threshold, and determine whether the predicted parameters exceed the threshold. If the predicted parameters at a certain moment exceed the normal range, send a warning message by means of email, text message, etc., indicating that the device may be about to have an abnormal operation, and notify relevant personnel to make preparations in advance. Obtain the real-time operating parameters of the device, compare them with the parameter sequence predicted by the ARIMA model in real time, and calculate the deviation value between the actual parameters and the predicted parameters. Set the deviation threshold. If the actual deviation exceeds the threshold, it is determined that the device operation is abnormal, trigger an abnormal alarm, and display the abnormal information through the human-machine interface. At the same time, feedback the abnormal parameters to the ARIMA prediction model, and use the deviation value to correct the relevant coefficient terms in the model so that the model can adapt to the changes in the device state. According to the adjusted and optimized ARIMA model parameters, re-predict the device operating parameters and continuously compare and monitor them with the device parameters obtained in real time until the device returns to the normal working state. Finally, use the optimized and corrected ARIMA model as the basis for predicting the device operating status and abnormal diagnosis, output the parameters of the ARIMA model, as well as the corrected device operating parameter sequence predicted by the model, to provide data support for the health management of the device.

[0022] Specifically, the ARIMA model is a commonly used time series forecasting method suitable for data with trends and seasonality. In equipment condition monitoring, ARIMA can effectively capture the patterns of parameter changes. Taking an industrial cooling system as an example, temperature is a key parameter. Suppose historical data shows that the temperature fluctuates between 20 - 25°C and exhibits a similar daily change pattern. The ARIMA model can learn this pattern and predict future temperature changes. For the problem of data transmission delay, the rolling prediction function of the ARIMA model is very useful. For example, if sensor data is transmitted every 5 minutes but there is actually a 10 - minute delay, the ARIMA model can predict the current temperature based on historical data. This prediction can fill the data gap caused by the delay and enable the monitoring module to respond in a timely manner. Setting the normal value range threshold is the key to the warning mechanism. Continuing with the cooling system example, if the normal operating temperature range is 22 - 26°C, 27°C can be set as the warning threshold. When the ARIMA model predicts that the temperature will reach 27°C within 30 minutes, a warning will be issued. This gives the operator enough time to take preventive measures, such as increasing the cooling water flow or checking the radiator. Real - time comparison of the deviation between the actual parameter and the predicted parameter is an important mechanism for model adaptation. Suppose the ARIMA model predicts a temperature of 24°C at a certain moment, but the actual measured value is 25.5°C. If the set deviation threshold is 1°C, this 1.5°C difference will trigger an anomaly alarm. At the same time, this deviation will be used to adjust the parameters of the ARIMA model to make future predictions more accurate. The continuous optimization of the model is crucial for long - term monitoring. For example, as the components of the cooling system age, its efficiency may gradually decline, resulting in a slight increase in the overall temperature. The optimized ARIMA model can adapt to this slow change and avoid frequent false alarms. The parameter sequence output by the ARIMA model provides comprehensive data support for equipment health management. For example, by analyzing the temperature change rate, it is possible to predict a decline in radiator efficiency. If it is found that the temperature is rising 20% faster than historical data, this may indicate that maintenance is required. This proactive maintenance strategy can significantly reduce the risk of equipment failure and improve reliability. In practical applications, the ARIMA model can be used in combination with other technologies. For example, the prediction results of ARIMA can be input into a machine learning model to further improve the accuracy of anomaly detection. This multi - model fusion method can simultaneously utilize the advantages of time series analysis and complex pattern recognition to provide a more comprehensive solution for equipment condition monitoring.

[0023] Step S104, based on the third processed data, combined with the road condition information, weather forecast data, and historical driving records obtained in real - time from the traffic management department, use a multi - factor regression model to dynamically evaluate the vehicle driving state and obtain the fourth processed data.

[0024] Through the API interface of the traffic management department, obtain real-time road congestion conditions and accident information, which serve as one of the input factors for vehicle driving state assessment. Access the public database of the meteorological bureau to obtain weather forecast data such as temperature, humidity, and visibility within a certain period in the future, and use it as another input factor for vehicle driving state assessment. Read the historical driving data of the vehicle from the in-vehicle OBD device, including driving speed, acceleration, braking frequency, etc., and upload it to the data center for storage and management, serving as the basic data for evaluating the vehicle driving state. Input multiple factors such as the obtained real-time road condition information, weather forecast data, and historical driving records into the vehicle driving state assessment model pre-trained based on neural network algorithms. The vehicle driving state assessment model calculates the driving state assessment score of the current vehicle through weight assignment and comprehensive analysis of various factors. Compare the vehicle driving state score calculated by the vehicle driving state assessment model with the preset safety threshold. If the score exceeds the threshold, it is determined that the current vehicle is in a dangerous driving state, and a voice and image safety warning is automatically sent to the driver, prompting the driver to take measures such as decelerating and stopping to avoid risks. Compare and analyze the vehicle driving state assessment score with the data in the vehicle's historical driving record database, and identify whether there is an abnormal change trend in the vehicle driving state through clustering algorithms. If the scores of consecutive multiple assessments show an upward trend, it is predicted that the driving risk of the vehicle will continue to increase in the future. According to the vehicle driving state risk prediction result, automatically adjust the driving parameters such as the cruise speed and following distance of the vehicle dynamically. When the predicted risk level is relatively high, control the vehicle to decelerate by 10% and increase the following distance to more than 100 meters to ensure driving safety. When the predicted risk level returns to normal, restore the cruise speed and following distance to the original set values to improve driving efficiency.

[0025] Specifically, the acquisition of real-time road congestion conditions and accident information is crucial for vehicle driving state assessment. For example, in a large city like Beijing, the traffic management department may provide API interfaces to return the congestion index and the latest accident reports for each major road. The data can be updated every 5 minutes, quantifying the congestion index of the Third Ring Road from 0 to 10, while recording the specific location and time of the accident. These information directly affect the driving state assessment, because congestion or accidents may cause the vehicle to accelerate and decelerate frequently, increasing the risk of collision. Meteorological conditions have a significant impact on driving safety. By accessing the database of the meteorological bureau, the hourly weather forecast for the next 24 hours can be obtained. For example, the forecast shows that the visibility will drop below 200 meters and the temperature will be close to 0°C in two hours, which means that icing may occur. Incorporating this information into the vehicle driving state assessment model can adjust the vehicle driving parameters in advance, such as reducing the recommended speed and increasing the safety distance. The historical driving data recorded by the in-vehicle OBD device provides a personalized basis for assessment. Suppose a vehicle has an average driving speed of 60 km / h, a maximum acceleration of 2 m / s², and brakes 3 times per kilometer in the past month. After these data are uploaded to the data center, they can be used to establish a baseline of the vehicle's normal driving mode, and any significant deviation from this baseline may be regarded as a potential risk. The vehicle driving state assessment model trained by neural network algorithms can comprehensively analyze multiple factors. The vehicle driving state assessment model may give 30% weight to real-time road conditions, 20% weight to weather forecasts, and 50% weight to historical driving data. For example, when it is detected that there is a traffic accident 2 kilometers ahead, the current visibility is low, and the vehicle's recent braking frequency is higher than the average level, the vehicle driving state assessment model may output a score from 0 to 100, and a score above 80 is considered a high-risk state. When the assessment score exceeds the safety threshold, the vehicle's automatic warning function is crucial. If the set safety threshold is 75 points, when the score reaches 78 points, the in-vehicle system will immediately issue an audible warning: "There is an accident 2 kilometers ahead, please slow down", and at the same time display the recommended deceleration amplitude and safe following distance on the central control screen. By analyzing historical data through clustering algorithms, abnormal driving state change trends can be identified. For example, if a vehicle's assessment score shows an upward trend in the past week, gradually rising from an average of 65 points to 75 points, it is predicted that without taking measures, the risk score of the next trip may break through the high-risk limit of 80 points. Based on the risk prediction results, the vehicle's driving parameters can be actively adjusted. When a high risk is predicted, if the vehicle's original cruise speed is set at 100 km / h, it will be automatically reduced to 90 km / h. At the same time, if the original following distance is set at 50 meters, it will be increased to more than 100 meters. These adjustments can effectively reduce the risk of collision and improve driving safety.

[0026] Step S105: Extract vehicle location distribution, cargo loading volume, and remaining fuel information from the fourth processed data. For the problem of setting status thresholds, adopt an adaptive threshold algorithm based on historical data and current operating conditions to dynamically adjust the warning rules, and obtain the fifth processed data.

[0027] According to the vehicle location distribution information, use the DBSCAN density clustering algorithm to divide the vehicles into regions to obtain the vehicle distribution density in different regions. For each region, obtain the cargo loading volume and remaining fuel data of the vehicles in that region. Through data cleaning and feature engineering, remove outliers and redundant data to obtain a normalized feature dataset. According to the vehicle historical operation data, define the judgment criteria for vehicle status, label the feature dataset, and obtain the label data of vehicle status. Using the label data and feature data, train the association model between cargo loading volume, remaining fuel, and vehicle status through the logistic regression algorithm to obtain the parameters and thresholds of the model. During the vehicle operation, collect the cargo loading volume and remaining fuel data of the vehicle in real time to construct time series data. Use the Kalman filter algorithm to filter the time series data to estimate the optimal values of the cargo loading volume and remaining fuel, and reduce the influence of measurement noise. Input the filtered cargo loading volume and remaining fuel data into the association model to judge the vehicle status. If the judgment result exceeds the threshold range, trigger a warning of the corresponding level. According to the warning level, match the preset vehicle scheduling strategy to dynamically adjust the transportation tasks of the warning vehicles, and reduce the risks of vehicle overloading and insufficient fuel. Output the adjusted vehicle scheduling strategy and transportation task allocation results to form an optimized vehicle operation plan to guide the safe operation and efficient scheduling of the vehicle. At the same time, transmit the vehicle operation plan to relevant departments and personnel to timely grasp the vehicle status and task progress for easy coordination and management.

[0028] Specifically, vehicle location distribution information is a key factor in logistics scheduling. The DBSCAN density clustering algorithm can effectively identify the vehicle distribution density in different regions, providing a basis for subsequent scheduling decisions. For example, in the Beijing area, the algorithm may divide the area within the Fifth Ring Road into high-density regions, while the suburbs are low-density regions. This division helps optimize vehicle allocation and improve transportation efficiency. Obtaining data on the cargo loading capacity and remaining fuel of vehicles within a region is an important basis for evaluating vehicle status. During the data cleaning and feature engineering process, it may be found that the reported loading capacity of some vehicles is abnormally high, such as exceeding 120% of the vehicle's maximum load. Such data should be regarded as outliers and excluded to ensure the accuracy of subsequent analysis. Defining vehicle status judgment criteria requires considering multiple factors. For example, vehicles with a loading capacity exceeding 90% and remaining fuel below 20% can be marked as "high-risk" status. Such criteria help identify vehicles that may face risks of overloading or fuel shortage in a timely manner. When training the association model between cargo loading capacity and remaining fuel and vehicle status using the logistic regression algorithm, it may be found that the influence weight of the loading capacity on vehicle status is 0.7, while the influence weight of the remaining fuel is 0.3. This means that changes in the loading capacity have a more significant impact when judging vehicle status. The Kalman filter algorithm can effectively reduce the influence of measurement noise when processing real-time collected data. Suppose the cargo loading capacity sensor of a vehicle is affected by vibration interference, resulting in large fluctuations in data within a short period. The Kalman filter can smooth these fluctuations and provide a more stable estimated value, thereby improving the accuracy of status judgment. The early warning mechanism triggered when the vehicle status judgment result exceeds the threshold range is the key to ensuring transportation safety. For example, when it is detected that the status score of a vehicle reaches 85 points (assuming the threshold is 80 points), a yellow early warning will be issued immediately. Such timely early warning can prevent problems before they occur and avoid potential safety hazards. After the early warning is triggered, the transportation tasks will be dynamically adjusted according to the preset scheduling strategy. For example, for vehicles about to be overloaded, some of the cargo may be transferred to nearby vehicles with lighter loads. Such dynamic adjustment can not only reduce risks but also improve the overall transportation efficiency. The optimized vehicle operation plan is the result of comprehensively considering various factors. It may include specific instructions such as suggesting that some vehicles refuel in advance and adjusting the driving route to avoid congested areas. These optimization measures can significantly improve the overall operation efficiency and safety of the fleet. Transmitting the operation plan to relevant departments and personnel helps achieve collaborative management. For example, the dispatching center can keep track of the status and task progress of each vehicle in real time and respond quickly when abnormal situations are found. Such an information sharing mechanism can greatly improve the flexibility of the logistics system and its ability to respond to emergencies.

[0029] Step S106, according to the fifth processed data, for the problem of false alarms and missed alarms, use the support vector machine classification algorithm to perform secondary screening on the warning signals to obtain the sixth processed data.

[0030] Extract the characteristics of the warning signal according to the fifth processed data, including signal strength, duration, frequency, etc., to form a warning signal feature vector. Using the labeled warning signal sample data, use the support vector machine algorithm to train a support vector machine classification model by adjusting parameters such as the kernel function and penalty coefficient. For the new warning signal, input its feature vector into the trained support vector machine classification model for binary classification prediction. The support vector machine classification model will classify the warning signal into two categories: normal signal or abnormal signal. For the warning signal predicted as an abnormal signal by the support vector machine classification model, mark it as a potential false alarm or missed alarm. Set a confidence threshold. When the confidence of the abnormal signal is lower than this confidence threshold, determine that the warning signal is a false alarm or missed alarm and filter it out. Retain the warning signals with a confidence higher than the confidence threshold as the sixth processed data. According to the classification results of the support vector machine classification model, screen out the warning signals with higher confidence from the sixth processed data as the final effective warning signals. Transmit the selected effective warning signals to the subsequent business processes for guiding relevant decisions and disposals. By using the support vector machine classification model to perform secondary classification and filtering on the warning signals, the false alarm rate and missed alarm rate can be effectively reduced, and the accuracy and reliability of the warning signals can be improved. At the same time, the support vector machine classification model can be retrained according to the continuously updated sample data to continuously optimize the classification effect and adapt to the changes in actual business requirements.

[0031] Specifically, the extraction of warning signal features is the key to achieving accurate classification. Signal intensity can reflect the severity of anomalies, duration indicates the persistence of problems, and frequency reveals the pattern of anomaly occurrences. For example, in a logistics system, the intensity of a vehicle overloading warning signal may be proportional to the overloading amount, the duration reflects the maintenance time of the overloaded state, and the frequency indicates the frequency of repeated overloading of the vehicle. The application of the support vector machine algorithm in warning signal classification has certain advantages. By adjusting the kernel function, such as selecting the radial basis function (RBF) kernel, it can effectively handle non-linearly separable warning signals. The adjustment of the penalty coefficient balances the complexity and classification accuracy of the support vector machine classification model. For example, in vehicle status warning, a large penalty coefficient may cause the support vector machine classification model to be too sensitive and misclassify normal minor fluctuations as anomalies; while a small penalty coefficient may ignore some potential risk signals. After the support vector machine classification model is trained, new warning signals will be input for classification. Suppose there is a new warning signal with a feature vector of [0.8, 30, 0.05], representing signal intensity, duration (in minutes), and frequency (times per hour) respectively. The support vector machine classification model will make a judgment based on these features and may classify it as an abnormal signal. For signals determined to be abnormal, further confidence evaluation is crucial. Setting a reasonable threshold, such as 0.85, can effectively filter out abnormal signals with low confidence. For example, an abnormal signal with a confidence of 0.92 will be retained, while a signal with a confidence of 0.78 will be regarded as a potential false alarm and filtered. This secondary classification and filtering mechanism greatly improves the reliability of the warning mechanism. In practical applications, it can effectively reduce false alarms caused by sensor failures or data transmission errors. For example, in a large logistics center, hundreds of warning signals may be generated every day. After being screened by the support vector machine classification model, only dozens of high-confidence signals may be retained, greatly reducing the burden on management personnel. The continuous optimization of the support vector machine classification model is also the key to ensuring the long-term effectiveness of the warning mechanism. As new sample data accumulates, the support vector machine classification model can be retrained regularly. This dynamic update mechanism enables the support vector machine classification model to adapt to changes in the business environment. For example, with the expansion of the fleet size and changes in road conditions, the original warning criteria may no longer apply. Through the retraining of the support vector machine classification model, these changes can be captured to maintain the efficiency of the warning mechanism. Finally, high-confidence effective warning signals will be used to guide actual operations. In the logistics system, this may mean timely adjusting vehicle routes, arranging maintenance, or reallocating goods. In this way, the warning mechanism not only improves the safety of operations but also optimizes resource utilization, ultimately enhancing the efficiency and reliability of the entire logistics network.

[0032] Step S107: Extract the vehicle scheduling requirements from the sixth processed data. For the problem of processing massive data generated by multiple vehicles, multiple routes, and multiple target points, use the Apache Spark distributed computing framework to parallelize and optimize the scheduling strategy to obtain the seventh processed data.

[0033] According to the vehicle scheduling requirements, the Scrapy distributed crawler framework is adopted to scrape relevant data such as vehicles, routes, and target points from a vast data source. Data cleaning and standardization processing are carried out through the Pandas and NumPy libraries to remove noise data and convert the data into a unified format. Key features of vehicles and routes are extracted, such as the position coordinates, speed, fuel consumption of vehicles, and the length, road conditions, speed limits of routes. An input vector for the vehicle scheduling optimization model is constructed through feature engineering. The XGBoost algorithm is used to train the vehicle scheduling strategy classification model. According to the vehicle and route features, the optimal scheduling strategy for the vehicle is predicted, and a preliminary scheduling plan is generated. The preliminary scheduling plan is loaded into the Spark distributed computing framework, and the RDD data structure is used to partition the scheduling plan to form multiple sub-tasks that can be executed in parallel. Through the Map and Reduce operations of Spark, the scheduling sub-tasks are executed in parallel on multiple nodes of the cluster to optimize the vehicle scheduling strategy. In the Map operation of Spark, for each scheduling sub-task, the tabu search algorithm is used for local optimization. Constraint conditions such as vehicle loading capacity, driving mileage, and time window are set, and a feasible scheduling plan that meets the constraints is searched. By setting the tabu step size and tabu list length, the search is prevented from falling into a local optimum. The optimization results of each sub-task are summarized through the Reduce operation to form a set of optimized candidate scheduling plans. For the set of candidate scheduling plans, the simulated annealing algorithm is used for global optimization. The scheduling strategies of each vehicle are regarded as particles. By setting parameters such as the initial temperature and cooling rate, the movement of particles in the state space is simulated, good solutions are accepted, and poor solutions are probabilistically accepted to jump out of the local optimum and find the global optimal vehicle scheduling strategy. The optimized vehicle scheduling strategy is converted into a scheduling instruction in JSON format and distributed through the Kafka message queue. Each vehicle terminal device subscribes to the scheduling instruction topic of Kafka, obtains the scheduling instruction in real time, and executes the corresponding scheduling task to achieve intelligent vehicle scheduling. During the process of the vehicle executing the scheduling task, information such as the vehicle's running trajectory, speed, and fuel consumption is collected through devices such as GPS and OBD, and the collected data is transmitted to the stream processing framework through Kafka. SparkStreaming is used to perform real-time analysis on the vehicle operation data. If abnormal situations such as the vehicle deviating from the predetermined route and speeding are found, warning information is generated in a timely manner to trigger the dynamic optimization of the scheduling strategy. The vehicle operation data and the optimized scheduling strategy are persistently stored in HDFS and HBase to form a historical dataset of vehicle scheduling. SparkSQL is used to analyze and mine the historical data, extract the key influencing factors of vehicle scheduling, optimize the scheduling strategy model, and continuously improve the efficiency and accuracy of vehicle scheduling.

[0034] Specifically, vehicle scheduling optimization is a crucial link in the logistics industry, and it can significantly improve efficiency through data-driven and intelligent algorithms. First, the Scrapy distributed crawler framework is used to obtain information such as vehicles and routes from multiple data sources. For example, it can scrape publicly available road condition data from the transportation department, weather forecasts from the meteorological department, and location information uploaded in real-time by on-vehicle devices. After being processed by Pandas and NumPy, these data form standardized feature vectors. In the feature engineering stage, composite features can be constructed. For instance, by comprehensively considering information such as the vehicle's current location, destination, road conditions, and weather, a feature of "estimated arrival time" can be generated. This feature can more comprehensively reflect the influencing factors of scheduling decisions. The XGBoost algorithm performs well in training the scheduling strategy classification model. It can effectively handle the interaction between features, such as the impact of the combination of vehicle load and road conditions on fuel consumption. By learning from historical data, the model can predict the optimal scheduling strategy under given conditions. After importing the preliminary scheduling plan into the Spark framework, the advantages of distributed computing can be fully utilized. For example, for a large-scale scheduling problem involving 100 vehicles and 20 delivery points, it can be split into multiple sub-tasks, and each sub-task is responsible for optimizing the scheduling plan for a specific area or time period. In the Map stage, the application of the tabu search algorithm can effectively avoid falling into local optima. Suppose when optimizing the route of a certain truck, the algorithm may temporarily accept a seemingly worse solution (such as detouring to avoid traffic congestion). This decision may increase the driving distance in the short term, but in the long run, it may be a better choice. In the global optimization stage, the simulated annealing algorithm is adopted, and its unique feature is that it can accept a certain degree of "bad" solutions during the search process. This characteristic enables the algorithm to have the opportunity to jump out of local optima and find better global solutions. For example, when optimizing the scheduling strategy of the entire fleet, it may temporarily accept non-optimal paths for some vehicles in exchange for an improvement in overall scheduling efficiency. The optimized scheduling strategy is distributed to each vehicle terminal through the Kafka message queue. This real-time communication mechanism enables the scheduling system to quickly respond to emergencies. For example, when a sudden traffic accident occurs on a certain road, the scheduling system can immediately send new route instructions to the affected vehicles. The real-time analysis function of SparkStreaming provides the basis for dynamic adjustment. For example, the scheduling system can monitor the driving speed and location of each vehicle in real-time. If it is found that a certain vehicle's speed is continuously lower than expected, it may indicate that the traffic conditions in that area have deteriorated, and the scheduling system will promptly adjust the route planning of subsequent vehicles. Finally, the scheduling data is stored in HDFS and HBase, and in-depth analysis is carried out through SparkSQL, and some potential optimization opportunities can be discovered. For example, by analyzing historical data, it may be found that the delivery efficiency during certain periods and on certain routes is particularly high, and these findings can guide future scheduling decisions and continuously improve the overall operational efficiency.

[0035] Step S108, according to the seventh processed data, for the problem of the computing power requirement of real-time scheduling, use the genetic algorithm to quickly generate an optimal scheduling plan and obtain the final scheduling result.

[0036] Randomly initialize a set of scheduling plans as the initial population. Each scheduling plan contains the execution order of tasks and the allocated resource information. For each scheduling plan in the population, calculate the fitness value of the plan by using the weighted summation method according to factors such as task priority, execution time, and resource utilization rate, as an index to evaluate the quality of the scheduling plan. The calculation formula of the fitness value is: F = w1×P + w2×(1 / T) + w3×R, where F is the fitness value, P is the task priority, T is the task execution time, R is the resource utilization rate, and w1, w2, w3 are the weight coefficients of each factor. Perform selection, crossover, and mutation operations on the population according to the fitness value to generate new scheduling plans. The selection operation uses the roulette wheel selection method to select individuals from the current population according to the size of the fitness value. The larger the fitness value, the higher the probability of being selected. The crossover operation uses the two-point crossover method, randomly selects two crossover points, and swaps the gene segments between the two selected individuals at the crossover points. The mutation operation uses the basic bit mutation method to invert some genes of an individual with a certain probability to introduce new gene combinations. Determine whether the maximum fitness value of the current population has not changed significantly after multiple consecutive iterations or has reached the preset maximum number of iterations. If satisfied, enter step 6; otherwise, add the newly generated scheduling plan to the population, repeat steps 2-4 for iterative optimization. Output the scheduling plan with the highest fitness value as the optimal plan, determine the execution order of tasks and the resource allocation situation, and generate the final scheduling result. Send the scheduling result to the task execution module to guide the scheduling execution of real-time tasks. At the same time, monitor the task execution situation and obtain feedback information such as task completion time and resource usage. Use the feedback information of task execution to dynamically adjust the parameters of the genetic algorithm, including population size, crossover probability, mutation probability, etc., so that the genetic algorithm can adapt to the dynamic changes of task execution and improve the efficiency and quality of subsequent scheduling. According to the adjusted genetic algorithm parameters, return to step 1 to start a new round of real-time scheduling tasks, and loop through the above steps to continuously optimize the scheduling plan.

[0037] Specifically, the application of genetic algorithms in vehicle scheduling optimization demonstrates its powerful search ability. In the initialization phase, a diverse set of scheduling plans can be generated based on historical data. For example, in an urban logistics distribution system, the initial population may consist of 10 different combinations of distribution routes, with each combination involving 20 vehicles and 100 delivery points. Fitness calculation is the key to evaluating the quality of the plans. In practical applications, task priority may reflect the urgency of orders, execution time corresponds to the total distribution time, and resource utilization rate reflects the loading efficiency of vehicles. Suppose a certain distribution plan has a task priority of 0.8, an expected execution time of 2 hours, and a resource utilization rate of 0.7, with weights of 0.4, 0.3, and 0.3 respectively. Then its fitness value is: 0.4×0.8 + 0.3×(1 / 2) + 0.3×0.7 = 0.62. The selection operation mimics the natural law of "survival of the fittest". When using the roulette wheel selection method, plans with higher fitness values have a greater probability of being retained. This ensures that excellent features can be inherited and developed in the population. The crossover operation combines the advantages of different plans. For example, two distribution routes may perform well in certain areas. Through two-point crossover, these excellent segments can be combined to create a better plan. The mutation operation introduces randomness to prevent the algorithm from falling into a local optimum. In vehicle scheduling, mutation may manifest as randomly adjusting the visiting order of a certain delivery point or replacing the vehicle executing a certain delivery task. This small-scale random change can sometimes bring unexpected results, such as discovering a more efficient route combination. The iterative optimization process continuously improves the quality of the plan. When the change amplitude of the maximum fitness value is less than a preset threshold (such as 0.001) for multiple consecutive iterations (such as 50 times), or when the maximum number of iterations (such as 1000 times) is reached, it is considered that the algorithm has converged to a relatively optimal solution. The output and execution of the optimal plan are the implementation steps of the algorithm. For example, the finally determined scheduling plan may specify that vehicle No. 1 delivers to point A first and then to point B, and vehicle No. 2 is responsible for delivering to points C, D, and E. These instructions are transmitted to the driver through the in-vehicle terminal to guide the actual distribution work. The feedback mechanism and dynamic adjustment enable the algorithm to adapt to changes in the actual situation. For example, if it is found that certain sections are often congested, resulting in task delays, the mutation probability can be appropriately increased to increase the chance of exploring new routes. Another example is that if the population diversity is found to decline, the population size can be increased to introduce more possible solutions. This cyclic optimization process enables the scheduling system to continuously learn and improve. By accumulating experience, the scheduling system may find that the distribution efficiency of certain routes is particularly high during certain periods, or that certain vehicle types are more suitable for specific types of distribution tasks. These findings can guide future scheduling decisions and continuously improve the overall operation efficiency.

[0038] Obviously, those skilled in the art can make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the embodiments of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A vehicle scheduling optimization method based on the Internet of Things, characterized in that: include: Acquire original data of vehicle status, and use Kalman filter algorithm to smooth the original data to obtain first processed data; According to the first processed data, fill in the missing values ​​using an interpolation algorithm to obtain second processed data; Extracting key state parameters from the second processed data, and modifying the parameters in real time using a time series prediction model to obtain third processed data; According to the third processed data, combined with road condition information, weather forecast data and historical driving records, Dynamically evaluating the vehicle driving state using a multi-factor regression model to obtain fourth processed data; Extracting the vehicle location distribution, cargo loading, and remaining fuel information from the fourth processed data, and dynamically adjusting the warning rules using an adaptive threshold algorithm to obtain fifth processed data; According to the fifth processed data, the warning signal is secondary screened using a support vector machine classification algorithm to obtain sixth processed data; Extracting the vehicle dispatching requirements from the sixth processed data, and optimizing the dispatching strategy in parallel using a distributed computing framework to obtain seventh processed data; According to the seventh processed data, a genetic algorithm is used to quickly generate an optimal scheduling solution to obtain a final scheduling result.

2. The method according to claim 1, characterized in that The acquiring of the original data of the vehicle state and smoothing the original data using a Kalman filter algorithm to obtain first processed data include: Obtain vehicle speed, acceleration and angular data collected by on-board sensors, establish a Kalman filter model, and set the process noise covariance matrix and the measurement noise covariance matrix; Construct the state transfer matrix, observation matrix and noise matrix of the Kalman filter; Inputting the raw data into the Kalman filter model, and recursively estimating the optimal value of the vehicle state at each moment through prediction and updating; In the prediction step, the state prior estimation and the error covariance matrix of the current moment are predicted based on the state estimation value and the state transfer matrix at the previous moment; In the update step, the Kalman gain is calculated, and the predicted value is corrected using the observed value at the current moment to obtain the state posterior estimate and the updated error covariance matrix; Iterate the prediction and update steps until all the raw data of vehicle status at all time points are processed; The vehicle state estimate after Kalman filter processing is evaluated, and if the error exceeds a preset threshold, the process noise covariance matrix and the measurement noise covariance matrix parameters of the Kalman filter are adjusted.

3. The method according to claim 1, characterized in that The method of filling missing values ​​using an interpolation algorithm according to the first processed data to obtain second processed data includes: According to the distribution characteristics and change trends of existing data, a data model is constructed to predict the value range of missing data; For missing continuous values, linear interpolation method is used to estimate and fill in; For missing categorical data, the nearest neighbor algorithm was used for interpolation; Perform outlier detection on the filled data, analyze the distribution characteristics of outliers, set anomaly thresholds, and remove abnormal data points introduced by the interpolation process; After removing outliers, the data size is expanded through data enhancement technology; For continuous values, new data samples are generated by adding random noise; For classified data, a synthetic minority class oversampling technique algorithm is used to synthesize new minority class samples to balance the category distribution of the data set. Divide the expanded data set into training set and test set; Different interpolation algorithms are used to fill in the missing data, and the prediction performance of each algorithm on the training set and test set is compared and evaluated. The interpolation model with the best filling effect is selected and used to estimate all missing values ​​in the original data set.

4. The method according to claim 1, characterized in that The step of extracting key state parameters from the second processed data and modifying the parameters in real time using a time series prediction model to obtain third processed data includes: Obtain key parameter values ​​that reflect the operating status of the equipment; Using the acquired historical parameter value sequence, an autoregressive integrated moving average model is established; The established autoregressive integrated moving average model is used to predict and correct the delayed state parameters at the current moment in real time according to the historical state parameter value sequence of the equipment, so as to obtain the real operating parameters of the equipment at the current moment; Substituting the corrected current state parameters into the autoregressive integrated moving average model, continuously rolling forecasting the equipment state parameters in a subsequent period of time is performed in a sliding time window manner to obtain a forecast parameter sequence for a period of time in the future; According to the predicted equipment operation parameter sequence, a normal value range threshold is set to determine whether the predicted parameter exceeds the threshold; If the forecast parameters at a certain moment are beyond the normal range, an early warning message will be issued; The real-time operating parameters of the equipment are obtained, and compared with the parameter sequence predicted by the autoregressive integrated moving average model in real time to calculate the deviation value between the actual parameters and the predicted parameters.

5. The method according to claim 1, characterized in that The fourth processed data is obtained by dynamically evaluating the vehicle driving state using a multi-factor regression model based on the third processed data in combination with road condition information, weather forecast data and historical driving records, including: Get real-time road congestion and accident information; Get temperature, humidity and visibility data for a period of time in the future; Reading historical driving data of the vehicle, wherein the historical driving data includes driving speed, acceleration, and braking frequency data; Inputting the acquired real-time traffic information, the weather forecast data and the historical driving record data into a vehicle driving state evaluation model pre-trained based on a neural network algorithm; The vehicle driving state evaluation model calculates the current vehicle driving state evaluation score by weighting and comprehensively analyzing various factors; Comparing the vehicle driving state score calculated by the vehicle driving state assessment model with a preset safety threshold; If the score exceeds the safety threshold, the current vehicle is judged to be in a dangerous driving state and a safety warning is automatically issued to the driver; The vehicle driving status assessment score is compared and analyzed with the data in the vehicle's historical driving record database, and a clustering algorithm is used to identify whether there is an abnormal change trend in the vehicle's driving status.

6. The method according to claim 1, characterized in that The extracting of the vehicle position distribution, cargo loading, and remaining fuel information from the fourth processed data, and dynamically adjusting the warning rules using an adaptive threshold algorithm to obtain the fifth processed data includes: According to the vehicle location distribution information, a density-based clustering algorithm is used to divide the vehicles into regions to obtain the vehicle distribution density in different regions; For each area, obtain the cargo loading and remaining fuel data of the vehicles in the area, remove outliers and redundant data, and obtain a normalized feature data set; According to the historical operation data of the vehicle, a judgment standard of the vehicle state is defined, and the characteristic data set is annotated to obtain label data of the vehicle state; Using the label data and the feature data, a correlation model between the cargo loading amount, the remaining fuel amount and the vehicle status is trained by a logistic regression algorithm to obtain parameters and thresholds of the correlation model; During the operation of the vehicle, the cargo load and remaining fuel data of the vehicle are collected in real time to construct time series data; The time series data is filtered using a Kalman filter algorithm to estimate optimal values ​​of the cargo loading amount and the remaining fuel amount; The filtered cargo loading and remaining fuel data are input into the association model to determine the vehicle status.

7. The method according to claim 1, characterized in that The method of performing secondary screening on the warning signal using a support vector machine classification algorithm according to the fifth processed data to obtain sixth processed data includes: Extract the features of the warning signal to form a warning signal feature vector; Using the labeled warning signal sample data and the support vector machine algorithm, the support vector machine classification model is trained by adjusting the kernel function and penalty coefficient parameters; For a new warning signal, its feature vector is input into the trained support vector machine classification model to perform binary classification prediction; The support vector machine classification model will classify the warning signal into two categories: normal signal or abnormal signal; For the warning signal predicted as an abnormal signal by the support vector machine classification model, mark it as a potential false positive or false negative; A confidence threshold is set. When the confidence of an abnormal signal is lower than the confidence threshold, the warning signal is determined to be a false alarm or a missed alarm and is filtered out. retaining the warning signal with a confidence level higher than the confidence level threshold as sixth processed data; According to the classification result of the support vector machine classification model, a warning signal with a higher confidence level is screened out from the sixth processed data as a final effective warning signal.

8. The method according to claim 1, characterized in that: The method of rapidly generating an optimal scheduling solution using a genetic algorithm according to the seventh processed data to obtain a final scheduling result includes: Initialize a set of scheduling schemes as the initial population, each of which contains the execution order of tasks and the allocated resource information; For each scheduling scheme in the population, the fitness value of the scheme is calculated by weighted summation based on the task priority, execution time and resource utilization factors, which is used as an indicator to evaluate the quality of the scheduling scheme; According to the fitness value, the population is selected, crossed and mutated to generate a new scheduling plan; the selection operation adopts the roulette selection method to select individuals from the current population according to the size of the fitness value. The larger the fitness value, the higher the probability of being selected; The crossover operation uses a two-point crossover method, randomly selecting two crossover points and exchanging the gene fragments between the crossover points of the two selected individuals; The mutation operation uses the basic bit mutation method to invert certain genes of individuals with a certain probability and introduce new gene combinations; Determine whether the maximum fitness value of the current population has not changed for multiple consecutive iterations, or has reached the preset maximum number of iterations; If the conditions are met, the scheduling scheme with the highest fitness value is output as the optimal scheme, the execution order of tasks and resource allocation are determined, and the final scheduling result is generated.

Citation Information

Patent Citations

  • Electric network synthetic disaster prevention system based on geographic information system

    CN101295172A

  • Taxi dispatching method and taxi dispatching system on basis of video monitoring system

    CN102867411A

  • Optimization method and device for riding order scheduling system

    CN111598307A

  • Task scheduling method for intelligent fire-fighting management service platform

    CN119623976A