Expressway area risk factor identification method and device
By acquiring and processing multi-dimensional data on highway roads, using neural network models to identify and correlate risk factors, and dynamically adjusting factors, the subjectivity and uncertainty of risk identification in traditional methods are solved, and more accurate and efficient risk assessment and management are achieved.
Patent Information
- Application Number
- CN202411963171.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Traditional highway road risk factor identification methods have subjectivity, uncertainty and inability to make full use of modern information technology, making it difficult to accurately identify risk factors and their relationships.
By obtaining multi-dimensional data such as traffic flow, weather conditions and road conditions, preprocessing and key feature extraction, using neural network models to identify risk factors and relevance mining, and dynamically adjusting factors to optimize risk assessment.
It realizes more accurate and efficient risk factor identification, reveals the mutual relationship between risk factors, improves the objectivity and real-time nature of risk assessment, and supports more accurate risk management strategies.
Smart Images

Figure CN119992819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and device for identifying highway road risk factors. Background Art
[0002] There are many risk factors on highways, such as bad weather, road damage, high traffic volume, etc. These factors directly affect driving safety and may even lead to traffic accidents. Therefore, it is particularly important to accurately identify and evaluate the risk factors on highways.
[0003] Some traditional methods for identifying highway risk factors mainly rely on manual inspections and regular testing. Although these methods can detect potential risks to a certain extent, they have the following defects:
[0004] For example, some traditional risk assessment methods rely on the experience and judgment of inspectors and lack unified and objective evaluation standards, which may lead to greater subjectivity and uncertainty in the assessment results. Some traditional methods fail to make full use of modern information technology, such as big data and artificial intelligence, to conduct in-depth analysis and mining of massive data, making it difficult to comprehensively and accurately identify risk factors and their interrelationships. Some traditional methods use static risk assessment models, which cannot dynamically adjust risk assessment results based on real-time data and are difficult to adapt to the rapid changes in risk factors in highway areas. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a method and device for identifying risk factors on a highway area, which can identify risk factors more accurately and efficiently.
[0006] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0007] In a first aspect, a method for identifying highway risk factors is provided, the method comprising:
[0008] Step 1, obtaining data related to highway risks, including traffic flow, weather conditions, and road conditions;
[0009] Step 2, preprocessing the data related to the highway road area risk to obtain preprocessed data;
[0010] Step 3, extracting key features from the preprocessed data, wherein the key features are used for risk factor identification;
[0011] Step 4, using the key features to train the neural network model to obtain a trained neural network model;
[0012] Step 5, identifying risk factors through the trained neural network model, and performing association mining on the risk factors to obtain the relationships between the risk factors;
[0013] Step 6, based on the relationship between risk factors, assign an importance score to each risk factor; calculate the dynamic adjustment factor, multiply the dynamic adjustment factor by the importance score of each risk factor to obtain the adjusted importance score; sort the risk factors according to the adjusted importance score to obtain the key factors affecting the highway area risk.
[0014] Furthermore, data related to highway risks are obtained, including traffic flow, weather conditions, and road conditions, including:
[0015] The selection and configuration of data collection points is considered as an optimization problem, with the goal of maximizing the efficiency and quality of data collection; decision variables are determined, including the location, type, and frequency of data collection of sensors;
[0016] Use binary coding to represent the solution and initialize a population containing multiple solutions;
[0017] Define a fitness function to evaluate the quality of each individual; use the roulette wheel selection method to select the corresponding individuals from the current population to enter the next generation; perform a crossover operation on the selected individuals to generate a new solution; perform a mutation operation on the newly generated individuals, and repeat the selection, crossover and mutation operations until the termination condition is met. When the termination occurs, output the final solution, that is, the final sensor configuration solution;
[0018] Deploy sensors according to the final sensor configuration plan to obtain data related to highway road risks.
[0019] Furthermore, key features are extracted from the preprocessed data, and the key features are used for risk factor identification, including:
[0020] Determine the name, definition, and expected meaning of each feature based on the data dictionary in the preprocessed data; quantify the value range and distribution of each feature through the mean; determine the target variable, that is, the risk factor indicator to be identified;
[0021] Calculate the Pearson correlation coefficient between each feature and the target variable; construct a correlation coefficient matrix based on the Pearson correlation coefficient; analyze the correlation coefficient matrix to identify the identification features related to the target variable;
[0022] According to the Pearson correlation coefficient, a screening threshold is set; according to the screening threshold, key features that meet the screening threshold conditions are screened out from the identification features.
[0023] Furthermore, the neural network model is trained using the key features to obtain a trained neural network model, including:
[0024] Initialize the weights and bias parameters of the convolutional neural network;
[0025] Define the loss function to measure the gap between the predicted value of the neural network model and the true value;
[0026] During the training process, the gradient descent algorithm is used to optimize the weights and bias parameters of the neural network;
[0027] For each training batch, calculate the gradient of the loss function with respect to the weights and biases;
[0028] According to the calculated gradient, update the weight and bias parameters and repeat the process until the stopping condition is met; set the learning rate and optimizer; divide the preprocessed data into training set and validation set;
[0029] Using the training set data, calculate the predicted value of the neural network model through forward propagation and calculate the value of the loss function; using the gradient descent algorithm and optimizer, update the weight and bias parameters of the neural network according to the gradient of the loss function; after each iteration cycle, evaluate the performance of the neural network model on the validation set to obtain the evaluation results;
[0030] The structure, learning rate and optimizer parameters of the neural network model are adjusted according to the evaluation results to obtain the trained neural network model.
[0031] Furthermore, the risk factors are identified by the trained neural network model, and the correlation mining of the risk factors is performed to obtain the relationship between the risk factors, including:
[0032] The data set to be analyzed is input into the trained neural network model; through the forward propagation process of the trained neural network model, the predicted value of the risk factor corresponding to each observation point is obtained;
[0033] Set the minimum support and minimum confidence thresholds for the Apriori algorithm; run the Apriori algorithm, input the transaction data set, minimum support, and minimum confidence;
[0034] The Apriori algorithm will output a series of association rules, each rule is represented as a combination of an antecedent and a consequent;
[0035] For each association rule, calculate its support, confidence and lift. Support indicates the frequency of the rule in the data, confidence indicates the reliability of the rule, and lift indicates the strength of the correlation between the antecedent and the consequent in the rule.
[0036] Sort and filter the association rules according to support, confidence and lift to find the corresponding key rules; analyze the key rules to obtain the correlation and dependency between different risk factors.
[0037] Furthermore, each risk factor is assigned an importance score based on the interrelationships between risk factors, including:
[0038] Determine the quantitative indicators of the relationship, including support, confidence and lift;
[0039] According to quantitative indicators, Calculate importance scores;
[0040] Among them, f r (A∪B) represents the frequency of transactions where item sets A and B appear simultaneously for association rule r; f r (A) and f r (B) represents the transaction frequency of item set A and item set B for association rule r; N represents the total number of transactions in the data set; w S 、w c and w L is the weight parameter; max r∈R N represents the maximum value of the unnormalized importance score in all association rule sets R; r represents the normalized importance score of association rule r.
[0041] Furthermore, the calculation formula of the dynamic adjustment factor is:
[0042]
[0043] Among them, F represents the dynamic adjustment factor; w represents the weight of traffic flow; V represents the current traffic flow; U represents the benchmark traffic flow; z represents the weight of visibility; a represents the peak value of the Gaussian function; S represents the current visibility; T represents the optimal visibility; and C represents the standard deviation of the Gaussian function related to visibility.
[0044] In a second aspect, a highway risk factor identification device includes:
[0045] An acquisition module, used to acquire data related to highway risks, including traffic flow, weather conditions, and road conditions;
[0046] A preprocessing module is used to preprocess the data related to highway road risks to obtain preprocessed data;
[0047] An extraction module, used to extract key features from the preprocessed data, wherein the key features are used for risk factor identification;
[0048] A training module, used to train the neural network model using key features to obtain a trained neural network model;
[0049] The identification module is used to identify risk factors through the trained neural network model and conduct association mining on risk factors to obtain the relationship between risk factors;
[0050] The calculation module is used to assign an importance score to each risk factor based on the relationship between the risk factors; calculate the dynamic adjustment factor, multiply the dynamic adjustment factor by the importance score of each risk factor to obtain the adjusted importance score; and sort the risk factors according to the adjusted importance score to obtain the key factors affecting the risk of the highway area.
[0051] According to a third aspect, a computing device includes:
[0052] one or more processors;
[0053] The storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described.
[0054] In a fourth aspect, a computer-readable storage medium stores a program, and when the program is executed by a processor, the method described is implemented.
[0055] The above solution of the present invention includes at least the following beneficial effects:
[0056] By acquiring multi-dimensional data (such as traffic flow, weather conditions, road conditions, etc.), this method can more comprehensively reflect the actual risk situation of the highway area. Combined with the deep learning ability of the neural network model, key features can be accurately extracted from complex data, thereby improving the accuracy of risk factor identification.
[0057] Compared with traditional manual inspection and data processing methods, this method uses automated data preprocessing and feature extraction steps to significantly improve the efficiency of data processing. This not only reduces the input of human resources, but also shortens the time cycle from data collection to risk factor identification, making risk response more timely.
[0058] This method can reveal the interrelationships between different risk factors through association mining technology. This understanding of the correlation helps to more comprehensively assess risks, formulate more accurate risk management strategies, and optimize resource allocation.
[0059] By dynamically adjusting the factors, the risk assessment can be adjusted in real time according to the actual situation. This dynamicity not only reflects the time-varying nature of highway risks, but also enhances the flexibility and adaptability of the risk assessment model.
[0060] By ranking risk factors, this method can identify the key factors that affect highway risks, which provides strong information support for decision makers and helps to formulate risk management measures with clear priorities and strong pertinence, thereby improving the safety and efficiency of highway operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 It is a flow chart of a method for identifying highway road area risk factors provided by an embodiment of the present invention.
[0062] Figure 2 It is a schematic diagram of a highway road area risk factor identification device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0063] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0064] like Figure 1 As shown, an embodiment of the present invention provides a method for identifying highway risk factors, the method comprising the following steps:
[0065] Step 1, obtaining data related to highway risks, including traffic flow, weather conditions, and road conditions;
[0066] Step 2, preprocessing the data related to the highway road area risk to obtain preprocessed data;
[0067] Step 3, extracting key features from the preprocessed data, wherein the key features are used for risk factor identification;
[0068] Step 4, using the key features to train the neural network model to obtain a trained neural network model;
[0069] Step 5, identifying risk factors through the trained neural network model, and performing association mining on the risk factors to obtain the relationships between the risk factors;
[0070] Step 6, based on the relationship between risk factors, assign an importance score to each risk factor; calculate the dynamic adjustment factor, multiply the dynamic adjustment factor by the importance score of each risk factor to obtain the adjusted importance score; sort the risk factors according to the adjusted importance score to obtain the key factors affecting the highway area risk.
[0071] In the embodiment of the present invention, by comprehensively collecting multi-dimensional data such as traffic flow, weather conditions, and road conditions, rich basic information is provided for subsequent risk factor identification and analysis. The collection of such comprehensive data helps to more comprehensively understand the operating status and potential risks of the highway area. Data preprocessing can clean and organize the original data, remove noise and outliers, and improve the quality and availability of the data. This step ensures the accuracy and reliability of subsequent analysis and provides a clean and standardized data set for the training of the neural network model. Feature extraction helps to extract key information directly related to risk factor identification from complex data sets. This not only reduces the dimension of the data, but also enables the neural network model to focus more on learning these key features, thereby improving the efficiency and accuracy of risk identification. By training the neural network model, it can learn the mapping relationship from key features to risk factors. The powerful learning ability of the neural network model can capture the nonlinear relationship in the data, further improving the accuracy and complexity of risk factor identification. Using the trained neural network model for risk factor identification can quickly and accurately locate potential risks. At the same time, association mining technology can reveal the inherent connection between risk factors, which helps to deeply understand the cause and propagation mechanism of risks. By assigning importance scores to risk factors and introducing dynamic adjustment factors for real-time adjustment, the impact of each risk factor on highway safety can be quantified more accurately.
[0072] In a preferred embodiment of the present invention, the above step 1, obtaining data related to highway road area risks, the data including traffic flow, weather conditions, road conditions, may include:
[0073] Step 11, treat the selection and configuration of data collection points as an optimization problem, with the goal of maximizing the efficiency and quality of data collection; determine the decision variables, including the location, type and collection frequency of sensors, specifically including: determining specific indicators of data collection efficiency and quality, such as the timeliness, completeness, accuracy, etc. of data collection; quantifying these indicators into measurable objective functions, for example, by defining data collection efficiency index and quality index; listing all possible sensor locations based on the geographical characteristics and key risk points of the highway; selecting sensor types suitable for data collection such as traffic flow, weather conditions, road conditions, etc.; setting the range of sensor collection frequency, considering the balance between data real-time and transmission cost; considering the physical limitations of sensor deployment, such as the availability of infrastructure such as power supply and communication, and setting budget constraints to ensure the economic feasibility of the overall solution.
[0074] Step 12, using binary coding to represent the solution, initialize a population containing multiple solutions, specifically including: using binary coding to represent the decision variables of sensor location, type and acquisition frequency; for example, the location can be represented by a series of binary bits to represent different candidate locations, the type can be represented by a few bits of binary to represent different sensor types, and the acquisition frequency can also be represented by binary coding to represent different frequency levels; randomly generate a certain number of initial solutions, each solution is a binary string, to ensure that the population size is large enough to cover the diversity of the search space.
[0075] Step 13, define a fitness function to evaluate the quality of each individual; use the roulette wheel selection method to select the corresponding individuals from the current population to enter the next generation; perform a crossover operation on the selected individuals to generate a new solution; perform a mutation operation on the newly generated individuals, and repeat the selection, crossover and mutation operations until the termination condition is met. When the termination occurs, the final solution, that is, the final sensor configuration solution, is output, which specifically includes:
[0076] According to the optimization goal, design a fitness function that can evaluate the quality of each individual (solution); this function should be able to comprehensively consider the efficiency and quality of data collection and give a quantitative score; use the roulette wheel selection method to select the probability of each individual entering the next generation according to its fitness value, and individuals with high fitness will have a greater chance of being selected and passed on to the next generation; randomly pair the selected individuals and select a crossover point for binary string crossover; through the crossover operation, generate new solutions, which combine the excellent characteristics of the parent individuals; randomly flip the binary bits of the newly generated individuals to introduce new genetic mutations. The mutation operation helps to increase the diversity of the population and prevent premature convergence. Set the termination condition of the genetic algorithm, such as reaching the maximum number of iterations or the fitness value reaching the preset threshold. When the termination condition is met, output the individual with the highest fitness in the current population as the final solution, that is, the optimal sensor configuration solution. Among them, the calculation formula of the fitness function is:
[0077]
[0078] Where F represents the fitness function value; α represents the weight adjustment coefficient of efficiency; β represents the scaling coefficient of efficiency term; γ represents the impact weight of sensor cost; δ represents the impact index of sensor frequency; N represents the number of sensors; c i represents the deployment cost coefficient of the i-th sensor; f i represents the acquisition frequency of the i-th sensor; M represents the number of collected data points; d mj represents the measured value of the jth data point; d aj represents the actual value of the jth data point; ∈ represents the influence weight of the data error; ζ represents the influence weight of the actual data value; η represents the exponential adjustment weight of the quality item.
[0079] Step 14, deploy sensors according to the final sensor configuration plan to obtain data related to highway road area risks, including: performing actual sensor deployment according to the sensor location, type and collection frequency determined by the final solution, ensuring that the sensors are installed correctly and establishing a stable communication connection with the data acquisition system; starting the sensors to start collecting data related to highway road area risks, and performing preliminary verification of the collected data to ensure that its accuracy and completeness meet the requirements; if data quality problems are found, promptly adjust the sensor configuration or perform necessary maintenance operations.
[0080] In the embodiment of the present invention, by optimizing the location, type and acquisition frequency of the sensor, effective data collection can be ensured in key areas and key time points, thereby avoiding unnecessary waste of resources and improving the overall efficiency of data collection. Reasonable sensor configuration can more accurately capture the relevant data of highway road risks. By selecting the appropriate sensor type and location, it can be ensured that the collected data is more representative and accurate. Optimizing the sensor configuration can avoid deploying too many sensors at unnecessary locations or times, thereby reducing hardware costs, maintenance costs and data processing costs. High-quality data collection provides a solid foundation for risk identification. Through accurate data, potential risk factors can be more accurately identified, and corresponding preventive and response measures can be taken in time. Optimizing sensor configuration is an important part of the construction of intelligent transportation systems (ITS). Through efficient data collection, richer and more real-time information can be provided for intelligent transportation systems, thereby promoting the further development of intelligent transportation technology. Based on the data collected by the optimized sensor configuration, road managers can more accurately grasp the road operation status, discover and deal with safety hazards in a timely manner, thereby improving the efficiency and effectiveness of road safety management.
[0081] In a preferred embodiment of the present invention, the above step 2, preprocessing the data related to the highway road area risk to obtain preprocessed data, may include:
[0082] Integrate the raw data collected from different data sources (such as traffic flow monitoring equipment, weather stations, road condition sensors, etc.) into a unified data storage system; perform preliminary checks on the data to confirm that the data is complete and formatted correctly. Identify and delete duplicate data records to ensure that each data is unique. Detect and handle outliers, such as sudden large or small values in traffic flow data, which may be caused by sensor errors or data transmission problems. For missing data, fill or interpolate according to the specific situation. If there are too many missing data, consider deleting the record.
[0083] Since data such as traffic flow, weather conditions, and road conditions may have different dimensions and ranges, in order to eliminate the impact of these differences on subsequent analysis, the data needs to be normalized. You can choose the minimum-maximum normalization method to scale the data to the range of [0,1], or use Z-score normalization to convert the data into a distribution with a mean of 0 and a standard deviation of 1.
[0084] In a preferred embodiment of the present invention, the above step 3, extracting key features from the preprocessed data, the key features used for risk factor identification, may include:
[0085] Step 31, determine the name, definition and expected meaning of each feature based on the data dictionary in the preprocessed data; quantify the value range and distribution of each feature through the mean; determine the target variable, that is, the risk factor indicator to be identified, specifically including: consulting the data dictionary attached to the preprocessed data, which is the first step to understand the data set.
[0086] The data dictionary contains information such as the name, data type, description, and possible value range of each feature. According to the data dictionary, sort out the name, definition (i.e., what the feature represents) and expected meaning (how the feature should be understood in the business logic) of each feature. For example, the definition of the feature "average speed" is the average speed of all vehicles passing through a certain section of road during a certain period of time, and its expected meaning is an important indicator reflecting the smoothness of road traffic. For numerical features, quantify their value range and distribution by calculating statistics such as mean, standard deviation, minimum value, and maximum value. For categorical features, count the frequency or frequency of each category to understand the distribution of different categories. Determine the target variable, which is the risk factor indicator that you want to predict or explain, such as accident rate, traffic congestion level, etc. According to business needs and the purpose of data analysis, clarify the target variable and ensure that the data set contains relevant information about this variable.
[0087] Step 32, calculate the Pearson correlation coefficient between each feature and the target variable; construct a correlation coefficient matrix based on the Pearson correlation coefficient; analyze the correlation coefficient matrix to identify the identification features related to the target variable, specifically including: calculating the Pearson correlation coefficient between each feature and the target variable, the Pearson correlation coefficient measures the strength and direction of the linear relationship between two variables; organizing the calculated Pearson correlation coefficient into a matrix form, in which rows and columns represent different features, and the elements in the matrix represent the correlation coefficients between the corresponding feature pairs. The correlation coefficient between the target variable and each feature can be extracted separately to form a column or a row; by observing the correlation coefficient matrix, especially the correlation coefficient between the target variable and each feature, the features with a strong correlation with the target variable can be preliminarily identified; the closer the absolute value of the correlation coefficient is to 1, the stronger the linear relationship between the two variables is.
[0088] Step 33, according to the Pearson correlation coefficient, a screening threshold is set; according to the screening threshold, key features that meet the screening threshold conditions are screened out from the identification features, specifically including:
[0089] According to business needs and data analysis experience, a screening threshold is set. This threshold is used to determine which features are strongly correlated with the target variable and are therefore considered key features. The threshold can be selected either absolutely (such as a correlation coefficient greater than 0.5 or less than -0.5) or relative (such as selecting features with a correlation coefficient ranking in the top 10%). The correlation coefficient calculated in step 32 is compared with the set screening threshold, and features whose correlation coefficients exceed (or are lower than, for negative correlations) the screening threshold are retained. These features are considered key features that are significantly correlated with the target variable.
[0090] In an embodiment of the present invention, by calculating the Pearson correlation coefficient between each feature and the target variable, those features that are highly correlated with the target variable can be accurately identified. These key features will play an important role in the subsequent risk factor identification model, thereby improving the recognition accuracy of the model. In feature engineering, removing features with low correlation with the target variable can effectively reduce the complexity of the model. This not only makes the model more concise and easy to understand, but also reduces the risk of overfitting and improves the generalization ability of the model. The screening process of key features is actually a dimensionality reduction process. By reducing the number of input features, the amount of calculation during model training can be significantly reduced, thereby improving the calculation efficiency. This is particularly important when processing large-scale data sets. By screening key features, unstable features that may be caused by noise or outliers can be removed. In this way, the constructed risk factor identification model will be more robust and have better tolerance for small changes in input data. The identification results of key features can also provide guidance for future data collection work. By clarifying which features are crucial to risk factor identification, more attention can be paid to these features in the subsequent data collection process, thereby optimizing data collection strategies and improving data quality.
[0091] In a preferred embodiment of the present invention, the above step 4, using the key features to train the neural network model to obtain the trained neural network model, may include:
[0092] Step 41, initialize the weight and bias parameters of the convolutional neural network, specifically including: using random values (such as normal distribution or uniform distribution) to initialize the weight parameters of the neural network. These weights are key parameters for neural network learning and they will be continuously adjusted during the training process. Initialize the bias parameter to zero or a decimal close to zero. The bias parameter is used to adjust the output of the neuron to make it easier to process through the activation function. According to the specific neural network structure and task requirements, select a suitable initialization method, such as He initialization, Xavier initialization, etc.
[0093] Step 42, define a loss function to measure the gap between the predicted value of the neural network model and the true value, wherein the calculation formula of the loss function is:
[0094]
[0095] Where n is the number of samples, yi is the true value, is the model prediction value.
[0096] Step 43, during the training process, use the gradient descent algorithm to optimize the weights and bias parameters of the neural network, specifically including: selecting a standard gradient descent algorithm, such as stochastic gradient descent (SGD), according to the requirements. Determine an appropriate learning rate to control the step size of parameter updates. Too large a learning rate may lead to unstable training, while too small a learning rate may lead to slow training. Configure the corresponding optimizer according to the selected gradient descent algorithm and learning rate, such as using the optimizer library in TensorFlow or PyTorch.
[0097] Step 44, for each training batch, calculate the gradient of the loss function with respect to the weights and biases, specifically including: inputting the training data into the neural network, and obtaining the predicted value of the model through calculations at each layer; calculating the loss value based on the predicted value and the true value using the loss function defined in step 42; using the chain rule to calculate the gradient of the loss function with respect to each weight and bias parameter, these gradients indicate how the parameters should be adjusted in order to reduce the loss.
[0098] Step 45, update the weight and bias parameters according to the calculated gradient, repeat the process until the stop condition is met; set the learning rate and optimizer; divide the preprocessed data into a training set and a validation set, specifically including: using the optimizer configured in step 43 and the calculated gradient, update the weight and bias parameters of the neural network; repeat steps 44 and 45, that is, continuously perform forward propagation, calculate loss, back propagation and parameter update until the stop condition is met (such as reaching a preset number of iterations or loss value convergence); before starting iterative training, ensure that the preprocessed data has been divided into a training set and a validation set. The training set is used to train the model, and the validation set is used to evaluate the performance of the model.
[0099] Step 46, using the training set data, calculate the predicted value of the neural network model through forward propagation, and calculate the value of the loss function; using the gradient descent algorithm and optimizer, update the weight and bias parameters of the neural network according to the gradient of the loss function; after each iteration cycle, evaluate the performance of the neural network model on the validation set to obtain the evaluation results, specifically including: after each iteration cycle, use the validation set data to evaluate the model. This helps to understand the generalization ability of the model on unseen data; select appropriate performance indicators (such as accuracy, recall rate, F1 score, etc.) according to the nature of the task, and calculate the values of these indicators; monitor the performance changes of the model on the validation set, and record the performance indicator values of each iteration cycle for subsequent analysis and comparison.
[0100] Step 47, adjusting the structure, learning rate and optimizer parameters of the neural network model according to the evaluation results to obtain the trained neural network model, specifically including: analyzing the performance of the model according to the performance indicator values recorded in step 46, identifying possible problems and room for improvement; based on the analysis results, trying to adjust the number of layers, number of neurons, activation function and other structural parameters of the neural network to improve the performance of the model; if the training speed of the model is slow or unstable, you can try to adjust the learning rate or replace other optimizers for trial; after adjusting the model structure and parameters, repeat training (steps 41 to 46) and evaluation (step 46) until a satisfactory performance level is achieved.
[0101] In an embodiment of the present invention, by training the model with key features, the neural network can more accurately capture the potential laws and patterns in the data. This helps to improve the accuracy of the model in predicting risk factors, thereby providing more reliable decision support. Using key features instead of all features for training helps to reduce the complexity of the model, thereby reducing the risk of overfitting. Overfitting means that the model performs well on the training data, but has poor generalization ability on new data. By screening key features, the model can focus more on learning the information that is really important for the prediction results. Since only key features are used for training, the input dimension of the neural network is reduced, which can significantly speed up the training speed and reduce the consumption of computing resources. The screening process of key features helps to remove noise and redundant information, making the model more robust. This means that when faced with small changes or outliers in the data, the model can maintain relatively stable prediction performance. Adjusting the structure, learning rate and optimizer parameters of the neural network model according to the evaluation results mentioned in step 47 helps to find the best configuration of the model. By continuously optimizing these hyperparameters, the prediction performance and generalization ability of the model can be further improved. Training with key features also helps to improve the interpretability of the model. Since there are fewer features and each feature is significantly associated with the target variable, it is easier to understand how the model makes predictions based on these features.
[0102] In a preferred embodiment of the present invention, the above step 5, identifying risk factors through the trained neural network model, and performing association mining on the risk factors to obtain the relationship between the risk factors, may include:
[0103] Step 51, input the data set to be analyzed into the trained neural network model; obtain the risk factor prediction value corresponding to each observation point through the forward propagation process of the trained neural network model, specifically including: ensuring that the data set to be analyzed has been properly preprocessed, including data cleaning, standardization or normalization, etc., to meet the input requirements of the neural network model; loading the previously trained neural network model to ensure that the model weights and structure have been correctly loaded; using the data set to be analyzed as input, and performing forward propagation calculations through the neural network model. This involves calculating the data through each layer of the model, and finally obtaining the predicted value of the output layer; obtaining the risk factor prediction value corresponding to each observation point from the output layer of the neural network, and these predicted values will be used for subsequent association rule mining.
[0104] Step 52, set the minimum support and minimum confidence thresholds of the Apriori algorithm; run the Apriori algorithm, input the transaction data set, minimum support and minimum confidence, specifically including: according to specific needs and data characteristics, set the minimum support (min_support) and minimum confidence (min_confidence) thresholds of the Apriori algorithm. These thresholds will be used to screen effective association rules; convert the risk factor prediction values obtained in step 51 into a transaction data set format suitable for Apriori algorithm processing, which requires discretizing the continuous risk factor prediction values into specific intervals or categories. Call the Apriori algorithm function or library, and input the converted transaction data set, minimum support and minimum confidence parameters to start the association rule mining process.
[0105] Step 53, the Apriori algorithm will output a series of association rules, each rule is represented by a combination of an antecedent and a consequent, specifically including: the Apriori algorithm traverses the transaction data set, finds itemsets that meet the minimum support and minimum confidence thresholds, and generates association rules based on these itemsets. Each association rule consists of an antecedent and a consequent. After the algorithm is completed, a series of association rules that meet the conditions will be output. These rules represent the potential associations between different risk factors.
[0106] Step 54, for each association rule, calculate its support, confidence and lift, support indicates the frequency of the rule appearing in the data, confidence indicates the reliability of the rule, lift indicates the correlation strength between the antecedent and the consequent in the rule, specifically including: for each association rule, calculate its support, confidence and lift. Support indicates the frequency of the rule appearing in the data; confidence indicates the reliability or conditional probability of the rule; lift indicates the degree of improvement of the correlation strength between the antecedent and the consequent in the rule relative to when they appear independently; traverse all the output association rules, and calculate the values of the above three indicators according to the definition.
[0107] Step 55, sort and screen the association rules according to support, confidence and lift, and find the corresponding key rules; analyze the key rules to obtain the correlation and dependency between different risk factors, specifically including: sorting all association rules according to the calculated support, confidence and lift values. This helps to identify the rules with the most statistical significance and business value. According to the sorting results and preset thresholds (such as the minimum values of support, confidence and lift), screen out the key rules, which represent the strongest and most meaningful associations in the data; conduct in-depth analysis of the screened key rules to reveal the correlation and dependency between different risk factors. This helps to understand the interaction mechanism between risk factors and provide support for subsequent risk management and decision-making.
[0108] In an embodiment of the present invention, the trained neural network model is used to forward propagate the data set to be analyzed, and the risk factor value corresponding to each observation point can be accurately predicted. This helps to identify potential risk factors more accurately and provide support for subsequent risk management and decision-making. The Apriori algorithm can mine the association rules between risk factors in the data set, which may be difficult to find in traditional statistical analysis. By revealing these hidden associations, the interaction and influence mechanism between risk factors can be more deeply understood. By calculating the support, confidence and lift of the association rules, the frequency, reliability and correlation strength of the evaluation rules can be quantified, which helps to more objectively evaluate the degree of association between different risk factors. Based on the key association rules mined, risk management strategies can be formulated more specifically. For example, by identifying highly correlated risk factor combinations, those factors that have the greatest impact on the overall risk can be prioritized, thereby improving the efficiency and effectiveness of risk management.
[0109] In a preferred embodiment of the present invention, the above step 6, assigning an importance score to each risk factor based on the relationship between the risk factors, may include:
[0110] Step 61, determining quantitative indicators of the mutual relationship, the quantitative indicators include support, confidence and lift;
[0111] Step 62, according to the quantitative indicators, Calculate importance scores;
[0112] Among them, f r (A∪B) represents the frequency of transactions where item sets A and B appear simultaneously for association rule r; F r (A) and f r (B) represents the transaction frequency of item set A and item set B for association rule r; N represents the total number of transactions in the data set; w S 、w C and w L is the weight parameter; max r∈R N represents the maximum value of the unnormalized importance score in all association rule sets R; r represents the normalized importance score of association rule r.
[0113] In an embodiment of the present invention, by assigning an importance score to each risk factor, the key risk factors in the data set can be identified more intuitively. This helps risk managers quickly locate risk points that need to be focused on and improves the pertinence and efficiency of risk management. The use of quantitative indicators such as support, confidence and lift enables quantitative evaluation of the relationships between risk factors. This quantitative evaluation method is more objective and accurate, and helps to deeply understand the complex relationships between risk factors. By introducing weight parameters, users can flexibly adjust the proportion of support, confidence and lift in the score calculation according to actual needs. This flexibility makes the evaluation process more in line with actual business scenarios and improves the practicality and operability of the evaluation results. The importance scores of all association rules are normalized so that the scores between different rules are comparable. This helps to sort and screen risk factors on a global scale, further simplifying the risk management decision-making process.
[0114] In a preferred embodiment of the present invention, the calculation formula of the dynamic adjustment factor is:
[0115]
[0116] Among them, F represents the dynamic adjustment factor; w represents the weight of traffic flow; V represents the current traffic flow; U represents the benchmark traffic flow; z represents the weight of visibility; a represents the peak value of the Gaussian function; S represents the current visibility; T represents the optimal visibility; and C represents the standard deviation of the Gaussian function related to visibility.
[0117] In the embodiment of the present invention, the traffic flow part in the formula It can reflect the proportional relationship between the current traffic flow and the benchmark traffic flow in real time. When the traffic flow changes, the dynamic adjustment factor will be adjusted accordingly, thereby achieving a rapid response to the traffic status. By introducing the Gaussian function The formula can accurately quantify the impact of visibility on the traffic environment. The characteristics of the Gaussian function make the adjustment factor change smoothly when visibility is close to the optimal value; and when visibility is far from the optimal value, the adjustment factor changes rapidly, thus more accurately reflecting the importance of visibility to traffic safety. The weight parameters in the formula and the parameters of the Gaussian function can be flexibly configured according to actual conditions. This enables the dynamic adjustment factor to adapt to different traffic scenarios and specific needs, improving its versatility and practicality. By comprehensively considering traffic flow and visibility, the dynamic adjustment factor can provide a scientific decision-making basis for the traffic management system. For example, in the case of traffic congestion or poor visibility, the control strategy of traffic lights can be adjusted in time to optimize the vehicle passage order, thereby improving traffic safety and efficiency.
[0118] In the embodiment of the present invention, the risk factors are sorted according to the adjusted importance scores to obtain the key factors affecting the highway area risk, which specifically include:
[0119] Using the adjusted importance score, all risk factors are ranked from high to low; the ranking results will reflect the relative impact of each risk factor on the highway area risk. Based on the ranking results, risk factors with higher scores (i.e., greater impact) are selected as key factors. The number of key factors can be determined based on actual conditions and needs, such as selecting the top 5% or top 10% of risk factors.
[0120] A highway risk factor identification device, comprising:
[0121] An acquisition module, used to acquire data related to highway risks, including traffic flow, weather conditions, and road conditions;
[0122] A preprocessing module is used to preprocess the data related to highway road risks to obtain preprocessed data;
[0123] An extraction module, used to extract key features from the preprocessed data, wherein the key features are used for risk factor identification;
[0124] A training module, used to train the neural network model using key features to obtain a trained neural network model;
[0125] The identification module is used to identify risk factors through the trained neural network model and conduct association mining on risk factors to obtain the relationship between risk factors;
[0126] The calculation module is used to assign an importance score to each risk factor based on the relationship between the risk factors; calculate the dynamic adjustment factor, multiply the dynamic adjustment factor by the importance score of each risk factor to obtain the adjusted importance score; and sort the risk factors according to the adjusted importance score to obtain the key factors affecting the risk of the highway area.
[0127] The embodiment of the present invention also provides a computer-readable storage medium storing instructions, which, when executed on a computer, enable the computer to execute the method described above. All implementations in the above method embodiment are applicable to this embodiment and can achieve the same technical effect.
[0128] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for identifying highway risk factors, characterized in that: The method comprises: Step 1, obtaining data related to highway risks, including traffic flow, weather conditions, and road conditions; Step 2, preprocessing the data related to the highway road area risk to obtain preprocessed data; Step 3, extracting key features from the preprocessed data, wherein the key features are used for risk factor identification; Step 4, using the key features to train the neural network model to obtain a trained neural network model; Step 5, identifying risk factors through the trained neural network model, and performing association mining on the risk factors to obtain the relationships between the risk factors; Step 6, based on the relationship between risk factors, assign an importance score to each risk factor; calculate the dynamic adjustment factor, multiply the dynamic adjustment factor by the importance score of each risk factor to obtain the adjusted importance score; sort the risk factors according to the adjusted importance score to obtain the key factors affecting the highway area risk.
2. The highway risk factor identification method according to claim 1 is characterized in that: Obtain data related to highway risks, including traffic flow, weather conditions, and road conditions, including: The selection and configuration of data collection points is considered as an optimization problem, with the goal of maximizing the efficiency and quality of data collection; decision variables are determined, including the location, type, and frequency of data collection of sensors; Use binary coding to represent the solution and initialize a population containing multiple solutions; Define a fitness function to evaluate the quality of each individual; use the roulette wheel selection method to select the corresponding individuals from the current population to enter the next generation; perform a crossover operation on the selected individuals to generate a new solution; perform a mutation operation on the newly generated individuals, and repeat the selection, crossover and mutation operations until the termination condition is met. When the termination occurs, output the final solution, that is, the final sensor configuration solution; Deploy sensors according to the final sensor configuration plan to obtain data related to highway road risks.
3. The highway risk factor identification method according to claim 2 is characterized in that: Key features are extracted from the preprocessed data, and the key features are used for risk factor identification, including: Determine the name, definition, and expected meaning of each feature based on the data dictionary in the preprocessed data; quantify the value range and distribution of each feature through the mean; determine the target variable, that is, the risk factor indicator to be identified; Calculate the Pearson correlation coefficient between each feature and the target variable; construct a correlation coefficient matrix based on the Pearson correlation coefficient; analyze the correlation coefficient matrix to identify the identification features related to the target variable; According to the Pearson correlation coefficient, a screening threshold is set; according to the screening threshold, key features that meet the screening threshold conditions are screened out from the identification features.
4. The highway risk factor identification method according to claim 3 is characterized in that: The neural network model is trained using key features to obtain a trained neural network model, including: Initialize the weights and bias parameters of the convolutional neural network; Define the loss function to measure the gap between the predicted value of the neural network model and the true value; During the training process, the gradient descent algorithm is used to optimize the weights and bias parameters of the neural network; For each training batch, calculate the gradient of the loss function with respect to the weights and biases; According to the calculated gradient, update the weight and bias parameters and repeat the process until the stopping condition is met; set the learning rate and optimizer; divide the preprocessed data into training set and validation set; Using the training set data, calculate the predicted value of the neural network model through forward propagation and calculate the value of the loss function; using the gradient descent algorithm and optimizer, update the weight and bias parameters of the neural network according to the gradient of the loss function; after each iteration cycle, evaluate the performance of the neural network model on the validation set to obtain the evaluation results; The structure, learning rate and optimizer parameters of the neural network model are adjusted according to the evaluation results to obtain the trained neural network model.
5. The highway risk factor identification method according to claim 4 is characterized in that: The risk factors are identified through the trained neural network model, and the correlation mining of risk factors is carried out to obtain the relationship between risk factors, including: The data set to be analyzed is input into the trained neural network model; through the forward propagation process of the trained neural network model, the predicted value of the risk factor corresponding to each observation point is obtained; Set the minimum support and minimum confidence thresholds for the Apriori algorithm; run the Apriori algorithm, input the transaction data set, minimum support, and minimum confidence; The Apriori algorithm will output a series of association rules, each rule is represented as a combination of an antecedent and a consequent; For each association rule, calculate its support, confidence and lift. Support indicates the frequency of the rule in the data, confidence indicates the reliability of the rule, and lift indicates the strength of the correlation between the antecedent and the consequent in the rule. Sort and filter the association rules according to support, confidence and lift to find the corresponding key rules; analyze the key rules to obtain the correlation and dependency between different risk factors.
6. The highway risk factor identification method according to claim 5 is characterized in that: Each risk factor is assigned an importance score based on the interrelationships between risk factors, including: Determine the quantitative indicators of the relationship, including support, confidence and lift; According to quantitative indicators, Calculate importance scores; Among them, f r (A∪B) represents the frequency of transactions where item sets A and B appear simultaneously for association rule r; f r (A) and f r (B) represents the transaction frequency of item set A and item set B for association rule r; N represents the total number of transactions in the data set; w S 、w C and w L is the weight parameter; max r∈R N represents the maximum value of the unnormalized importance score in all association rule sets R; r represents the normalized importance score of association rule r.
7. The highway risk factor identification method according to claim 6 is characterized in that: The calculation formula of the dynamic adjustment factor is: Among them, F represents the dynamic adjustment factor; w represents the weight of traffic flow; V represents the current traffic flow; U represents the benchmark traffic flow; z represents the weight of visibility; a represents the peak value of the Gaussian function; S represents the current visibility; T represents the optimal visibility; and C represents the standard deviation of the Gaussian function related to visibility.
8. A highway risk factor identification device, the system implementing the method according to any one of claims 1 to 7, characterized in that: include: An acquisition module, used to acquire data related to highway risks, including traffic flow, weather conditions, and road conditions; A preprocessing module is used to preprocess the data related to highway road risks to obtain preprocessed data; An extraction module, used to extract key features from the preprocessed data, wherein the key features are used for risk factor identification; A training module, used to train the neural network model using key features to obtain a trained neural network model; The identification module is used to identify risk factors through the trained neural network model and conduct association mining on risk factors to obtain the relationship between risk factors; a calculation module for assigning an importance score to each risk factor based on the interrelationships between the risk factors; Calculate the dynamic adjustment factor and multiply the dynamic adjustment factor by the importance score of each risk factor to obtain an adjusted importance score; According to the adjusted importance scores, the risk factors are ranked to obtain the key factors affecting the highway area risk.
9. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Highway safety hazard road section identification method and device
CN106935030A
Traffic risk assessment method, system and device
CN117475631A
Expressway area risk prediction method and device
CN118015839A
Real-time expressway traffic operation risk identification method and identification system under ice and snow conditions based on dynamic multilayer fuzzy logic
CN119068680A
Systems and methods for generating project plans from predictive project models
WO2014043338A1
Cited By
Traffic safety support system and learning method executable by the same
US12431011B2