Data processing method and device for risk category analysis, equipment and medium

By constructing vehicle features and training a category recognition model, and using deep learning methods to extract features from vehicle data and perform K-fold cross-validation, the problem of refining vehicle risk category recognition is solved, enabling rapid and accurate identification and management of vehicle risk categories.

CN121834450APending Publication Date: 2026-04-10PCI TECH GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies cannot achieve refined classification and identification of vehicle risk categories, making them unsuitable for the actual needs of risk source classification and hierarchical management.

Method used

By constructing vehicle features and training a category recognition model, deep learning methods are used to extract features from vehicle data and perform K-fold cross-validation to generate a feature dataset and deploy the category recognition model to identify the specific risk category of the vehicle.

Benefits of technology

It enables rapid and accurate identification of vehicle risk categories, supports refined classification and hierarchical management, and timely prevention and control of potential risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834450A_ABST
    Figure CN121834450A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device for risk category analysis, equipment and a medium, relates to the technical field of computers, and solves the problem that the risk category of a vehicle is difficult to accurately position in the related art. The risk category of the target vehicle is identified, so that the specific risk category of the vehicle is further accurately identified, and fine classification and level-to-level management is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a data processing method and device for risk category analysis, equipment and medium. BACKGROUND

[0002] As the main participant in the transportation process, accurately identifying various high-risk vehicles and managing the risk source is an important means to eliminate traffic safety hazards. In related technologies, vehicle traffic safety risk identification is usually based on one or more types of data such as vehicle data and driving environment data, and methods such as clustering, decision tree, logistic regression, and deep learning are used to quantitatively evaluate or grade the risk level of a single vehicle. However, in related technologies, whether based on traditional decision trees, analytic hierarchy process, logistic regression, or machine learning methods, the above solutions focus on static or dynamic quantitative evaluation of vehicle risk to identify high-risk vehicles, and cannot achieve fine classification and identification. However, from the perspective of vehicles, the risk hidden dangers and solutions of risk vehicles driven by different behaviors may be very different. For example, illegal transport vehicles and high-risk vehicles may both be high-risk vehicles, but the actual risks they contain are obviously different. Therefore, the identification scheme of related technologies that classifies all high-risk vehicles into one category cannot meet the actual needs of risk source classification and grading management. SUMMARY

[0003] The present application provides a data processing method, device, equipment and medium for risk category analysis, which solves the problem of accurately positioning the risk category of vehicles in related technologies. The present application establishes a category identification model by constructing vehicle features and training, and realizes the identification of the risk category of target vehicles, thereby accurately identifying the specific risk category to which the vehicle belongs, which helps to achieve classification and grading management.

[0004] In a first aspect, the present application provides a data processing method for risk category analysis, which includes: Performing outlier detection on the obtained vehicle data corresponding to a plurality of target vehicles respectively to determine sample data as valid samples; Based on a preset risk feature type, performing feature extraction on the sample data and determining feature items corresponding to each risk feature type; According to the data type corresponding to the feature item, the feature item is feature data to generate a feature data set; Based on the feature data set, the pre-constructed identification model is trained using a K-fold cross-validation method to obtain a category identification model, which is used to identify the corresponding risk category of the input vehicle data; deploy the category recognition model and perform risk category analysis on the input vehicle data through the category recognition model to determine the risk category of the vehicle.

[0005] In a second aspect, the present application also provides a data processing apparatus for risk category analysis, comprising: a data preprocessing module configured to perform outlier detection on the acquired vehicle data corresponding to a plurality of target vehicles respectively to determine sample data as valid samples; a feature extraction module configured to perform feature extraction on the sample data based on a preset risk feature type and determine feature items corresponding to each risk feature type; a feature dataization module configured to perform feature dataization on the feature items according to a data type corresponding to the feature items to generate a feature data set; a model training module configured to perform model training on a pre-constructed recognition model based on the feature data set using a K-fold cross-validation method to obtain a category recognition model, the category recognition model being used to recognize a corresponding risk category of input vehicle data; a model output module configured to deploy the category recognition model and perform risk category analysis on the input vehicle data through the category recognition model to determine the risk category of the vehicle.

[0006] In a third aspect, the present application also provides an electronic device, comprising: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method for risk category analysis of the present application.

[0007] In a fourth aspect, the present application also provides a storage medium storing computer executable instructions, when the computer executable instructions are executed by a processor, the computer executable instructions are used to execute the data processing method for risk category analysis of the present application.

[0008] The present application scheme fuses vehicle related data, constructs features of the vehicle, and trains a model capable of recognizing the risk category of the vehicle through sample data by using a deep learning method, so as to quickly and accurately evaluate and analyze the risk category of a specific key vehicle. Not only can the vehicle be identified as high risk, but also the specific risk category to which the vehicle belongs can be accurately identified, which helps to achieve fine classification and grading management, and helps to guide the actual vehicle risk source management work, and is beneficial to timely prevention and control of risk hidden dangers. BRIEF DESCRIPTION OF DRAWINGS

[0009] Figure 1 A step schematic diagram of the data processing method for risk category analysis provided by an embodiment of the present application.

[0010] Figure 2 A step schematic diagram of feature processing provided for an embodiment of the present application.

[0011] Figure 3 A step schematic diagram of model training provided for an embodiment of the present application.

[0012] Figure 4 A structure schematic diagram of a data processing apparatus for risk category analysis provided for an embodiment of the present application.

[0013] Figure 5 A structure schematic diagram of an electronic device provided for an embodiment of the present application. DETAILED DESCRIPTION

[0014] The embodiments of the present application will be further described below in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the embodiments of the present application, and not to limit the embodiments of the present application. In addition, it should be noted that, in order to facilitate description, only parts related to the embodiments of the present application are shown in the drawings, and those skilled in the art can think that, as long as the technical features are not contradictory, any combination of technical features can constitute an optional embodiment after reading the description of the present application.

[0015] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be exchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually a class, not limited to the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally represents an "or" relationship between the front and rear associated objects. In the description of the present application, "multiple" means two or more, and "several" means one or more.

[0016] As the most important participant in the transportation process, accurately identifying various high-risk vehicles and managing the risk source is an important means to eliminate traffic safety hazards. In related technologies, big data, cloud computing, artificial intelligence and other digital technologies provide new detection means for the management of road traffic safety risks. The identification of vehicle traffic safety risks is usually based on one or more types of data such as vehicle data and driving environment data, and methods such as clustering, decision tree, logistic regression, and deep learning are used to quantitatively evaluate or grade the risk level of a single vehicle.

[0017] However, whether based on traditional decision tree, analytic hierarchy process, logistic regression and other mathematical models, or based on machine learning methods, the above solutions are focused on the quantitative evaluation of static or dynamic vehicle risks to identify high-risk vehicles, and cannot achieve fine classification and identification. But from the perspective of vehicles, the risk hidden dangers and their solutions of risk vehicles driven by different behaviors may be very different. For example, illegal transport vehicles and high-risk vehicles may be high-risk vehicles, but the actual risks they contain are obviously different. Therefore, the identification scheme of the related art that classifies all high-risk vehicles into one category cannot meet the actual needs of risk source classification and hierarchical management.

[0018] The present application provides a data processing method for risk category analysis, Figure 1 The steps of the data processing method for risk category analysis provided by an embodiment of the present application are shown in the schematic diagram. The method can learn features based on various parameters and information carried in vehicle data, and then learn the nonlinear, high-dimensional and complex correlation between the data of target vehicles such as typical accident risk vehicles or specific risk type vehicles and the risk categories corresponding to the vehicles through a deep learning method, to establish a class category identification model for identifying the risk categories of vehicles, and to realize the analysis and identification of the risk categories of vehicles, thereby helping to guide the risk hidden danger monitoring and management of vehicles, and being beneficial to timely prevention and control of risk hidden dangers. The specific steps include steps S110-S150.

[0019] Step S110, performing outlier detection on the obtained vehicle data corresponding to a plurality of target vehicles respectively, to determine sample data as valid samples.

[0020] The obtained vehicle data exist in multiple and correspond to different target vehicles respectively. It can be conceived that the target vehicles are vehicles as training samples. Exemplarily, vehicle data corresponding to vehicles belonging to a specific risk category such as vehicles involved in road traffic accidents or vehicles involved in traffic violations can be selected. Of course, multiple vehicle data can also be configured according to a plurality of risk categories, so that there are multiple vehicle data under each risk category. After receiving the vehicle data, data preprocessing is performed on the vehicle data to realize data screening and cleaning, so as to facilitate subsequent data processing. For this purpose, the plurality of vehicle data are subjected to outlier detection to screen out outliers in the data, thereby leaving valid samples. It can be conceived that in the case where a plurality of types of parameter information are included in the same vehicle data, different screening strategies are adopted for different types of parameter information.

[0021] Optionally, the vehicle data includes vehicle basic information, vehicle passing detection information, accident and illegal information, and preset risk category information. The vehicle basic information is used to record the basic attribute parameters of the vehicle, such as the license plate of the vehicle, the use nature of the vehicle (such as a commercial vehicle or a non-commercial vehicle), the total weight, the number of seats, and other basic information related to the vehicle. The vehicle passing detection information is used to record the detection time and the detection position when the vehicle passes through the road monitoring portal. It is conceivable that the road monitoring portal can be a position corresponding to a road toll station, a monitoring intersection, and the like, and the passing vehicle can be detected through a monitoring camera arranged at the position. The vehicle passing detection information is associated with the vehicle and the road monitoring portal, that is, the road monitoring portal that detects the vehicle in the vehicle data generates corresponding vehicle passing detection information, and the time when the vehicle passes through the road monitoring portal and the position of the road monitoring portal can be determined through the vehicle passing detection information. The accident and illegal information is used to record the number of violations and the number of accidents of the vehicle in each time period, such as recording the number of violations and the number of accidents of the vehicle in the month. The preset risk category information is used to record the preconfigured risk category of the vehicle. In this regard, corresponding risk categories are configured in the sample data corresponding to different vehicle data to train the model.

[0022] Specifically, the vehicle data carrying empty vehicle basic information or preset risk category information is removed from the plurality of vehicle data. That is, for the vehicle basic information and the preset risk category information, the empty information is removed, that is, the vehicle data is removed when any of the vehicle basic information or the preset risk category information of the vehicle data is empty. Moreover, based on the detection position and the detection time in the passing vehicle detection information, the vehicle speed value is determined and the vehicle data with a vehicle speed value exceeding a preset vehicle speed is removed. It is conceivable that the detection position recorded in the passing vehicle detection information is associated with the position of the road monitoring portal, and the vehicle speed value can be determined based on the distance between the plurality of road monitoring portals and the detection time of the vehicle passing through, and the vehicle data with a vehicle speed value exceeding a preset vehicle speed is removed. By setting the preset vehicle speed, the data with abnormal speed is removed to avoid its interference with subsequent model training. In addition, the box plot corresponding to the number of violations and the number of accidents recorded in the accident violation information in all vehicle data is determined, and the vehicle data corresponding to the abnormal data points located outside the box plot is removed. The box plot (also known as the box plot) is a statistical chart used to display the dispersion of a set of data, which can represent the maximum value, the minimum value, the median value and the upper and lower quartiles in the set of data. By determining the box plot corresponding to the number of violations and the box plot corresponding to the accident, the data points outside the edge line of the box plot are deleted, so that the abnormal data is extracted. Furthermore, after the screening operation according to the vehicle basic information, the passing vehicle detection information, the accident violation information and the preset risk category information, the remaining vehicle data is used as sample data. It is conceivable that the screening operation of the above different parameter information can be performed synchronously or in any order, and the screening can be achieved. Therefore, by screening the abnormal data, the present scheme can construct more accurate sample data to facilitate better model training and improve the recognition accuracy of the model.

[0023] In step S120, based on the preset risk feature type, the feature extraction is performed on the sample data to determine the feature item corresponding to each risk feature type.

[0024] After obtaining the sample data, the feature extraction is performed on the data to determine the features of the target vehicle. It is conceivable that the feature extraction of the sample data is used to construct the feature item associated with different parameter information. In the present embodiment, the risk feature types corresponding to different risk feature types correspond to different feature items, for example, the preset risk feature types include a first type associated with vehicle basic information, a second type associated with passing vehicle detection information and a third type associated with accident violation information. For different types, the feature items constructed therefrom are determined based on different parameter information, such as the feature items of the first type based on the vehicle basic information for feature extraction.

[0025] Step S130, according to the data type corresponding to the feature item, the feature data of the feature item is generated to generate a feature data set.

[0026] It can be understood that in the constructed feature item, there are different data types, wherein the data types include category type and numerical type. For this, the feature item includes a first feature item as a category type feature and a second feature item as a numerical type feature, wherein the first feature item is a feature represented in a non-numeric form, and the second feature item is a feature represented in a numeric form. It is conceivable that the feature data processing of the feature item is used to convert the feature item into a form recognizable by the model, so as to input the pre-constructed recognition model for model training. In the feature data processing, different processing methods are used for different data types in the embodiments.

[0027] Optionally, after the construction of the feature item is completed, the first feature item as a category type feature and the second feature item as a numerical type feature in all feature items are determined. In the constructed initial data set, for the first feature item, the first feature item is numerized based on label encoding or one-hot encoding and stored in the initial data set. Wherein, the label encoding is used to map each category to an integer value, such as incrementing from 0. The one-hot encoding is used to create a new binary feature for each category of each discrete attribute. For example, in some embodiments, the label encoding is used for features with strong mutual exclusivity and small number of categories, while the one-hot encoding is used for features with large number of categories and no hierarchical relationship, so as to avoid false order relationship between categories.

[0028] And for the second feature item, the second feature item is normalized according to the Z-score standardization method and stored in the initial data set. Wherein, the Z-score is also called the standard score, and the Z-score standardization method is used to divide the difference between a number and the average number by the standard deviation, which can truly reflect the relative standard distance of a score from the average number. By the Z-score standardization method, the present scheme can eliminate the influence of the dimension difference between different feature items on the model training.

[0029] Furthermore, after traversing all feature items, for the constructed initial dataset, the Pearson correlation coefficients between each feature in the initial dataset are determined, and features with absolute Pearson correlation coefficients greater than a preset value are removed to output the feature dataset. The Pearson correlation coefficient (PPMCC or PCCs) measures the degree of correlation (linear correlation) between two variables X and Y, with a value between -1 and 1. Therefore, the Pearson correlation coefficient is used to analyze the linear correlation between features, and by comparing the absolute value of the Pearson correlation coefficient with a preset value, features with absolute values ​​greater than the preset value are removed, thus finally obtaining the feature dataset. Each data point in this feature dataset corresponds to a feature vector and risk category label for a single vehicle. This scheme, by digitizing the feature items, forms a feature dataset with a unified structure and optimized dimensions, which helps to better optimize the recognition model and thus improve the accuracy of vehicle risk category recognition.

[0030] Step S140: Based on the feature dataset, the pre-built recognition model is trained using the K-fold cross-validation method to obtain the category recognition model.

[0031] The category recognition model is used to identify the corresponding risk category from the input vehicle data. The feature dataset serves as samples, and is used to train the pre-built recognition model. During model training, a K-fold cross-validation method is employed. This method divides the dataset into K mutually exclusive subsets for K iterations. In each iteration, the model is trained and evaluated based on the dataset, and the overall performance of the model is determined through the evaluation results. This method can more comprehensively evaluate the model's generalization ability, reduce evaluation fluctuations caused by the randomness of data partitioning, and thus provide more stable performance metrics, contributing to obtaining a more accurate category recognition model.

[0032] Step S150: Deploy the category recognition model and perform risk category analysis on the input vehicle data using the category recognition model to determine the risk category of the vehicle.

[0033] The category recognition model is deployed on electronic devices such as the devices and servers to facilitate identification. The input vehicle data serves as the data to be identified, is associated with the vehicle to be identified, and the risk category of the vehicle is determined through the category recognition model. Optionally, during model deployment, the model can be optimized for the specific electronic device, such as through pruning and quantization to reduce its size and adapt to the hardware of the electronic device. Furthermore, corresponding API interfaces can be configured to output the recognition results.

[0034] As can be seen from the above scheme, this scheme integrates vehicle-related data to construct vehicle features, and uses deep learning methods to train a model that can identify risk categories using sample data. This enables rapid and accurate assessment and analysis of the risk categories of specific key vehicles. It can not only identify whether a vehicle is high-risk, but also accurately identify the specific risk category to which the vehicle belongs. This helps to achieve refined classification and hierarchical management, and also helps to provide targeted guidance for the actual management of vehicle risk sources, which is conducive to timely prevention and control of potential risks.

[0035] In one embodiment, the risk characteristic types include a first type, a second type, and a third type. The vehicle basic information includes vehicle usage type, vehicle age, number of seats, and total weight. The first type of feature includes a percentage value and a load factor. Based on the vehicle usage type and vehicle age recorded in the sample data, multiple different target combinations and their percentage values ​​in all vehicle data are determined, and these percentage values ​​are used as the first type of feature. The target combination is a parameter combination of the target vehicle usage type and a preset vehicle age. Different target combinations are associated with different risk categories. By determining the percentage of different target combinations, derived features of the corresponding vehicle basic attributes can be constructed. For example, if the configured target combination is an operating vehicle and the vehicle is older than 5 years, the percentage value of this parameter combination is found in the vehicle data, and this percentage value is used as the first type of feature. Additionally, the ratio of the maximum load capacity to the total weight corresponding to the number of seats in the vehicle basic information is used as the corresponding load factor feature in the first type. It is conceivable that this load factor is the ratio of the maximum load capacity to the total weight, and the maximum load capacity is the product of the number of seats and the preset load capacity.

[0036] For the second type, which is associated with vehicle detection information, the driving trajectory information is reconstructed based on the vehicle detection information in the sample data, and the corresponding travel characteristics, range concentration, and abnormal travel patterns in the second type are determined. Travel characteristics are features associated with the travel range corresponding to the driving trajectory information. For example, the travel range covered by the driving trajectory information can be used to determine whether the vehicle frequently crosses cities or provinces, thus constructing the corresponding travel characteristics. Range concentration is used to represent the number of target areas covered by the travel range. For example, if the target areas are divided into different regions according to cities, the number of cities covered can be used to determine the number of cities the vehicle is active in, or to determine whether the number of cities meets a preset requirement. Abnormal travel mode refers to the mode where there are abnormal trajectories in the driving trajectory information. It can be conceivable that the abnormal trajectory can be determined by setting corresponding parameters to determine whether the driving trajectory is abnormal. For example, a trajectory distance threshold can be configured to determine that the current driving trajectory is a long-distance travel behavior when the driving trajectory exceeds the trajectory distance threshold. And the corresponding average daily travel time can be determined by the detection time in the vehicle passing detection information. Combining the long-distance travel behavior and the average daily travel time, it can be determined whether there is an abnormal trajectory. For example, if there is a long-distance travel and the average daily travel time is less than 2 hours, it can be determined that there is an abnormal behavior.

[0037] Based on the accident and violation information in the sample data, the accident rate (corresponding to the number of accidents and total mileage), the month-on-month growth rate of the corresponding number of violations, and the type of violation are determined as the third type of feature items. The accident and violation information records the number of violations and accidents for each time period, such as recording violations and accidents monthly. The accident rate is expressed as the ratio of the number of accidents to the total mileage, i.e., the accident rate per unit of travel mileage. The month-on-month growth rate can be determined by comparing the number of violations in two adjacent months, or optionally by comparing the number of violations in two adjacent quarters. The comparison period can be set according to the actual application requirements.

[0038] To address this, this solution extracts features from the sample data to construct feature terms, which then characterize the vehicle data, facilitating the construction of a feature dataset for model training. Furthermore, by extracting features from various parameter information within the vehicle data, the feature dataset constructed by this solution more closely reflects the actual risk categories of vehicles, thereby improving the accuracy of model identification.

[0039] Optionally, Figure 2 This is a schematic diagram illustrating the feature processing steps provided in an embodiment of this application, as shown below. Figure 2 As shown, for vehicle detection information, the reconstructed driving trajectory information can, in one embodiment, be determined by path planning based on the detection time and location recorded in the vehicle detection information when passing through the road monitoring checkpoint, and by calculating the probability of each path. Specific steps include S210-S220: Step S210: Determine several vehicle detection information corresponding to the same vehicle, and based on the Dijkstra method, determine multiple shortest paths between different road monitoring checkpoints corresponding to the several vehicle detection information to generate a candidate path set.

[0040] Step S220: Based on Bayesian theory algorithm, determine the probability of passage of all paths in the candidate path set, and select the target path with the highest probability of passage in the candidate path set as the driving trajectory in the driving trajectory information.

[0041] Understandably, Dijkstra's method (i.e., Dijkstra's algorithm) uses a greedy algorithm strategy, starting from a starting point, to iterate through the nearest unvisited vertex's adjacent nodes until it reaches the endpoint, thus determining the shortest path from one vertex to all other vertices. It's conceivable that by using vehicle detection information, we can identify the two road monitoring checkpoints that detected the same vehicle, i.e., determine the upstream detection time (i.e., the detection time of the upstream road monitoring checkpoint), the downstream detection time (i.e., the detection time of the downstream road monitoring checkpoint), the upstream detection segment, and the downstream detection segment. Dijkstra's method can then be used to find multiple shortest paths, thereby constructing a candidate path set.

[0042] Bayesian theory algorithms describe the relationship between the conditional probabilities of two events using Bayes' theorem, such as the start time. End time Starting point and finish line The probability of passage for the corresponding path can be calculated using the following formula:

[0043] in Indicates the path between two points. Indicates road segment, The first part of the above formula (i.e. The expression represents the prevalence of choosing path p from the starting point to the ending point. Optionally, it can be further decomposed into the probability of the starting segment and the transition probability of subsequent segments, such as obtaining the corresponding probability statistics through Dirichlet priors or historical GPS vehicle trajectories, and using these as probability values. The second part of the above equation (i.e. This represents the consistency between actual and expected driving events, such as dividing time into 24 hourly intervals and calculating the average speed for each time interval of each road segment. Optionally, matrix factorization algorithms can be used to address the problem of data scarcity. Based on this, the probability of passage for all paths in the candidate path set can be determined, and then the target path with the highest probability of passage can be selected as the driving trajectory in the driving trajectory information. To this end, this scheme analyzes and extracts travel features from vehicle data by reconstructing the driving trajectory, thereby constructing a feature dataset that better represents the risk category. This feature dataset is then used for model training, thereby improving the accuracy of risk category identification.

[0044] Figure 3 This is a schematic diagram illustrating the model training steps provided in one embodiment of the present application. In one embodiment, a pre-built recognition model is trained using the K-fold cross-validation method to obtain a category recognition model. Specifically, in this embodiment, the K-fold cross-validation method can be used to partition the feature training set to achieve model training, and the specific steps include S310-S350.

[0045] Step S310: Perform stratified sampling of the feature dataset according to a preset ratio to divide the feature dataset into a training set and a test set.

[0046] Step S320: Divide the training set into multiple mutually exclusive subsets on an average basis according to the number of partitions preset in the K-fold cross-validation method.

[0047] Step S330: Select several mutually exclusive subsets as training subsets and use the remaining mutually exclusive subsets as verification subsets.

[0048] Step S340: Perform multiple full training sessions based on the selected training subset each time, and record the classification accuracy and F1 score of the validation subset in each training session.

[0049] Step S350: When both the target accuracy and the target score reach the corresponding thresholds, the category recognition model is determined.

[0050] Understandably, the feature dataset is partitioned to distinguish between training and test sets. The training set is used to train the model, while the test set is used to validate model performance. Specifically, stratified sampling is used during the partitioning process to ensure that the sample proportions of each risk category in the training and test sets are consistent with those in the original dataset, avoiding model bias caused by class imbalance. The training set is then further partitioned into multiple mutually exclusive subsets. This number of partitions corresponds to the pre-set K value in the K-fold cross-validation method. For example, when K is 5, the training set is divided into an average of 4 training subsets and 1 validation subset. Furthermore, the selected training and validation subsets are different each time, for K iterations. It can be expected that in the 5 mutually exclusive subsets, one mutually exclusive subset is selected as the validation subset in each iteration, and the validation subset is different in each iteration.

[0051] Therefore, in each training iteration, the model is fully trained on the training subset, and the classification accuracy and F1 score of the validation subset are recorded in each training iteration. Classification accuracy represents the percentage of correctly predicted results out of the total samples, while the F1 score is related to precision and recall, and is typically twice the ratio of the product of precision and recall to their sum. Precision is the probability that a sample predicted as positive is actually positive among all samples predicted as positive, and recall is the probability that a sample actually positive is predicted as positive. Classification accuracy reflects overall classification correctness, precision and recall measure the accuracy and completeness of positive class predictions, respectively, and the F1 score balances both aspects to optimize model performance. Therefore, classification accuracy and F1 score are selected as evaluation metrics, and corresponding thresholds are set for each, for example, a threshold of 90% for classification accuracy and a threshold of 0.85 for F1 score. Furthermore, the evaluation results are averaged before comparison to determine the target accuracy and target score. The target accuracy is the average of multiple classification accuracies, and the target score is the average of multiple F1 scores. Then, when both the target accuracy and target score reach their respective thresholds—for example, referring to the above example, when the target accuracy is greater than or equal to 90% and the target score is greater than or equal to 0.85—the model is considered successfully trained, ultimately yielding a category recognition model that meets the needs of practical applications. It is conceivable that retraining would be performed if the target accuracy and target score did not meet the thresholds.

[0052] In response, by using the above-mentioned scheme for model training, this scheme can more comprehensively evaluate the model's generalization ability, reduce evaluation fluctuations caused by the randomness of data partitioning, and thus provide more stable performance indicators, which helps to obtain a category recognition model with more accurate recognition accuracy.

[0053] In some embodiments, the pre-built recognition model includes a CNN feature extraction module and an XGBoost classification prediction module. The CNN feature extraction module is used to extract features from the feature vectors in the feature dataset through at least three convolutional layers, perform data standardization through batch normalization layers, introduce non-linear features through the ReLU activation function, and compress features through max pooling layers to output a high-dimensional feature vector. The XGBoost classification prediction module is used to predict the high-dimensional feature vector through multiple decision trees to determine the probability values ​​of the corresponding risk categories, and take the target risk category with the highest probability value as the risk category of the high-dimensional feature vector.

[0054] Understandably, the identification model employs a feature extraction and classification prediction architecture, combining the local feature capture capability of Convolutional Neural Networks (CNNs) with the classification prediction optimization of eXtreme Gradient Boosting (XGBoost) to achieve risk category identification. The CNN feature extraction module includes three convolutional layers, a batch normalization layer, and a max pooling layer. For example, the first convolutional layer has 64 kernels with a kernel size of 3, the second layer has 128 kernels with a kernel size of 2, and the third layer has 256 kernels with a kernel size of 1. Specifically, the first convolutional layer uses 64 kernels for feature capture, the batch normalization layer standardizes the data, the ReLU activation function introduces non-linear features, and finally, the max pooling layer compresses the feature dimensionality. Then it enters the next convolutional layer to capture features again through 128 convolutional kernels in the second convolutional layer. The above operations are repeated through batch normalization layer and max pooling layer. Finally, after three layers of processing, a high-dimensional feature vector is output.

[0055] The XGBoost classification prediction module uses multiple decision trees. Optionally, the maximum depth of these decision trees can range from 3 to 5, and the corresponding learning rate can be set from 0.01 to 0.1. The hyperparameters for each training iteration are randomly selected within this range, and the cross-entropy loss function is used as the optimization function. These hyperparameters are parameters that need to be pre-set before model training in machine learning. The XGBoost classification prediction module then uses multiple decision trees to make predictions, obtaining the prediction results of each decision tree (i.e., the probability of each risk category). The target risk category with the highest probability value is then selected as the risk category of the high-dimensional feature vector.

[0056] The category recognition model constructed in this scheme can combine the local feature capture capability of convolutional neural networks with the classification prediction advantage of extreme gradient boosting trees to accurately identify the risk category to which a vehicle belongs, achieving high-precision risk category recognition, thereby contributing to the refined implementation of classification and hierarchical management.

[0057] Figure 4 This is a schematic diagram of the structure of a data processing device for risk category analysis provided in an embodiment of this application. The device is used to execute the data processing method for risk category analysis provided in the above embodiment, and the device also has a functional module for executing the method and beneficial effects. As shown in the figure, the data processing device for risk category analysis includes a data preprocessing module 401, a feature extraction module 402, a feature data conversion module 403, a model training module 404, and a model output module 405.

[0058] The data preprocessing module 401 is configured to perform outlier detection on the acquired vehicle data corresponding to multiple target vehicles to determine the sample data that can be used as valid samples. The feature extraction module 402 is configured to extract features from the sample data based on preset risk feature types and determine the feature items corresponding to each risk feature type; The feature datafication module 403 is configured to perform feature datafication on the feature items according to the data type corresponding to the feature items to generate a feature dataset; The model training module 404 is configured to train the pre-built recognition model using the K-fold cross-validation method based on the feature dataset to obtain the category recognition model. The category recognition model is used to identify the corresponding risk category of the input vehicle data. The model output module 405 is configured to deploy a category recognition model and perform risk category analysis on the input vehicle data to determine the risk category of the vehicle.

[0059] Based on the above embodiments, the vehicle data includes basic vehicle information, vehicle detection information, accident and violation information, and preset risk category information. Basic vehicle information records the vehicle's basic attribute parameters; vehicle detection information records the detection time and location when the vehicle passes through a road monitoring checkpoint; accident and violation information records the number of violations and accidents for the vehicle in each time period; and preset risk category information records the risk categories pre-configured for the vehicle. The data preprocessing module 401 is specifically configured as follows: Remove vehicle data from multiple vehicle datasets where the basic vehicle information or preset risk category information is null. Based on the detection location and detection time in the vehicle detection information, determine the vehicle speed value and remove vehicle data whose speed value exceeds the preset speed. Determine the box plot of the number of violations and accidents recorded in the accident violation information of all vehicle data, and remove the vehicle data corresponding to the abnormal data points located outside the box plot; After filtering out vehicles based on their basic information, vehicle inspection information, accident and violation information, and preset risk category information, the remaining vehicle data is used as sample data.

[0060] Based on the above embodiments, the risk feature types include a first type associated with vehicle basic information, a second type associated with vehicle detection information, and a third type associated with accident and violation information. The vehicle basic information includes vehicle usage type, vehicle age, number of seats, and total weight. The feature extraction module 402 is specifically configured as follows: Based on the vehicle usage type and age recorded in the sample data, we determine multiple different target combinations and the proportion of each target combination in all vehicle data. The proportion is used as the first type of feature. The target combination is a parameter combination of the target vehicle usage type and the preset vehicle age. Different target combinations are associated with different risk categories. The ratio of the number of seats in the vehicle's basic information to the load capacity and total weight is used as the feature item of the corresponding load factor in the first type. Based on the vehicle detection information in the sample data, the driving trajectory information is restored and the corresponding travel characteristics, range concentration, and abnormal travel patterns in the second type are determined. The travel characteristics are the features associated with the travel range corresponding to the driving trajectory information. The range concentration is used to represent the number of target areas covered by the travel range. The abnormal travel patterns are the patterns in the driving trajectory information where there are abnormal trajectories. Based on the accident and violation information in the sample data, the accident rate of the corresponding number of accidents and total mileage, the month-on-month growth rate of the corresponding number of violations, and the violation type are determined as the feature items of the third type.

[0061] Based on the above embodiments, the feature extraction module 402 is further configured as follows: Several vehicle detection information corresponding to the same vehicle are identified, and multiple shortest paths between different road monitoring checkpoints in the several vehicle detection information are determined based on the Dijkstra method to generate a candidate path set; Based on Bayesian theory algorithms, the probability of passage of all paths in the candidate path set is determined, and the target path with the highest probability of passage is selected as the driving trajectory in the driving trajectory information.

[0062] Based on the above embodiments, the feature data generation module 403 is specifically configured as follows: Identify the first feature term that is a categorical feature and the second feature term that is a numerical feature among all feature terms; Based on label encoding or one-hot encoding, the first feature term is numericalized and stored in the initial dataset; The second feature term is normalized using the Z-score standardization method and stored in the initial dataset; Determine the Pearson correlation coefficients among the features in the initial dataset, and remove features whose absolute Pearson correlation coefficients are greater than a preset value to output a feature dataset. Each data point in the feature dataset corresponds to a feature vector and risk category label for a single vehicle.

[0063] Based on the above embodiments, the model training module 404 is specifically configured as follows: The feature dataset is stratified and sampled according to a preset ratio to divide the feature dataset into a training set and a test set. The training set is used to train the model, and the test set is used to verify the model performance. The training set is divided into multiple mutually exclusive subsets on an average basis according to the pre-set number of partitions in the K-fold cross-validation method. Select several mutually exclusive subsets as training subsets and use the remaining mutually exclusive subsets as validation subsets; Multiple full training sessions are performed based on the selected training subset each time, and the classification accuracy and F1 score of the validation subset are recorded in each training session. When both the target accuracy and the target score reach the corresponding thresholds, the category recognition model is determined. The target accuracy is the average of multiple classification accuracies, and the target score is the average of multiple F1 scores.

[0064] Based on the above embodiments, the pre-built recognition model includes a CNN feature extraction module and an XGBoost classification prediction module. The CNN feature extraction module is used to extract features from the feature vectors in the feature dataset through at least three convolutional layers, perform data standardization through batch normalization layers, introduce nonlinear features through the ReLU activation function, and compress features through max pooling layers to output high-dimensional feature vectors. The XGBoost classification prediction module is used to predict the high-dimensional feature vectors through multiple decision trees to determine the probability values ​​of the corresponding risk categories, and take the target risk category with the highest probability value as the risk category of the high-dimensional feature vector.

[0065] It is worth noting that in the embodiments of the above-mentioned device, the modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each module are only for easy differentiation and are not used to limit the protection scope of the embodiments of this application.

[0066] Figure 5This is a schematic diagram of an electronic device provided in an embodiment of this application. The device is used to execute the data processing method for risk category analysis provided in the above embodiment, and has corresponding functional modules and beneficial effects for executing the method. As shown in the figure, the device includes a processor 501, a memory 502, an input device 503, and an output device 504. The number of processors 501 can be one or more; one processor 501 is shown as an example in the figure. The processor 501, memory 502, input device 503, and output device 504 can be connected via a bus or other means; a bus connection is shown as an example in the figure. The memory 502, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the data processing method for risk category analysis in the embodiments of this application. The processor 501 executes various corresponding functional applications and data processing by running the software programs, instructions, and modules stored in the memory 502, thereby implementing the above-mentioned data processing method for risk category analysis.

[0067] The memory 502 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data recorded or created during use. Furthermore, the memory 502 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 502 may further include memory remotely configured relative to the processor 501, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof. The input device 503 can be used to input corresponding numerical or character information to the processor 501 and to generate key signal inputs related to user settings and function control of the device; the output device 504 can be used to send or display key signal outputs related to user settings and function control of the device.

[0068] This application also provides a storage medium storing computer-executable instructions, which, when executed by a processor, are used to perform relevant operations in the data processing method for risk category analysis provided in any embodiment of this application.

[0069] Computer-readable storage media include both permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0070] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0071] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.

Claims

1. A data processing method for risk category analysis, characterized in that, include: Outlier detection is performed on the acquired vehicle data corresponding to multiple target vehicles to determine the sample data that can be used as valid samples. Based on preset risk feature types, feature extraction is performed on the sample data and feature items corresponding to each risk feature type are determined; Based on the data type corresponding to the feature item, the feature item is digitized to generate a feature dataset; Based on the feature dataset, the pre-built recognition model is trained using the K-fold cross-validation method to obtain a category recognition model, which is used to identify the corresponding risk category of the input vehicle data. The category recognition model is deployed and used to perform risk category analysis on the input vehicle data to determine the risk category of the vehicle.

2. The data processing method for risk category analysis according to claim 1, characterized in that, The vehicle data includes basic vehicle information, vehicle detection information, accident and violation information, and preset risk category information. The basic vehicle information is used to record the basic attribute parameters of the vehicle. The vehicle detection information is used to record the detection time and location when the vehicle passes through the road monitoring checkpoint. The accident and violation information is used to record the number of violations and accidents of the vehicle in each time period. The preset risk category information is used to record the risk categories pre-configured for the vehicle. The step of performing outlier detection on multiple vehicle data sets corresponding to different target vehicles to determine the sample data as valid samples includes: Remove vehicle data from multiple vehicle data sets where the vehicle's basic information or the preset risk category information is null. Based on the detection location and detection time in the vehicle detection information, the vehicle speed value is determined and vehicle data with speed values ​​exceeding a preset speed are removed. Determine the box plot of the number of violations and accidents recorded in the accident violation information in all vehicle data, and remove the vehicle data corresponding to abnormal data points located outside the box plot; After filtering out vehicles based on their basic information, vehicle detection information, accident and violation information, and preset risk category information, the remaining vehicle data is used as the sample data.

3. The data processing method for risk category analysis according to claim 2, characterized in that, The risk characteristic types include a first type associated with the vehicle's basic information, a second type associated with the vehicle detection information, and a third type associated with the accident and violation information. The vehicle's basic information includes vehicle usage type, vehicle age, number of seats, and total weight. The step of extracting features from the sample data based on preset risk feature types and determining feature items corresponding to each risk feature type includes: Based on the vehicle usage type and vehicle age recorded in the sample data, multiple different target combinations and the proportion of the target combination in all vehicle data are determined, and the proportion is used as a feature item of the first type. The target combination is a parameter combination of the target vehicle usage type and the preset vehicle age. Different target combinations are associated with different risk categories. The ratio of the number of seats in the vehicle's basic information to the load capacity and total weight is used as a feature item of the corresponding load factor in the first type. Based on the vehicle detection information in the sample data, the driving trajectory information is restored and the corresponding travel features, range concentration, and abnormal travel patterns in the second type are determined. The travel features are features associated with the travel range corresponding to the driving trajectory information. The range concentration is used to represent the number of target areas covered by the travel range. The abnormal travel patterns are patterns in which the driving trajectory information has abnormal trajectories. Based on the accident and violation information in the sample data, the accident rate of the corresponding number of accidents and total mileage, the month-on-month growth rate of the corresponding number of violations, and the violation type are determined as the feature items of the third type.

4. The data processing method for risk category analysis according to claim 3, characterized in that, The process of reconstructing the driving trajectory information based on the vehicle detection information in the sample data includes: Several vehicle detection information corresponding to the same vehicle are identified, and multiple shortest paths between different road monitoring checkpoints in the several vehicle detection information are determined based on the Dijkstra method to generate a candidate path set; Based on Bayesian theory algorithms, the probability of passage of all paths in the candidate path set is determined, and the target path with the highest probability of passage is selected as the driving trajectory in the driving trajectory information.

5. The data processing method for risk category analysis according to claim 1, characterized in that, The step of datafiing the feature terms according to their corresponding data types to generate a feature dataset includes: Identify the first feature term that is a categorical feature and the second feature term that is a numerical feature among all feature terms; The first feature term is numericalized based on label encoding or one-hot encoding and stored in the initial dataset; The second feature term is normalized according to the Z-score standardization method and stored in the initial dataset; The Pearson correlation coefficients among the features in the initial dataset are determined, and features whose absolute values ​​of the Pearson correlation coefficients are greater than a preset value are removed to output the feature dataset. Each data point in the feature dataset corresponds to a feature vector and risk category label for a single vehicle.

6. The data processing method for risk category analysis according to claim 1, characterized in that, The step of training a pre-built recognition model using K-fold cross-validation based on the feature dataset to obtain a category recognition model includes: The feature dataset is stratified and sampled according to a preset ratio to divide the feature dataset into a training set and a test set. The training set is used to train the model, and the test set is used to verify the model performance. The training set is divided into multiple mutually exclusive subsets on an average basis according to the preset number of partitions in the K-fold cross-validation method. Select several mutually exclusive subsets as training subsets and use the remaining mutually exclusive subsets as validation subsets; Multiple full training sessions are performed based on the selected training subset each time, and the classification accuracy and F1 score of the validation subset are recorded in each training session. When both the target accuracy and the target score reach the corresponding thresholds, the category recognition model is determined, where the target accuracy is the average of multiple classification accuracies and the target score is the average of multiple F1 scores.

7. The data processing method for risk category analysis according to any one of claims 1-6, characterized in that, The pre-built recognition model includes a CNN feature extraction module and an XGBoost classification prediction module. The CNN feature extraction module is used to extract features from the feature vectors in the feature dataset through at least three convolutional layers, perform data standardization through batch normalization layers, introduce non-linear features through the ReLU activation function, and compress features through max pooling layers to output a high-dimensional feature vector. The XGBoost classification prediction module is used to predict the high-dimensional feature vector through multiple decision trees to determine the probability values ​​of each risk category, and take the target risk category with the highest probability value as the risk category of the high-dimensional feature vector.

8. A data processing apparatus for risk category analysis, characterized in that, include: The data preprocessing module is configured to perform outlier detection on the acquired vehicle data corresponding to multiple target vehicles to determine the sample data that can be used as valid samples. The feature extraction module is configured to extract features from the sample data based on preset risk feature types and determine the feature items corresponding to each risk feature type; The feature datafication module is configured to perform feature datafication on the feature items according to the data type corresponding to the feature items to generate a feature dataset; The model training module is configured to train the pre-built recognition model using the K-fold cross-validation method based on the feature dataset to obtain a category recognition model, which is used to identify the corresponding risk category of the input vehicle data. The model output module is configured to deploy the category recognition model and perform risk category analysis on the input vehicle data through the category recognition model to determine the risk category of the vehicle.

9. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the data processing method for risk category analysis as described in any one of claims 1-7.

10. A storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a processor, are used to perform the data processing method for risk category analysis as described in any one of claims 1-7.