Electric vehicle transportation risk scene construction method and system based on multi-source data clustering

By integrating multi-source data and performing cluster analysis, the problem of lagging dynamic risk identification during electric vehicle transportation was solved, enabling the automatic construction and real-time updating of fine-grained risk scenarios, thereby improving the accuracy and efficiency of risk management.

CN121504160APending Publication Date: 2026-02-10RES INST OF HIGHWAY MINIST OF TRANSPORT
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511661180.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies are insufficient to fully reflect the dynamic and interactive risk status under the coupled effects of multiple factors during electric vehicle transportation, resulting in delayed risk identification and insufficient ability to discover new or complex risk patterns.

Method used

By collecting multi-source heterogeneous data, performing preprocessing, feature extraction, and cluster analysis, a risk scenario model for electric vehicle transportation is constructed, including the fusion and cluster analysis of vehicle operating status, weather and road conditions, real-time traffic flow, and risk event data.

Benefits of technology

It enables automatic and accurate identification of risks in electric vehicle transportation and fine-grained scenario construction, improving the accuracy and efficiency of risk management and supporting real-time dynamic updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121504160A_ABST
    Figure CN121504160A_ABST
Patent Text Reader

Abstract

The invention discloses an electric vehicle transportation risk scene construction method and system based on multi-source data clustering, and the method comprises the following steps: collecting a multi-source heterogeneous data source, and obtaining original data in the transportation process of an electric vehicle; preprocessing the original data to obtain a standardized data set; performing feature extraction on the standardized data set to form a feature vector set; performing clustering analysis on the feature vector set by using a clustering algorithm to obtain a plurality of data clusters; and constructing an electric vehicle transportation risk scene model based on the data clustering cluster. According to the invention, through multi-source data fusion and clustering analysis, a potential risk mode in electric vehicle transportation can be automatically identified, a fine-grained risk scene is constructed, and the accuracy and efficiency of risk management are improved. In the embodiment, data quality is ensured by data preprocessing, dynamic risk factors are captured by feature extraction, natural data grouping is realized by clustering analysis, and interpretable risk classification is provided by a risk scene model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electric vehicle transportation risk management, specifically to a method and system for constructing electric vehicle transportation risk scenarios based on multi-source data clustering. Background Technology

[0002] With the rapid popularization and commercial application of electric vehicle technology, its transportation scale in urban logistics, public transportation, and personal travel continues to expand. However, in the complex actual operating environment, the transportation process of electric vehicles faces multi-dimensional and dynamically changing risk factors, including the battery and power system status of the vehicle itself, constantly changing external traffic conditions, the impact of weather and road conditions, and various uncertainties such as driver behavior. These factors are intertwined, making transportation risks multi-sourced, heterogeneous, and highly contextualized.

[0003] Currently, in the field of electric vehicle transportation risk management, existing methods mostly rely on the analysis of a single data source or rule-based judgments based on fixed thresholds. For example, some solutions trigger alarms simply by monitoring whether battery voltage or temperature exceeds preset limits using onboard sensors; others mainly rely on static statistics from historical accident reports to summarize risk types. While these methods have some practicality, they struggle to comprehensively reflect the true risk state under the coupled effects of multiple factors during operation, especially lacking the ability to characterize dynamic and interactive risk scenarios. Due to the failure to effectively integrate vehicle operation data, environmental data, and traffic flow data, existing technologies often lag in identifying potential risks and lack the ability to discover new or complex risk patterns, thus limiting the accuracy and real-time nature of risk assessments.

[0004] Therefore, there is an urgent need in this field for a technical solution that can comprehensively utilize multi-source heterogeneous data and use intelligent analysis methods to automatically and accurately construct risk scenarios for electric vehicle transportation, so as to improve the foresight and systematic nature of risk identification. Summary of the Invention

[0005] To address the technical problems mentioned above, this invention provides a method for constructing electric vehicle transportation risk scenarios based on multi-source data clustering, comprising the following steps:

[0006] Collect data from multiple heterogeneous data sources to obtain raw data during the transportation of electric vehicles;

[0007] The original data is preprocessed to obtain a standardized dataset;

[0008] Feature extraction is performed on the standardized dataset to form a feature vector set;

[0009] Clustering algorithms are used to perform cluster analysis on the feature vector set to obtain several data clusters;

[0010] Based on the data clusters, a risk scenario model for electric vehicle transportation is constructed.

[0011] Preferably, the multi-source heterogeneous data includes: vehicle operating status data, weather and road condition data, real-time traffic flow data, and risk event data; wherein, the method for obtaining the raw data includes:

[0012] Based on the vehicle sensor system, acquire vehicle operating status data;

[0013] Obtain weather and road condition data from the environmental monitoring platform;

[0014] Obtain real-time traffic flow data from the traffic management database;

[0015] Obtain risk event data based on historical accident records.

[0016] Preferably, the preprocessing of the raw data includes:

[0017] The Z-score method is used to remove outliers, and linear interpolation is used to handle missing values.

[0018] A timestamp-based alignment method is used to merge multi-source data;

[0019] Min-max normalization is used to convert the data to a uniform scale.

[0020] Preferably, the method for forming the feature vector set includes:

[0021] The sliding window method is used to calculate the mean, standard deviation, and trend slope of the data within the window in order to extract dynamic features;

[0022] The mean was calculated using the arithmetic mean method, the variance was calculated using the second central moments, and the extreme value characteristics were determined using the comparison and ranking method.

[0023] Based on domain knowledge, specific features related to transportation risks, including battery charge / discharge rate, braking frequency, and number of rapid accelerations, are extracted.

[0024] Preferred methods for obtaining data clusters include:

[0025] Based on the feature vector set, the distance between each feature vector is calculated using the Euclidean distance metric to form a similarity matrix;

[0026] Based on the similarity matrix, a density-based clustering algorithm is applied to group data objects.

[0027] Based on the grouping results, a cluster label is assigned to each data object to generate a data cluster.

[0028] Preferably, methods for grouping data objects using clustering algorithms include:

[0029] Based on the similarity matrix, the k-means algorithm is used to initialize the cluster centers;

[0030] Based on the initialized cluster centers, the updated cluster centers are obtained through iterative optimization;

[0031] Based on the updated cluster centers, the cluster quality assessment results are obtained through silhouette coefficient analysis.

[0032] Based on the clustering quality assessment results, the final clustering grouping scheme is obtained.

[0033] Preferred methods for constructing risk scenario models for electric vehicle transportation include:

[0034] Based on the characteristic distribution of clusters, risk patterns are identified by analyzing the statistical distribution characteristics of each characteristic variable;

[0035] Based on the identified risk patterns, risk scenario categories are defined using a rule-based risk classification method.

[0036] Based on the defined risk scenario categories, risk scenario descriptions containing key feature values ​​are generated using scenario parameterization methods.

[0037] The present invention also provides a system for constructing risk scenarios for electric vehicle transportation based on multi-source data clustering. The system is used to implement the above method and includes: a collection module, a processing module, an extraction module, a clustering module, and a construction module.

[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0039] This invention, through multi-source data fusion and cluster analysis, can automatically identify potential risk patterns in electric vehicle transportation, construct fine-grained risk scenarios, and improve the accuracy and efficiency of risk management. In the embodiments, data preprocessing ensures data quality, feature extraction captures dynamic risk factors, cluster analysis achieves natural data grouping, and the risk scenario model provides interpretable risk classification. In practical applications, this method can be dynamically updated in conjunction with real-time data streams, providing continuous support for the safety of electric vehicle transportation. Attached Figure Description

[0040] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention. Detailed Implementation

[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0043] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0044] Example 1

[0045] like Figure 1 The diagram shown is a schematic representation of the method flow in this embodiment, and the steps include:

[0046] S1. Collect data from multiple heterogeneous data sources to obtain raw data during the transportation of electric vehicles.

[0047] First, multi-source heterogeneous data sources are collected to obtain raw data during the electric vehicle transportation process. This data includes vehicle operating status data, weather and road condition data, real-time traffic flow data, and risk event data. Specifically, vehicle operating status data is acquired through vehicle sensor systems, including parameters such as vehicle speed, battery voltage, current, and temperature; weather and road condition data is acquired through an environmental monitoring platform, including temperature, humidity, precipitation, and road surface slippage; real-time traffic flow data is obtained from a traffic management database, including traffic volume, average speed, and congestion index; and risk event data is extracted from historical accident records, including accident time, location, type, and severity. This data is stored in time-series format with timestamps to ensure temporal consistency.

[0048] S2. Preprocess the raw data to obtain a standardized dataset.

[0049] After obtaining the raw data, it needs to be preprocessed to eliminate noise, handle missing values, and standardize the data scale to form a standardized dataset. The preprocessing process includes several steps: First, outliers are removed using the Z-score method. The Z-score calculation formula is:

[0050]

[0051] Where x is the original data point, μ is the mean, and is the standard deviation; when |z|>3, the data point is considered an outlier and is removed.

[0052] Secondly, for missing values, a linear interpolation method is used to handle them, assuming that adjacent data points in the time series... and Then the missing point t k The value x at that location k The calculation is as follows:

[0053] .

[0054] Then, a timestamp-based alignment method is used to merge multi-source data, aligning data from different sources according to a unified timeline to ensure that data points are consistent in time.

[0055] Finally, min-max normalization is used to convert the data to a uniform scale. The normalization formula is:

[0056]

[0057] in, and These represent the minimum and maximum values ​​of the data, respectively. After normalization, all data values ​​fall within the range of [0,1]. Through preprocessing, the original data is transformed into a high-quality, standardized dataset, laying the foundation for subsequent analysis.

[0058] S3. Extract features from the standardized dataset to form a feature vector set.

[0059] Feature extraction aims to extract dynamic and static features related to transportation risks from time-series data. First, a sliding window method is used to calculate the statistical characteristics of the data within a window. The window size can be set according to the data sampling frequency, for example, a 5-minute window. For each window, the mean, standard deviation, and trend slope of the data are calculated. The trend slope is fitted to the relationship between the data points within the window and time using linear regression. The formula for calculating the slope m is as follows:

[0060]

[0061] in, and , where represents time and the average value of the data, respectively; n represents the total number of data points; and i represents the index of the data point.

[0062] Secondly, the variance is calculated using the second-order central moments, i.e.:

[0063] .

[0064] The algorithm employs a comparative ranking method to determine extreme value features, such as extracting the maximum and minimum values ​​within a window. Furthermore, based on domain knowledge, specific features related to transportation risks are extracted, including battery charge / discharge rate (calculated as the ratio of current to rated capacity), braking frequency (number of braking actions per unit time), and rapid acceleration frequency (number of times acceleration exceeds a threshold per unit time). These features collectively constitute a feature vector set, with each feature vector representing a data summary within a time window.

[0065] S4. Use clustering algorithms to perform cluster analysis on the feature vector set to obtain several data clusters;

[0066] Cluster analysis first calculates the distance between each feature vector using the Euclidean distance metric, forming a similarity matrix. The similarity matrix describes the similarity between feature vectors; a smaller distance indicates a higher similarity. Next, the k-means algorithm is used for clustering, specifically including:

[0067] Based on the similarity matrix, the k-means algorithm is used to initialize cluster centers, typically by randomly selecting k feature vectors as initial centers. Based on these initialized cluster centers, updated cluster centers are obtained through iterative optimization. During the iteration process, each data point is assigned to the nearest cluster center, and the center point of each cluster is recalculated, updated to the mean of all points within the cluster. This iteration continues until the change in center points is less than a threshold or the maximum number of iterations is reached. Based on the updated cluster centers, the cluster quality assessment result is obtained through silhouette coefficient analysis. The silhouette coefficient s(i) is calculated as follows:

[0068] .

[0069] Where a(i) is the average distance between point i and other points in the same cluster, and b(i) is the average distance between point i and the nearest neighbor cluster point.

[0070] The silhouette coefficient ranges from [-1, 1], and a larger value indicates a better clustering effect. Based on the clustering quality assessment results, the clustering parameters (such as the k value or neighborhood radius) are adjusted to obtain the final clustering grouping scheme, and a cluster label is assigned to each data object to generate data clusters.

[0071] S5. Construct a risk scenario model for electric vehicle transportation based on data clustering.

[0072] First, risk patterns are identified by analyzing the statistical distribution characteristics of each feature variable based on the characteristic distribution of clusters. For example, the mean, variance, and quantiles of feature variables in each cluster are calculated to identify high-risk clusters (such as clusters with high battery charge / discharge rates and high braking frequencies). Second, risk scenario categories are defined based on the identified risk patterns using a rule-based risk classification method. Rules can be based on domain expert knowledge or historical data. For example, a "high-risk scenario" is defined as a cluster with battery temperature exceeding 50°C and more than 5 rapid accelerations per minute; a "medium-risk scenario" is a cluster with high braking frequency but normal battery parameters; and a "low-risk scenario" is a cluster where all characteristics are within safe limits. Finally, risk scenario descriptions containing key feature values ​​are generated using scenario parameterization methods based on the defined risk scenario categories. For example, for each risk scenario, representative feature values ​​(such as average vehicle speed, maximum battery temperature, and average braking frequency) are extracted, and a scenario description document is generated for subsequent risk analysis and early warning.

[0073] Example 2

[0074] This embodiment also provides a system for constructing risk scenarios for electric vehicle transportation based on multi-source data clustering, including: a data acquisition module, a processing module, an extraction module, a clustering module, and a construction module.

[0075] The following will describe in detail, with reference to this embodiment, how the present invention solves the technical problems in practical work.

[0076] First, a data acquisition module is used to collect raw data from multiple heterogeneous data sources during the electric vehicle transportation process. This data includes vehicle operating status data, weather and road condition data, real-time traffic flow data, and risk event data. Specifically, vehicle operating status data is acquired through the vehicle's sensor system, including parameters such as vehicle speed, battery voltage, current, and temperature; weather and road condition data is acquired through an environmental monitoring platform, including temperature, humidity, precipitation, and road surface slippage; real-time traffic flow data is obtained from a traffic management database, including traffic volume, average speed, and congestion index; and risk event data is extracted from historical accident records, including accident time, location, type, and severity. This data is stored in time-series format with timestamps to ensure temporal consistency.

[0077] After obtaining the raw data, the processing module preprocesses it to eliminate noise, handle missing values, and standardize the data scale, forming a standardized dataset. The preprocessing process includes several steps: First, outliers are removed using the Z-score method. The Z-score calculation formula is:

[0078]

[0079] Where x is the original data point, μ is the mean, and is the standard deviation; when |z|>3, the data point is considered an outlier and is removed.

[0080] Secondly, for missing values, a linear interpolation method is used to handle them, assuming that adjacent data points in the time series... and Then the missing point t k The value x at that location k The calculation is as follows:

[0081] .

[0082] Then, a timestamp-based alignment method is used to merge multi-source data, aligning data from different sources according to a unified timeline to ensure that data points are consistent in time.

[0083] Finally, min-max normalization is used to convert the data to a uniform scale. The normalization formula is:

[0084]

[0085] in, and These represent the minimum and maximum values ​​of the data, respectively. After normalization, all data values ​​fall within the range of [0,1]. Through preprocessing, the original data is transformed into a high-quality, standardized dataset, laying the foundation for subsequent analysis.

[0086] The extraction module aims to extract dynamic and static features related to transportation risks from time-series data. First, a sliding window method is used to calculate the statistical characteristics of the data within a window. The window size can be set according to the data sampling frequency, for example, a 5-minute window. For each window, the mean, standard deviation, and trend slope of the data are calculated. The trend slope is fitted using linear regression to the relationship between the data points within the window and time. The formula for calculating the slope m is as follows:

[0087]

[0088] in, and , where represents time and the average value of the data, respectively; n represents the total number of data points; and i represents the index of the data point.

[0089] Secondly, the variance is calculated using the second-order central moments, i.e.:

[0090] .

[0091] The algorithm employs a comparative ranking method to determine extreme value features, such as extracting the maximum and minimum values ​​within a window. Furthermore, based on domain knowledge, specific features related to transportation risks are extracted, including battery charge / discharge rate (calculated as the ratio of current to rated capacity), braking frequency (number of braking actions per unit time), and rapid acceleration frequency (number of times acceleration exceeds a threshold per unit time). These features collectively constitute a feature vector set, with each feature vector representing a data summary within a time window.

[0092] The clustering module first calculates the distance between each feature vector using the Euclidean distance metric, forming a similarity matrix. The similarity matrix describes the similarity between feature vectors; a smaller distance indicates a higher similarity. Next, the k-means algorithm is used for clustering, specifically including:

[0093] Based on the similarity matrix, the k-means algorithm is used to initialize cluster centers, typically by randomly selecting k feature vectors as initial centers. Based on these initialized cluster centers, updated cluster centers are obtained through iterative optimization. During the iteration process, each data point is assigned to the nearest cluster center, and the center point of each cluster is recalculated, updated to the mean of all points within the cluster. This iteration continues until the change in center points is less than a threshold or the maximum number of iterations is reached. Based on the updated cluster centers, the cluster quality assessment result is obtained through silhouette coefficient analysis. The silhouette coefficient s(i) is calculated as follows:

[0094] .

[0095] Where a(i) is the average distance between point i and other points in the same cluster, and b(i) is the average distance between point i and the nearest neighbor cluster point.

[0096] The silhouette coefficient ranges from [-1, 1], and a larger value indicates a better clustering effect. Based on the clustering quality assessment results, the clustering parameters (such as the k value or neighborhood radius) are adjusted to obtain the final clustering grouping scheme, and a cluster label is assigned to each data object to generate data clusters.

[0097] Finally, the module builds a risk scenario model for electric vehicle transportation based on data clustering.

[0098] First, risk patterns are identified by analyzing the statistical distribution characteristics of each feature variable based on the characteristic distribution of clusters. For example, the mean, variance, and quantiles of feature variables in each cluster are calculated to identify high-risk clusters (such as clusters with high battery charge / discharge rates and high braking frequencies). Second, risk scenario categories are defined based on the identified risk patterns using a rule-based risk classification method. Rules can be based on domain expert knowledge or historical data. For example, a "high-risk scenario" is defined as a cluster with battery temperature exceeding 50°C and more than 5 rapid accelerations per minute; a "medium-risk scenario" is a cluster with high braking frequency but normal battery parameters; and a "low-risk scenario" is a cluster where all characteristics are within safe limits. Finally, risk scenario descriptions containing key feature values ​​are generated using scenario parameterization methods based on the defined risk scenario categories. For example, for each risk scenario, representative feature values ​​(such as average vehicle speed, maximum battery temperature, and average braking frequency) are extracted, and a scenario description document is generated for subsequent risk analysis and early warning.

[0099] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for constructing electric vehicle transportation risk scenarios based on multi-source data clustering, characterized in that, Includes the following steps: Collect data from multiple heterogeneous data sources to obtain raw data during the transportation of electric vehicles; The original data is preprocessed to obtain a standardized dataset; Feature extraction is performed on the standardized dataset to form a feature vector set; Clustering algorithms are used to perform cluster analysis on the feature vector set to obtain several data clusters; Based on the data clusters, a risk scenario model for electric vehicle transportation is constructed.

2. The method for constructing electric vehicle transportation risk scenarios based on multi-source data clustering according to claim 1, characterized in that, The multi-source heterogeneous data includes: vehicle operating status data, weather and road condition data, real-time traffic flow data, and risk event data; wherein, the methods for obtaining the raw data include: Based on the vehicle sensor system, acquire vehicle operating status data; Obtain weather and road condition data from the environmental monitoring platform; Obtain real-time traffic flow data from the traffic management database; Obtain risk event data based on historical accident records.

3. The method for constructing electric vehicle transportation risk scenarios based on multi-source data clustering according to claim 1, characterized in that, The preprocessing of the raw data includes: The Z-score method is used to remove outliers, and linear interpolation is used to handle missing values. A timestamp-based alignment method is used to merge multi-source data; Min-max normalization is used to convert the data to a uniform scale.

4. The method for constructing electric vehicle transportation risk scenarios based on multi-source data clustering according to claim 1, characterized in that, The methods for forming the feature vector set include: The sliding window method is used to calculate the mean, standard deviation, and trend slope of the data within the window in order to extract dynamic features; The mean was calculated using the arithmetic mean method, the variance was calculated using the second central moments, and the extreme value characteristics were determined using the comparison and ranking method. Based on domain knowledge, specific features related to transportation risks, including battery charge / discharge rate, braking frequency, and number of rapid accelerations, are extracted.

5. The method for constructing electric vehicle transportation risk scenarios based on multi-source data clustering according to claim 1, characterized in that, Methods for obtaining data clusters include: Based on the feature vector set, the distance between each feature vector is calculated using the Euclidean distance metric to form a similarity matrix; Based on the similarity matrix, a density-based clustering algorithm is applied to group data objects. Based on the grouping results, a cluster label is assigned to each data object to generate data clusters.

6. The method for constructing electric vehicle transportation risk scenarios based on multi-source data clustering according to claim 5, characterized in that, Methods for grouping data objects using clustering algorithms include: Based on the similarity matrix, the k-means algorithm is used to initialize the cluster centers; Based on the initialized cluster centers, the updated cluster centers are obtained through iterative optimization; Based on the updated cluster centers, the cluster quality assessment results are obtained through silhouette coefficient analysis. Based on the clustering quality assessment results, the final clustering grouping scheme is obtained.

7. The method for constructing electric vehicle transportation risk scenarios based on multi-source data clustering according to claim 1, characterized in that, Methods for constructing risk scenario models for electric vehicle transportation include: Based on the characteristic distribution of clusters, risk patterns are identified by analyzing the statistical distribution characteristics of each characteristic variable; Based on the identified risk patterns, risk scenario categories are defined using a rule-based risk classification method. Based on the defined risk scenario categories, risk scenario descriptions containing key feature values ​​are generated using scenario parameterization methods.

8. A system for constructing risk scenarios for electric vehicle transportation based on multi-source data clustering, the system being used to implement the method described in any one of claims 1-7, characterized in that, include: The module consists of a data acquisition module, a processing module, an extraction module, a clustering module, and a construction module.

Citation Information

Patent Citations

  • Commercial vehicle driving behavior risk level identification method based on Internet of Vehicles data

    CN113657432A

  • Vehicle condition data exception processing system based on clustering analysis

    CN120544294A

  • Multi-dimensional data fusion system for operating truck risk rating

    CN120822193A

  • System and method of vehicle risk analysis

    US20230055238A1