Intelligent connected vehicle anomaly detection system and method based on big data analysis
By acquiring and analyzing metadata in intelligent connected vehicles, building and detecting abnormal traffic, combining fingerprint feature analysis and clustering technology of 5G network slices, the complexity of abnormal detection in heterogeneous data is solved, efficient collaborative detection and feature consistency are achieved, and the accuracy and reliability of abnormal traffic detection are improved.
Patent Information
- Application Number
- CN202510409847.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-02
AI Technical Summary
In the abnormal detection of intelligent connected vehicles, how to establish a unified abnormal traffic baseline in heterogeneous data, especially in the face of the complexity brought about by the diversity and dynamics of 5G network slices, to achieve efficient collaborative detection and feature consistency.
By obtaining the metadata of the vehicle CAN bus communication with the cloud, analyzing the load characteristics, building a load characteristic matrix and normalizing it. The Marshall distance algorithm is used to detect abnormal traffic, and the transmission link fingerprint features are extracted from 5G network slices, and feature drift is identified through K-mean clustering analysis. Then, the weighted fusion algorithm is used to correct the features, and the abnormal flow detection results are correlated with the feature drift to generate the final abnormality detection report.
Effectively respond to security threats in complex Internet of Vehicles environments, improve the accuracy and reliability of abnormal flow detection, and ensure efficient coordinated detection and fusion characteristics in a multi-slicing environment.
Smart Images

Figure CN119946640B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to an intelligent connected vehicle anomaly detection system and method based on big data analysis. Background Art
[0002] In the anomaly detection of intelligent connected vehicles, the metadata of the vehicle CAN bus and cloud communication constitutes the basis of a multi-level security situation awareness network. Due to the huge amount of data generated by vehicles during driving, traditional centralized data processing methods are difficult to cope with real-time requirements. Therefore, it is inevitable to use a distributed streaming computing engine to process massive audit logs. However, the load distribution characteristics of packets of different protocol types are different. How to establish a unified abnormal traffic baseline in these heterogeneous data has become a technical difficulty. Although the abnormal traffic baseline construction method based on Mahalanobis distance can effectively identify anomalies, in practical applications, due to the diversity and complexity of vehicle communication protocols, the accuracy and stability of the baseline model are challenged.
[0003] Furthermore, the transmission link fingerprint features in 5G network slicing provide a new means for real-time determination of distributed denial of service attacks or forged service request pulse attacks.
[0004] However, the introduction of 5G network slicing technology also brings new problems: How to achieve efficient collaborative detection between different network slices? Since the transmission link fingerprint characteristics of each network slice are different, how to ensure the consistency of these characteristics in cross-slice scenarios and how to efficiently integrate these characteristics in a multi-slice environment are key technical issues that need to be solved.
[0005] In addition, due to the dynamic nature of 5G network slicing technology, the transmission link fingerprint characteristics may drift as the network slices change, which further increases the complexity of anomaly detection. Summary of the invention
[0006] The purpose of the present invention is to overcome the above-mentioned shortcomings. The present invention provides an intelligent connected vehicle anomaly detection system and method based on big data analysis, which can effectively deal with security threats in complex vehicle networking environments and improve the accuracy and reliability of abnormal traffic detection.
[0007] The first aspect of the present invention provides a method for detecting abnormalities of intelligent connected vehicles based on big data analysis, which specifically comprises the following steps:
[0008] Obtain metadata from the communication between the vehicle CAN bus and the cloud, including protocol type and load characteristics; parse the load characteristics according to the protocol type and extract the key fields of the data packet;
[0009] Construct a load feature matrix based on key fields, normalize heterogeneous data, and obtain a feature data set in a unified format;
[0010] The Mahalanobis distance algorithm is used to calculate the distance between each data point in the feature data set and the preset baseline. If the distance exceeds the preset threshold, it is determined to be abnormal traffic and an abnormal traffic mark list is generated;
[0011] Obtain the fingerprint features of the transmission link from the 5G network slices, standardize the fingerprint features of different slices, and generate a consistent fingerprint feature dataset across slices;
[0012] Analyze the fingerprint feature data set. If the fingerprint features of the same category show significant differences in different slices, it is determined that the feature drift is caused by the dynamic nature of the network slice, and a feature drift marker list is generated.
[0013] According to the feature drift marker list, the fingerprint features of different slices are weightedly fused. If the Euclidean distance between the fused feature and the preset standard feature is less than the preset threshold, the collaborative detection result is generated;
[0014] The collaborative detection results are correlated with the abnormal traffic mark list for analysis. If the same vehicle has abnormal traffic and feature drift in different slices at the same time, a final anomaly detection report containing the vehicle ID, timestamp, anomaly type and detection results is generated.
[0015] The second aspect of the present invention further provides an intelligent connected vehicle abnormality detection system based on big data analysis, comprising:
[0016] The first data acquisition module acquires metadata in the communication between the vehicle CAN bus and the cloud, the metadata including the protocol type and the load characteristics; according to the protocol type, the load characteristics are parsed to extract the key fields of the data packet;
[0017] The data processing module builds a load feature matrix based on key fields, normalizes heterogeneous data, and obtains a feature data set in a unified format;
[0018] The traffic mark construction module uses the Mahalanobis distance algorithm to calculate the distance between each data point in the feature data set and the preset baseline. If the distance exceeds the preset threshold, it is determined to be abnormal traffic and a list of abnormal traffic marks is generated;
[0019] The second data acquisition module obtains the fingerprint features of the transmission link from the 5G network slice, standardizes the fingerprint features of different slices, and generates a fingerprint feature data set that is consistent across slices;
[0020] The drift mark construction module analyzes the fingerprint feature data set. If the fingerprint features of the same category show significant differences in different slices, it is determined that the feature drift is caused by the dynamic nature of the network slice, and a feature drift mark list is generated;
[0021] The feature drift fusion module performs weighted fusion on the fingerprint features of different slices according to the feature drift mark list. If the Euclidean distance between the fused feature and the preset standard feature is less than the preset threshold, a collaborative detection result is generated;
[0022] The collaborative detection module associates and analyzes the collaborative detection results with the abnormal traffic mark list. If the same vehicle has abnormal traffic and feature drift in different slices at the same time, a final anomaly detection report containing vehicle ID, timestamp, anomaly type and detection results is generated.
[0023] The present invention has the following beneficial effects:
[0024] The present invention targets security threats in the communication between the vehicle CAN bus and the cloud, obtains metadata and parses load characteristics, and constructs a feature matrix for normalization processing. The Mahalanobis distance algorithm is used to detect abnormal traffic, and the transmission link fingerprint features are extracted from the 5G network slice. K-means clustering is used to analyze the fingerprint features of different slices to identify the feature drift caused by the dynamic nature of the network slice. The features are corrected by a weighted fusion algorithm, and the abnormal traffic detection results are correlated with the feature drift, and finally an abnormal detection report containing the vehicle ID, timestamp, abnormal type and detection results is generated; thereby, it can effectively respond to security threats in complex Internet of Vehicles environments and improve the accuracy and reliability of abnormal traffic detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 The present invention is a flow chart of the method for detecting abnormalities of intelligent connected vehicles based on big data analysis. DETAILED DESCRIPTION
[0026] The following will describe the technical solutions in the embodiments of the present invention in detail in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention.
[0027] like Figure 1 As shown, in this embodiment, a method for detecting abnormalities of intelligent connected vehicles based on big data analysis may specifically include the following steps:
[0028] Step S101: Obtain metadata in the communication between the vehicle CAN bus and the cloud, where the metadata includes the protocol type and load characteristics. According to the protocol type, the load characteristics are parsed to extract key fields of the data packet.
[0029] Specifically, the protocol type of the metadata is identified through a regular expression parsing method; based on the identified protocol type, the load characteristics are analyzed, and the structured information of the data packet is extracted; through a clustering analysis method, the position and format of the key fields in the data packet are determined, for example, the position and format of the key fields are determined by bit offset and length, such as the ID field is located at bits 0-11, and the data segment is located at bits 12-63; based on the position of the key fields, the string interception technology is used to obtain the specific content of the key fields, for example, 8 bits are intercepted starting from the 15th bit to obtain the specific content of the key fields such as the vehicle speed value; the extracted key fields are matched with a preset metadata template to generate structured data, for example, the vehicle speed value is matched with the "SPEED" field in the template; if the field content matches the template, the structured data is stored in the local database; if it does not match, the abnormal field is marked and a log is recorded; the structured data is transmitted to the cloud storage system through the HTTP protocol interface.
[0030] Specifically, when using regular expressions to identify metadata protocol types, the data packet header features can be analyzed. For example, for sensor data packets, temperature sensor data may start with "TEMP:" and humidity sensor data may start with "HUM:". The data type can be quickly determined using the preset regular expression pattern.
[0031] When analyzing load characteristics, the data structures of different protocols are different. For example, temperature sensor data may use the format of "device number | timestamp | temperature value | status code", while the video data stream contains fields such as frame header, timestamp, pixel data, etc. The corresponding structure information needs to be extracted according to the protocol characteristics.
[0032] During cluster analysis, the sliding window method can be used to locate key fields. For example, in temperature data, numerical fields are usually floating point types with a fixed number of digits. Similar fields can be clustered according to data characteristics to determine the locations of fields such as temperature values and timestamps.
[0033] When obtaining the content of key fields, extract the fields at the determined positions. For example, in the temperature sensor data, if it is determined that the temperature value is located in the third field and is five digits long, the string at the corresponding position can be directly intercepted. String interception can avoid complex parsing processes and improve processing efficiency. In the metadata template matching stage, the format specifications of various sensor data need to be defined in advance. For example, the temperature value range should be between minus fifty degrees and one hundred and fifty degrees. If the extracted temperature value exceeds this range, it is marked as abnormal. Template matching can ensure the legitimacy and integrity of the data. When storing structured data, normal data can be classified and stored according to dimensions such as device type and time for subsequent query and analysis. Abnormal data needs to be marked with reasons, such as "value out of bounds" and "format error", etc., to facilitate subsequent processing.
[0034] During the data transmission phase, HTTP protocol requests can be used to send structured data in batches to the cloud. Data can be packaged by minute or hour to ensure real-time performance and avoid frequent requests. After receiving the data, the cloud stores it persistently and provides query services through the interface. Taking the temperature sensor as an example, the data packet may be in the format of "TS001|202401011230|25.6|0". It is identified as temperature data through regular expressions, and the fields such as device number, time, and temperature value are parsed to verify that the temperature value is within a reasonable range. Finally, the structured data is stored in the cloud database.
[0035] This processing method not only ensures the accuracy of the data, but also provides an exception handling mechanism, laying the foundation for subsequent data analysis.
[0036] In this embodiment, according to the identified protocol type, the load characteristics are analyzed and the structured information of the data packet is extracted, which may specifically include the following steps:
[0037] According to the protocol type, a corresponding decoding strategy is selected to decode the load characteristics to obtain a protocol identifier and a data segment; the protocol header field is extracted according to the protocol identifier, and the key field in the data segment is identified in combination with the protocol type; the key field is matched with the preset metadata mapping rule to generate a structured metadata table; the storage address of the key field is obtained according to the structured metadata table, and the storage address is mapped with the protocol type; a protocol load parsing template is generated based on the mapping relationship, and the parsing template contains a mapping relationship between the protocol type and the field name; the protocol load is parsed for the CAN bus traffic according to the parsing template, the key field content is extracted, and the parsing result is matched and verified with the preset metadata feature library; the metadata attributes of the fields in the protocol load parsing template are updated according to the matching result.
[0038] Specifically, when selecting a decoding strategy for a protocol type, you can distinguish between different types of network layers and application layers. For example, sensor data uses a lightweight protocol, while audio and video streams use a streaming protocol. A typical scenario is that temperature sensor data is transmitted in the form of a string, the protocol identifier is "TEMP", and the data segment contains a specific temperature value; while the humidity sensor uses "HUM" as an identifier. The protocol type can be quickly identified through preset rules.
[0039] The key to extracting protocol header fields is to understand the structural characteristics of different protocols. Taking the temperature sensor as an example, the protocol header contains two main fields: device number and timestamp, and the data segment contains temperature value and status code. The data segment position can be located by the protocol type. For example, the temperature data structure is "device number | timestamp | temperature value | status", where the temperature value is located in the third field. When generating a structured metadata table, it is necessary to establish a mapping relationship between the field and the physical meaning. Taking sensor data as an example, the device number is mapped to the specific location information, and the temperature value is mapped to the actual measurement range. This mapping can help with subsequent data analysis and anomaly detection, such as timely detection of temperature values that exceed the normal range.
[0040] In the process of storage address mapping, a hierarchical storage strategy can be adopted. Sensor data is stored in chronological order, and the address format is "device type / year / month / hour / minute / second", which is convenient for retrieval by time dimension. Data of different protocol types are stored in independent areas to avoid mutual interference. The generation of parsing templates needs to consider protocol characteristics. The temperature sensor template contains field definitions such as "device number: string, timestamp: integer, temperature value: floating point number, status: integer". Through the template, each field in the data packet can be quickly located and parsed. When verifying the parsing results, it is necessary to combine specific business scenarios. Temperature data is usually in the range of minus fifty degrees to one hundred and fifty degrees. If it exceeds the range, an abnormality is recorded. The status code reflects the working status of the sensor. Zero means normal and non-zero means abnormal. It can be used for device monitoring. Metadata attribute updates need to be based on actual data characteristics. If it is found that temperature data often has negative values, the temperature range definition can be adjusted; if a certain type of status code appears frequently, the corresponding processing logic can be added. This dynamic update mechanism makes the system more adaptable and can better handle various types of data.
[0041] In this embodiment, the method of identifying the protocol type of metadata by using a regular expression parsing method may specifically include the following steps:
[0042] Acquire the communication data stream and intercept the protocol header information; extract the protocol feature fields and generate a feature vector set; use the regular expression matching algorithm to compare the preset protocol feature library; determine the protocol category based on the matching results; if the match is successful, output the protocol type; if the match fails, use the machine learning algorithm to classify again; train the protocol feature data set through the support vector machine model and establish a protocol classifier; use the classifier to judge the category of unidentified protocols and obtain the final protocol type.
[0043] Specifically, during the communication data collection phase, a cache queue can be used to temporarily store data streams to ensure that continuous data packets are completely intercepted. For sensor scenarios, the interception of data packets needs to consider the header length and identifier. For example, the protocol header of a temperature sensor usually contains a device identifier and a timestamp, while a humidity sensor contains additional information such as air quality. During the feature field extraction process, attention should be paid to the structural characteristics of the data packet. For example, the device number in the temperature data packet is encoded in eight digits, the timestamp is sixteen digits, and the temperature value is in floating point format. By analyzing these features, a feature vector can be constructed, which includes dimensions such as the protocol header length, identifier type, and data segment format.
[0044] When matching regular expressions, matching rules can be customized for different protocols. The matching mode of temperature sensor data includes device number format verification and value range verification, such as the temperature value must be between -50 degrees and 150 degrees. Humidity data must match the percentage format, and the value range is between 0 and 100. The protocol classification process adopts a multi-level judgment strategy. First, quick classification is performed through preset rules. If the temperature data contains a specific identifier, it can be directly determined as a temperature protocol.
[0045] For complex situations, machine learning methods are used for in-depth analysis. The training data set of the support vector machine model contains typical features of various protocols, such as header structure, data format, etc. When new sensors are connected, the protocol library can be dynamically expanded. For example, when adding an air quality sensor, its data features are first analyzed, including the particle concentration format, air pressure value range, etc. These features are added to the training set and the classifier model is updated. This method enables the system to have protocol expansion capabilities and can adapt to different sensor types. In practical applications, protocol classification also needs to consider data integrity. Some sensors may experience data loss or damage, and fault tolerance is required at this time. The integrity of the data packet can be judged by checksum, and incomplete data packets can be reorganized or discarded. At the same time, an update mechanism for the protocol feature library is established to dynamically adjust the matching rules and classification models according to the actual data features to improve the system recognition accuracy.
[0046] This multi-level protocol recognition scheme can effectively process all types of sensor data. The system first attempts to quickly match, and then uses machine learning methods if it fails, which ensures both processing efficiency and recognition accuracy. At the same time, through the dynamic update mechanism, the system can adapt to new sensor types and has good scalability.
[0047] Step S102: construct a load feature matrix according to key fields, normalize the heterogeneous data, and obtain a feature data set in a unified format.
[0048] Specifically, the key fields are received to generate an initial feature matrix, the initial feature matrix is normalized to obtain an intermediate feature matrix; feature vectors in heterogeneous data are extracted according to the intermediate feature matrix, the feature vectors are normalized to obtain a standard feature matrix; a feature weight matrix is calculated according to the standard feature matrix, and a final feature matrix is obtained by performing a dot multiplication of the feature weight matrix and the standard feature matrix; and a feature data set containing a unified format is output according to the final feature matrix.
[0049] Specifically, the generation of the initial feature matrix requires the fusion of multi-source data, such as vehicle speed, engine speed, throttle opening and other sensor data. These data have different dimensions and value ranges, so they need to be normalized. Taking vehicle driving data as an example, the vehicle speed ranges from 0 to 200 kilometers per hour, the engine speed ranges from 800 to 6,000 revolutions per minute, and the throttle opening ranges from 0 to 100 percent. Through maximum and minimum value normalization, the data can be uniformly mapped to the range of zero to one to form an intermediate feature matrix.
[0050] The extraction of feature vectors needs to consider the correlation between data, such as the positive correlation between engine speed and throttle opening, and the negative correlation between vehicle speed and braking depth. By analyzing these correlations, feature vectors that can characterize the vehicle state can be extracted. For example, in abnormal steering detection, data such as steering wheel angle, yaw rate, and lateral acceleration form feature vectors, which are normalized to form a standard feature matrix.
[0051] The calculation of the feature weight matrix needs to consider the contribution of different features to anomaly detection. In the detection of brake system anomalies, the importance of features such as brake pressure, wheel speed, and brake pedal position is different. The weights can be determined based on expert experience or historical data analysis, such as brake pressure weight 0.4, wheel speed weight 0.3, and pedal position weight 0.3. After the feature weight matrix is multiplied by the standard feature matrix, the final feature matrix reflecting the importance of each feature is obtained.
[0052] The output of feature data sets in a unified format needs to consider the time series and relevance of the data. Taking engine fault warning as an example, feature data such as intake pressure, fuel pressure, and coolant temperature are sampled at fixed time intervals to form a time series data set. Each record contains fields such as timestamp, feature value, and abnormal flag. This format facilitates the training and verification of subsequent machine learning models.
[0053] In the detection of power system anomalies, parameters such as engine speed, torque, and power can be used as features, and dimensional differences can be eliminated through normalization. When allocating weights, the engine speed weight is 0.35, the torque weight is 0.35, and the power weight is 0.3, forming a feature weight matrix. The final feature matrix reflects the comprehensive impact of each parameter on the state of the power system, which helps to detect power output anomalies in a timely manner. In the detection of transmission anomalies, it is necessary to combine features such as gear position, input shaft speed, and output shaft speed. Through normalization and weighted calculation, faults such as gear shifting anomalies and slipping can be identified. The final output feature data set contains information such as fault type and severity, providing a basis for maintenance decisions.
[0054] This feature extraction method not only ensures the standardization of data, but also improves the accuracy of anomaly detection.
[0055] Step S103: Use the Mahalanobis distance algorithm to calculate the distance between each data point in the feature data set and the preset baseline. If the distance exceeds a preset threshold, it is determined to be abnormal traffic and an abnormal traffic mark list is generated.
[0056] Specifically, all data points in the feature data set are obtained, and for each data point, the Mahalanobis distance algorithm is used to calculate the distance value between the data point and the preset baseline, which is determined by the mean of the historical normal traffic data; the calculated distance value is compared with the preset threshold, which is determined by the standard deviation of the historical normal traffic data. If the distance value is greater than the preset threshold, the traffic corresponding to the data point is determined to be abnormal traffic; all data points determined as abnormal traffic are summarized to generate a preliminary abnormal traffic list; the isolation forest algorithm is used to perform secondary verification on each data point in the preliminary abnormal traffic list, and this secondary verification mechanism can effectively reduce the false alarm rate. The training set of the isolation forest algorithm is trained with historical normal traffic data, specifically including constructing a training sample set with historical normal traffic data, and using the training sample set to train the isolation forest algorithm to ensure the accuracy of the secondary verification, and updating the preliminary abnormal traffic list according to the secondary verification results to obtain the final abnormal traffic list; the density-based clustering algorithm is used to group and analyze the data points in the final abnormal traffic list, and the parameters of the clustering algorithm are determined by optimizing the silhouette coefficient to obtain abnormal traffic features of different categories, and the features of each category are represented by the mean of the data points of the category.
[0057] Specifically, the Mahalanobis distance algorithm identifies anomalies by calculating the difference between data points and the baseline. This method takes into account the correlation between features. Taking the engine monitoring of smart connected vehicles as an example, there is an inherent correlation between engine temperature and speed. When the engine speed increases, the temperature usually increases. From the historical normal operation data, it can be seen that the mean values of engine temperature and speed are 90 degrees and 3,000 revolutions per minute, respectively. This set of data can be used as a preset baseline. If the engine temperature is detected to be 120 degrees and the speed is only 2,000 revolutions per minute at a certain moment, this data that does not conform to the conventional correlation will produce a larger Mahalanobis distance value. The preset threshold is determined based on the standard deviation of the historical data and can be adjusted according to the actual application scenario.
[0058] In vehicle braking system monitoring, the standard deviations of brake pedal pressure and braking distance are 0.5 MPa and 1.5 meters respectively, and the corresponding abnormality determination threshold is set accordingly. When it is detected that the brake pedal pressure is normal but the braking distance is significantly extended, the system will mark it as a potential abnormality.
[0059] The Isolation Forest algorithm evaluates the degree of abnormality of data points through data segmentation. For example, in a vehicle battery management system, the voltage, current, temperature and other parameters during historical normal operation are used as training sets. When the battery temperature is normal but the voltage fluctuates abnormally, the data point is quickly separated, indicating that it is an abnormal point. This secondary verification mechanism can effectively reduce the false alarm rate.
[0060] Density-based clustering analysis can identify different types of abnormal patterns. Taking the vehicle communication system as an example, when network anomalies occur, they may be characterized by high packet loss rate and increased latency. Through density clustering, these anomalies can be divided into different categories such as communication interruption and network congestion. The silhouette coefficient is used to evaluate the clustering effect, and the optimal classification result is obtained when the cluster radius is set to 0.3.
[0061] In actual applications, an abnormality occurred during the operation of a certain intelligent connected vehicle. The system first detected multiple suspicious data points through Mahalanobis distance, such as abnormal increase in engine temperature but normal load. After verification by the isolation forest algorithm, it was confirmed that these points did reflect abnormal conditions. Density clustering analysis showed that these abnormal points were mainly concentrated in the cooling system fault category, characterized by continuous increase in coolant temperature and decreased circulation efficiency. This multi-level anomaly detection method can accurately identify the type of fault and provide a decision-making basis for timely maintenance. Through this systematic anomaly detection process, the safety of vehicle operation can be effectively improved and the risk of failure can be reduced.
[0062] The isolation forest algorithm identifies anomalies by randomly dividing the space, and anomalies are often more easily isolated. Taking network traffic monitoring as an example, a data center normally processes about one thousand requests per second. When there are five thousand abnormal burst requests, the average path length of the data point in the random tree is significantly lower than the normal point and is confirmed as abnormal traffic. The density clustering algorithm divides clusters based on the density of data points and can find abnormal clusters of irregular shapes. High-frequency, low-amplitude vibration is manifested as the number of vibrations per minute exceeding the normal value by twice, and the amplitude is within the normal range, which may indicate insufficient bearing lubrication. Low-frequency, high-amplitude vibration is manifested as normal vibration frequency, but the amplitude exceeds the threshold by 50%, which may be caused by bearing wear.
[0063] The silhouette coefficient is used to evaluate the clustering effect. The value range is from negative one to one. The closer to one, the better the clustering effect. Taking the abnormal detection of the air-conditioning system as an example, when the silhouette coefficient is 0.8, the abnormal data can be accurately divided into different categories such as refrigerant leakage and compressor failure. The abnormal characteristics of refrigerant leakage are manifested as a continuous decrease in cooling capacity, while compressor failure is manifested as a sudden increase in energy consumption. Through cluster analysis, the fault type can be quickly located and the formulation of maintenance plans can be guided.
[0064] In this embodiment, the following steps are also included: for the abnormal traffic mark list, a sliding window mechanism is used to dynamically update the baseline model. If the abnormal traffic ratio in the window exceeds a preset threshold, the Mahalanobis distance of the baseline model is recalculated to obtain an updated baseline model.
[0065] Specifically, a flow data set in a window is received, the flow data set includes flow feature vectors in a time window; a sliding window mechanism is adopted to obtain a flow data set of a current window according to a preset time interval; the number of abnormal flow marks in the flow data set is obtained from a flow detection node, and the abnormal flow ratio of the current window is determined according to the ratio of the number of abnormal flow marks to the total flow; an abnormal ratio threshold is preset, and the abnormal flow ratio of the current window is compared with the abnormal ratio threshold; if the abnormal flow ratio of the current window exceeds the abnormal ratio threshold, a baseline model update process is triggered; the flow feature vector set of the current window is filtered according to the abnormal flow mark to obtain a feature vector subset for updating; a covariance matrix of the flow feature vector subset is calculated, and the covariance matrix is used to describe the linear relationship between feature vectors; the Mahalanobis distance algorithm parameters of the flow feature vector subset are calculated based on the covariance matrix, and the Mahalanobis distance algorithm parameters are used to measure the distribution difference of feature vectors; the baseline model parameters are updated according to the Mahalanobis distance algorithm parameters to generate an updated baseline model parameter set; the updated baseline model parameter set is synchronized to the flow detection node, and the flow detection node is triggered to load the new baseline model parameters.
[0066] Specifically, when the Mahalanobis distance algorithm is used in traffic monitoring, it is necessary to continuously track the distribution changes of traffic data. In the field of traffic monitoring, a typical time window may be hourly or daily, such as collecting network traffic data every hour. Assume that the traffic data collected within a certain time window contains features such as packet size and transmission rate, which constitute a feature vector. When the number of abnormal traffic markers detected within a certain time window exceeds the set threshold, the baseline model needs to be updated.
[0067] To illustrate this process, assume that in a network monitoring, a time window is set every hour, and the threshold for the abnormal traffic ratio is set to 5%. In a certain time window, if 600 out of 10,000 packets are detected and marked as abnormal, then the abnormal ratio is 6%, which exceeds the preset threshold, and the update process of the baseline model will be triggered. When calculating the covariance matrix, the correlation between features needs to be considered. For example, in network traffic, there may be a linear relationship between packet size and transmission time. By calculating the covariance between these features, a matrix describing their mutual relationship can be obtained. This matrix can reflect the degree of association between different features and help to more accurately measure the distance between data points.
[0068] In practical applications, the calculation of the parameters of the Mahalanobis distance algorithm needs to take into account the distribution characteristics of the data. For example, within a certain time window, if the mean of the traffic feature is found to have changed significantly, or the correlation between the features has changed significantly, the calculation parameters of the Mahalanobis distance need to be adjusted accordingly. This adjustment can make anomaly detection more accurately adapt to the dynamic changes of network traffic. When the baseline model parameters are updated, these new parameters need to be synchronized to the traffic detection node in a timely manner. For example, when multiple detection nodes may work at the same time, the updated parameters need to ensure that all nodes can obtain and use the new detection standards in a timely manner. This can ensure the consistency of the detection standards of the entire system and avoid the situation where different nodes use different detection standards. When loading new baseline model parameters on the traffic detection node, smooth transition also needs to be considered. For example, a progressive update strategy can be adopted to use the new and old parameters for detection in a short period of time, and the effectiveness of the new parameters can be verified by comparing the differences in the detection results. This method can reduce the risks brought by parameter updates and ensure the stable operation of the system.
[0069] Step S104: Obtain the fingerprint features of the transmission link from the 5G network slice, standardize the fingerprint features of different slices, and generate a fingerprint feature data set that is consistent across slices.
[0070] Specifically, the original data of the transmission link is obtained from the 5G network slice, and the fast Fourier transform is used to extract fingerprint features for different slice types; the extracted fingerprint features are normalized according to the slice type, and if the eigenvalue exceeds the preset range, the logarithmic transformation is used for data compression to obtain standardized feature data; the standardized feature data is input into the principal component analysis method to determine the dimension with the largest variance in the feature vector, and obtain the feature subset after dimensionality reduction; the feature subset after dimensionality reduction is grouped by the K-means algorithm, and if the distance within the group is less than the preset threshold, it is merged into the same category to generate a feature data set with category labels; the feature data set is sorted according to the category label The rows are reorganized. If there is consistency across slices, the feature data is retained to obtain a fingerprint feature data set that is consistent across slices; the similarity of the fingerprint feature data set that is consistent across slices is calculated using the cosine similarity algorithm to determine the cosine similarity between the feature vectors, and the calculated cosine similarity values are filled into the preset matrix structure in row and column order to obtain a feature similarity matrix; the feature similarity matrix is input into a multilayer perceptron for training, and cross entropy is used as the loss function to calculate the difference between the model output and the true label. The weight parameters of the multilayer perceptron are adjusted according to the loss value through the back propagation algorithm to obtain a classification model that can identify different slice types.
[0071] Specifically, the original data collection of the transmission link needs to consider the characteristics of different slice types. For example, the in-vehicle entertainment system uses high-bandwidth slices, and its data contains the characteristics of streaming media transmission; the autonomous driving system uses low-latency slices, and the data reflects the properties of real-time control. Fourier transform of the collected data can extract frequency domain features and convert time domain signals into spectrum features. For example, in vehicle communication data, normal communication traffic presents stable periodicity in the frequency domain, while abnormal traffic is manifested as spectrum mutations. When normalizing, different normalization methods are used for different slice types. The normalization range of safety control data is set to zero to one, and data outside the range is compressed by logarithmic transformation. For example, the braking command data of a vehicle originally ranges from zero to one thousand milliseconds, which is compressed to the range of zero to one through logarithmic transformation to maintain the relative relationship of the data.
[0072] Principal component analysis is used to reduce feature dimensions and retain the most representative features. For example, vehicle status monitoring data contains multiple dimensions such as speed, acceleration, and direction. Principal component analysis can identify the most important feature combinations for anomaly detection, such as the coordinated change characteristics of speed and acceleration. Feature grouping uses a distance measurement method to group similar features together. In vehicle operation status monitoring, related parameters such as engine speed and throttle opening are close and can be classified into the power system feature group. When an abnormality is detected in a group of features, the faulty system can be quickly located.
[0073] Cross-slice consistency analysis ensures that the characteristics of different slice types are comparable. For example, when autonomous driving data is transmitted in different network slices, abnormal transmission status can be identified by analyzing the consistency of its latency characteristics. Under normal circumstances, the transmission latency of control instructions should remain stable, and sudden latency fluctuations indicate network abnormalities.
[0074] The feature similarity matrix reflects the degree of association between different features. In vehicle communication systems, there is a correlation between parameters such as network throughput and signal strength, and this correlation can be quantified by cosine similarity. When the similarity between certain features suddenly decreases, it may indicate a system failure. The multilayer perceptron model identifies anomalies by learning the patterns of the feature similarity matrix. Labeled normal and abnormal data are used during training, and the model gradually adjusts parameters to improve recognition accuracy. For example, if a smart connected car has a communication anomaly during operation, the model can quickly determine the type of anomaly based on the learned feature pattern and provide decision support for fault handling.
[0075] In this embodiment, similarity calculation is performed on the fingerprint feature data set that is consistent across slices using a cosine similarity algorithm, which may specifically include the following steps:
[0076] Obtain consistent feature vectors across slices from the fingerprint feature data set and construct a feature vector matrix; for the feature vector matrix, calculate the dot product value of each pair of feature vectors to obtain the dot product result matrix; calculate the modulus of each feature vector and generate a modulus vector for subsequent similarity calculation; based on the dot product result matrix and the modulus vector, use the cosine similarity formula to calculate the similarity value of each pair of feature vectors; if the similarity value is greater than a preset threshold, the pair of feature vectors is judged to be a high-similarity pair and marked as a candidate matching pair; perform cluster analysis on the marked candidate matching pairs to generate fingerprint feature clustering results that are consistent across slices; output the clustering results as the final data for cross-slice fingerprint feature similarity analysis.
[0077] Specifically, when obtaining consistent feature vectors across slices from the fingerprint feature data set, multiple dimensions of the vehicle's operating status need to be considered. For example, the vehicle's power system characteristics include parameters such as engine speed and throttle opening, which form feature vectors that reflect the vehicle's power performance status. Safety control system characteristics include values such as brake pressure and wheel speed, which constitute another set of feature vectors for monitoring driving safety.
[0078] When constructing the eigenvector matrix, these vectors are arranged in time sequence to form a dynamic monitoring data stream. When calculating the dot product of the eigenvector, the potential failure mode is revealed by analyzing the correlation between different system parameters. For example, if the dot product value of the engine speed and fuel consumption data is large, it indicates that the two have a strong correlation. When this correlation changes significantly, it may indicate abnormal engine operation. The modulus calculation of the eigenvector reflects the overall change amplitude of each parameter and provides a benchmark for subsequent similarity analysis. When the cosine similarity is used to calculate the similarity of the eigenvectors, the collaborative anomalies between different systems can be found. For example, when the steering system fails, the similarity between the steering wheel angle and the tire steering angle will be significantly reduced, while the similarity between the two is usually maintained above 0.9 during normal driving. The similarity threshold is set to 0.85, and the feature pairs above this value are marked as candidate matching pairs.
[0079] Cluster analysis of candidate matching pairs can identify systemic faults. For example, when a smart connected car is driving, the similarity of the feature vectors formed by multiple sensor data of the braking system decreases at the same time. Through cluster analysis, these anomalies can be classified as brake system faults. At the same time, by analyzing the control instruction data transmitted in different slices, if the transmission delay of multiple control parameters is found to increase suddenly, the clustering results will show network transmission anomalies. The final data of the cross-slice fingerprint feature similarity analysis can reflect the overall status of the system.
[0080] In actual applications, the autonomous driving function of a certain intelligent connected car relies on multiple network slices to transmit different types of data. By analyzing the similarity changes of these data features, potential risks can be discovered in a timely manner. For example, when steering control data is transmitted in a low-latency slice, if the similarity with historical normal data is observed to decrease, the system will immediately issue an early warning to ensure driving safety. This analysis method achieves all-round monitoring of vehicle systems by establishing feature similarity associations across slices.
[0081] In this embodiment, cross entropy is used as the loss function to calculate the difference between the model output and the true label. The weight parameters of the multilayer perceptron are adjusted through the back propagation algorithm according to the loss value to obtain a classification model that can identify different slice types. Specifically, the following steps may be included:
[0082] Obtain a feature similarity matrix, which contains feature data of different slice types; input the matrix into a multilayer perceptron, which contains at least one hidden layer, and the number of neurons in the hidden layer is determined according to the feature dimension; use cross entropy as the loss function to calculate the difference between the model output and the true label; adjust the weight parameters of the multilayer perceptron through the back propagation algorithm according to the loss value; if the loss value is less than the preset threshold, stop training, otherwise continue to iterate and optimize the model parameters; after training, obtain a classification model, which can identify the slice type according to the input features; use a test data set to evaluate the performance of the classification model, and calculate the classification accuracy and recall rate.
[0083] Specifically, the feature similarity matrix contains vehicle status information of multiple slice types. For example, high-priority slices record real-time control parameters such as steering wheel angle and vehicle speed, medium-priority slices save vehicle camera data, and low-priority slices store vehicle diagnostic information. When constructing the matrix, these features are organized by dimension to form input data in a unified format. The design of a multi-layer perceptron needs to consider the complexity of data features. Assuming that vehicle status monitoring involves fifty parameters, fifty neurons are set in the input layer. According to experience, two hidden layers are set, with one hundred neurons in the first layer and fifty in the second layer. Such a network structure can learn nonlinear relationships between features.
[0084] For example, the correspondence between engine speed and throttle opening, or the correlation between vehicle speed and brake pressure. The cross entropy loss function measures the difference between the model prediction result and the actual slice type. When the prediction is completely correct, the loss is zero, and when the prediction is wrong, the loss value is large. In this way, the model can learn the characteristic patterns of different slice data. For example, data of high-priority slices usually have smaller timestamp differences, while data of low-priority slices are updated less frequently. During the backpropagation process, the model weight parameters are continuously adjusted to reduce the loss value. Set the loss threshold to 0.01, and when the loss is less than this value, the model is considered to have converged.
[0085] In actual applications, the data acquisition system of a certain intelligent connected car generates about 10,000 records per hour. After 5,000 iterations of training, the model loss dropped to 0.008, achieving the expected effect. The trained classification model can accurately identify the slice type to which the data belongs. For example, when the status data of the vehicle steering system is input, the model will classify it as a high-priority slice, which is consistent with the actual needs of the vehicle control system. At the same time, for data with low non-real-time requirements such as vehicle positioning and road condition recognition, the model will assign it to the corresponding low-priority slice. In the testing phase, the model performance is evaluated using data that did not participate in the training. In a test, one thousand sets of vehicle status data were used, and the model correctly identified the slice type of 930 sets of data, with a classification accuracy of 93%. The recognition recall rate for high-priority slices reached 95%, which ensures that key control data can be processed in a timely manner.
[0086] In this embodiment, the raw data of the transmission link is obtained from the 5G network slice, and the fingerprint features are extracted by using fast Fourier transform for different slice types, which may specifically include the following steps:
[0087] The raw data is classified and stored according to different slice types; the link data in the raw data is preprocessed according to the slice type, the noise is removed and the data format is standardized to obtain standardized link data; the standardized link data is converted into the frequency domain using the fast Fourier transform, and the frequency domain feature vector is extracted as the preliminary fingerprint feature; the preliminary fingerprint features are clustered according to the preset feature classification rules to determine the feature distribution pattern of different slice types; if the feature distribution pattern meets the preset distinction conditions, the clustering result is used as the final fingerprint feature; if not, the feature classification rules are adjusted to re-cluster; the feature analysis module is used to reduce the dimension of the final fingerprint feature to generate a low-dimensional feature representation for subsequent matching and identification; the low-dimensional feature representation is compared with the pre-established fingerprint feature library through the feature processing module to complete the identification and verification of the slice type.
[0088] Specifically, the raw data is first classified and stored. For example, chassis control-related data is stored in high-priority slices, including key parameters such as braking torque and wheel speed; in-vehicle entertainment system data is stored in low-priority slices, such as audio playback status, screen brightness and other information. This classified storage method is helpful for subsequent data processing and analysis. The link data preprocessing stage mainly eliminates signal interference and outliers. The steering angle sensor data of a certain model has large fluctuations before filtering. After the abnormal peaks are removed by median filtering, the data curve is smoother. Standardization processing unifies data of different dimensions to the same scale, such as mapping the vehicle speed from zero to 200 kilometers per hour to the interval from zero to one. Fast Fourier transform can reveal the periodic characteristics of data. For example, after the engine speed data is transformed, a significant frequency peak appears near 20 Hz, which reflects the inherent vibration characteristics of the engine. These frequency domain features constitute the preliminary fingerprint feature vector. In the feature clustering analysis stage, preset rules are used to group fingerprint features, for example, features with a similarity of more than 85% are classified into one category.
[0089] For example, in one analysis, the engine-related sensor data formed an obvious cluster, while the body electronic equipment data formed another independent cluster. In feature distinguishability verification, if the feature overlap of different slice types is less than 10%, the clustering result is considered to meet the distinguishing condition. Otherwise, the classification rules need to be adjusted, such as increasing the similarity threshold or adding feature dimensions.
[0090] In actual applications, a certain brand of cars achieved an ideal clustering effect after three rounds of rule adjustments. The dimensionality reduction processing stage maps high-dimensional features to low-dimensional space to maintain the key information of the data. For example, the 50-dimensional sensor data is reduced to 10 dimensions through principal component analysis, and 95% of the variance information is still retained. This not only reduces the storage space, but also improves the efficiency of subsequent matching. The final feature matching stage compares the processed features with the pre-established feature library.
[0091] In this embodiment, the normalization process of the extracted fingerprint features according to the slice type may specifically include the following steps:
[0092] Obtain fingerprint feature data and determine the slice type; extract feature vectors based on the slice type; analyze feature distribution and determine feature range; transform feature values to achieve feature mapping; use normalization method to process feature dimensions; obtain normalized feature data; determine feature transformation results and complete the feature processing flow.
[0093] Specifically, multi-dimensional data is first obtained through the vehicle sensor network, including vehicle operating status, environmental perception, control instructions, etc. Slice types are divided according to the importance of the business. For example, safety-critical data such as steering system and braking system are classified into high-priority slices, while comfort function data such as air conditioning and audio are classified into low-priority slices.
[0094] The feature vector extraction stage focuses on the time domain and frequency domain characteristics of the data. Taking the engine vibration data as an example, by analyzing the amplitude, frequency, and phase of the vibration waveform, the characteristic fingerprint of the engine working state can be identified. When the engine is in normal working condition, the vibration frequency is usually in the range of 20 Hz to 50 Hz, while there will be obvious frequency deviation under abnormal conditions. The feature distribution analysis adopts probabilistic statistical methods. The steering system data of a certain brand of car shows that the distribution of steering angular velocity under normal conditions presents a typical Gaussian distribution, and 95% of the data falls within the range of plus or minus two standard deviations of the mean. This statistical characteristic can be used as an important basis for judging the abnormality of the steering system.
[0095] The feature transformation process needs to consider the physical meaning of the data. The linear interval of the wheel speed signal from zero to 200 kilometers per hour can be compressed by logarithmic transformation, making the small changes in the high-speed section easier to detect. The body posture data is suitable for Fourier transformation to extract periodic features such as vehicle bumps and rolls. Feature normalization ensures that data of different dimensions can be compared uniformly. For example, physical quantities such as brake pressure, wheel speed, and steering angle are mapped to the interval from zero to one. The original data of a new energy vehicle braking system shows that the brake pressure range is from zero to ten megapascals. After normalization by minimum and maximum values, it can be comprehensively analyzed with other normalized features. The determination of the transformation results needs to consider the physical meaning of the features and the actual application requirements. The chassis control system of a certain model uses twelve key features. After principal component analysis and dimensionality reduction, five main features are retained, which can still explain 95% of the data variance. This dimensionality reduction process not only retains the core information of the data, but also improves the real-time performance of anomaly detection. The integrity of the feature processing process directly affects the accuracy of anomaly detection.
[0096] In this embodiment, the step of grouping the feature subsets after dimension reduction by using the K-means algorithm may specifically include the following steps:
[0097] In the intelligent connected vehicle anomaly detection system, the number of clusters is determined from the preset K value range, and the cluster center of the K-means algorithm is initialized; the distance between each feature subset and the cluster center is iteratively calculated, and the feature subset is assigned to the nearest cluster center; the position of the cluster center is updated, and if the change of the cluster center is less than the preset threshold, the iteration is stopped; based on the final cluster center, the grouping result of each feature subset is determined; the clustering effect is evaluated using the silhouette coefficient, and if the silhouette coefficient is lower than the preset value, the K value is adjusted to re-group; the grouping result and the corresponding cluster center are output to complete the grouping process of the feature subset.
[0098] Specifically, in cluster analysis, determining the appropriate number of clusters is a key step. Taking fingerprint feature analysis as an example, when the preset number of clusters ranges from three to eight categories, it is necessary to make a preliminary judgment by observing the data distribution characteristics. For example, in the fingerprint texture direction feature, there may be three main types: spiral, bow and tent. At this time, you can first select three cluster centers for trial. The initialization of the cluster center has an important impact on the final result. Suppose there is a set of fingerprint feature data, which contains three hundred sample points, each of which contains features such as texture direction and ridge density. Three sample points can be randomly selected as the initial cluster centers, or three points that are far away from each other can be selected as the initial centers, which helps to avoid local optimal solutions. Calculating the distance between the feature subset and the cluster center is the core of the iterative process. For each sample point, it is necessary to calculate its Euclidean distance to the three cluster centers. For example, if the texture direction of a sample point is ninety degrees and the ridge density is ten per centimeter, and the distance values obtained after calculating the distance with the three cluster centers are 0.5, 1.2 and 0.8 respectively, then this sample point will be assigned to the first cluster center with the smallest distance.
[0099] The position update of the cluster center reflects the adaptability of the algorithm. When all sample points are assigned to the nearest cluster center, the center position of each category needs to be recalculated. If a category contains one hundred sample points, the new cluster center will be the average value of these one hundred points in each feature dimension. When the position change between the new and old cluster centers is less than 0.01, it can be considered that the algorithm has converged. The silhouette coefficient is an important indicator for evaluating the clustering effect. For each sample point, calculate its average distance from other samples of the same type and its average distance from the nearest samples of other classes. If the average distance between a sample point and samples of the same type is 0.3 and the average distance to the nearest other class is 0.8, the silhouette coefficient of the point is 0.625. If the overall silhouette coefficient is lower than 0.5, it means that the clustering effect is not ideal and the number of clusters needs to be adjusted and tried again. The final grouping result reflects the inherent structure of the data.
[0100] Through cluster analysis, it may be found that fingerprint features can be divided into several clear groups, each of which has similar characteristic attributes. This grouping result can be used for subsequent fingerprint matching and recognition to improve the accuracy and efficiency of the system. The cluster center represents the typical characteristics of each category and can be used as a standard template for that category.
[0101] Step S105: Use the K-means clustering algorithm to analyze the fingerprint feature data set. If the fingerprint features of the same category show significant differences in different slices, it is determined that the feature drift is caused by the dynamic nature of the network slice, and a feature drift marker list is generated.
[0102] Specifically, after obtaining the fingerprint feature data set, the K-means clustering algorithm is used to divide the data set to obtain a cluster category set; for each category in the cluster category set, its fingerprint feature value in different network slices is extracted, and the feature difference value is calculated; the feature difference value is obtained by calculating the Euclidean distance of the feature value of the same category in different network slices; according to the pre-set difference threshold, it is judged whether the feature difference value exceeds the difference threshold, and if it exceeds, a feature drift mark is generated; after obtaining the feature drift mark set, combined with the network slice dynamic data, through comparative analysis, it is determined whether the drift cause is the network slice dynamics; the principal component analysis method is used to reduce the dimension of the feature drift mark set, and the high-dimensional features are reduced to three dimensions to obtain the drift feature principal component set; according to the drift feature principal component set, the support vector machine algorithm is used to construct a drift classification model, and the radial basis function is set as the kernel function to obtain the drift classification result; combined with the drift classification result and the network slice dynamic data, a feature drift mark list containing the drift category and the corresponding network slice information is generated to complete the feature drift judgment.
[0103] Specifically, different types of network slices carry different business traffic. By clustering analysis of fingerprint feature data sets, data groups with similar features can be identified. For example, the data traffic of vehicle entertainment systems and navigation systems shows obvious differences in bandwidth requirements and latency requirements, and can be divided into different categories through clustering. For each clustering category, extract and compare its feature values in different slices. For example, when the autonomous driving control signal is transmitted in a high-reliability slice and a low-latency slice, the feature value should be relatively stable under normal circumstances. The feature difference is quantified by calculating the Euclidean distance. For example, if the transmission delay difference of the steering control data of a vehicle in two slices exceeds the preset threshold, the system will generate a feature drift marker.
[0104] Feature drift may be caused by the dynamic adjustment of network slices. For example, when a vehicle is driving at high speed, the network slice may reallocate resources due to load balancing, resulting in changes in the transmission characteristics of certain services. By analyzing the temporal relationship between slice adjustment logs and feature drift, it can be determined whether the drift is caused by the dynamic nature of the slice. For high-dimensional feature drift data, principal component analysis can be used for dimensionality reduction to highlight key features. For example, the drift characteristics of the vehicle communication system may include multiple dimensions such as delay, packet loss rate, and signal strength. Through dimensionality reduction, the three most representative principal components can be retained for subsequent analysis.
[0105] The drift classification model uses the support vector machine algorithm to achieve nonlinear mapping of the feature space through the radial basis kernel function. For example, vehicle communication anomalies are classified into categories such as network congestion, signal interference, and equipment failure. Labeled drift data is used for model training, such as a communication anomaly sample when a vehicle passes through a tunnel. The final generated feature drift tag list contains detailed drift information. For example, the autonomous driving control data of a vehicle is abnormal in a low-latency slice. After analysis, it is found that the temporary performance degradation is caused by the dynamic adjustment of slice resources. This information can be used for the optimization management of network slices, such as reserving sufficient bandwidth for key services when adjusting resources.
[0106] By analyzing the drift characteristics in different time periods and under different road conditions, the drift pattern can be summarized. For example, in traffic-intensive areas, due to the large number of vehicles competing for network resources, the characteristics of some non-critical services drift more frequently. These analysis results can guide the dynamic optimization of network slices and improve the service quality of the system.
[0107] Step S106: According to the feature drift mark list, a weighted average algorithm is used to perform weighted fusion on the fingerprint features of different slices. If the Euclidean distance between the fused feature and the preset standard feature is less than a preset threshold, a collaborative detection result is generated.
[0108] Specifically, the weight value of the slice fingerprint feature is obtained from the feature drift marker list, and the weight value is used as input to train the historical data using the PCA method in Python's Scikit-learn library to generate the preset standard feature; according to the obtained weight value, the fingerprint feature is weighted and merged using the weighted average method to generate a fused feature vector; the dimension of the fused feature vector is consistent with the preset standard feature; the Euclidean distance between the fused feature vector and the preset standard feature is calculated, and the Euclidean distance formula is the square root of the sum of the squares of the numerical differences of the corresponding dimensions of the two vectors. If the Euclidean distance is lower than the preset threshold, a collaborative detection result is generated; the preset threshold is determined according to the Euclidean distance distribution of normal samples in the historical data.
[0109] Specifically, the weight value of the slice fingerprint feature reflects the importance of different features. For example, in the vehicle's automatic driving control data, the weight value of the steering wheel angle data is 0.8, and the weight value of the vehicle speed data is 0.6, indicating that steering control is more important in driving safety. These weight values can be used in the subsequent feature fusion process. When performing principal component analysis on historical data, the most representative feature combination can be extracted. For example, the communication data of a vehicle in a highway scenario contains multiple dimensions such as delay, throughput, and signal strength. Through principal component analysis, it can be reduced to three main features as preset standard features. These standard features can reflect the network performance indicators under normal driving conditions.
[0110] When the weighted average method is used for feature fusion, the multi-dimensional features collected in real time are multiplied by the corresponding weights and summed. For example, in the network slice monitoring data of a vehicle, the delay feature value is ten milliseconds, the weight is 0.5, the bandwidth utilization is seventy percent, the weight is 0.3, and the signal quality is negative seventy decibels, the weight is 0.2. After fusion, a normalized feature vector is obtained. This fusion method not only retains the influence of important features, but also reduces the interference of secondary features.
[0111] Euclidean distance calculation can quantify the degree of deviation between the current state and the standard state. For example, when an autonomous driving vehicle passes through a tunnel, the Euclidean distance between the fusion feature vector and the preset standard feature is 0.8, which exceeds the preset threshold of 0.5, indicating that the network performance is obviously abnormal. This anomaly may be caused by signal attenuation caused by the tunnel environment. The determination of the preset threshold needs to consider the statistical characteristics of historical data. If the Euclidean distance distribution of normal driving data within one month is analyzed and it is found that 95% of the sample distances are less than 0.5, then 0.5 can be set as the anomaly detection threshold. This threshold setting method based on data statistics can effectively control the false alarm rate.
[0112] In practical applications, different scenarios may require different threshold standards. For example, in urban road environments, due to more building occlusion and signal interference, the threshold can be appropriately relaxed to 0.6, while in open environments such as highways, the threshold can be tightened to 0.4 to improve detection sensitivity. This dynamic threshold mechanism can adapt to complex and changing driving environments.
[0113] The collaborative detection results can be used to guide the dynamic adjustment of network slices. For example, if a logistics fleet detects that the communication quality has decreased but is still within the threshold during the delivery process, the system can pre-adjust the slice resource allocation to ensure business continuity. This prediction-based resource regulation method can improve network service quality and reduce the risk of business interruption.
[0114] Step S107: perform correlation analysis on the collaborative detection results and the abnormal traffic mark list. If the same vehicle has abnormal traffic and feature drift in different slices at the same time, a final abnormal detection report including vehicle ID, timestamp, abnormal type and detection results is generated.
[0115] Specifically, a set of vehicle traffic features corresponding to different network slice identifiers is obtained, which includes protocol features and traffic timing features; the distribution difference of protocol features in multiple slices is detected, and if the distribution difference exceeds the preset feature drift threshold, the feature drift mark is triggered; the fluctuation frequency in the traffic timing features is extracted, and the alarm time window is obtained from the abnormal traffic mark list, and the alarm time window and the fluctuation frequency are matched by a sliding window; when there is a continuous overlapping interval in the sliding window matching calculation result, the association rule between the vehicle identifier and the feature drift mark is activated, and a detection report containing the protocol feature offset and the traffic fluctuation trajectory is generated; a real-time traffic data stream containing the vehicle unique identifier and timestamp is received, and the traffic features in the real-time traffic data are extracted Sequence and associated time window; use K-means algorithm to cluster traffic feature sequence according to preset conditions to generate multiple traffic clustering clusters; calculate the feature drift distance of traffic clustering cluster in the time window, and the feature drift distance is obtained by Euclidean distance calculation; when the feature drift distance exceeds the preset threshold, generate an abnormal feature vector containing vehicle identification, offset period and drift amount; query the matching records in the abnormal traffic mark list according to the abnormal feature vector to obtain all feature drift events within the corresponding time range; associate and match traffic anomaly events with overlapping time windows with feature drift events. When the same vehicle identification appears in both feature drift events and traffic anomaly events, generate a detection report containing the vehicle unique identification, anomaly type and detailed detection parameters.
[0116] Specifically, the vehicle traffic characteristics corresponding to the network slice identifier can be divided into protocol characteristics and traffic timing characteristics. Protocol characteristics include data packet size distribution, protocol type ratio, etc. For example, in a certain vehicle communication, the data packet size is bimodal, concentrated in about 100 bytes and 1,000 bytes respectively, and the transmission control protocol accounts for 85%.
[0117] Traffic timing features include timing information such as fluctuation period and peak value change. When vehicles travel through different network slices, these features will have significant differences due to changes in network environment and business requirements. Feature drift detection is achieved by comparing the distribution of protocol features in different slices.
[0118] The feature distribution in the baseline slice is set as a reference. If the distribution difference of the corresponding features in other slices exceeds the preset threshold, the feature drift flag is triggered. For example, the variance of the packet size distribution of a vehicle in the baseline slice is 0.5, while it becomes 2.0 in the target slice. The difference exceeds the preset threshold of 1.0, and the system will generate a feature drift flag.
[0119] Traffic fluctuation frequency analysis focuses on the time characteristics of abnormal traffic. The system obtains the alarm time window from the abnormal traffic mark list, such as detecting that the traffic fluctuation frequency of a vehicle suddenly increases from the normal 0.2 Hz to 0.8 Hz within five minutes.
[0120] Through sliding window matching calculation, the system can find the temporal correlation between abnormal traffic and feature drift. In real-time traffic monitoring, the system continuously receives data streams with vehicle unique identification and timestamps. Through cluster analysis, traffic feature sequences can be divided into different clusters, such as dividing vehicle communication traffic into high-speed mobile clusters, stationary clusters, and low-speed mobile clusters.
[0121] The feature drift distance reflects the degree of change of the cluster center over time. When the drift distance exceeds the preset threshold, the system will generate an abnormal feature vector.
[0122] Abnormal correlation analysis identifies abnormal events with overlapping time windows by matching records in the abnormal traffic tag list. For example, if a vehicle experiences protocol feature drift and a sudden increase in traffic within ten minutes, the system will associate these two types of abnormalities and generate a detection report containing complete abnormal information. By combining the analysis results of the two, the source and nature of the abnormality can be more accurately determined, and collaborative detection can be achieved. This multi-dimensional abnormal correlation analysis helps improve detection accuracy and reduce false alarms.
[0123] Finally, when performing feature fusion based on the feature drift marker list, a weighted average algorithm is used, and the features of different slices are given different weights according to their reliability. If the Euclidean distance between the fused feature and the preset standard feature is less than the threshold, it is confirmed as a valid collaborative detection result. This method can effectively balance the feature differences of each slice and provide more accurate detection results.
[0124] This embodiment further generates a final anomaly detection report by combining the feature analysis of 5G network slicing with the existing abnormal traffic judgment. The two are used to monitor the abnormal conditions of intelligent connected vehicles from different dimensions, complementing each other and forming a more comprehensive and three-dimensional anomaly detection system.
[0125] The embodiment of the present invention further provides a smart connected vehicle abnormality detection system based on big data analysis, which may specifically include:
[0126] The first data acquisition module acquires metadata in the communication between the vehicle CAN bus and the cloud, the metadata including the protocol type and the load characteristics; according to the protocol type, the load characteristics are parsed to extract the key fields of the data packet;
[0127] The data processing module builds a load feature matrix based on key fields, normalizes heterogeneous data, and obtains a feature data set in a unified format;
[0128] The traffic mark construction module uses the Mahalanobis distance algorithm to calculate the distance between each data point in the feature data set and the preset baseline. If the distance exceeds the preset threshold, it is determined to be abnormal traffic and a list of abnormal traffic marks is generated;
[0129] The second data acquisition module obtains the fingerprint features of the transmission link from the 5G network slice, standardizes the fingerprint features of different slices, and generates a fingerprint feature data set that is consistent across slices; the drift mark construction module analyzes the fingerprint feature data set. If the fingerprint features of the same category show significant differences in different slices, it is determined that the feature drift is caused by the dynamic nature of the network slice, and a feature drift mark list is generated;
[0130] The feature drift fusion module performs weighted fusion on the fingerprint features of different slices according to the feature drift mark list. If the Euclidean distance between the fused feature and the preset standard feature is less than the preset threshold, a collaborative detection result is generated;
[0131] The collaborative detection module associates and analyzes the collaborative detection results with the abnormal traffic mark list. If the same vehicle has abnormal traffic and feature drift in different slices at the same time, a final abnormal detection report containing vehicle ID, timestamp, abnormal type and detection results is generated. The above is only a preferred embodiment of the present invention, so any equivalent changes or modifications made according to the structure, features and principles described in the scope of the patent application of the present invention are included in the protection scope of the patent application of the present invention.
Claims
1. A method for detecting abnormalities in intelligent connected vehicles based on big data analysis, characterized in that: include: Obtain metadata from the vehicle CAN bus and cloud communications, including protocol type and load characteristics; According to the protocol type, the load characteristics are parsed and the key fields of the data packet are extracted; Construct a load feature matrix based on key fields, normalize heterogeneous data, and obtain a feature data set in a unified format; The Mahalanobis distance algorithm is used to calculate the distance between each data point in the feature data set and the preset baseline. If the distance exceeds the preset threshold, it is determined to be abnormal traffic and an abnormal traffic mark list is generated; Obtain the fingerprint features of the transmission link from the 5G network slices, standardize the fingerprint features of different slices, and generate a consistent fingerprint feature dataset across slices; Analyze the fingerprint feature data set. If the fingerprint features of the same category show significant differences in different slices, it is determined that the feature drift is caused by the dynamic nature of the network slice, and a feature drift marker list is generated. According to the feature drift mark list, the fingerprint features of different slices are weighted fused. If the Euclidean distance between the fused feature and the preset standard feature is less than the preset threshold, the collaborative detection result is generated; The collaborative detection results are correlated with the abnormal traffic mark list for analysis. If the same vehicle has abnormal traffic and feature drift in different slices at the same time, a final anomaly detection report containing the vehicle ID, timestamp, anomaly type and detection results is generated.
2. The method for detecting abnormality of an intelligent connected vehicle according to claim 1, characterized in that: The Mahalanobis distance algorithm is used to calculate the distance between each data point in the feature data set and the preset baseline. If the distance exceeds the preset threshold, it is determined to be abnormal traffic, and an abnormal traffic mark list is generated, which specifically includes: Obtain all data points in the feature data set, and use the Mahalanobis distance algorithm to calculate the distance between each data point and the preset baseline, which is determined by the mean of historical normal traffic data; The calculated distance value is compared with a preset threshold value, which is determined by the standard deviation of historical normal traffic data. If the distance value is greater than the preset threshold value, the traffic corresponding to the data point is determined to be abnormal traffic; All data points determined to be abnormal traffic are aggregated to generate a preliminary abnormal traffic list; The isolation forest algorithm is used to perform secondary verification on each data point in the preliminary abnormal traffic list. The training set of the isolation forest algorithm uses historical normal traffic data. The preliminary abnormal traffic list is updated according to the secondary verification results to obtain the final abnormal traffic list.
3. The method for detecting abnormality of an intelligent connected vehicle according to claim 2, characterized in that: Also includes: For the abnormal traffic mark list, a sliding window mechanism is used to dynamically update the baseline model. If the abnormal traffic ratio in the window exceeds the preset threshold, the Mahalanobis distance of the baseline model is recalculated to obtain an updated baseline model.
4. The method for detecting abnormality of an intelligent connected vehicle according to claim 3, characterized in that: The dynamic updating of the baseline model by using a sliding window mechanism specifically includes: A flow data set within a receiving window, wherein the flow data set includes a flow feature vector within the time window; Adopting the sliding window mechanism, the flow data set of the current window is obtained according to the preset time interval; Acquire the number of abnormal traffic marks in the traffic data set from the traffic detection node, and determine the abnormal traffic proportion of the current window according to the ratio of the number of abnormal traffic marks to the total traffic; Preset an abnormal ratio threshold, and compare the abnormal traffic ratio of the current window with the abnormal ratio threshold; If the abnormal traffic ratio of the current window exceeds the abnormal ratio threshold, the baseline model update process is triggered; Filter the flow feature vector set of the current window according to the abnormal flow mark to obtain a feature vector subset for updating; Calculating a covariance matrix of the flow feature vector subset, wherein the covariance matrix is used to describe a linear relationship between feature vectors; Calculating Mahalanobis distance algorithm parameters of the subset of flow feature vectors based on the covariance matrix, wherein the Mahalanobis distance algorithm parameters are used to measure the distribution difference of the feature vectors; Updating the baseline model parameters according to the Mahalanobis distance algorithm parameters to generate an updated baseline model parameter set; The updated baseline model parameter set is synchronized to the traffic detection node, triggering the traffic detection node to load the new baseline model parameters.
5. The method for detecting abnormality of an intelligent connected vehicle according to claim 1, characterized in that: The method of obtaining the fingerprint features of the transmission link from the 5G network slice, standardizing the fingerprint features of different slices, and generating a fingerprint feature data set that is consistent across slices includes: Obtain the original data of the transmission link from the 5G network slice, and use fast Fourier transform to extract fingerprint features for different slice types; The extracted fingerprint features are normalized according to the slice type. If the feature value exceeds the preset range, logarithmic transformation is used to compress the data to obtain standardized feature data; The standardized feature data is input into the principal component analysis method to determine the dimension with the largest variance in the feature vector and obtain the feature subset after dimensionality reduction; The feature subsets after dimensionality reduction are grouped using the K-means algorithm. If the distance within the group is less than the preset threshold, they are merged into the same category to generate a feature data set with category labels. The feature data set is reorganized according to the category label. If there is consistency across slices, the feature data is retained to obtain a fingerprint feature data set that is consistent across slices.
6. The method for detecting abnormality of an intelligent connected vehicle according to claim 5, characterized in that: Also includes: The similarity of the fingerprint feature data set that is consistent across slices is calculated by using the cosine similarity algorithm to determine the cosine similarity between feature vectors, and the calculated cosine similarity values are filled into a preset matrix structure in row and column order to obtain a feature similarity matrix; The feature similarity matrix is input into the multilayer perceptron for training. The cross entropy is used as the loss function to calculate the difference between the model output and the true label. The weight parameters of the multilayer perceptron are adjusted according to the loss value through the back propagation algorithm to obtain a classification model that can identify different slice types.
7. The method for detecting abnormality of an intelligent connected vehicle according to claim 1, characterized in that: The K-means clustering algorithm is used to analyze the fingerprint feature data set. If the fingerprint features of the same category show significant differences in different slices, it is determined that the feature drift is caused by the dynamic nature of the network slice, and a feature drift mark list is generated, including: After obtaining the fingerprint feature data set, the K-means clustering algorithm is used to divide the data set to obtain a cluster category set; For each category in the cluster category set, extract its fingerprint feature values in different network slices and calculate the feature difference value; The feature difference value is obtained by calculating the Euclidean distance of the feature values of the same category in different network slices; According to the preset difference threshold, determine whether the feature difference value exceeds the difference threshold, and if so, generate a feature drift mark; After obtaining the feature drift marker set, combined with the network slice dynamic data, a comparative analysis is performed to determine whether the drift cause is the network slice dynamics; The principal component analysis method is used to reduce the dimension of the feature drift label set, reducing the high-dimensional features to three dimensions and obtaining the drift feature principal component set; According to the drift feature principal component set, the drift classification model is constructed using the support vector machine algorithm, and the radial basis function is set as the kernel function to obtain the drift classification result; Combining the drift classification results and the network slice dynamic data, a feature drift tag list containing drift categories and corresponding network slice information is generated to complete the feature drift determination.
8. The method for detecting abnormality of an intelligent connected vehicle according to claim 1, characterized in that: According to the feature drift mark list, the weighted average algorithm is used to perform weighted fusion on the fingerprint features of different slices. If the Euclidean distance between the fused feature and the preset standard feature is less than a preset threshold, a collaborative detection result is generated, including: Get the weight value of the slice fingerprint feature from the feature drift marker list; According to the obtained weight value, the fingerprint features are weighted and merged using the weighted average method to generate a fused feature vector; the dimension of the fused feature vector is consistent with the preset standard feature; Calculate the Euclidean distance between the fused feature vector and the preset standard feature. The Euclidean distance formula is the square root of the sum of the squares of the numerical differences of the corresponding dimensions of the two vectors. If the Euclidean distance is lower than a preset threshold, a collaborative detection result is generated.
9. The method for detecting abnormality of an intelligent connected vehicle according to claim 1, characterized in that: The collaborative detection results are associated with the abnormal traffic mark list. If the same vehicle has abnormal traffic and feature drift in different slices at the same time, a final abnormal detection report containing vehicle ID, timestamp, abnormal type and detection results is generated, including: Obtain a set of vehicle traffic features corresponding to different network slice identifiers, which includes protocol features and traffic timing features; Detect the distribution difference of the protocol features in multiple slices. If the distribution difference exceeds the preset feature drift threshold, the feature drift flag is triggered; Extract the fluctuation frequency in the traffic time series characteristics, obtain the alarm time window from the abnormal traffic mark list, and perform sliding window matching calculation on the alarm time window and the fluctuation frequency; When there are continuous overlapping intervals in the calculation results, the association rules between the vehicle identifier and the characteristic drift mark are activated to generate a detection report containing the protocol characteristic deviation and the flow fluctuation trajectory; Receive a real-time traffic data stream containing a unique vehicle identifier and a timestamp, and extract a traffic feature sequence and an associated time window from the real-time traffic data; Clustering the traffic feature sequences to generate multiple traffic clusters; Calculate the characteristic drift distance of traffic clusters within the time window; When the characteristic drift distance exceeds a preset threshold, an abnormal characteristic vector including the vehicle identification, the offset period and the drift amount is generated; Query matching records in the abnormal traffic mark list according to the abnormal feature vector to obtain all feature drift events within the corresponding time range; The traffic anomaly events with overlapping time windows are associated and matched with the feature drift events. When the same vehicle identification appears in both the feature drift event and the traffic anomaly event, a detection report containing the vehicle ID, timestamp, anomaly type and detection results is generated.
10. An intelligent connected vehicle anomaly detection system based on big data analysis, characterized in that: include: A first data acquisition module acquires metadata in the communication between the vehicle CAN bus and the cloud, the metadata including the protocol type and load characteristics; According to the protocol type, the load characteristics are parsed and the key fields of the data packet are extracted; The data processing module builds a load feature matrix based on key fields, normalizes heterogeneous data, and obtains a feature data set in a unified format; The traffic mark construction module uses the Mahalanobis distance algorithm to calculate the distance between each data point in the feature data set and the preset baseline. If the distance exceeds the preset threshold, it is determined to be abnormal traffic and a list of abnormal traffic marks is generated; The second data acquisition module obtains the fingerprint features of the transmission link from the 5G network slice, standardizes the fingerprint features of different slices, and generates a fingerprint feature data set that is consistent across slices; The drift mark construction module analyzes the fingerprint feature data set. If the fingerprint features of the same category show significant differences in different slices, it is determined that the feature drift is caused by the dynamic nature of the network slice, and a feature drift mark list is generated; The feature drift fusion module performs weighted fusion on the fingerprint features of different slices according to the feature drift mark list. If the Euclidean distance between the fused feature and the preset standard feature is less than the preset threshold, a collaborative detection result is generated; The collaborative detection module associates and analyzes the collaborative detection results with the abnormal traffic mark list. If the same vehicle has abnormal traffic and feature drift in different slices at the same time, a final anomaly detection report containing vehicle ID, timestamp, anomaly type and detection results is generated.
Citation Information
Patent Citations
Feature extraction method for signals of electronic tongue based on manifold learning
CN106018515A
5G slice network anomaly detection method based on virtual network flow analysis
CN114401516A