Model training method, abnormal data identification method, device, equipment and medium

CN118013402BActive Publication Date: 2026-09-25GREAT WALL MOTOR CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311800179.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2026-09-25
Estimated Expiration
2043-12-25

AI Technical Summary

Technical Problem

这些因素都可能导致驾驶员无法正常驾驶车辆,从而增加了交通事故的风险

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118013402B_ABST
    Figure CN118013402B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a model training method, an abnormal data identification method, device, electronic equipment and computer readable storage medium, which are used for training of an abnormal identification model including a random forest model and a time series model. The method comprises the following steps: obtaining a data set including a plurality of data groups, each data group including sample data of a plurality of dimensions, sample data belonging to the same dimension in the data set being collected at different time points; inputting each data group into the random forest model to obtain first identification data with an abnormal data category; obtaining sample data belonging to the same dimension as the first identification data from the data set to obtain target sample data; inputting the target sample data into the time series model to obtain second identification data with an abnormal data category; and training the abnormal identification model by using the data category and the label information of the first and second identification data respectively. The application can improve the accuracy and stability of abnormal data detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a model training method, anomaly data identification method, apparatus, electronic device, and computer-readable storage medium in the field of data processing technology. Background Technology

[0002] In recent years, with the rapid increase in car ownership, the number of traffic accidents on my country's highways has also continued to rise. This trend has attracted widespread public attention, and relevant departments have stepped up measures to address it. According to investigations and research by relevant institutions, vehicle malfunctions are a major factor causing traffic accidents. This further exacerbates our concerns about highway traffic safety.

[0003] Vehicle malfunctions can be caused by a variety of factors, including driver discomfort, fatigue, and driving under the influence of alcohol, as well as issues with the vehicle itself such as insufficient battery power, tire slippage, or low tire pressure. These factors can all prevent the driver from operating the vehicle properly, thus increasing the risk of traffic accidents. If these abnormal vehicles are not detected and managed in a timely manner, it can lead to anything from increased traffic congestion and reduced road capacity to serious accidents causing personal injury, death, and significant economic losses.

[0004] Therefore, identifying and intervening in abnormal situations that occur during vehicle operation, and reducing traffic accidents caused by vehicle malfunctions, is of great significance for improving road safety. Summary of the Invention

[0005] This application provides a model training method, an anomaly data identification method, device, electronic device, and computer-readable storage medium. The method can identify anomaly data in vehicle data, thereby locating the abnormal vehicle components in a timely manner based on the anomaly data, which helps to avoid accidents caused by vehicle component malfunctions.

[0006] Firstly, a model training method is provided for training an anomaly detection model. The anomaly detection model includes a pre-trained random forest model and a time series model, with the output layer of the random forest model connected to the input layer of the time series model. The model training method includes: acquiring sample datasets corresponding to multiple vehicles; wherein each vehicle's sample dataset includes multiple sample data groups, each sample data group includes sample data with multiple data dimensions, and each sample data has annotation information indicating whether the data category is normal or abnormal; sample data belonging to the same data dimension in each vehicle's sample dataset are collected at different times; for each sample data group in each vehicle's sample dataset, the sample data group is input into the random forest model, and the random forest model identifies anomalies in the sample data group based on the multi-dimensional characteristics of the input data. The data is identified, and a first identification result is output. The first identification result includes first identified data and a first data category of the first identified data, where the first data category is an anomaly category. Sample data belonging to the same data dimension as the first identified data is obtained from the sample dataset corresponding to the sample data group, resulting in target sample data. The target sample data and the first identified data are input into the time series model, which identifies anomaly data in the target sample data based on the time series characteristics of the input data in the same data dimension, and outputs a second identification result. The second identification result includes second identified data and a second data category of the second identified data, where the second data category is an anomaly category. The anomaly identification model is trained based on the first data category, the second data category, the annotation information of the first identified data, and the annotation information of the second identified data.

[0007] In the above technical solution, this application first fuses a pre-trained random forest model and a time series model to construct an anomaly identification model for anomaly data identification. This is achieved by acquiring sample datasets corresponding to multiple vehicles. For each sample data group in the sample dataset corresponding to each vehicle, the sample data group is input into the random forest model. The random forest model identifies anomaly data in the sample data group based on the multi-dimensional characteristics of the input data, outputting a first identification result including data of the first anomaly category. Sample data belonging to the same data dimension as the first identification data is obtained from the sample dataset corresponding to the sample data group, resulting in target sample data. The target sample data and the first identification data are input into the time series model. The time series model identifies anomaly data in the target sample data based on the time series characteristics of the input data of the same data dimension, outputting a second identification result including data of the second anomaly category. The anomaly identification model is trained based on the data category of the first identification data, the data category of the second identification data, the annotation information of the first identification data, and the annotation information of the second identification data. This technical solution realizes the training of an anomaly identification model obtained by fusing the random forest model and the time series model. Applying the anomaly identification model can achieve the identification of anomaly data. Since the anomaly detection model is a fusion of the random forest model and the time series model, it combines the advantages of both models, which helps to improve the accuracy and stability of anomaly detection.

[0008] In conjunction with the first aspect, in some possible implementations, obtaining the sample datasets corresponding to each of the multiple vehicles includes: for each of the multiple vehicles, obtaining the sample operation data generated by each vehicle at different times, thus obtaining the sample operation data for each vehicle at multiple times; setting annotation information for the sample operation data for each vehicle at multiple times; determining whether the difference between the first quantity of the first data and the second quantity of the second data in the sample operation data for each vehicle at multiple times is less than a preset threshold; wherein, the first data includes sample operation data whose data category is normal in the sample operation data for each vehicle at multiple times, and the second data includes sample operation data whose data category is abnormal in the sample operation data for each vehicle at multiple times; if so, then for the sample operation data for each time corresponding to each vehicle, according to the preset threshold... The sample running data at each time moment is categorized according to multiple data dimensions to obtain a sample data group corresponding to each time moment. Based on the sample data group corresponding to each time moment, a sample dataset corresponding to each vehicle is generated to obtain sample datasets corresponding to multiple vehicles. If not, the sample running data at multiple time moments corresponding to each vehicle is preprocessed to adjust the difference between the first quantity of the first data and the second quantity of the second data in the sample running data at multiple time moments corresponding to each vehicle to be less than a preset threshold. The preprocessing includes sampling processing or undersampling processing. The preprocessed sample running data at each time moment corresponding to each vehicle is categorized according to multiple preset data dimensions to obtain a sample data group corresponding to each time moment. Based on the sample data group corresponding to each time moment, a sample dataset corresponding to each vehicle is generated to obtain sample datasets corresponding to multiple vehicles.

[0009] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, training the anomaly recognition model based on the first data category, the second data category, the annotation information of the first identification data, and the annotation information of the second identification data includes: determining whether the first identification data and the second identification data are the same; if so, fusing the first identification result and the second identification result to obtain a fused identification result; wherein, the fused identification result includes third identification data and the third data category of the third identification data, the third data category being an anomaly category, and the third identification data being the same as the first identification data; and training the anomaly recognition model based on the difference information between the annotation information of the third data category and the first identification data.

[0010] In combination with the first aspect and the above implementation methods, in some possible implementation methods, the step of fusing the first identification result and the second identification result to obtain the fused identification result includes: determining the average value of the first identification result and the second identification result to obtain the fused identification result; or, performing a weighted average processing on the first identification result and the second identification result to obtain the fused identification result.

[0011] In combination with the first aspect and the above implementation methods, in some possible implementation methods, training the anomaly recognition model based on the first data category, the second data category, the annotation information of the first identification data, and the annotation information of the second identification data includes: determining whether the first identification data and the second identification data are the same; if not, training the anomaly recognition model based on the difference information between the annotation information of the first data category and the first identification data and the difference information between the annotation information of the second data category and the second identification data.

[0012] In conjunction with the first aspect and the above implementation methods, in some possible implementation methods, the model training method further includes: for any one of the multiple vehicles, acquiring data generated by the vehicle at different times and belonging to the same data dimension to obtain time series sample data; verifying the stationarity of the time series sample data; if the stationarity verification of the time series sample data fails, performing differencing on the time series sample data to obtain the differencing order of the seasonal autoregressive moving average model and the stationary time series sample data corresponding to the time series sample data; drawing an autocorrelation plot and a partial autocorrelation plot based on the stationary time series sample data; determining the autoregressive order and the moving average order of the seasonal autoregressive moving average model based on the autocorrelation plot and the partial autocorrelation plot; constructing the seasonal autoregressive moving average model based on the differencing order, the autoregressive order, and the moving average order to obtain the time series model.

[0013] Secondly, a vehicle anomaly identification method is provided, comprising: acquiring target data groups generated by a target vehicle at multiple set times to obtain multiple target data groups; wherein each target data group includes data to be identified in multiple data dimensions; inputting the multiple target data groups into an anomaly identification model, and outputting anomaly identification results by the anomaly identification model; wherein the anomaly identification model is trained according to the above-mentioned model training method, and the anomaly identification results include data to be identified in the multiple target data groups whose data category is anomaly; acquiring vehicle components corresponding to the data to be identified in the anomaly identification results; and identifying the vehicle components as anomaly components of the target vehicle.

[0014] In the above technical solution, the vehicle anomaly identification method provided in this application obtains multiple target data groups generated by the target vehicle at multiple set times, each target data group including data to be identified in multiple data dimensions. These multiple target data groups are input into an anomaly identification model, which outputs anomaly identification results for the data in the multiple target data groups that fall under the anomaly category. The method then identifies the vehicle components corresponding to the data to be identified in the anomaly identification results, thus identifying the abnormal components of the target vehicle. This technical solution uses the anomaly identification model to identify abnormal data in the vehicle's operational data and locates abnormal components within the vehicle. Since the anomaly identification model is a fusion of a random forest model and a time series model, it combines the advantages of both models, improving not only the accuracy of anomaly data identification but also the accuracy of locating abnormal vehicle components. This helps users promptly detect potential anomalies or malfunctions in the vehicle, thus helping to prevent accidents caused by abnormal vehicle components.

[0015] Thirdly, a model training apparatus is provided for training an anomaly detection model, the anomaly detection model including a pre-trained random forest model and a time series model, wherein the output layer of the random forest model is connected to the input layer of the time series model; the model training apparatus includes:

[0016] The sample acquisition module is used to acquire sample datasets corresponding to multiple vehicles. Each sample dataset for each vehicle includes multiple sample data groups, each sample data group includes sample data of multiple data dimensions, and each sample data has annotation information. The annotation information is used to indicate whether the data category of the sample data is normal or abnormal. The sample data belonging to the same data dimension in the sample dataset for each vehicle are collected at different times.

[0017] The first identification module is used to input each sample data group in the sample dataset corresponding to each vehicle into the random forest model, and the random forest model identifies abnormal data in the sample data group based on the multi-dimensional characteristics of the input data, and outputs a first identification result; wherein, the first identification result includes first identification data and a first data category of the first identification data, and the first data category is an abnormal category;

[0018] The data selection module is used to obtain sample data belonging to the same data dimension as the first identification data from the sample dataset corresponding to the sample data group, so as to obtain the target sample data;

[0019] The second identification module is used to input the target sample data and the first identification data into the time series model, and the time series model identifies abnormal data in the target sample data based on the time series characteristics of the input data in the same data dimension, and outputs a second identification result; wherein, the second identification result includes the second identification data and the second data category of the second identification data, and the second data category is an anomaly category;

[0020] The model training module is used to train the anomaly recognition model based on the first data category, the second data category, the annotation information of the first recognition data, and the annotation information of the second recognition data.

[0021] In conjunction with the third aspect, in some possible implementations, the sample acquisition module is specifically used for: acquiring sample operation data generated by each vehicle at different times for each of the multiple vehicles, thus obtaining sample operation data for each vehicle at multiple times; setting annotation information for the sample operation data for each vehicle at multiple times; determining whether the difference between the first quantity of the first data and the second quantity of the second data in the sample operation data for each vehicle at multiple times is less than a preset threshold; wherein, the first data includes sample operation data whose data category is normal in the sample operation data for each vehicle at multiple times, and the second data includes sample operation data whose data category is abnormal in the sample operation data for each vehicle at multiple times; if so, then for the sample operation data for each vehicle at each time, according to the preset multiple... The data dimensions are used to categorize the sample running data at each time point to obtain the sample data group corresponding to each time point; based on the sample data group corresponding to each time point, a sample dataset corresponding to each vehicle is generated to obtain the sample dataset corresponding to each vehicle; otherwise, the sample running data at multiple time points corresponding to each vehicle is preprocessed to adjust the difference between the first quantity of the first data and the second quantity of the second data in the sample running data at multiple time points corresponding to each vehicle to be less than a preset threshold; wherein, the preprocessing includes sampling processing or undersampling processing; the preprocessed sample running data at each time point corresponding to each vehicle is categorized according to preset multiple data dimensions to obtain the sample data group corresponding to each time point; based on the sample data group corresponding to each time point, a sample dataset corresponding to each vehicle is generated to obtain the sample dataset corresponding to each vehicle.

[0022] In combination with the third aspect and the above implementation methods, in some possible implementations, the model training module includes:

[0023] The first training unit is used to determine whether the first identification data and the second identification data are the same; if so, the first identification result and the second identification result are fused to obtain a fused identification result; wherein, the fused identification result includes third identification data and a third data category of the third identification data, the third data category is an anomaly category, and the third identification data is the same as the first identification data; the anomaly identification model is trained based on the difference information between the annotation information of the third data category and the first identification data.

[0024] In conjunction with the third aspect and the above implementation methods, in some possible implementation methods, the first training unit, in fusing the first recognition result and the second recognition result to obtain the fused recognition result, is specifically used to: determine the average value of the first recognition result and the second recognition result to obtain the fused recognition result; or, perform weighted averaging on the first recognition result and the second recognition result to obtain the fused recognition result.

[0025] In conjunction with the third aspect and the above implementation methods, in some possible implementations, the model training module further includes:

[0026] The second training unit is used to determine whether the first identification data and the second identification data are the same; if not, the anomaly identification model is trained based on the difference information between the annotation information of the first data category and the first identification data and the difference information between the annotation information of the second data category and the second identification data.

[0027] In conjunction with the third aspect and the above implementation methods, in some possible implementations, the model training device further includes:

[0028] The model building unit is used to acquire, for any one of multiple vehicles, data generated by that vehicle at different times and belonging to the same data dimension, to obtain time series sample data; verify the stationarity of the time series sample data; if the stationarity verification of the time series sample data fails, perform differencing on the time series sample data to obtain the differencing order of the seasonal autoregressive moving average model and the stationary time series sample data corresponding to the time series sample data; draw autocorrelation plots and partial autocorrelation plots based on the stationary time series sample data; determine the autoregressive order and moving average order of the seasonal autoregressive moving average model based on the autocorrelation plots and partial autocorrelation plots; and construct the seasonal autoregressive moving average model based on the differencing order, the autoregressive order, and the moving average order to obtain the time series model.

[0029] Fourthly, an abnormal data identification device is provided, the abnormal data identification device comprising:

[0030] The data acquisition module is used to acquire target data sets generated by the target vehicle at multiple set time periods, resulting in multiple target data sets; each target data set includes data to be identified in multiple data dimensions;

[0031] A data recognition module is used to input the plurality of target data groups into an anomaly recognition model, and the anomaly recognition model outputs anomaly recognition results; wherein, the anomaly recognition model is trained according to the above-mentioned model training method, and the anomaly recognition results include the data to be identified in the plurality of target data groups whose data category is anomaly;

[0032] The anomaly determination module is used to obtain the vehicle component corresponding to the data to be identified in the anomaly identification result, and to identify the vehicle component as an abnormal component of the target vehicle.

[0033] Fifthly, an electronic device is provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, causing the electronic device to execute the model training method in the first aspect or any possible implementation of the first aspect, or to execute the vehicle anomaly recognition method in the implementation of the second aspect.

[0034] In a sixth aspect, a computer program product is provided, comprising: computer program code, which, when executed on a computer, causes the computer to execute the model training method in the first aspect or any possible implementation thereof, or to execute the vehicle anomaly recognition method in the implementation thereof.

[0035] In a seventh aspect, a computer-readable storage medium is provided, which stores computer program code that, when executed on a computer, causes the computer to perform the model training method in the first aspect or any possible implementation thereof, or to perform the vehicle anomaly recognition method in the implementation thereof. Attached Figure Description

[0036] Figure 1 A schematic flowchart of a model training method provided in an embodiment of this application is shown;

[0037] Figure 2 A schematic diagram of the anomaly detection model is shown.

[0038] Figure 3A schematic flowchart of a vehicle anomaly identification method provided in an embodiment of this application is shown;

[0039] Figure 4 This paper shows a schematic diagram of the structure of a model training device provided in an embodiment of this application;

[0040] Figure 5 This illustration shows a structural schematic diagram of an abnormal data identification device provided in an embodiment of this application;

[0041] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0042] The technical solutions in this application will be clearly and thoroughly described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0043] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0044] The following is an embodiment of a model training method provided in this application.

[0045] Figure 1 A schematic flowchart of a model training method provided in an embodiment of this application is shown, such as... Figure 1 As shown, the model training method provided in this application embodiment is applied to an electronic device with computing power. This model training method is used for training an anomaly recognition model, such as... Figure 2 As shown, Figure 2 The diagram illustrates the structure of an anomaly detection model, which includes a pre-trained Random Forest model 100 and a pre-trained time series model 200. In other words, the anomaly detection model can be understood as a fusion model. Figure 2 As shown in the figure within the dashed box on the left, the output layer of the random forest model 100 is connected to the input layer of the time series model 200, that is, the output of the random forest model 100 is used as the input of the time series model 200.

[0046] The above model training methods include the following schemes:

[0047] S110: Obtain the sample datasets corresponding to each of the multiple vehicles.

[0048] In one exemplary embodiment, the sample dataset corresponding to each vehicle is obtained after processing vehicle-to-everything (V2X) data, which includes vehicle performance data, sensor data, location data, etc. The sample dataset corresponding to each vehicle includes multiple sample data groups, and each sample data group includes sample data with multiple data dimensions. For example, if the data with multiple data dimensions is 9, then each sample data group includes sample data with 9 data dimensions, and the number of data dimensions included in each sample data group is the same.

[0049] Multiple data dimensions include: time feature dimension, vehicle status feature dimension, sensor feature dimension, sensor data statistical feature dimension, location feature dimension, specific event feature dimension, statistical feature dimension, rate of change feature dimension, historical feature dimension, etc.

[0050] 1. Time-related characteristics, including:

[0051] Hours, minutes, and seconds: These represent the specific time information extracted from the timestamp.

[0052] Day of the week: This indicates the day of the week calculated based on the timestamp to capture the cyclical changes each week.

[0053] Is it a working day? This indicates whether it is a working day based on the date, which may affect the vehicle's usage mode.

[0054] 2. Vehicle status feature dimensions include:

[0055] Speed: Represents the vehicle's speed information.

[0056] Acceleration: This indicates the vehicle's acceleration information, which helps detect abnormal situations such as rapid acceleration or deceleration.

[0057] Steering angle: This indicates the vehicle's steering angle and may be related to abnormal driving behavior.

[0058] 3. Sensor feature dimensions, including:

[0059] Temperature, pressure, humidity, and other sensor data are used to monitor the vehicle's environmental conditions. Abnormal data may indicate equipment malfunction or abnormal conditions.

[0060] 4. Sensor data statistical feature dimensions: such as maximum value, minimum value, average value, etc., used to capture abnormal situations in data distribution.

[0061] 5. Location feature dimension, including:

[0062] Longitude and latitude: indicates the geographic location information of the vehicle, which may be related to abnormal conditions in specific areas.

[0063] Position change speed: used to calculate the change speed of the vehicle position, and abnormal speed may indicate abnormal driving.

[0064] 6. Specific event feature dimension, including:

[0065] Engine fault code: used to map the engine fault code into a numerical feature, which may affect the state of the vehicle.

[0066] Number of sharp brakes and number of sharp turns: the number of sharp brakes and sharp turns is calculated based on sensor data, which is used to detect abnormal driving behavior.

[0067] 7. Statistical feature dimension, including:

[0068] Sliding window statistical features: used to calculate statistical features such as average value and standard deviation within a period of time, so as to capture dynamic changes of data.

[0069] 8. Change rate feature dimension, including:

[0070] Data change rate: used to calculate the change rate between adjacent data points, and is used to detect abnormal sudden changes.

[0071] 9. Historical feature dimension, including:

[0072] Data value at the previous moment: used to take the data at the previous moment as a feature, and is used to capture the trend change of data.

[0073] For each sample data in the sample dataset corresponding to each vehicle, each sample data has annotation information. The annotation information is used to indicate that the data category of the sample data is a normal category or an abnormal category. For example, annotation information "1" indicates that the data category of the sample data is a normal category, and annotation information "0" indicates that the data category of the sample data is an abnormal category. Sample data belonging to the same data dimension in the sample dataset corresponding to each vehicle is collected at different times. For example, the sample data under the historical feature dimension in the sample dataset corresponding to vehicle A includes: data 1, data 2, data 3, ..., data 10, and the respective corresponding times of data 1, data 2, data 3, ..., data 10 are t1, t2, t3, ..., t10, t1<t2<t3, ..., t9<t10, that is, data 1, data 2, data 3, ..., data 10 belong to time series data.

[0074] S120: For each sample data group in the sample dataset corresponding to each vehicle, the sample data group is input into the random forest model, and the random forest model identifies abnormal data in the sample data group based on the multi-dimensional characteristics of the input data, and outputs the first identification result.

[0075] We obtain sample datasets for multiple vehicles. For each vehicle's sample dataset, we divide it into a training set, a test set, or a validation set according to a preset ratio, such as dividing it into a training set and a test set, with a preset ratio of 8:2. Each vehicle's sample dataset comprises 80% training data and 20% testing data.

[0076] like Figure 2 As shown in the flowchart within the dashed box on the right, after dividing the training set and the test set, for each sample data group in the training set corresponding to each vehicle, each sample data group is used as the input of the random forest model, i.e., the first input. The random forest model identifies abnormal data in the input sample data group based on the multi-dimensional characteristics of the input data, and outputs the first identification result. The first identification result includes the first identification data and the first data category of the first identification data. The first data category is the abnormal category. The first data category is represented by a probability value. If the probability value of the first data category is greater than a preset value, it indicates that the category is abnormal. The first identification data is one of the multiple sample data included in the input sample data group.

[0077] S130: Obtain sample data belonging to the same data dimension as the first identification data from the sample dataset corresponding to the sample data group to obtain the target sample data.

[0078] After obtaining the first recognition result, sample data belonging to the same data dimension as the first recognition data is obtained from the sample dataset corresponding to the sample data group input to the random forest model. The first recognition data and the obtained sample data belonging to the same data dimension as the first recognition data are used as the target sample data. For example, if the sample dataset corresponding to the sample data group input to the random forest model is sample dataset B corresponding to car B, and the data dimension of the first recognition data is the rate of change feature dimension, then sample data with the rate of change feature dimension is obtained from sample dataset B. The obtained sample data with the rate of change feature dimension and the first recognition data are used as the target sample data.

[0079] S140: Input the target sample data and the first identification data into the time series model, and the time series model identifies abnormal data in the target sample data based on the time series characteristics of the input data of the same data dimension, and outputs a second identification result.

[0080] like Figure 2As shown in the flowchart within the dashed box on the right, target sample data is obtained and used as the input to the time series model, i.e., the second input. The time series model identifies abnormal data in the target sample data based on the time series characteristics of the input data, thereby outputting a second identification result. The second identification result includes the second identification data and the second data category of the second identification data. The second data category is an anomaly category, which is represented by a probability value. If the probability value of the second data category is greater than a preset value, it indicates that the category is an anomaly. The second identification data is a data point in the target sample data.

[0081] S150: Train the anomaly recognition model based on the first data category, the second data category, the annotation information of the first identification data, and the annotation information of the second identification data.

[0082] After obtaining the first and second identification results, the annotation information of the first and second identification data is obtained, and the anomaly identification model is iteratively trained based on the first data category, the second data category, the annotation information of the first and second identification data, until the anomaly identification model converges.

[0083] In one possible implementation, training the anomaly detection model based on the first data category, the second data category, the annotation information of the first identification data, and the annotation information of the second identification data includes the following schemes:

[0084] Determine whether the first identification data and the second identification data are the same;

[0085] If so, the first identification result and the second identification result are fused to obtain a fused identification result;

[0086] The anomaly recognition model is trained based on the difference between the annotation information of the third data category and the first identification data.

[0087] After obtaining the first and second identification results, it is determined whether the first and second identification data are the same sample data. If they are, it indicates that the random forest model and the time series model have identified the same data. The first and second identification results are then fused. Specifically, this fusion involves combining the probability values ​​of the first and second data categories to obtain a fused identification result. The fused identification result includes third identification data and its third category. The probability value of the third data category is the fusion of the probability values ​​of the first and second data categories. If the probability value of the third data category is greater than a preset value, it indicates that the third data category is an anomaly category, and the third identification data is the same as either the first or second identification data.

[0088] In one possible implementation, training the anomaly detection model based on the first data category, the second data category, the annotation information of the first identification data, and the annotation information of the second identification data includes the following schemes:

[0089] The average value of the first recognition result and the second recognition result is determined to obtain the fused recognition result; or,

[0090] The first identification result and the second identification result are weighted and averaged to obtain the fused identification result.

[0091] One approach is to calculate the average of the probability values ​​of the first and second data categories to obtain the probability value of the third data category. Another approach pre-sets the first fusion weight value corresponding to the random forest model and the second fusion weight value corresponding to the time series model, where the first fusion weight value + second fusion weight value = 1, and the probability value of the third data category = (probability value of the first data category × first fusion weight value + probability value of the second data category × second fusion weight value) / 2.

[0092] After obtaining the probability value of the third data category, the first difference between the annotation information of the third data category and the first identification data is calculated using a pre-designed first loss function. This first difference serves as the difference information between the annotation information of the third data category and the first identification data. It is determined whether the first difference is less than or equal to a first preset difference. If the first difference is greater than the first preset difference, the anomaly identification model is iteratively trained until the first difference is less than or equal to the first preset difference. At this point, the training of the model is stopped, and the anomaly identification model training is completed.

[0093] In one possible implementation, training the anomaly detection model based on the first data category, the second data category, the annotation information of the first identification data, and the annotation information of the second identification data includes the following schemes:

[0094] Determine whether the first identification data and the second identification data are the same;

[0095] If not, the anomaly recognition model is trained based on the differences between the annotation information of the first data category and the first identification data and the differences between the annotation information of the second data category and the second identification data.

[0096] After obtaining the first and second identification results, it is determined whether the first and second identification data are the same sample data. If not, it indicates that the data identified by the random forest model and the time series model are not the same data. Then, a second difference is calculated using a second loss function between the first data category and the labeled information of the first identification data. This second difference represents the discrepancy between the first data category and the labeled information of the first identification data. A third difference is also calculated using the second loss function between the second data category and the labeled information of the second identification data. This third difference represents the discrepancy between the second data category and the labeled information of the second identification data. It is then determined whether the second difference is less than or equal to a second preset difference and whether the third difference is less than or equal to a third preset difference. If both the second and third differences are greater than the second and third preset differences, iterative training of the anomaly identification model continues until both the second and third differences are less than or equal to the second and third preset differences. At this point, model training stops, and the anomaly identification model training is complete.

[0097] After the anomaly detection model is trained, it undergoes evaluation and parameter tuning. Model evaluation involves using a test set to assess the model, with metrics including accuracy, precision, recall, and F1 score. A baseline comparison is established by comparing the results of the random forest model with those of the anomaly detection model, and also by comparing the results of the time series model with those of the anomaly detection model. The comparison results determine the performance improvement achieved by the anomaly detection model. Parameter tuning involves adjusting the model's weights, voting numbers, averaging methods, etc., based on the evaluation results. If the anomaly detection model's performance does not meet expectations, the output features of the random forest model can be reselected, as certain features may negatively impact the fusion effect of the anomaly detection model.

[0098] The anomaly detection model is evaluated and its parameters are tuned to bring it to the desired standard, thus completing the training. The model is then saved and can be used in real-world scenarios, such as for identifying abnormal vehicle data.

[0099] This application integrates a pre-trained random forest model and a time series model to construct an anomaly identification model. It acquires sample datasets corresponding to multiple vehicles, and for each sample data group in the sample dataset for each vehicle, inputs the sample data group into the random forest model. The random forest model identifies anomalies in the sample data group based on the multi-dimensional characteristics of the input data, outputting a first identification result including data categorized as an anomaly. Then, it obtains target sample data from the sample dataset corresponding to the sample data group, which belongs to the same data dimension as the first identification data. The target sample data and the first identification data are input into the time series model, which identifies anomalies in the target sample data based on the time series characteristics of the input data in the same data dimension, outputting a second identification result including data categorized as an anomaly. The application of this anomaly identification model trains the model based on the data categories of the first and second identification data, the annotation information of the first and second identification data, and the annotation information of the second identification data. This technical solution realizes the training of an anomaly identification model obtained by integrating the random forest model and the time series model, and the application of this model enables the identification of anomalies. Since the anomaly detection model is a fusion of the random forest model and the time series model, it combines the advantages of both models, which helps to improve the accuracy and stability of anomaly detection.

[0100] The advantages of anomaly detection models include:

[0101] 1. It comprehensively considers the multi-dimensional features of the data, meaning the random forest model can handle feature information from multiple data dimensions, such as vehicle status, operational behavior, and environmental data. By comprehensively considering various features, it can more comprehensively capture data with anomalies, thereby improving the sensitivity of detection.

[0102] 2. It can capture time trends and periodicity; that is, time series models can analyze the trends and periodicity of time series data, helping to identify seasonal and trend anomalies in the data. Combining the time series analysis capabilities of time series models with the comprehensive feature extraction capabilities of random forest models can more accurately capture anomalous data.

[0103] 3. It can improve the robustness of the model. By fusing the random forest model and the time series model, the robustness of the anomaly detection model can be improved, and the risk of overfitting can be reduced. Because the random forest model and the time series model use different methods to identify anomalous data, they have strong capabilities in different aspects. Fusing them together can improve the stability of the overall model.

[0104] 4. It can improve the reliability of model output results. By fusing the random forest model and the time series model, model fusion can reduce the prediction bias of a single model and improve the reliability of the final output results. By combining the opinions of multiple models, more consistent and reliable results of identifying abnormal data can be obtained.

[0105] 5. Expanded application scenarios allow the model to be applied to more complex scenarios. For example, in complex vehicle systems, there may be multiple abnormal patterns and factors that are difficult to cover completely with a single model. The anomaly recognition model generated by the fusion of random forest model and time series model can cope with complex scenarios and identify different types of anomalies.

[0106] In one possible implementation, obtaining the sample datasets corresponding to each of the multiple vehicles includes the following schemes:

[0107] For each of the multiple vehicles, obtain the sample operation data generated by each vehicle at different times to obtain the sample operation data of each vehicle at multiple times;

[0108] Labeling information is set for the sample operation data of each vehicle at multiple time points;

[0109] Determine whether the difference between the first quantity of the first data and the second quantity of the second data in the sample running data at multiple times corresponding to each vehicle is less than a preset threshold; wherein, the first data includes sample running data of the normal category in the sample running data at multiple times corresponding to each vehicle, and the second data includes sample running data of the abnormal category in the sample running data at multiple times corresponding to each vehicle.

[0110] If so, then for the sample running data of each vehicle at each time moment, the sample running data at each time moment is classified according to multiple preset data dimensions to obtain the sample data group corresponding to each time moment;

[0111] Based on the sample data group corresponding to each time moment, generate a sample dataset for each vehicle to obtain sample datasets for multiple vehicles.

[0112] If not, the sample running data for each vehicle at multiple time points are preprocessed to adjust the difference between the first quantity of the first data and the second quantity of the second data in the sample running data for each vehicle at multiple time points to be less than a preset threshold; wherein, the preprocessing includes sampling processing or undersampling processing;

[0113] The preprocessed sample operation data for each vehicle at each time point are categorized according to multiple preset data dimensions to obtain the sample data group corresponding to each time point.

[0114] Based on the sample data group corresponding to each time moment, generate a sample dataset for each vehicle to obtain sample datasets for each of the multiple vehicles.

[0115] The process of generating the sample dataset is as follows:

[0116] Multiple vehicles are selected in advance. These vehicles can be different models from the same brand, or vehicles from different brands, etc. This application does not specify any particular limitation. Each of the multiple vehicles is referred to as vehicle i. Sample operation data generated by vehicle i at different times in the past are obtained, thereby obtaining sample operation data for vehicle i at multiple times. This sample operation data comes from the aforementioned vehicle network data.

[0117] We obtain sample running data for vehicle i at multiple time points and set annotation information for each sample running data. That is, if the sample running data is abnormal, the data category is set to the abnormal category, and if the sample running data is normal, the data category is set to the normal category.

[0118] After setting the annotation information for the sample running data at multiple time points corresponding to vehicle i, determine whether the difference between the first quantity of the first data and the second quantity of the second data in the sample running data at multiple time points corresponding to vehicle i is less than a preset threshold.

[0119] If so, it means that the number of sample operation data belonging to the normal category is balanced with the number of sample operation data belonging to the abnormal category. Then, the sample operation data of vehicle i at multiple time points are classified according to multiple preset data dimensions (such as time feature dimension, vehicle state feature dimension, sensor feature dimension, etc.). That is, the sample operation data of vehicle i at multiple time points are grouped into a sample data group for sample operation data of multiple data dimensions collected at the same time. For example, the sample operation data of vehicle i at multiple time points that belongs to the time feature dimension and collected at time t1 are grouped into a sample data group. The operational data, sample operational data of the vehicle state feature dimension, and sample operational data of the sensor feature dimension are divided into a sample data group. Among the sample operational data of multiple time points corresponding to vehicle i, the sample operational data of the time feature dimension, the sample operational data of the vehicle state feature dimension, and the sample operational data of the sensor feature dimension collected at time t2 are divided into a sample data group. This process is repeated to obtain the sample data group for each time point corresponding to vehicle i. The sample dataset corresponding to vehicle i is generated by the sample data group for each time point corresponding to vehicle i. In this way, the sample datasets corresponding to multiple vehicles can be obtained.

[0120] If not, it means that the number of sample data in the normal category is unbalanced with the number of sample data in the abnormal category. There may be too much sample data in the normal category and too little sample data in the abnormal category, or there may be too much sample data in the abnormal category and too little sample data in the abnormal category. Therefore, it is necessary to adjust the balance between the number of sample data in the normal category and the number of sample data in the abnormal category.

[0121] The adjustment process includes: preprocessing (sampling or undersampling) the sample running data at multiple time points corresponding to vehicle i, so that the difference between the first quantity of the first data and the second quantity of the second data in the sample running data at multiple time points corresponding to vehicle i is less than a preset threshold. This balances the number of sample running data in the normal category and the number of sample running data in the abnormal category. Then, the preprocessed sample running data at each time point corresponding to vehicle i is categorized according to multiple preset data dimensions to obtain sample data sets for each time point corresponding to vehicle i. The categorization process here is the same as the one described above and will not be repeated here. The sample data sets for each time point corresponding to vehicle i are obtained, and a sample dataset for vehicle i is generated from these sample data sets. This results in sample datasets for multiple vehicles, where the number of sample running data in the normal category and the number of sample running data in the abnormal category are balanced, ensuring better model performance in identifying abnormal data.

[0122] In one possible implementation, the model training method further includes: constructing a random forest model, wherein the construction of the random forest model includes:

[0123] This application obtains operational data belonging to the aforementioned multiple data dimensions collected at different times from vehicle network data. The time of the obtained operational data is the same as that of the sample operational data mentioned above in terms of dimensions, but the time of collection is different. For ease of distinction, this application refers to the operational data used to construct the random forest model as the first model sample data. By obtaining multiple first model sample data from vehicle network data, a first model sample dataset for constructing the random forest model is generated. The first model sample dataset has a balanced quantity of first model sample data of normal category and first model sample data of abnormal category, which can ensure the quality and consistency of the data. The first model sample dataset is divided into features (i.e., model input) and target variables (i.e. model output).

[0124] Training samples are randomly drawn from the first model sample dataset. For each decision tree, sampling with replacement is performed from the first model sample dataset to construct a random sample set. In this way, the training data for each decision tree is slightly different to improve the diversity of the model. Each decision tree corresponds to a random sample set, that is, there are multiple random sample sets.

[0125] For each random sample set, an independent decision tree is constructed. The construction of the decision tree includes:

[0126] Feature selection: This means selecting a subset of features at each node for optimal segmentation;

[0127] Data segmentation: Divide the data into two subsets based on selected features and segmentation criteria;

[0128] Recursive construction: Continue to recursively split each subset until a termination condition (such as the number of leaf nodes, depth, etc.) is met.

[0129] Ensemble decision trees: Multiple decision trees are combined into a random forest model. For classification tasks, voting can be used to select the final prediction result; for regression tasks, the average or weighted average can be calculated as the final prediction result.

[0130] Feature importance assessment: Random forest models can provide an importance score for each feature, which measures the impact of each feature on model performance; the importance score can help users understand which features in the data have the greatest influence on the prediction results.

[0131] Model hyperparameter tuning: Adjust the hyperparameters of the random forest model, such as the number of decision trees, maximum depth, minimum number of sample splits, etc., to obtain better performance. After the model hyperparameter tuning is completed, the random forest model is built.

[0132] In one possible implementation, the model training method further includes constructing a time series model, the construction process of which includes:

[0133] For any one of the multiple vehicles, obtain the data generated by that vehicle at different times that belong to the same data dimension to obtain time series sample data;

[0134] Verify the stationarity of the time series sample data;

[0135] If the stationarity verification of the time series sample data fails, the time series sample data is differentially processed to obtain the difference order of the seasonal autoregressive moving average model and the stationary time series sample data corresponding to the time series sample data.

[0136] Autocorrelation and partial autocorrelation plots were drawn based on the stationary time series sample data.

[0137] Based on the autocorrelation plot and the partial autocorrelation plot, determine the autoregression order and the moving average order of the seasonal autoregressive composite moving average model;

[0138] The seasonal autoregressive composite moving average model is constructed based on the difference order, the autoregression order, and the moving average order to obtain the time series model.

[0139] Let any one of the multiple vehicles be referred to as vehicle j. In chronological order, retrieve the data generated by vehicle j at different times in the past, which belong to the same data dimension. For example, retrieve data 1 generated by vehicle j at time t1 that belongs to the vehicle state feature dimension, data 2 generated at time t2 that belongs to the vehicle state feature dimension, data 2 generated at time t3 that belongs to the vehicle state feature dimension, ..., data n generated at time tn that belongs to the vehicle state feature dimension.

[0140] To ensure the data is arranged in chronological order, if any data point is missing, interpolation or padding is performed to maintain temporal continuity. For example, the final time series sample data corresponding to t1-tn might be data 1, data 2, data 3, ..., data (n-1), where t1... <t2<t3、..、t(n-1)<tn。

[0141] The construction of time series models requires ensuring that the time series data is stationary, meaning it does not exhibit a clear trend or seasonality over time. Therefore, after obtaining time series sample data corresponding to multiple time points, an autocorrelation function (ACF) and a partial autocorrelation function (PACF) are plotted based on the time series sample data. The autocorrelation and partial autocorrelation of the time series sample data are determined through the autocorrelation and partial autocorrelation, and the stationarity of the time series sample data is verified based on the autocorrelation and partial autocorrelation.

[0142] If the time series sample data fails the stationarity test through autocorrelation and partial autocorrelation, then the time series sample data is differentially processed, that is, the time series sample data is differentially processed to transform the time series sample data into a stationary time series; among which, the differential operation includes first-order difference and multiple-order difference.

[0143] After performing differencing operations on the time series sample data, we can obtain the stationary time series sample data transformed from the time series sample data, as well as the difference order of the stationary time series sample data transformed from the time series sample data. This is also the difference order of the Autoregressive Integrated Moving Average model (ARIMA).

[0144] Then, based on stationary time series sample data, autocorrelation plots and partial autocorrelation plots are drawn, and the autoregression order and moving average order of the ARIMA model are obtained from the autocorrelation plots and partial autocorrelation plots obtained from the stationary time series sample data. Then, the ARIMA model is constructed according to the difference order, autoregression order and moving average order, and the constructed ARIMA model is used as a time series model.

[0145] After obtaining the time series model, operational data belonging to the same data dimension collected at different times are obtained from the vehicle network data to form the second model sample dataset. The time series model is trained using the time series data of the same data dimension in the second model sample dataset until the time series model converges, thus completing the training of the time series model.

[0146] The following is an embodiment of a vehicle anomaly identification method provided in this application.

[0147] Figure 3 A schematic flowchart of a vehicle anomaly identification method provided in an embodiment of this application is shown, such as... Figure 3 As shown, the vehicle anomaly identification method provided in this application is applied to electronic devices with computing power, such as computers. The vehicle anomaly identification method includes the following schemes:

[0148] S210: Obtain target data sets generated by the target vehicle at multiple set times to obtain multiple target data sets;

[0149] S220: Input the plurality of target data sets into the anomaly detection model, and have the anomaly detection model output the anomaly detection result;

[0150] S230: Obtain the vehicle component corresponding to the data to be identified in the anomaly identification result;

[0151] S240: The vehicle component is identified as an abnormal component of the target vehicle.

[0152] The target vehicle is the vehicle that needs to be detected for anomalies, and the set time refers to a historical moment. The target vehicle will generate operational data during its previous operation. From the operational data generated by the target vehicle during its previous operation, target data groups generated by the target vehicle at multiple set times are obtained, resulting in multiple target data groups. Each target data group includes data to be identified in multiple data dimensions. The data to be identified generated at different set times in the multiple target data groups but belonging to the same data dimension forms time series data.

[0153] After obtaining multiple target data sets, the multiple target data sets are input into the anomaly recognition model trained by the above model training method. The anomaly recognition model identifies the data with anomalies in the multiple target data sets and obtains the anomaly recognition results. The anomaly recognition results include the data to be identified in the multiple target data sets whose data category is anomaly.

[0154] The identification information of vehicle components is pre-associated with the corresponding operational data. When an anomaly occurs in the operational data, the vehicle component experiencing the anomaly can be located using the identification information corresponding to the abnormal operational data. Therefore, after obtaining the anomaly identification result, we can obtain the data to be identified that falls under the anomaly category from multiple target data groups. In other words, we obtain the data to be identified that is abnormal data from multiple target data groups. Then, we determine the identification information corresponding to the data to be identified in the anomaly identification result. The vehicle component corresponding to the determined identification information is the abnormal component of the target vehicle.

[0155] The vehicle anomaly identification method provided in this application obtains multiple target data sets generated by a target vehicle at multiple set time periods, each containing data to be identified across multiple dimensions. These multiple target data sets are input into an anomaly identification model, which outputs anomaly identification results for the data in the multiple target data sets that fall under the anomaly category. The method then identifies the vehicle components corresponding to the data in the anomaly identification results and identifies these components as abnormal parts of the target vehicle. This technical solution uses the anomaly identification model to identify abnormal data in the vehicle's operational data and locates abnormal components within the vehicle. Since the anomaly identification model is a fusion of a random forest model and a time series model, it combines the advantages of both, improving not only the accuracy of anomaly data identification but also the accuracy of locating abnormal vehicle components. This helps users promptly detect potential anomalies or malfunctions in the vehicle, helping to prevent accidents caused by component malfunctions. When an abnormal component is detected in the vehicle, the user can be notified promptly for timely vehicle repair or intervention. This vehicle anomaly identification method can be configured into the vehicle after-sales department's maintenance system, which can increase revenue for businesses, thereby increasing customer volume and preventing customer churn.

[0156] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0157] Figure 4 This application provides a schematic diagram of the structure of a model training device according to an embodiment of the present application. Figure 4 As shown, a model training device 400 is used for training an anomaly detection model. The anomaly detection model includes a pre-trained random forest model and a time series model, with the output layer of the random forest model connected to the input layer of the time series model. The model training device 400 includes:

[0158] The sample acquisition module 410 is used to acquire sample datasets corresponding to multiple vehicles. Each sample dataset for each vehicle includes multiple sample data groups, each sample data group includes sample data of multiple data dimensions, and each sample data has annotation information. The annotation information is used to indicate whether the data category of the sample data is normal or abnormal. The sample data belonging to the same data dimension in the sample dataset for each vehicle are collected at different times.

[0159] The first identification module 420 is used to input each sample data group in the sample dataset corresponding to each vehicle into the random forest model, and the random forest model identifies abnormal data in the sample data group based on the multi-dimensional characteristics of the input data, and outputs a first identification result; wherein, the first identification result includes first identification data and a first data category of the first identification data, and the first data category is an abnormal category;

[0160] The data selection module 430 is used to obtain sample data belonging to the same data dimension as the first identification data from the sample dataset corresponding to the sample data group, so as to obtain the target sample data;

[0161] The second identification module 440 is used to input the target sample data and the first identification data into the time series model, and the time series model identifies abnormal data in the target sample data based on the time series characteristics of the input data in the same data dimension, and outputs a second identification result; wherein, the second identification result includes the second identification data and the second data category of the second identification data, and the second data category is an anomaly category;

[0162] The model training module 450 is used to train the anomaly recognition model based on the first data category, the second data category, the annotation information of the first recognition data, and the annotation information of the second recognition data.

[0163] In one possible implementation, the sample acquisition module 410 is specifically used for: acquiring sample operation data generated by each vehicle at different times for each of the multiple vehicles, thus obtaining sample operation data for each vehicle at multiple times; setting annotation information for the sample operation data for each vehicle at multiple times; determining whether the difference between the first quantity of the first data and the second quantity of the second data in the sample operation data for each vehicle at multiple times is less than a preset threshold; wherein, the first data includes sample operation data whose data category is normal in the sample operation data for each vehicle at multiple times, and the second data includes sample operation data whose data category is abnormal in the sample operation data for each vehicle at multiple times; if so, then for the sample operation data for each vehicle at each time, according to the preset multiple data... The sample running data at each time point is categorized according to dimensions to obtain the sample data group corresponding to each time point; based on the sample data group corresponding to each time point, a sample dataset corresponding to each vehicle is generated to obtain the sample dataset corresponding to each vehicle; otherwise, the sample running data at multiple time points corresponding to each vehicle is preprocessed to adjust the difference between the first quantity of the first data and the second quantity of the second data in the sample running data at multiple time points corresponding to each vehicle to be less than a preset threshold; wherein, the preprocessing includes sampling processing or undersampling processing; the preprocessed sample running data at each time point corresponding to each vehicle is categorized according to preset multiple data dimensions to obtain the sample data group corresponding to each time point; based on the sample data group corresponding to each time point, a sample dataset corresponding to each vehicle is generated to obtain the sample dataset corresponding to each vehicle.

[0164] In one possible implementation, the model training module 450 includes:

[0165] The first training unit is used to determine whether the first identification data and the second identification data are the same; if so, the first identification result and the second identification result are fused to obtain a fused identification result; wherein, the fused identification result includes third identification data and a third data category of the third identification data, the third data category is an anomaly category, and the third identification data is the same as the first identification data; the anomaly identification model is trained based on the difference information between the annotation information of the third data category and the first identification data.

[0166] In one possible implementation, the first training unit, in fusing the first recognition result and the second recognition result to obtain the fused recognition result, is specifically configured to: determine the average value of the first recognition result and the second recognition result to obtain the fused recognition result; or, perform a weighted average processing on the first recognition result and the second recognition result to obtain the fused recognition result.

[0167] In one possible implementation, the model training module 450 further includes:

[0168] The second training unit is used to determine whether the first identification data and the second identification data are the same; if not, the anomaly identification model is trained based on the difference information between the annotation information of the first data category and the first identification data and the difference information between the annotation information of the second data category and the second identification data.

[0169] In one possible implementation, the model training device 400 further includes:

[0170] The model building unit is used to acquire, for any one of multiple vehicles, data generated by that vehicle at different times and belonging to the same data dimension, to obtain time series sample data; verify the stationarity of the time series sample data; if the stationarity verification of the time series sample data fails, perform differencing on the time series sample data to obtain the differencing order of the seasonal autoregressive moving average model and the stationary time series sample data corresponding to the time series sample data; draw autocorrelation plots and partial autocorrelation plots based on the stationary time series sample data; determine the autoregressive order and moving average order of the seasonal autoregressive moving average model based on the autocorrelation plots and partial autocorrelation plots; and construct the seasonal autoregressive moving average model based on the differencing order, the autoregressive order, and the moving average order to obtain the time series model.

[0171] It should be noted that the model training device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the model training method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the model training device and the model training method embodiments provided in the above embodiments belong to the same concept. Therefore, for details not disclosed in the device embodiments of this application, please refer to the above embodiments of the model training method of this application, which will not be repeated here.

[0172] Figure 5 This paper shows a schematic diagram of the structure of an abnormal data identification device provided in an embodiment of this application, such as... Figure 5 As shown, the abnormal data identification device 500 includes:

[0173] The data acquisition module 510 is used to acquire target data sets generated by the target vehicle at multiple set times, resulting in multiple target data sets; wherein, each target data set includes data to be identified in multiple data dimensions;

[0174] The data recognition module 520 is used to input the plurality of target data groups into the anomaly recognition model, and the anomaly recognition model outputs anomaly recognition results; wherein, the anomaly recognition model is trained by the above-mentioned model training method, and the anomaly recognition results include the data to be identified in the plurality of target data groups whose data category is anomaly category;

[0175] The anomaly determination module 530 is used to obtain the vehicle component corresponding to the data to be identified in the anomaly identification result, and to determine the vehicle component as the abnormal component of the target vehicle.

[0176] It should be noted that the abnormal data identification device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the abnormal data identification method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the abnormal data identification device and the abnormal data identification method embodiments provided in the above embodiments belong to the same concept. Therefore, for details not disclosed in the device embodiments of this application, please refer to the embodiments of the abnormal data identification method of this application, which will not be repeated here.

[0177] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0178] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown.

[0179] For example, such as Figure 6 As shown, the electronic device 600 includes a memory 601 and a processor 602. The memory 601 stores executable program code 6011, and the processor 602 is used to call and execute the executable program code 6011 to perform a model training method or a vehicle anomaly recognition method.

[0180] This embodiment can divide the electronic device into functional modules according to the above method example. For example, each module can correspond to a separate functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware. It should be noted that the module division in this embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0181] When each functional module is divided according to its corresponding function, the electronic device may include: a sample acquisition module, a first recognition module, a data selection module, a second recognition module, a model training module, a data acquisition module, a data recognition module, an anomaly determination module, etc. It should be noted that all relevant content of each step involved in the above method embodiments can be referenced to the functional description of the corresponding functional module, and will not be repeated here.

[0182] The electronic device provided in this embodiment is used to execute the above-described model training method or vehicle anomaly recognition method, and thus can achieve the same effect as the above-described implementation method.

[0183] When using integrated units, the electronic device may include a processing module and a storage module. The processing module is used to control and manage the operation of the electronic device. The storage module is used to support the execution of program code and data by the electronic device.

[0184] The processing module may be a processor or a controller, which can implement or execute various exemplary logic blocks, modules, and circuits as disclosed in this application. The processor may also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and microprocessors, etc., and the storage module may be a memory.

[0185] This embodiment also provides a computer-readable storage medium storing computer program code. When the computer program code is run on a computer, the computer executes the above-described related method steps to implement a model training method or a vehicle anomaly recognition method in the above embodiment.

[0186] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned related steps to implement a model training method or a vehicle anomaly recognition method as described in the above embodiment.

[0187] In addition, the electronic device provided in the embodiments of this application may specifically be a chip, component or module. The electronic device may include a connected processor and a memory. The memory is used to store instructions. When the electronic device is running, the processor may call and execute the instructions to make the chip execute a model training method or a vehicle anomaly recognition method in the above embodiments.

[0188] In this embodiment, the electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding model training method or vehicle anomaly recognition method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding model training method or vehicle anomaly recognition method provided above, and will not be repeated here.

[0189] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0190] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0191] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be defined by the scope of the claims.

Claims

1. A model training method, characterized in that, The anomaly detection model is used for training, and the anomaly detection model includes a pre-trained random forest model and a time series model, wherein the output layer of the random forest model is connected to the input layer of the time series model. The model training method includes: Obtain sample datasets corresponding to multiple vehicles; wherein, the sample dataset corresponding to each vehicle includes multiple sample data groups, each sample data group includes sample data of multiple data dimensions, each sample data has annotation information, the annotation information is used to indicate whether the data category of the sample data is normal category or abnormal category, and the sample data belonging to the same data dimension in the sample dataset corresponding to each vehicle are collected at different times; For each sample data group in the sample dataset corresponding to each vehicle, the sample data group is input into the random forest model. The random forest model identifies abnormal data in the sample data group based on the multi-dimensional characteristics of the input data and outputs a first identification result. The first identification result includes first identified data and a first data category of the first identified data, where the first data category is an abnormal category. The target sample data is obtained by acquiring sample data belonging to the same data dimension as the first identification data from the sample dataset corresponding to the sample data group; The target sample data and the first identification data are input into the time series model. The time series model identifies abnormal data in the target sample data based on the time series characteristics of the input data in the same data dimension, and outputs a second identification result. The second identification result includes second identification data and a second data category of the second identification data, and the second data category is an anomaly category. Determine whether the first identification data and the second identification data are the same; If so, the first identification result and the second identification result are fused to obtain a fused identification result; wherein, fusing the first identification result and the second identification result is to fuse the probability value of the first data category and the probability value of the second data category to obtain the fused identification result, the fused identification result includes third identification data and a third data category of the third identification data, the third data category is an anomaly category, and the third identification data is the same as the first identification data; The anomaly recognition model is trained based on the difference information between the annotation information of the third data category and the first identification data. If not, the anomaly recognition model is trained based on the differences between the annotation information of the first data category and the first identification data and the differences between the annotation information of the second data category and the second identification data.

2. The model training method according to claim 1, characterized in that, The process of obtaining the sample datasets corresponding to each of the multiple vehicles includes: For each of the multiple vehicles, obtain the sample operation data generated by each vehicle at different times to obtain the sample operation data of each vehicle at multiple times; Labeling information is set for the sample operation data of each vehicle at multiple time points; Determine whether the difference between the first quantity of the first data and the second quantity of the second data in the sample running data at multiple times corresponding to each vehicle is less than a preset threshold; wherein, the first data includes sample running data of the normal category in the sample running data at multiple times corresponding to each vehicle, and the second data includes sample running data of the abnormal category in the sample running data at multiple times corresponding to each vehicle. If so, then for the sample running data of each vehicle at each time moment, the sample running data at each time moment is classified according to multiple preset data dimensions to obtain the sample data group corresponding to each time moment; Based on the sample data group corresponding to each time moment, generate a sample dataset for each vehicle to obtain sample datasets for multiple vehicles. If not, the sample running data for each vehicle at multiple time points are preprocessed to adjust the difference between the first quantity of the first data and the second quantity of the second data in the sample running data for each vehicle at multiple time points to be less than a preset threshold; wherein, the preprocessing includes sampling processing or undersampling processing; The preprocessed sample operation data for each vehicle at each time point are categorized according to multiple preset data dimensions to obtain the sample data group corresponding to each time point. Based on the sample data group corresponding to each time moment, generate a sample dataset for each vehicle to obtain sample datasets for each of the multiple vehicles.

3. The model training method according to claim 1, characterized in that, The step of fusing the first recognition result and the second recognition result to obtain the fused recognition result includes: The average value of the first recognition result and the second recognition result is determined to obtain the fused recognition result; or, The first identification result and the second identification result are weighted and averaged to obtain the fused identification result.

4. The model training method according to claim 1, characterized in that, The model training method also includes: For any one of the multiple vehicles, obtain the data generated by that vehicle at different times that belong to the same data dimension to obtain time series sample data; Verify the stationarity of the time series sample data; If the stationarity verification of the time series sample data fails, the time series sample data is differentially processed to obtain the difference order of the seasonal autoregressive moving average model and the stationary time series sample data corresponding to the time series sample data. Autocorrelation and partial autocorrelation plots were drawn based on the stationary time series sample data. Based on the autocorrelation plot and the partial autocorrelation plot, determine the autoregression order and the moving average order of the seasonal autoregressive composite moving average model; The seasonal autoregressive composite moving average model is constructed based on the difference order, the autoregression order, and the moving average order to obtain the time series model.

5. A method for identifying vehicle anomalies, characterized in that, The vehicle anomaly identification method includes: The target data sets generated by the target vehicle at multiple set time periods are obtained, resulting in multiple target data sets; each target data set includes data to be identified in multiple data dimensions; The plurality of target data groups are input into the anomaly detection model, and the anomaly detection model outputs anomaly detection results; wherein, the anomaly detection model is trained by the model training method according to any one of claims 1 to 4, and the anomaly detection results include the data to be identified in the plurality of target data groups whose data category is anomaly; Obtain the vehicle component corresponding to the data to be identified in the anomaly identification results; The vehicle component is identified as an abnormal component of the target vehicle.

6. A model training device, characterized in that, The anomaly detection model is used for training, and the anomaly detection model includes a pre-trained random forest model and a time series model, wherein the output layer of the random forest model is connected to the input layer of the time series model. The model training device includes: The sample acquisition module is used to acquire sample datasets corresponding to multiple vehicles. Each sample dataset for each vehicle includes multiple sample data groups, each sample data group includes sample data of multiple data dimensions, and each sample data has annotation information. The annotation information is used to indicate whether the data category of the sample data is normal or abnormal. The sample data belonging to the same data dimension in the sample dataset for each vehicle are collected at different times. The first identification module is used to input each sample data group in the sample dataset corresponding to each vehicle into the random forest model, and the random forest model identifies abnormal data in the sample data group based on the multi-dimensional characteristics of the input data, and outputs a first identification result; wherein, the first identification result includes first identification data and a first data category of the first identification data, and the first data category is an abnormal category; The data selection module is used to obtain sample data belonging to the same data dimension as the first identification data from the sample dataset corresponding to the sample data group, so as to obtain the target sample data; The second identification module is used to input the target sample data and the first identification data into the time series model, and the time series model identifies abnormal data in the target sample data based on the time series characteristics of the input data in the same data dimension, and outputs a second identification result; wherein, the second identification result includes the second identification data and the second data category of the second identification data, and the second data category is an anomaly category; The model training module is used to determine whether the first identification data and the second identification data are the same; if so, the first identification result and the second identification result are fused to obtain a fused identification result; wherein, fusing the first identification result and the second identification result is to fuse the probability value of the first data category and the probability value of the second data category to obtain the fused identification result, the fused identification result includes third identification data and a third data category of the third identification data, the third data category being an anomaly category, and the third identification data being the same as the first identification data; the anomaly identification model is trained based on the difference information between the annotation information of the third data category and the first identification data; if not, the anomaly identification model is trained based on the difference information between the annotation information of the first data category and the first identification data and the difference information between the annotation information of the second data category and the second identification data.

7. An abnormal data identification device, characterized in that, The abnormal data identification device includes: The data acquisition module is used to acquire target data sets generated by the target vehicle at multiple set time periods, resulting in multiple target data sets; each target data set includes data to be identified in multiple data dimensions; A data recognition module is used to input the plurality of target data groups into an anomaly recognition model, and the anomaly recognition model outputs anomaly recognition results; wherein, the anomaly recognition model is trained according to the model training method according to any one of claims 1 to 4, and the anomaly recognition results include the data to be identified in the plurality of target data groups whose data category is anomaly category; The anomaly determination module is used to obtain the vehicle component corresponding to the data to be identified in the anomaly identification result, and to identify the vehicle component as an abnormal component of the target vehicle.

8. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable program code; A processor is configured to call and run the executable program code from the memory, causing the electronic device to perform the model training method as described in any one of claims 1 to 4 or the vehicle anomaly recognition method as described in claim 5.

Citation Information

Patent Citations

  • Data type recognition, model training and risk recognition method and device and equipment

    CN107391569A

  • Time series data anomaly detection method fusing LSTM and GAN

    CN110598851A