A Dynamic Driver Identity Recognition Method and System Based on Incremental Learning
By using incremental learning and sliding window technology in the driver identity recognition system to extract and process CAN bus data, the problem of low driver identity recognition efficiency in dynamic scenarios in the prior art is solved, and fast and accurate identity recognition and efficient model retraining are achieved.
Patent Information
- Application Number
- CN202210611924.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-05-31
AI Technical Summary
The existing driver identity recognition scheme is not efficient in dynamic scenarios, and the existing methods have efficiency and resource occupancy problems in the driver feature extraction and model retraining.
The dynamic driver identity recognition method based on incremental learning is adopted, and the CAN bus data is obtained from the OBD-II interface, the driver behavior characteristics are extracted using the BIR model, and an identity recognition model based on sliding window and incremental learning is constructed. The model adjustment and retraining are used to use replay, context threshold and knowledge distillation.
It realizes the rapid and accurate identification of driver identity in dynamic scenarios, improves the efficiency and resource utilization of model retraining, and is suitable for applications in dynamic scenarios.
Smart Images

Figure CN114987504B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of identity recognition, and particularly relates to a dynamic driver identity recognition method and system based on incremental learning. Background Art
[0002] Most of the existing driver identity recognition schemes are static single - recognition schemes based on keys or driver biometrics. When the data is leaked, it is vulnerable to attacks. Some researchers have noticed this and proposed some driver identity recognition schemes based on deep learning. However, these existing schemes have two defects: the extraction of driver biometrics is not obvious, and the existing schemes are not efficient in dynamic scenarios where new drivers need to be continuously added to the recognition system. In terms of driver feature extraction, the existing methods use data from the OBD - II interface or simulator as input, which contains a large number of irrelevant fields and cannot fully reflect the driver's behavior characteristics. Secondly, when the nth new driver needs to be added to the recognition system, the existing schemes need to retrain the entire model containing n drivers. As the number of added drivers increases, the retraining time of the recognition system gradually becomes longer, the occupied system resources become more and more, and the efficiency becomes worse, which is not suitable for application in dynamic scenarios. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a dynamic driver identity recognition method and system based on incremental learning in view of the above - mentioned deficiencies in the prior art. By using the BIR model, CAN bus data is obtained from the OBD - II interface, and feature extraction is performed on this data to achieve driver identity recognition in dynamic scenarios.
[0004] The present invention adopts the following technical solutions:
[0005] A dynamic driver identity recognition method based on incremental learning of the present invention includes the following steps:
[0006] S1. Collect CAN bus data generated during driving, convert the CAN bus data into decimal data to obtain driver behavior feature data containing m features; collect all data transmitted by the sensor within a time unit, take the average value of the collected data, and integrate the collected driver behavior feature data into t pieces of data to obtain a driver feature dataset of [t×m].
[0007] S2. Using a sliding window method, intercept continuous data within a period of time from the driver behavior feature dataset of size [t×m] obtained in step S1 to obtain driver behavior feature data slices; construct a driver identity recognition model based on incremental learning, and input the driver behavior feature data slices into the driver identity recognition model based on incremental learning to identify the driver's identity.
[0008] Specifically, in step S1, the sensors include a brake pedal, engine torque, steering wheel, lateral acceleration, throttle valve, yaw rate, and accelerator pedal.
[0009] Specifically, in step S2, the size of the driver behavior feature data slice is [300×m], and the step size of the sliding window is 60.
[0010] Specifically, in step S2, the identification of the driver is specifically as follows:
[0011] The driver identity recognition model based on incremental learning includes a main model M and a generator model G;
[0012] When a new driver numbered n intends to be added to the driver identity recognition model based on incremental learning that can recognize n - 1 drivers, the generator model G n-1 First generates the original replay data
[0013] Then, the generator model G n-1 Utilizes the original replay data And the feature data D of the driver numbered n who intends to be added to the recognition model n To retrain the generator model to obtain a new generator model G n ;
[0014] Next, use the new generator model G n To generate new replay data At this time, use a context threshold for gating, and select some new replay data according to a set probability Use the method of knowledge distillation to label the prediction results of the new replay data The knowledge distillation uses the prediction probabilities of all classes for labeling; use the new replay data And the feature data D of the driver numbered n who intends to be added to the recognition model n To retrain the main model M n-1 When the main model determines the classification result during the retraining process, use the method of knowledge distillation to label the result. Finally, obtain a new main model M that can recognize n drivers n Thus, the retraining process of the driver identity recognition model based on incremental learning is completed.
[0015] Further, in the convolutional kernel, the size of the filter is [21×m], the stride size is 1, and the number of channels is 256.
[0016] Further, the main model is a classifier with a basic neural network structure, and the training result of the main model is represented by the loss function L total Quantitatively represented, the loss function L total is:
[0017]
[0018] where, W n represents the loss weight of the driver labeled n, L current is the loss generated by the new driver behavior feature data to be recognized, L replay is the loss generated by the replay data.
[0019] Further, the generator model uses a variational autoencoder network for deep learning and shares the same convolutional layer and fully connected layer with the main model.
[0020] Further, in the generator model, x is used to represent the driver task to be added to the recognition model, P(x) represents the probability that the driver is correctly recognized, and the maximum value is determined by taking the logarithm of P(x). logP(x) is:
[0021] logP(x) = L b + KL(q(z|x)||P(z|x)
[0022]
[0023] where, z is a latent variable, representing the encoding result of the VAE network encoder used by the generator model, L b is the lower bound of logP(x), KL(q(z|x)||P(z|x)) is the KL divergence, q(z|x) represents the probability that the encoder outputs z when the input is x, P(x|z) is the probability that the decoder outputs x when the input is z, and P(z) is the probability that the latent variable z is correctly recognized.
[0024] Further, the loss L D during the distillation of driver n is:
[0025]
[0026] where, θ refers to the hyperparameter when the nth driver is training, and T is the temperature variable is the probability that the main model correctly recognizes the nth driver when the knowledge distillation temperature is set to T, is the conditional probability when the main model pre-defines the temperature as T.
[0027] In a second aspect, an embodiment of the present invention provides a dynamic driver identity recognition system based on incremental learning, including:
[0028] A data module that collects CAN bus data generated during driving, converts the CAN bus data into decimal data, and obtains driver behavior feature data containing m features; collects all data transmitted by sensors within a time unit, takes the average of the collected data, and integrates the collected driver behavior feature data into t pieces of data to obtain a driver feature dataset of [t×m];
[0029] An identification module that uses a sliding window method to intercept continuous data within a period of time from the driver behavior feature dataset of size [t×m] obtained by the data module to obtain a driver behavior feature data slice; constructs a driver identity recognition model based on incremental learning, inputs the driver behavior feature data slice into the driver identity recognition model based on incremental learning, and conducts identity recognition on the driver.
[0030] Compared with the prior art, the present invention has at least the following beneficial effects:
[0031] A dynamic driver identity recognition method based on incremental learning. Compared with existing driver identity recognition methods, this method uses an incremental learning algorithm to establish an open driver identity recognition model. Due to the openness of the incremental learning algorithm, in a dynamic scenario where new drivers to be identified continuously join the driver identity recognition model, the model can be adjusted in ways such as replay, context threshold, and knowledge distillation. The model does not need to store the driver data that has been added to the driver identity recognition model, and when a new driver to be identified joins the driver identity recognition model, the retraining process of the driver identity recognition model can be quickly completed, improving time and space utilization.
[0032] Furthermore, driver behavior features are reflected in CAN bus data, and different sensors have different weights for reflecting driver behavior. Through literature and experiments, seven types of sensor data with high weights for reflecting driver behavior features are selected, and these data are collected to extract driver behavior features and establish a driver behavior feature dataset. The driver identity recognition result is more accurate, the driver behavior feature dataset occupies less storage space, but can more fully reflect driver behavior features.
[0033] Furthermore, due to the continuity of driver behavior features, a section of driver behavior feature data needs to be selected for feature extraction. Therefore, a sliding window method can be used to extract driver behavior feature data slices. Through experiments, it is found that the size of the driver behavior feature data slice is [300×m], and the step size of the sliding window sliding is 60. A suitable sliding window size can fully reflect driver behavior features and obtain the best recognition effect.
[0034] Furthermore, the driver identity recognition model based on incremental learning is divided into two parts. The first part is the main model M, and the second part is the generator model G. The driver identity recognition model based on incremental learning is retrained by means of replay, context threshold, and knowledge distillation, without the need to save the driver data that has been added to the driver identity recognition model. And when a new driver to be recognized is added to the driver identity recognition model, the retraining process of the driver identity recognition model can be completed quickly, improving the utilization rate of time and space.
[0035] Furthermore, during the convolution process, the convolution kernel size needs to be set. During the collection and extraction of driver behavior feature data, 21 columns of data are obtained as driver behavior feature data. Therefore, the filter size is set to [21×m]. Then, the stride size is set to 1 and the number of channels is 256 to ensure the continuity of driver behavior features. Setting the filter size according to the driver behavior feature data results facilitates data preprocessing.
[0036] Furthermore, different weights W are used n to dynamically adjust the calculation result of the loss function, which can improve the recognition accuracy of the model.
[0037] Furthermore, the VAE network trains the encoder and decoder, which can learn the continuous representation of the input driver behavior feature data in the latent space, matching the continuity of driver behavior. The main model and the generator model share the data of the same convolutional layer and fully connected layer, which can ensure the consistency of data usage.
[0038] Furthermore, in the generator model, x is used to represent the driver task to be added to the recognition model, and P(x) represents the probability that the driver is correctly recognized by the generator model. The larger P(x) is, the better the recognition result is. Therefore, by taking the logarithm of P(x), its maximum value can be found; the probability that the driver is correctly recognized by the generator model is quantitatively represented.
[0039] Furthermore, knowledge distillation is to multiply a temperature variable T when training the soft target to generate labels, softening the target probability. Driver feature data often has a certain degree of similarity. Therefore, there is sometimes almost no probability difference in the prediction results. Previous prediction methods use a certain probability value to represent the prediction result, and this method is often called hard target prediction. Hard target prediction often only selects the result with the largest probability, ignoring the influence caused by the remaining results. Therefore, the prediction result can be represented by a prediction probability vector containing all possible categories, and this method is often called soft target prediction.
[0040] In summary, the present invention uses an incremental learning algorithm, which can accurately identify the driver's identity, complete the retraining process of the driver identity recognition model faster, improve the time utilization rate of the driver identity recognition solution, and use the incremental learning algorithm to improve the space utilization rate of the driver identity recognition solution.
[0041] Next, through the accompanying drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0042] Figure 1 is the flowchart of incremental learning;
[0043] Figure 2 is the flowchart of the dynamic driver identity recognition model based on incremental learning;
[0044] Figure 3 is the confusion matrix diagram of the recognition results of the dynamic driver identity recognition model based on incremental learning;
[0045] Figure 4 is the experimental result diagram of the comparison of similar incremental learning algorithms;
[0046] Figure 5 is the experimental result diagram of the comparison with the existing driver identity recognition solution, where (a) is vehicle V1 and (b) is vehicle V2. Detailed Embodiments
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] In the description of the present invention, it should be understood that the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0049] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0050] It should be further understood that the term "and / or" as used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally indicates an "or" relationship between the associated objects before and after.
[0051] It should be understood that although terms such as first, second, third, etc. may be used to describe preset ranges in the embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0052] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0053] Various structural schematic diagrams according to the disclosed embodiments of the present invention are shown in the drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures and their relative sizes and positional relationships are only exemplary, and may actually deviate due to manufacturing tolerances or technical limitations, and those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.
[0054] The present invention provides a dynamic driver identity recognition method based on incremental learning, which collects CAN bus data generated during driving, converts the CAN bus data into decimal data, and obtains driver behavior feature data containing m features; collects all data transmitted by sensors within a time unit, takes the average value of the collected data, integrates the collected driver behavior feature data into t pieces of data, and obtains a driver feature data set of [t×m]; intercepts continuous data from the driver behavior feature data set to obtain driver behavior feature data slices; constructs a driver identity recognition model based on incremental learning, inputs the driver behavior feature data slices into the driver identity recognition model based on incremental learning, and performs identity recognition on the driver. The present invention can accurately identify the driver's identity, complete the retraining process of the driver identity recognition model faster, and improve the time utilization rate of the driver identity recognition solution.
[0055] Please refer to Figure 1 , a dynamic driver identity recognition method based on incremental learning of the present invention includes the following steps:
[0056] S1. Extraction of feature data and generation of feature data set
[0057] S101. Extract feature data
[0058] First, n volunteers are recruited as drivers to drive a car for experiments. The CAN bus data generated during driving is collected through the OBD interface, and these obtained data are analyzed and processed using the CANalyse automotive analysis tool to convert the hexadecimal CAN bus data into decimal data.
[0059] Table Ⅰ 7 kinds of sensors
[0060]
[0061] According to experiments and previous work, it is found that the data availability of different sensors is different. Table Ⅰ lists the 7 sensors with the strongest availability. Each sensor has a unique ID and frequency, and the data segment lengths are also different. The data that fully represents the driver's behavior characteristics is identified and extracted according to the ID, so that the input data can better reflect the driver's behavior characteristics. Therefore, in this way, driver behavior feature data containing m features is obtained.
[0062] S102. Generate a sample set
[0063] It is found from Table I that the transmission frequencies of the seven selected sensor data in the CAN bus are different. For example, the data transmission frequency of the brake pedal is 20 ms / time, and the data transmission frequency of the engine is 125 ms / time. Therefore, 0.5 seconds is selected as a time unit to collect all the data transmitted by the sensors within one time unit and take the mean of these data. In this way, the collected driver behavior characteristic data can be integrated into t pieces of data. Finally, a driver characteristic data set of [t×m] can be obtained.
[0064] S2. Build an incremental learning model
[0065] S201. Data input based on a sliding window
[0066] Drivers' habits are often hidden in the data of continuous time, and the data at a single time point cannot fully reflect drivers' behavior habits. Therefore, in the way of using a sliding window, continuous data within a period of time are intercepted from the driver behavior characteristic data set of size [t×m] obtained, to get driver behavior characteristic data slices. Then, an appropriate sliding step of the sliding window is selected, which should not only ensure that each data slice contains the data of the previous data slice to make the data change more continuous, but also improve the recognition efficiency. A large number of experiments show that when the size of the driver behavior characteristic data slice is [300×m] and the sliding step of the sliding window is 60, the driver behavior characteristics can be fully reflected and the best recognition effect can be obtained.
[0067] S202. Build a driver identity recognition model based on incremental learning
[0068] A driver identity recognition scheme based on incremental learning is proposed, named G-DriverID. Compared with previous work, G-DriverID is more suitable for dynamic scenarios that need to continuously add new drivers to the model. G-DriverID uses methods such as knowledge distillation and context threshold to solve the problem of catastrophic forgetting and improve the model retraining efficiency.
[0069] Driver identity recognition model based on incremental learning
[0070] Please refer to Figure 1 , incremental learning is an open machine learning algorithm. Different from previous driver identity recognition schemes, when a new driver to be recognized needs to be added to the model, common driver identity recognition schemes often need to collect all the data of target drivers and use these data to retrain the entire model. Incremental learning is to use the data of the newly added driver to adjust the existing driver identity recognition model, and finally make the model able to recognize all drivers.
[0071] Driver identity recognition model recognition process based on incremental learning
[0072] See also Figure 2 , when a new driver numbered n is to be added to the driver identification model based on incremental learning that can identify n-1 drivers, the main model M n-1 Use the new replay data generated by the generator model and the data of driver number n to retrain the new main model M n , and finally all n drivers can be identified. The specific process is as follows:
[0073] In the input layer, the input 300×m driver feature data slice is first normalized and converted into a tensor;
[0074] Then, these tensors are put into the convolution layer for convolution operation; in the convolution kernel, the filter size is [21×m], the step size is 1, and the number of channels is 256. Through the activation function ReLU(·), a group of C1 neurons of [280×1] are finally obtained. In the fully connected layer, the features of C1 neurons are further extracted, and finally the F2 neurons of size [128×1] are obtained. At this point, the driver data D to be added to the driver identification model numbered n based on incremental learning is n All ready.
[0075] The driver identification model based on incremental learning is divided into two parts, the first part is the main model M, and the second part is the generator model G. When a new driver numbered n is to be added to the driver identification model based on incremental learning that can identify n-1 drivers, the generator model G n-- First generate the original replay data Next, the generator model G n-1 Using original playback data and the driver characteristic data D with the number n to be added to the recognition model n Retrain the generator model to obtain a new generator model G n ; Next, use the new generator model G n Generate new replay data At this time, the context threshold is used for gating, and part of the new replay data is selected according to the set probability. Use these parts to replay the data and the driver characteristic data D with the number n to be added to the recognition model n Retrain the main model M n-1 When the main model determines the classification results during the training process, it uses knowledge distillation to mark the results. Knowledge distillation does not directly mark the input driver to be identified as the driver with the highest probability based on the main model results, but uses the predicted probabilities of all categories to mark it. Finally, a new main model M that can identify n drivers is obtained.n , thus far, the retraining process of the driver identity recognition model based on incremental learning is completed.
[0076] Main model
[0077] The main model is a classifier with a basic neural network structure. The training result of the main model is represented by the loss function L total quantitatively, L total consists of two parts. One part of the loss is generated by the new driver behavior feature data to be recognized, and this data is recorded as L current , and the other part of the loss is generated by the generated replay data, denoted as L replay . The weights of these two losses are determined by the number of input tasks.
[0078] L total is expressed as:
[0079]
[0080] where, W n represents the weight of the driver marked as n.
[0081] In the main model, the mean squared error loss is selected as the function used in the loss functions L current and L replay , and the loss function L can be expressed as:
[0082]
[0083] where, y i represents the true value of the data of the i-th driver. y' i represents the predicted value of the data of the i-th driver.
[0084] Generator model
[0085] The generator model uses a variational autoencoder (VAE) network for deep learning and shares the same convolutional layer and fully connected layer with the main model. The loss function of the generator model is also similar to that of the main model, that is:
[0086]
[0087] The theory of the VAE network comes from the Gaussian mixture model. The input driver behavior feature data is defined as x, and z is the output of the encoder.
[0088] The probability that the driver's behavior feature data x is correctly recognized by the VAE network is P(x). The larger it is, the better the recognition effect. The P(x) can be quantitatively calculated by calculating the set loss function L, and it can be expressed as:
[0089]
[0090] P(x) = ∫ z P(z)P(x|z)dz
[0091] The VAE network mainly consists of two parts: an encoder and a decoder. These two parts are fully connected networks with two hidden layers, containing 400 non-linear units and ReLU(·). q(z|x) represents the probability that the encoder outputs z when the input is x, and P(x|z) represents the probability that the decoder outputs x when the input is z. The encoder outputs two sets of codes, μ and σ, to control the degree of noise interference.
[0092] The formula is transformed into:
[0093] logP(x) = L b + KL(q(z|x) || P(z|x)
[0094]
[0095] Among them, KL(q(z|x) || P(z|x)) is the KL divergence, which is used to represent the asymmetry between the two probability distributions P(z|x) and q(z|x). The KL divergence is always greater than 0.
[0096] Therefore, L b is the lower bound of logP(x); from the L b formula, it can be seen that by adjusting the encoder to minimize q(z|x), and then adjusting the decoder to maximize P(x|z), L b is maximized, making the predicted value closest to the true value.
[0097] Knowledge Distillation
[0098] Driver characteristic data often has a certain similarity. Therefore, the prediction results sometimes have almost no probability difference. Previous prediction methods use a certain probability value to represent the prediction result. This way is often called hard target prediction. Hard target prediction often only selects the result with the highest probability, ignoring the impact caused by the remaining results. Therefore, the prediction result can be represented by a prediction probability vector containing all possible categories. This way is often called soft target prediction. Knowledge distillation is to multiply a temperature variable T when generating labels for soft target training to soften the target probability. During testing, the temperature T is changed back to the original value. The loss L D during the distillation of driver n during training is expressed as:
[0099]
[0100] Among them, θ refers to the hyperparameter of the nth driver during training, and T is the temperature variable When the knowledge distillation temperature is set to T, the probability that the main model correctly identifies the nth driver is the conditional probability when the temperature is predefined as T for the main model.
[0101] Context threshold
[0102] In order to replay new data in the generator model The generation part implements context-sensitive processing. For each driver to be identified, a certain proportion of unit subsets are randomly selected from each hidden layer in the decoder network to achieve full gating. Using a preset hyperparameter, when conditional replay is executed and replay samples are generated, which class to use is determined according to the selected parameter. When there is no conditional replay, the class used when generating replay samples will be randomly selected from the classes seen so far. Finally, new replay data after context selection is obtained.
[0103] In another embodiment of the present invention, a dynamic driver identity recognition system based on incremental learning is provided. This system can be used to implement the above-mentioned dynamic driver identity recognition method based on incremental learning. Specifically, the dynamic driver identity recognition system based on incremental learning includes a data module and an identification module.
[0104] Among them, the data module collects CAN bus data generated during driving, converts the CAN bus data into decimal data, and obtains driver behavior feature data containing m features; collects all data transmitted by the sensor within a time unit, takes the average value of the collected data, and integrates the collected driver behavior feature data into t pieces of data to obtain a driver feature dataset of [t×m].
[0105] The identification module uses a sliding window method to intercept continuous data within a period of time from the driver behavior feature dataset of size [t×m] obtained by the data module to obtain driver behavior feature data slices; constructs a driver identity recognition model based on incremental learning, and inputs the driver behavior feature data slices into the driver identity recognition model based on incremental learning to identify the driver's identity.
[0106] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Usually, the components of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0107] Experimental environment
[0108] To verify the robustness of G-DriverID, experimental studies were conducted in two cars, V1 and V2. In addition, 15 volunteers were recruited to drive these two cars to collect CAN bus data during driving. Among these 15 volunteers, there were 10 male volunteers and 5 female volunteers, aged between 24 and 50 years old, with a driving experience of more than two years. During driving, an in-vehicle diagnostic tool, CANalyse, was used to collect and analyze CAN bus data from the OBD-II port.
[0109] To reduce the interference of external factors on the experimental results, a specific driving route needs to be formulated. The formulated route mainly includes three road conditions: congested road conditions, winding road conditions, and straight road conditions. In a specific experiment, the designed route is about 9 kilometers long as a whole. According to road regulations, the driving speed is lower than 70 kilometers per hour, and it takes about 20 minutes to drive one week. Driving data on 45 sunny days were collected to prevent the interference of weather and pedestrians. During this period, each driver needed to drive two experimental vehicles on the designated route for 2 laps every day. The collected data was stored in a database, and the data was processed using the method in the solution. Finally, data that can fully represent the driver's behavior characteristics was extracted, and a driver behavior dataset including 37,500 training data and 7,500 test data was established for each driver; 23 and 27 feature data were selected from car V1 and car V2 respectively.
[0110] Experimental results
[0111] (1) Driver identification results
[0112] A model based on incremental learning was used to implement driver identification. The training data was sourced from the prepared driver behavior feature dataset. The method of sliding window was used to extract driver behavior feature data slices, these data slices were transformed into tensors, and the model was trained.
[0113] First, the driving data of 5 drivers to be recognized is used to obtain an initialized model.
[0114] Then, new driver data is added to the model one by one. After training, the accuracy of each driver in the new model is obtained. As the number of recognized drivers in the model increases, the average recognition accuracy of the drivers also changes.
[0115] Finally, when the data of 15 drivers are all added to the model, the recognition results are represented in the form of a confusion matrix, as Figure 3 shown. The number above each space represents the recognition result of vehicle V1, and the number below each space represents the recognition result of vehicle V2. Among them, T represents the true label of the test data, D represents the prediction result of the test data, and the diagonal represents the proportion of accurately predicted driver data.
[0116] It can be seen from Figure 3 that in V1, the highest recognition rate of the drivers can reach 97.46%, and the lowest recognition rate is 90.19%. In V2, the highest recognition rate of the drivers can reach 96.55%, and the lowest recognition rate is 90.25%.
[0117] In addition, in the latest model, the accuracy of newly added authorized drivers is higher than that of previous authorized drivers. This is because during retraining, the catastrophic forgetting problem will reduce the accuracy of previous authorized drivers.
[0118] (2) Comparison experiment results of the same type of incremental learning algorithm
[0119] Please refer to Figure 4 , and conduct experimental comparisons with two other representative incremental learning models (Generative Replay (GR) and Elastic Weight Consolidation (EWC)) using the same training / test data set. First, these models are initialized with 5 authorized drivers, and the current average recognition accuracy is recorded. Then, the other 10 authorized drivers are added to these models one by one, and the average accuracy each time is recorded. It can be seen from Figure 5 that as more and more drivers are added to these models, the driver recognition accuracy will slowly decrease. However, the average accuracy of the G-DriverID model is higher than that of the GR and EWC models. In the experiment of V1, the average accuracy rates of 15 authorized drivers in the G-DriverID, GR, and EWC models are 93.51%, 91.08%, and 76.25% respectively. In the experiment of V2, the average accuracy rates of 15 authorized drivers in these three models are 92.84%, 90.4%, and 76.41% respectively. Thus, it can be seen that the present invention has superiority.
[0120] Although incremental learning models are highly efficient in retraining in dynamic scenarios, these models suffer from the problem of catastrophic forgetting. As the number of drivers increases, the accuracy of identifying the identities of the first drivers added to the model decreases. From Figure 5 It can be seen that in V1, for the drivers with the first three digits of the number, the accuracies of the G-DriverID model are 90.9%, 91.9%, and 90.1% respectively. The accuracies of the GR model are 87.7%, 88.3%, and 88.7% respectively. In the EWC model, the accuracies are 59.7%, 61.1%, and 61.3% respectively. In V2, the accuracies of these three drivers in the G-DriverID model are 91.8%, 91.9%, and 90.6% respectively. In the GR model, the ratios are 86.6%, 87.3%, and 87.4% respectively. In the EWC model, the ratios are 60.8%, 62.4%, and 64.7% respectively. The experimental results show that compared with other incremental learning models, the recognition effect of G-DriverID is generally good.
[0121] (3) Comparison experimental results with existing driver identity recognition schemes
[0122] Compared with existing driver identity recognition schemes, when the model needs to be extended, the scheme has significant advantages. Among existing driver identity recognition schemes, two typical models, SVM and CNN, were selected for experimental comparison.
[0123] First, these models were trained using the data of five drivers. When new drivers to be identified need to be added to these models, CNN and SVM need to retrain the entire model, but the G-DriverID scheme only needs to perform incremental retraining on the previous model.
[0124] Add 15 drivers to these models in turn according to the above method, and record the running time of model retraining. From Figure 5 it can be seen that in two cars, V1 and V2, as the number of authorized drivers increases, the advantages of G-DriverID become more and more prominent, and the retraining running time is much lower than that of the CNN and SVM models. Therefore, in actual scenarios, G-DriverID has better application value than other models.
[0125] In summary, a dynamic driver identity recognition method and system based on incremental learning according to the present invention has the following characteristics:
[0126] First, the present invention conducts experiments in real cars, and the data acquisition and driver identity recognition results are true and reliable; the test results show that the method provided by the present invention can accurately identify the driver's identity, and has strong usability and practicality.
[0127] Second, the present invention further analyzes and processes the collected CAN bus data, extracts the CAN bus data that can better reflect the driver's characteristics, reduces the influence caused by irrelevant data, and makes the data source more representative.
[0128] Third, compared with the driver identity recognition solutions used in the existing industry, the present invention implicitly recognizes the driver's identity through the driver characteristic data, and the recognition result is more accurate, difficult to be attacked, and has strong security.
[0129] Fourth, compared with other machine learning driver identity recognition solutions, the present invention can not only accurately recognize the driver's identity, but also significantly improve the efficiency of the system re-training when a new driver joins the driver identity recognition system.
[0130] Fifth, the present invention conducts experiments in two cars produced by different manufacturers, considering the influence of different vehicles on the recognition result, and the experimental result has robustness.
[0131] Sixth, the present invention can be applied to long-distance buses, armored cars and other special vehicles with frequent driver changes, improve the efficiency of the recognition system when a new driver joins, and improve driving safety.
[0132] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0133] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0134] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks.
[0135] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 one or more of the processes and / or blocks Figure 1 specified in one or more of the blocks.
[0136] The above is only to illustrate the technical idea of the present invention and should not be used to limit the protection scope of the present invention. Any modifications made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A dynamic driver identity recognition method based on incremental learning, characterized in that, Including the following steps: S1. Collect the CAN bus data generated during driving, convert the CAN bus data into decimal data, and obtain driver behavior characteristic data containing m characteristics. Collect all the data transmitted by the sensor within a time unit, take the mean of the collected data, and integrate the collected driver behavior characteristic data into t pieces of data to obtain a t×m driver characteristic data set; S2, using the sliding window method, from step S1 obtained [ t×m ] size of driver behavior characteristic data set to intercept continuous data within a period of time to obtain driver behavior characteristic data slices; Construct a driver identity recognition model based on incremental learning, input the driver behavior feature data slices into the driver identity recognition model based on incremental learning, and perform driver identity recognition. The specific process of performing driver identity recognition is as follows: At the input layer, first, the input 300×m driver feature data slice is normalized and transformed into a tensor; then the tensor is put into the convolutional layer for convolutional operation; In the convolutional kernel, the filter size is 21×m , the stride size is 1, and the number of channels is 256; it is activated by the activation function ReLU(·), and finally a group of 280×1 C1 neurons are obtained; In the fully connected layer, extract the features of the C1 neurons, and finally obtain the F2 neurons with a size of 128×1 , which are to be added to the driver data of the driver identification model numbered n based on incremental learning Preparation completed; The driver identity recognition model based on incremental learning includes a main model M and a generator model G; When a new driver numbered n intends to be added to the incremental learning-based driver identity recognition model that can recognize n - 1 drivers, the generator model first generates the original replay data ; Generator model Using the original replay data and the driver feature data numbered n to be added to the recognition model Retrain the generator model to obtain a new generator model ; Use a new generator model Generate new replay data , at this time, use the context threshold for gating, and select part of the new replay data according to the set probability , use knowledge distillation to label the new replay data 's prediction results, and knowledge distillation uses the prediction probabilities of all categories for labeling; use the new replay data and the driver feature data numbered n to be added to the recognition model Retrain the main model , when the main model determines the classification result during the retraining process, use knowledge distillation to label the result Finally, a new main model for identifying n drivers is obtained. , and the retraining of the driver identification model based on incremental learning is completed.
2. The dynamic driver identity recognition method based on incremental learning according to claim 1, characterized in that, In step S1, the sensors include a brake pedal, engine torque, a steering wheel, lateral acceleration, a throttle valve, yaw rate, and an accelerator pedal.
3. The dynamic driver identity recognition method based on incremental learning according to claim 1, characterized in that, In step S2, the size of the driver behavior feature data slice is 300×m , and the step size for the sliding window to slide is 60.
4. The dynamic driver identity recognition method based on incremental learning according to claim 1, characterized in that, The main model is a classifier with a basic neural network structure, and the training result of the main model is represented by a loss function in a quantized manner, and the loss function is as follows: Among them, represents the loss weight of the driver marked as n , the loss generated for the new driver behavior feature data to be recognized, and is the loss generated for the replay data.
5. The dynamic driver identity recognition method based on incremental learning according to claim 1, characterized in that, The generator model uses a variational autoencoder network for deep learning and shares the same convolutional layer and fully connected layer with the main model.
6. The dynamic driver identity recognition method based on incremental learning according to claim 1, characterized in that, In the generator model, represents the driver task to be added to the recognition model, represents the probability that the driver is correctly recognized. By finding the logarithm to determine the maximum value, which is: Among them, is a latent variable, representing the encoding result of the VAE network encoder used by the generator model, is the lower limit of is the KL divergence, represents the probability that the encoder outputs z when the input is x, is the probability that the decoder outputs x when the input is z, is the probability that the latent variable z is correctly recognized.
7. The dynamic driver identity recognition method based on incremental learning according to claim 1, characterized in that, Loss distilled during driver n training is as follows: Among them, θ refers to the hyperparameter when the nth driver is in training, is the temperature variable is the probability that the main model correctly identifies the nth driver when the knowledge distillation temperature is set to T, is the conditional probability when the main model pre-defines the temperature as T.
8. A dynamic driver identity recognition system based on incremental learning, characterized in that, Including: The data module collects CAN bus data generated during driving, converts the CAN bus data into decimal data, and obtains driver behavior feature data containing m features; Collect all the data transmitted by the sensor within a time unit, take the mean of the collected data, and integrate the collected driver behavior characteristic data into t pieces of data to obtain a t×m driver characteristic data set; The recognition module uses a sliding window method to intercept continuous data within a period of time from the driver behavior feature dataset of t×m size obtained from the data module, and obtains a driver behavior feature data segment; Construct a driver identity recognition model based on incremental learning, input the driver behavior feature data slices into the driver identity recognition model based on incremental learning, and perform driver identity recognition. The specific process of performing driver identity recognition is as follows: At the input layer, first, the input 300×m driver feature data slice is normalized and converted into a tensor; then the tensor is put into the convolutional layer for convolutional operation; In the convolutional kernel, the filter size is 21×m , the stride size is 1, and the number of channels is 256; it is activated by the activation function ReLU(·), and finally a set of 280×1 C1 neurons are obtained; In the fully connected layer, extract the features of the C1 neurons, and finally obtain the F2 neurons with a size of 128×1 , which are to be added to the driver data of the driver identity recognition model numbered n based on incremental learning Preparation completed; The driver identity recognition model based on incremental learning includes a main model M and a generator model G; When a new driver with number n intends to be added to the incremental learning-based driver identification model that can identify n - 1 drivers, the generator model first generates the original replay data ; Generator model Using the original replay data and the driver feature data numbered n to be added to the recognition model Retrain the generator model to obtain a new generator model ; Use a new generator model Generate new replay data At this time, use the context threshold for gating and select some of the new replay data according to the set probability Mark the new replay data using knowledge distillation The prediction results of, knowledge distillation are marked using the prediction probabilities of all classes; use the new replay data And the driver feature data numbered n to be added to the recognition model Retrain the main model When the main model determines the classification result during retraining, mark the result using knowledge distillation Finally, a new main model for identifying n drivers is obtained , and the retraining of the driver identification model based on incremental learning is completed.
Citation Information
Patent Citations
Driver identity authentication method based on convolutional neural network and support vector domain description
CN110598734A
Image big data-oriented class increment classification method, system and device and medium
CN112990280A