Mileage robust calculation method based on LSTM anomaly identification and compensation

By building an LSTM anomaly identification and compensation model, the problem of inaccurate mileage estimation caused by location data anomalies is solved, and efficient and accurate mileage calculation is achieved, which is suitable for free-flow toll collection systems.

CN120632709APending Publication Date: 2025-09-12SOUTHEAST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510619669.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing location data anomaly detection methods cannot effectively capture time series characteristics, resulting in inaccurate mileage estimation and a lack of compensation mechanism for abnormal data, which affects the accuracy and continuity of the free-flow toll collection system.

Method used

An LSTM neural network training set is used to generate vehicle driving route data. Through data enhancement, normalization and time series formatting preprocessing, an LSTM anomaly identification and compensation model is constructed. The forget gate, input gate and output gate structures are used to capture time series features. Anomaly thresholds are set for detection and data compensation. Finally, mileage is calculated through Gauss-Krüger projection transformation.

Benefits of technology

It improves the accuracy and robustness of mileage calculation, enhances the generalization ability of the model, reduces computational complexity, meets real-time application requirements, and ensures the continuity and accuracy of location data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632709A_ABST
    Figure CN120632709A_ABST
Patent Text Reader

Abstract

The invention provides a mileage robust calculation method based on LSTM (Long Short Term Memory) anomaly identification and compensation, which comprises the following steps of: firstly, creating a driving track enhanced data set, then preprocessing the data, respectively determining characteristic attributes of a training set, and carrying out normalization and time sequence formatting on the data, so as to construct an LSTM-based anomaly data identification and compensation model; the model is trained to detect abnormal data in the longitude and latitude position data, compensation is performed by using a position time sequence predicted value to generate a continuous and accurate position data sequence, and the Euclidean distance is calculated after Gaussian-Kruger projection conversion is adopted on the basis of the sequence, so that the accurate driving mileage is calculated. Abnormal data generated in the positioning process can be effectively processed, and the method is suitable for the fields of navigation, automatic driving, free flow charging and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a mileage robust estimation method based on LSTM anomaly identification and compensation, and belongs to the technical field of vehicle navigation. Background Art

[0002] With the widespread adoption of technologies like the Global Positioning System (GPS), mileage calculation based on location data has become increasingly important in areas such as free-flow tolling. This next-generation system, which uses positioning technology to calculate vehicle mileage charges, is based on internet-based toll collection. However, existing location data collected can contain anomalies. These anomalies can be caused by position calculation errors in complex road environments such as tunnel entrances and exits and in cities; they can also be caused by aging BDS receiving equipment; or even by human negligence. The presence of anomalies in the location time series can lead to significant errors in mileage calculations.

[0003] Traditional anomaly detection algorithms are mostly implemented using multiple iterations, which requires a large amount of computation and takes a long time when encountering large amounts of data. Existing anomaly detection methods, such as statistical thresholding methods or simple machine learning models, often fail to fully capture the time series characteristics of location data, resulting in poor detection results. In addition, these methods generally do not provide a compensation mechanism for abnormal data and cannot guarantee the continuity and accuracy of the calculation results. Therefore, it is necessary to develop a new method that can effectively detect and compensate for anomalies in location data, thereby providing more accurate location data for subsequent mileage calculations. Designing a reliable location anomaly data identification and compensation algorithm is of great significance to the accuracy of free-flow charging. Summary of the Invention

[0004] This paper provides a robust mileage estimation method based on LSTM anomaly identification and compensation, aiming to address the issue of inaccurate mileage estimation caused by location data anomalies. It provides a new solution for accurate and reliable mileage estimation and for free-flow toll collection in smart transportation.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] Step S1: Data generation

[0007] The LSTM neural network training set contains the longitude and latitude data of all location points on the vehicle's driving route. The training data set is generated based on the Baidu Map Developer Platform.

[0008] Sub-step S11: Create a driving instance

[0009] Create a driving navigation instance and use the search function to retrieve the shortest path based on the latitude and longitude coordinates of the starting and ending points.

[0010] Sub-step S12: Data encoding

[0011] Based on the trajectory data created by the search function, separate the individual longitude and latitude data with "|", and separate the longitude and latitude data of different points with "," to achieve data encoding;

[0012] Sub-step S13: Data enhancement and segmentation

[0013] Random noise is added to the trajectory data generated by data encoding to replicate multiple trajectories for data enhancement. Gaussian noise with a mean of 0 and a standard deviation of 0.001 degrees is added to generate training set trajectories and test set trajectories in a 9:1 ratio.

[0014] Step S2: Data preprocessing

[0015] Sub-step S21: Determine the characteristic attributes of the training set

[0016] According to the relationship between vehicle location prediction and vehicle data information, the most basic timestamp, longitude, and latitude are first selected as the feature attribute set; the data collection frequency is set according to application requirements; the N trajectory position points based on the timestamp, longitude, and latitude feature attribute set are represented as

[0017] P i =<t i ,lon i ,lat i >,i=1,2,…,N (1)

[0018] where t i Indicates the timestamp attribute corresponding to the i-th position, lon i ,lat i They represent the longitude and latitude attributes corresponding to the i-th position respectively; the default imported data is arranged by timestamp, so the location data dimension is (N, 2);

[0019] Sub-step S22: Data normalization

[0020] The “min-max normalization” algorithm is used, as shown in the following formula, to perform linear changes on the original data set;

[0021]

[0022] Where max(x) and min(x) represent the maximum and minimum values ​​in the position sequence x respectively;

[0023] Sub-step S23: Time series formatting

[0024] The timestep position data points are passed into the input layer as training samples, and the timestep+1 position data points are used as training labels. The prediction error is calculated to update the model parameters. The sliding window data is segmented to generate training samples with a dimension of (N-timestep, timestep, 2) and training labels with a dimension of (N-timestep, 2).

[0025] Obtain N-timestep trajectory segments, each of which consists of timestep longitude and latitude data points, forming the above three-dimensional training set; the N-timestep trajectory segments obtained by segmentation are shuffled and randomly arranged, and the N-timestep training labels are also shuffled in the same order as the training samples; after the above sequence segmentation, the training set P = [p1, p2, ..., p i ,…p N-timestep ], the training set after random permutation is expressed as P'=[p N-timestep ,p i ,…,p2,…p1]; input the training set P' into the LSTM trajectory prediction model for training;

[0026] Step S3: LSTM anomaly identification and compensation model construction

[0027] Sub-step S31: Construct input layer

[0028] The input layer receives fixed-length time series information. The transmitted data is timestep vehicle trajectory points, each of which contains a set of vehicle attributes at that point. Each time a sample is input, the latitude is (timestep, 2), and there are N samples in total.

[0029] Sub-step S32: Construct hidden layer 1

[0030] Hidden layer 1 adds three forget gates, input gate and output gate structures;

[0031] Sub-step S33: Construct dropout layer

[0032] Temporarily discard some neurons in the LSTM neural network from the network;

[0033] Sub-step S34: Construct hidden layer 2

[0034] Hidden layer 2 adds three forget gates, input gate and output gate structures;

[0035] Sub-step S35: Construct output layer

[0036] The output layer calculates the predicted point position by connecting the weights with the last hidden layer. The number of output neurons is 2, and the predicted label dimension is (1, 2). The error between the output predicted position and the actual position observation value is used to determine whether the position data is abnormal, and the abnormal data is compensated by the predicted value.

[0037] Step S4: LSTM anomaly identification and compensation model training

[0038] Sub-step S41: Forward propagation

[0039] (2) Forget gate (f t ): The forget gate calculates the result through the state value of the neuron at the previous moment, the output value at the previous moment and the input value at the current moment;

[0040]

[0041] Among them, x t 、h t-1 They represent the input neurons at time t and the hidden neurons at time t-1 respectively; W fx 、W fh and b f Respectively represent the forget gate corresponding to x t 、h t-1 and the weight matrix of the bias vector; σ is the sigmoid function;

[0042] (2) Input gate (i t ): After the calculation of formula (4), it is decided whether to discard the input information or input it into the neuron;

[0043]

[0044] Among them, W ix 、W ih and b i Respectively represent the input gate corresponding to x t 、h t-1 and the weight matrix of the bias vector; g t represents the cell state at time t, W gx 、W gh and b g Respectively represent the current moment input corresponding to x t 、h t-1 and the weight matrix of the bias vector; tanh is the hyperbolic tangent function; calculate the unit state C at the current moment t

[0045] C t =C t-1 ⊙f t +g t ⊙i t(6)

[0046] where represents the Hadamard product;

[0047] (3) Output gate (o t ): The following operation is used to determine whether to output information to the next neuron;

[0048]

[0049] Among them, W ox 、W oh and b o Respectively represent the output gate corresponding to x t 、h t-1 and the weight matrix of the bias vector; calculate the hidden layer output h t

[0050] h t =o t ⊙tanh(c t ) (8)

[0051] Sub-step S42: Backward propagation

[0052] The prediction error is propagated to the previous moment along time and to the previous layer along the neural network, and the calculation of each weight gradient and bias gradient is derived:

[0053]

[0054] Among them, η∈{f,i,g,o}, E is the cost function; f,t , δ i,t , δ g,t , δ o,t The BPTT error calculation process is as follows:

[0055]

[0056] Step S5: Identification and compensation of abnormal position data

[0057] Sub-step S51: abnormal data generation

[0058] An increment of ±0.01 is randomly added to the longitude and latitude data as a trajectory segment containing abnormal position data;

[0059] Sub-step S52: Setting the abnormality threshold

[0060] Based on the prediction error distribution of training data and abnormal data, set the threshold for anomaly detection; calculate the Euclidean distance between each predicted point and the true point, and then calculate the mean μ and standard deviation σ of the error, and set the threshold to μ+2σ;

[0061] Sub-step S53: Abnormal data identification and compensation process

[0062] The preprocessed new data is fed into the trained LSTM model to generate predicted longitude and latitude for each data point. The error between the actual location data and the predicted value is calculated, and the error for each data point is compared with the preset anomaly threshold. If the error is greater than the threshold, the data point is marked as an anomaly; otherwise, it is marked as normal. For detected anomaly data points, the predicted value of the LSTM model is used to compensate.

[0063] Step S6: Robust mileage calculation

[0064] Sub-step S61: Gauss-Krüger projection conversion

[0065] The compensated latitude and longitude data are converted into plane Cartesian coordinates (x, y) using the Gauss-Krüger projection;

[0066] Sub-step S62: Calculate the distance between adjacent data points

[0067] For two adjacent data points (i, i+1) in the time series, the Euclidean distance is calculated in the plane coordinate system:

[0068]

[0069] Among them, (x i ,y i ) and (x i +1,yi+1) are the Cartesian coordinates of data points i and i+1;

[0070] Sub-step S63: Accumulate total mileage

[0071] Add up the distances of all adjacent data points to get the total mileage:

[0072]

[0073] Where N is the total number of data points.

[0074] The present invention achieves the following beneficial effects through the above technical solutions:

[0075] 1. This invention improves the accuracy of mileage calculation. It uses an LSTM neural network to identify anomalies in longitude and latitude position data and uses predicted values ​​to compensate for them, effectively eliminating data deviations caused by signal interference or sensor failures, ensuring the continuity and accuracy of position data, and thus significantly improving the accuracy of mileage calculation.

[0076] 2. Enhanced robustness of the method: By using preprocessing techniques such as data augmentation, normalization, and time series formatting, combined with the LSTM model's ability to capture time series features, this method can adapt to data fluctuations in complex driving scenarios, improving the model's generalization and noise resistance.

[0077] 3. Efficient data processing: Through sliding window segmentation and random permutation of training data, along with optimized LSTM model training parameters, the system reduces computational complexity while ensuring anomaly identification accuracy, meeting the needs of real-time applications. This beneficial effect is achieved through the synergy of data generation, preprocessing, model building and training, anomaly identification and compensation, and mileage calculation, significantly improving the overall performance of location data processing and mileage estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 Flowchart of the robust mileage estimation steps;

[0079] Figure 2 formatting process diagrams for time series;

[0080] Figure 3 This is the framework diagram of the LSTM abnormal data identification and compensation model;

[0081] Figure 4 This is a flowchart of location abnormal data identification and compensation based on LSTM. DETAILED DESCRIPTION

[0082] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0083] like Figure 1 As shown, the present invention provides a mileage robust estimation method based on LSTM anomaly identification and compensation. First, a driving trajectory enhanced dataset is created, and then the data is preprocessed to determine the characteristic attributes of the training set and normalize and format the data in time series. Then, an LSTM-based abnormal data identification and compensation model is constructed. The model is trained to detect abnormal data in longitude and latitude position data, and the position time series prediction value is used for compensation to generate a continuous and accurate position data sequence. The Euclidean distance is calculated based on the sequence after Gauss-Krüger projection transformation, so as to estimate the accurate mileage. The present invention can effectively process abnormal data generated in the positioning process and is applicable to fields such as free-flow charging. The implementation process is divided into six main steps, each of which contains detailed sub-steps to ensure the clarity and operability of the method.

[0084] Step S1: Data generation

[0085] The LSTM neural network training set must include the latitude and longitude data of all location points on the vehicle's driving route. The training dataset is generated based on the Baidu Map Developer Platform.

[0086] Sub-step S11: Create a driving instance

[0087] Create a driving navigation instance and use the search function to initiate a search to obtain the shortest path based on the latitude and longitude coordinates of the starting and ending points.

[0088] Sub-step S12: Data encoding

[0089] Based on the trajectory data created by the search function, separate the individual longitude and latitude data with "|" and separate the longitude and latitude data of different points with "," to achieve data encoding.

[0090] Sub-step S13: Data enhancement and segmentation

[0091] Based on the generated trajectory data, random noise is added to replicate multiple trajectories to achieve data augmentation and improve the network's generalization ability. Gaussian noise with a mean of 0 and a standard deviation of 0.001 degrees is added, and training and test trajectories are generated in a 9:1 ratio. For example, 90 trajectories with random noise are generated as a training set for LSTM network training, and 10 trajectories with random noise are generated as a test set for LSTM prediction accuracy testing.

[0092] Step S2: Data preprocessing

[0093] The LSTM training data generated above needs to be further preprocessed to ensure that the model can learn the normal time series pattern of the location data. The specific steps are as follows:

[0094] Sub-step S21: Determine the characteristic attributes of the training set

[0095] Before training the neural network model, it is necessary to determine the feature attributes used in the training set. Based on the relationship between vehicle position prediction and vehicle data information, the most basic timestamp, longitude, and latitude are first selected as the feature attribute set. These data should reflect the movement trajectory of the vehicle or equipment in a normal environment, such as driving data on urban roads or highways. The data collection frequency can be set according to application requirements, for example, once per second. The N trajectory position points based on the timestamp, longitude, and latitude feature attribute set can be expressed as

[0096] P i =<t i ,lon i ,lat i >,i=1,2,…,N (1)

[0097] where ti Indicates the timestamp attribute corresponding to the i-th position, lon i ,lat i They represent the longitude and latitude attributes corresponding to the i-th position respectively. By default, the imported data is arranged by timestamp, so the location data dimension is (N, 2).

[0098] Sub-step S22: Data normalization

[0099] To improve the generalization ability of the LSTM neural network in identifying data anomalies and accelerate the convergence of neural network training, the training data needs to be normalized. The "min-max normalization" algorithm, as shown in the following formula, is used to linearly transform the original data set so that each normalized data point falls between 0 and 1.

[0100]

[0101] Where max(x) and min(x) represent the maximum and minimum values ​​in the position sequence x, respectively.

[0102] Sub-step S23: Time series formatting

[0103] Since the LSTM neural network model is trained based on historical time series data, it is necessary to segment the sequence data through a sliding window with a timestamp step of timestep to generate training samples and corresponding labels for model training. That is, the timestep position data points are passed into the input layer as training samples, and the timestep+1 position data points are used as training labels to calculate the prediction error and update the model parameters. Assuming timestep = 10, it means that the position information of the 11th data point is predicted using 1 to 10 position data points; and the sliding window is continuously moved, and the position information of the 12th data point is predicted using 2 to 11 position data points. And so on, until it stops at the N-timestep position point. After the sliding window data segmentation, training samples with dimensions of (N-timestep, timestep, 2) and training labels with dimensions of (N-timestep, 2) can be generated. The sliding window dynamic training set extraction process is shown in the attached figure. Figure 2 shown.

[0104] Through the above sliding window sequence segmentation, N-timestep trajectory segments are obtained. Each trajectory segment consists of timestep longitude and latitude data points, forming the above three-dimensional training set. In order to improve the robustness of the training LSTM trajectory prediction model, the N-timestep trajectory segments obtained by segmentation will be shuffled and randomly arranged, and the N-timestep training labels will also be shuffled in the same order as the training samples. For example, the training set P = [p1, p2, ..., pi ,…p N-timestep ], then the training set after random arrangement can be expressed as P'=[p N-timestep ,p i ,…,p2,…p1]. Inputting the training set P' into the LSTM trajectory prediction model for training can greatly improve the robustness of trajectory prediction.

[0105] Step S3: LSTM anomaly identification and compensation model construction

[0106] After preparing the training data, we start to build the LSTM neural network model for anomaly detection and data prediction. The LSTM abnormal data identification and compensation model framework is shown in the attached figure. Figure 3 As shown in the figure, LSTM adds a special gate structure to capture the long-term and short-term dependencies of time series, which has great advantages in processing traffic trajectory data with time series characteristics. The specific steps are as follows:

[0107] Sub-step S31: Construct input layer

[0108] The input layer receives fixed-length time series information. The data it transmits is timestep vehicle trajectory points, each of which contains a set of vehicle attributes. Therefore, the number of neurons in the input layer is equal to the number of attributes (for longitude and latitude data only, the number of neurons is 2). Since the LSTM neural network time step is timestep, each input sample has a latitude of (timestep, 2), for a total of N samples.

[0109] Sub-step S32: Construct hidden layer 1

[0110] Hidden layer 1 adds three forget gates, input gate and output gate structures to complete the learning and memory of sequence time features.

[0111] Sub-step S33: Construct dropout layer

[0112] The dropout layer sets a certain probability to temporarily discard some neurons in the LSTM neural network from the network, so that the model has strong nonlinear learning ability without overfitting.

[0113] Sub-step S34: Construct hidden layer 2

[0114] Three forget gates, input gates, and output gate structures are added to the hidden layer 2 to further enhance the learning and memory of sequence time features.

[0115] Sub-step S35: Construct output layer

[0116] The output layer calculates the predicted point position by connecting the weights with the last hidden layer. The number of output neurons is 2, and the predicted label dimension is (1, 2). The error between the output predicted position and the actual position observation is used to determine whether the position data is abnormal, and the predicted value is used to compensate for the abnormal data.

[0117] Step S4: LSTM anomaly identification and compensation model training

[0118] Sub-step S41: Forward propagation

[0119] The memory module information is controlled through three gate structures to coordinate the retention and forgetting of information in different neurons:

[0120] (3) Forget gate (f t ): This gate primarily determines whether a neuron retains or forgets its previous historical data. The calculation process is shown in the following equation. The forget gate calculates the result using the neuron's previous state, output, and current input. The result ranges from 0 to 1, with 1 indicating information retention and 0 indicating information forgetfulness.

[0121]

[0122] Among them, x t 、h t-1 They represent the input neurons at time t and the hidden neurons at time t-1 respectively; W fx 、W fh and b f Respectively represent the forget gate corresponding to x t 、h t-1 and the weight matrix of the bias vector; σ is the sigmoid function.

[0123] (2) Input gate (i t ): It mainly determines the input of the neuron. The calculation of formula (4) determines whether the input information is discarded or input into the neuron. Therefore, the input gate determines which information is input. Its activation function is the same as that of the forget gate, so the output range is also 0 to 1. 1 means that information is input, and 0 means that the information is discarded and not input into the neuron.

[0124]

[0125] Among them, W ix 、W ih and b i Respectively represent the input gate corresponding to x t 、h t-1 and the weight matrix of the bias vector; g t represents the cell state at time t, W gx 、W gh and bg Respectively represent the current moment input corresponding to x t 、h t-1 and the weight matrix of the bias vector; tanh is the hyperbolic tangent function. The current cell state C can be calculated t

[0126] C t =C t-1 ⊙f t +g t ⊙i t (6)

[0127] where ⊙ represents the Hadamard product.

[0128] (3) Output gate (o t ) is primarily responsible for determining the neuron's current output. The calculation process is shown in the following equation. The following equation determines whether to output information to the next neuron. The activation function remains the same, so the output range is also 0 to 1. 1 indicates output, while 0 indicates no output.

[0129]

[0130] Among them, W ox 、W oh and b o Respectively represent the output gate corresponding to x t 、h t-1 and the weight matrix of the bias vector. The hidden layer output h can be calculated t

[0131] h t =o t ⊙tanh(c t ) (8)

[0132] Sub-step S42: Backward propagation

[0133] Memory unit training uses the time series backpropagation algorithm to calculate the gradients of each weight, and then combines it with the gradient descent algorithm to train the LSTM model. By propagating the prediction error to the previous moment in time and to the previous layer of the neural network, the gradient calculations of each weight and bias term can be derived:

[0134]

[0135] Among them, η∈{f,i,g,o}, E is the cost function; f,t , δ i,t , δ g,t , δ o,t The BPTT error calculation process is as follows:

[0136]

[0137]

[0138] Table 1 LSTM anomaly identification and compensation model training parameters

[0139] parameter value optimizer adam Activation Function relu monitor val_loss loss mse timestep 49 dropout 0.2 learning_rate 0.001 epsilon (learning rate reduction detection threshold) 0.0001 patience (learning rate reduction tolerance threshold) 1 factor (learning rate factor) 0.1 batch_size (batch processing size) 64 epoch (number of training times) 60 validation_split (ratio of validation set to training set) 0.1

[0140] Step S5: Identification and compensation of abnormal position data

[0141] Sub-step S51: abnormal data generation

[0142] Increments of ±0.01 are randomly added to the latitude and longitude data as trajectory segments containing abnormal location data.

[0143] Table 2 Normal position data and abnormal position data

[0144]

[0145] Sub-step S52: Setting the abnormality threshold

[0146] Based on the distribution of prediction errors between the training data and the anomaly data, we set a threshold for anomaly detection. We calculate the prediction error for each data point during training. We then calculate the Euclidean distance between each predicted point and the true point, and then calculate the mean μ and standard deviation σ of the error. We then set a threshold of μ + 2σ. This threshold is used in subsequent steps to determine whether a data point is an anomaly.

[0147] Sub-step S53: Abnormal data identification and compensation process

[0148] When new location data is input, the new data is preprocessed to ensure that the data format is consistent with the training data; the preprocessed new data is then input into the trained LSTM model to generate the predicted longitude and latitude for each data point; the error between the actual location data and the predicted value is calculated, and the error of each data point is compared with the preset abnormality threshold. If the error is greater than the threshold, the data point is marked as abnormal, otherwise it is marked as normal; for the detected abnormal data points, the predicted value of the LSTM model is used to compensate to generate a continuous and accurate location data sequence. The flow chart is shown in the attached figure. Figure 4 shown.

[0149] Step S6: Robust mileage calculation

[0150] Sub-step S61: Gauss-Krüger projection conversion

[0151] The compensated latitude and longitude data are converted into plane Cartesian coordinates (x, y) using the Gauss-Krüger projection. The specific formula can be referred to the standard geographic information system software.

[0152] Sub-step S62: Calculate the distance between adjacent data points

[0153] For two adjacent data points (i, i+1) in the time series, the Euclidean distance is calculated in the plane coordinate system:

[0154]

[0155] Among them, (x i ,y i ) and (x i +1,yi+1) are the Cartesian coordinates of data points i and i+1.

[0156] Sub-step S63: Accumulate total mileage

[0157] Add up the distances of all adjacent data points to get the total mileage:

[0158]

[0159] Where N is the total number of data points.

Claims

1. A mileage robust dead reckoning method based on LSTM anomaly identification and compensation, characterized by: The following steps are involved: Step S1: Data generation The LSTM neural network training set contains the longitude and latitude data of all location points on the vehicle's driving route. The training data set is generated based on the Baidu Map Developer Platform. Sub-step S11: Create a driving instance Create a driving navigation instance and use the search function to retrieve the shortest path based on the latitude and longitude coordinates of the starting and ending points. Sub-step S12: Data encoding Based on the trajectory data created by the search function, separate the individual longitude and latitude data with "|" and separate the longitude and latitude data of different points with "", to achieve data encoding; Sub-step S13: Data enhancement and segmentation Random noise is added to the trajectory data generated by data encoding to replicate multiple trajectories for data enhancement. Gaussian noise with a mean of 0 and a standard deviation of 0.001 degrees is added to generate training set trajectories and test set trajectories in a 9:1 ratio. Step S2: Data preprocessing Sub-step S21: Determine the characteristic attributes of the training set Based on the relationship between vehicle location prediction and vehicle data information, the most basic timestamp, longitude, and latitude were first selected as the feature attribute set; The data collection frequency is set according to the application requirements; the N trajectory location points based on the timestamp, longitude, and latitude feature attribute set are represented as P i =<t i ,lon i ,lat i >,i=1,2,…,N (1) where t i Indicates the timestamp attribute corresponding to the i-th position, lon i ,lat i They represent the longitude and latitude attributes corresponding to the i-th position respectively; the default imported data is arranged by timestamp, so the location data dimension is (N, 2); Sub-step S22: Data normalization The "min-max normalization" algorithm is used, as shown in the following formula, to perform linear changes on the original data set; Where max(x) and min(x) represent the maximum and minimum values ​​in the position sequence x respectively; Sub-step S23: Time series formatting The timestep position data points are passed into the input layer as training samples, and the timestep+1 position data points are used as training labels. The prediction error is calculated to update the model parameters. The sliding window data is segmented to generate training samples with a dimension of (N-timestep, timestep, 2) and training labels with a dimension of (N-timestep, 2). Obtain N-timestep trajectory segments, each of which consists of timestep longitude and latitude data points, forming the above three-dimensional training set; the N-timestep trajectory segments obtained by segmentation are shuffled and randomly arranged, and the N-timestep training labels are also shuffled in the same order as the training samples; after the above sequence segmentation, the training set P = [p1, p2, ..., p i ,…p N-timestep ], the training set after random permutation is expressed as P'=[p N-timestep ,p i ,…,p2,…p1]; input the training set P' into the LSTM trajectory prediction model for training; Step S3: LSTM anomaly identification and compensation model construction Sub-step S31: Construct input layer The input layer receives fixed-length time series information. The transmitted data is timestep vehicle trajectory points, each of which contains a set of vehicle attributes at that point. Each time a sample is input, the latitude is (timestep, 2), and there are N samples in total. Sub-step S32: Construct hidden layer 1 Hidden layer 1 adds three forget gates, input gate and output gate structures; Sub-step S33: Construct dropout layer Temporarily discard some neurons in the LSTM neural network from the network; Sub-step S34: Construct hidden layer 2 Hidden layer 2 adds three forget gates, input gate and output gate structures; Sub-step S35: Construct output layer The output layer calculates the predicted point position by connecting the weights with the last hidden layer. The number of output neurons is 2, and the predicted label dimension is (1, 2). The error between the output predicted position and the actual position observation value is used to determine whether the position data is abnormal, and the abnormal data is compensated by the predicted value. Step S4: LSTM anomaly identification and compensation model training Sub-step S41: Forward propagation (1) Forget gate (f t ): The forget gate calculates the result through the state value of the neuron at the previous moment, the output value at the previous moment and the input value at the current moment; Among them, x t 、h t-1 They represent the input neurons at time t and the hidden neurons at time t-1 respectively; W fx 、W fh and b f Respectively represent the forget gate corresponding to x t 、h t-1 and the weight matrix of the bias vector; σ is the sigmoid function; (2) Input gate (i t ): After the calculation of formula (4), it is decided whether to discard the input information or input it into the neuron; Among them, W ix 、W ih and b i Respectively represent the input gate corresponding to x t 、h t-1 and the weight matrix of the bias vector; g t represents the cell state at time t, W gx 、W gh and b g Respectively represent the current moment input corresponding to x t 、h t-1 and the weight matrix of the bias vector; tanh is the hyperbolic tangent function; calculate the unit state C at the current moment t C t =C t-1 ⊙f t +g t ⊙i t (6) where represents the Hadamard product; (3) Output gate (o t ): The following operation is used to determine whether to output information to the next neuron; Among them, W ox 、W oh and b o Respectively represent the output gate corresponding to x t 、h t-1 and the weight matrix of the bias vector; calculate the hidden layer output h t h t =o t ⊙tanh(c t ) (8) Sub-step S42: Backward propagation The prediction error is propagated to the previous moment along time and to the previous layer along the neural network, and the calculation of each weight gradient and bias gradient is derived: Among them, η∈{f,i,g,o}, E is the cost function; f,t , δ i,t , δ g,t , δ o,t The BPTT error calculation process is as follows: Step S5: Identification and compensation of abnormal position data Sub-step S51: abnormal data generation An increment of ±0.01 is randomly added to the longitude and latitude data as a trajectory segment containing abnormal position data; Sub-step S52: Setting the abnormality threshold Based on the prediction error distribution of training data and abnormal data, set the threshold for anomaly detection; calculate the Euclidean distance between each predicted point and the true point, and then calculate the mean μ and standard deviation σ of the error, and set the threshold to μ+2σ; Sub-step S53: Abnormal data identification and compensation process The preprocessed new data is fed into the trained LSTM model to generate predicted longitude and latitude for each data point. The error between the actual location data and the predicted value is calculated, and the error for each data point is compared with the preset anomaly threshold. If the error is greater than the threshold, the data point is marked as an anomaly; otherwise, it is marked as normal. For detected anomaly data points, the predicted value of the LSTM model is used to compensate. Step S6: Robust mileage calculation Sub-step S61: Gauss-Krüger projection conversion The compensated latitude and longitude data are converted into plane Cartesian coordinates (x, y) using the Gauss-Krüger projection; Sub-step S62: Calculate the distance between adjacent data points For two adjacent data points (i, i+1) in the time series, the Euclidean distance is calculated in the plane coordinate system: Among them, (x i ,y i ) and (x i +1,yi+1) are the Cartesian coordinates of data points i and i+1; Sub-step S63: Accumulate total mileage Add up the distances of all adjacent data points to get the total mileage: Where N is the total number of data points.

Citation Information

Cited By

  • Vehicle following target CIPV identification method and system for navigation automatic driving

    CN121448433A