Strip mine truck cross-scale speed prediction method based on space-time masking pre-training
Through space-time masking pre-training and space-time dual-span-scale Transformer prediction module, the problem of space-time heterogeneity distinction in the prediction of open-pit mine velocity in the prior art is solved, and the prediction accuracy and model adaptability are improved.
Patent Information
- Application Number
- CN202510092680.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The existing open-pit mine speed prediction technology is difficult to effectively distinguish between space-time heterogeneity, resulting in poor adaptability and low prediction accuracy when applied to new data.
The span scale velocity prediction method of open-pit mine card based on spatiotemporal masking pre-training is adopted. The spatiotemporal features and road network structure information are extracted through the spatiotemporal masking pre-training process, and the prediction is made using the spatiotemporal double-span scale Transformer prediction module.
It improves the generalization ability of the model and the accuracy of the temporal and spatial heterogeneity distinction, enhances the prediction accuracy of future speed data of road network nodes, and improves the efficiency and safety of open-pit mine card transportation.
Smart Images

Figure CN120013000A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of open-pit mine truck intelligent scheduling, and in particular to an open-pit mine truck cross-scale speed prediction method based on spatiotemporal masking pre-training suitable for open-pit mine truck intelligent scheduling. Background Art
[0002] In the intelligent dispatching of open-pit mine trucks, the speed prediction of mine trucks plays a vital role. This prediction is not only related to the production efficiency of the mine, but also directly affects the operation safety, fuel consumption and overall operating costs of the mine trucks. Therefore, accurate speed prediction of open-pit mine trucks is the key prerequisite for safe dispatching of open-pit mine trucks.
[0003] Compared with typical multivariate time series data, mining truck data exhibits spatiotemporal heterogeneity. Spatiotemporal heterogeneity refers to the unevenness and complexity of the distribution of spatiotemporal data. It describes the significant differences between the variables of the same sensor at different times and different sensors at the same time. When the data scale is small, this heterogeneity is clearly visible; however, when faced with data accumulated by hundreds of sensors over several months, spatial heterogeneity and temporal heterogeneity are highly mixed. Therefore, accurately distinguishing temporal heterogeneity is a huge challenge. Existing models are mostly trained in an end-to-end manner. Due to the high complexity of the model, its input data is usually limited to a shorter value (usually 12 steps). This will cause the model to be unable to adapt faster and make accurate predictions when applied to new, unseen data. Therefore, it is particularly important to improve the generalization ability of the model and accurately distinguish spatiotemporal heterogeneity.
[0004] In addition, although graph neural networks (GNNs) have shown significant results in dealing with spatial correlation and temporal dependency, they are still limited by their over-reliance on topological regularization patterns. In the process of node message propagation, GNNs mainly focus on capturing information strictly constrained by graph topology, which leads to a relatively narrow range of data representation and may ignore those non-topological patterns that cannot be directly captured by the graph structure, thus limiting its application scope and prediction accuracy. Summary of the invention
[0005] Technical problem: The purpose of the present invention is to provide a cross-scale speed prediction method for open-pit mine trucks based on spatiotemporal masking pre-training in response to the limitations of existing open-pit mine truck speed prediction technology. Through the spatiotemporal masking pre-training process, the spatiotemporal characteristics of long time series and the road network structure information are effectively extracted, and then these feature information is used through the spatiotemporal dual cross-scale Transformer prediction module to accurately predict the speed data of road network nodes for a period of time in the future, thereby improving the efficiency and safety of open-pit mine truck transportation.
[0006] Technical solution: To achieve the above-mentioned purpose, the present invention provides a method for cross-scale speed prediction of open-pit mine trucks based on spatiotemporal masking pre-training, which is characterized by: designing a speed prediction model based on the speed data of the open-pit mine truck, pre-processing the GPS trajectory data of the open-pit mine truck, and obtaining a speed sequence sample set and a road network adjacency matrix; selecting a long time series from the speed sequence sample set, processing it together with the adjacency matrix through a high-dimensional spatiotemporal data embedding module and a spatiotemporal position encoding module, and inputting it into a spatiotemporal coupling masking pre-training module to obtain the parameters of the spatiotemporal encoder in the spatiotemporal masking autoencoder, and using the parameters to configure a spatiotemporal dual-cross-scale Transformer prediction module; selecting a short time series from the speed sequence sample set, processing it together with the adjacency matrix through a high-dimensional spatiotemporal data embedding module, and inputting it into a spatiotemporal dual-cross-scale Transformer prediction module to predict the speed data of the road network nodes for a period of time in the future;
[0007] The specific steps are as follows:
[0008] Step 1: pre-process the GPS trajectory data of the open-pit mine truck, use the existing road network automatic generation algorithm, and generate a road network adjacency matrix based on the pre-processed GPS trajectory data; represent each node of the road network as a speed sequence sample, and use the speed sequence samples of all nodes to construct a speed sequence sample set X;
[0009] Step 2: Use random sampling method to select long time series data X from the speed series sample set X obtained in step 1 long , the long time series data X long Together with the adjacency matrix A, it is input into the high-dimensional spatiotemporal data embedding module to realize the mapping of data from the original feature space to the high-dimensional feature space, and obtain the high-dimensional spatiotemporal representation of the sequence data;
[0010] Step 3: Represent the obtained sequence data in high-dimensional space-time e Input includes time position code E tp , spatial position encoding E sp The spatiotemporal position encoding module is used to obtain the final embedding representation E of the spatiotemporal coupling masking pre-training module;
[0011] Step 4: Input the final embedding representation E of the spatiotemporal coupled masking pre-training module into the spatiotemporal data masking unit of the spatiotemporal coupled masking pre-training module, and perform masking using a randomly sampled spatiotemporal masking block generation method to obtain masked data X;
[0012] Step 5, design a pre-training loss function, input the masked data X obtained in step 4 into the spatiotemporal masked autoencoder of the spatiotemporal coupled masked pre-training module for pre-training, and obtain the parameters of the temporal multi-head attention layer and the spatial multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder; the spatiotemporal masked autoencoder of the spatiotemporal coupled masked pre-training module includes three core parts: a spatiotemporal encoder, a masked token replacement module, and a spatiotemporal reconstruction decoder;
[0013] Step 6: Select short time series data X from the speed series sample set obtained in step 1 by random sampling. short , the short time series data X short The high-dimensional spatiotemporal data embedding module is input together with the adjacency matrix to obtain a high-dimensional spatiotemporal representation. The high-dimensional spatiotemporal representation and the parameters of the temporal multi-head attention layer and the spatial multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder obtained in step 5 are used in the spatiotemporal dual-cross-scale Transformer prediction module to predict the speed of road network nodes in the future.
[0014] Step 7: Design the loss function of the spatiotemporal dual-cross-scale Transformer prediction module, perform iterative training, optimize the learning parameters of the prediction module, and use a variety of evaluation functions to evaluate the prediction performance of the prediction module.
[0015] In step 1, the GPS trajectory data of the open-pit mine truck is preprocessed, including the generation of the road network adjacency matrix and the construction of the speed sequence sample set;
[0016] The road network adjacency matrix generation process is:
[0017] First, collect GPS track data of open-pit mining trucks. These data record the driving information of mining trucks in the mining area, including mining truck ID, latitude and longitude coordinates, speed and timestamp;
[0018] Next, the GPS trajectory data is preprocessed, including removing noise, outliers, and duplicate data;
[0019] Then, the existing automatic road network generation algorithm is used to analyze the preprocessed GPS trajectory data to identify key locations in the trajectory, such as road intersections, turning points, etc. These locations will be set as nodes of the road network V = {v1, v2, ... v N}; After determining the road network nodes, the algorithm forms a road network adjacency matrix A∈R according to the connectivity of the trajectory data N×N , forming a road network G;
[0020] The constructed speed sequence sample set:
[0021] First, for each node v in the road network i, extract the speed data of all mining trucks passing through the node, including the speed value and the corresponding timestamp;
[0022] Then, based on the extracted speed values and timestamps, a speed sequence sample x is constructed for each node in the road network at a time interval of 5 minutes. i ∈R T×C ;
[0023] Finally, the speed sequence samples of all road network nodes are aggregated to construct the speed sequence sample set X = {x1, x2, …x N}∈R T×N×C ; Where T represents time, N represents the number of road network nodes, C represents node speed, and C=1 represents a single speed dimension.
[0024] In step 2, the long time series data X long Together with the adjacency matrix generated in step 1, it is input into the high-dimensional spatiotemporal data embedding module for processing:
[0025] The process of randomly extracting long time series data:
[0026] First, set the long time series extraction ratio to
[0027] Then, determine the length of the long time series
[0028] Finally, a long time series data with a time length of
[0029] The high-dimensional spatiotemporal data embedding module includes time dimension embedding, space dimension embedding, and period embedding;
[0030] The time dimension embedding process is:
[0031] First, the long time series data X long Divide into L time periods of length T P =T long / L time blocks, and splice along the feature dimension to obtain the spliced data
[0032] Then, the concatenated data X p Through a linear layer, we get the time embedding vector
[0033] E t =W p X p +b p
[0034] Among them, Wp is the weight matrix, b p is the bias vector;
[0035] The spatial dimension embedding process is:
[0036] First, the normalized Laplacian matrix L∈R is calculated from the adjacency matrix A N×N :
[0037] L=I n -D -1 / 2 AD -1 / 2 =U T ΛU
[0038] Among them, I n is the identity matrix, D is the degree matrix, and U is the eigenvector matrix;
[0039] Then, in the eigenvector matrix U, select the K smallest non-zero eigenvectors and project the K smallest non-zero eigenvectors to the D-dimensional features to generate the spatial embedding vector
[0040] The cycle embedding includes daily cycle embedding and weekly cycle embedding, and the embedding process is:
[0041] First, the timestamp corresponding to the speed sequence sample is converted into the time proportion T of the sample in a day d and the proportion of time in a week T w , each of which passes through the daily embedding linear layer and the weekly embedding linear layer to obtain the daily period embedding vector and the periodic embedding vector The processing process is:
[0042] E d =W d T d +b d
[0043] E w =W w T w +b w
[0044] Among them, W d and b d is the weight matrix and bias vector of the embedding linear layer, W w and b w The weight matrix and bias vector for the week embedding linear layer;
[0045] Finally, the above embedding vectors are summed to obtain the output E of the high-dimensional spatiotemporal data embedding module e :
[0046] Ee =E t +E s +E d +E w
[0047] Among them, E t is the time embedding vector, E s is the spatial embedding vector, E d is the daily period embedding vector, E w is the periodic embedding vector, It is a high-dimensional spatiotemporal representation of sequence data;
[0048] In step 3, the high-dimensional spatiotemporal representation E of the sequence data e Input includes time position code E tp , spatial position encoding E sp The spatiotemporal position encoding module process is as follows:
[0049] First, the high-dimensional spacetime is represented by E e Perform time position encoding E tp , the formula is as follows:
[0050] E tp [t,n,2i]=sin(t / 10000 2i / D )
[0051] E tp [t,n,2i+1]=cos(t / 10000 2i / D )
[0052] Where t represents the time step, n represents the length of the time series, i represents the index in the time embedding dimension D, sin(·) represents the sine function, and cos(·) represents the cosine function. Indicates time position coding;
[0053] Then, the high-dimensional space-time representation E e Spatial position encoding sp , the formula is as follows:
[0054] E sp [t,n,2i]=sin(n / 10000 2i / D )
[0055] E sp [t,n,2i+1]=cos(n / 10000 2i / D )
[0056] Where t represents the time step, n represents the length of the time series, i represents the index in the spatial embedding dimension D, sin(·) represents the sine function, and cos(·) represents the cosine function. Represents spatial position encoding;
[0057] Finally, the high-dimensional space-time is represented by E e , time position coding E tp and spatial position encoding E sp Add them together to obtain the output E of the spatiotemporal position encoding module, and the formula is as follows:
[0058] E=E e +E tp +E sp
[0059] in, The final embedding representation of the pre-trained module is masked for spatiotemporal coupling.
[0060] In step 4, the processing process of the spatiotemporal data masking unit of the spatiotemporal coupled masking pre-training module is as follows:
[0061] First, the final embedding representation E is input into the spatiotemporal data masking unit, and the masking ratio is determined to be α;
[0062] Then, the final embedding representation E is split into patches as follows:
[0063] N patch =T p ND / l m
[0064] Among them, N patch is the total number of patches, T p is the time step of the final embedding representation E, N is the number of nodes, D is the feature dimension of the final embedding representation E, l m The length of each patch block is based on the total number of patch blocks N patch Set the patch block index set N = {1, 2, ..., N patch};
[0065] Next, according to the masking ratio α and the total number of patches N patch Calculate the number of masked patches N m =αN patch ; Randomly select N from the patch block index set N m index, set the mask patch index set
[0066] Next, create a masking matrix with all initial values 1 Using the masked patch index set P, set the values of the patches corresponding to these indexes in the masking matrix M to 0 to generate a new masking matrix
[0067] Finally, the masked data is calculated Masking data As input to the spatiotemporal masked autoencoder;
[0068]
[0069] Where E is the final embedding representation of the spatiotemporal coupled masked pre-training module, M′ is the masking matrix, To mask data;
[0070] In step 5, the masked data X is input into the spatiotemporal masked autoencoder of the spatiotemporal coupled masked pre-training module for pre-training. During the pre-training process, a pre-training loss function is designed to evaluate the performance of the model in the reconstruction task and guide the model to continuously optimize its parameters during the training process; the spatiotemporal encoder includes a temporal multi-head attention layer and a spatial multi-head attention layer, the masked token replacement module replaces the masked part of the spatiotemporal encoder output data with a specific masked token, and the spatiotemporal reconstruction decoder includes an attention layer and a fully connected layer;
[0071] The pre-training process of spatiotemporal masked autoencoder is:
[0072] First, mask the data Input to the spatiotemporal encoder, masking data After passing through the time multi-head attention layer and the space multi-head attention layer respectively, the outputs of the two attention layers are merged to obtain the output X of the spatiotemporal encoder enc-out ; The calculation process of the space-time encoder is as follows:
[0073]
[0074] X temp-att =Attention(Q1,K1,V1)
[0075]
[0076] X spatial-att =Attention(Q2,K2,V2)
[0077] X enc-out =X temp-att +X spatial-att
[0078] in, To mask the data, Q1 is the query vector of the temporal multi-head attention layer, K1 is the key vector of the temporal multi-head attention layer, and V1 is the value vector of the temporal multi-head attention layer. is the weight matrix of Q1, is the weight matrix of K1, is the weight matrix of V1, X temp-att is the output of the temporal multi-head attention layer; Q2 is the query vector of the spatial multi-head attention layer, K2 is the key vector of the spatial multi-head attention layer, and V2 is the value vector of the temporal multi-head attention layer. is the weight matrix of Q2, is the weight matrix of K2, is the weight matrix of V2, X spatial-att is the output of the spatial multi-head attention layer; Attention represents the attention mechanism calculation, X enc-out is the output of the space-time encoder;
[0079] Second, the output X of the space-time encoder enc-out Passed to the masked token replacement module;
[0080] The process of masked token replacement is as follows: First, construct an output X enc-out Masked token vector T of the same shape m ; Next, construct an output X that is consistent with the spatiotemporal encoder enc-out A mask tensor M of the same shape p , the mask tensor M p The position value specified by the masked patch index set P is 1, and the other position values are 0; finally, apply M p T m With X enc-out Merge to generate masked token replacement result X dec-in ; The merging process is as follows:
[0081] X dec-in =M p ⊙T m +(1-M p )⊙X enc-out
[0082] Among them, M p is the mask tensor, T m is the masked token vector, X enc-out is the output of the spatiotemporal encoder, ⊙ represents element-by-element multiplication;
[0083] Third, replace the masked token with the result X dec-in Input into the space-time reconstruction decoder, decode and reconstruct the data, and obtain the output representation X of the space-time reconstruction decoder dec-out ; The reconstruction decoder process is as follows:
[0084] X dec-out =FC(Attention(X dec-in ))
[0085] Among them, Attention represents the attention layer, FC(·) represents the fully connected layer;
[0086] Fourth, calculate the masked reconstruction part after training and mask the true value Q∈R T×N×D , the calculation process is as follows:
[0087]
[0088] Among them, X dec-in is the masked token replacement result, M′ is the masking matrix, To mask the data, ⊙ represents element-by-element multiplication;
[0089] Fifth, based on the masked true value Q and the masked reconstruction part Design a pre-training loss function and adjust the model parameters through optimization strategy; the calculation formula of the pre-training loss function is as follows:
[0090]
[0091] Among them, L is the loss value, T is the time step, N is the number of nodes, and D is the hidden dimension. is the masked reconstruction part after pre-training, Q is the masked true value, t is the time index, n is the node index, and d is the hidden dimension index;
[0092] Sixth, after pre-training, obtain the weight matrix of the temporal multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder Weight matrix of the spatial multi-head attention layer
[0093] In step 6, the spatiotemporal dual-scale Transformer prediction module involves the division of spatial scale and temporal scale. The division process is:
[0094] The short time series data and the road network adjacency matrix are processed by the high-dimensional spatiotemporal data embedding module, and the obtained high-dimensional spatiotemporal representation H t-T:t ∈R T×N×D It is expressed as follows:
[0095]
[0096] in, Represents the value of the Nth node in the time period tT to t;
[0097] The spatial scale division refers to dividing the geographic space area into multiple grid units according to the pre-set length and width standards; the size of these grid units will vary due to the use of different length and width standards, thus producing different spatial scale divisions; then, the nodes in the road network will be assigned to the corresponding grid units according to their geographical locations, and the nodes in the same grid unit will be aggregated as the representation of the grid unit; the first s Aggregate representation of the mth grid at each spatial scale as follows:
[0098]
[0099] Where LN(·) represents the normalization operation, Indicates that the first s All values of the mth grid at the spatial scale are aggregated, represents the value of the i-th node in the time period tT to t, represents the weight matrix, represents the bias vector;
[0100] The time scale division refers to dividing a time period with a fixed length into several small time segments according to a preset unit time length. When different unit times are used, different time scale divisions will be generated. Assuming that the first t In the time scale, the unit time length is Then the input data will be split into time segments, where T is the length of the original time segment; the i-th node is in the l-th t The jth time segment of the time scale represents as follows:
[0101]
[0102] in, represents the value of the i-th node at time t, j represents the time segment index, Indicates the unit time length;
[0103] Will Input to the fully connected layer to get the lth t Aggregate representation of the jth time segment of a time scale
[0104]
[0105] Where LN(·) represents the normalization operation, Indicates the first t The j-th time segment of a time scale is represented by represents the weight matrix, represents the bias vector;
[0106] In step 6, the spatiotemporal dual-cross-scale Transformer prediction module includes a temporal Transformer unit and a spatial Transformer unit; the prediction process is as follows:
[0107] First, randomly select a speed sequence sample set X with a length of T short =θT long Short time series data The short time series data X short It is input into the high-dimensional spatiotemporal data embedding module together with the road network adjacency matrix A to obtain the high-dimensional spatiotemporal representation of the sequence data.
[0108] H t-T:t =FC(X short +X d +X w )+X s
[0109] Among them, FC(·) represents the fully connected layer, X d and X w represents the daily and weekly period embedding, X s represents the spatial embedding of the adjacency matrix after Laplace eigentransformation;
[0110] Then, the high-dimensional spatiotemporal data is embedded into the high-dimensional spatiotemporal representation H of the module. t-T:t Input to the temporal Transformer unit; the temporal Transformer unit has a total of l t Layer, divide the temporal Transformer unit into time scales, and the final output is expressed as The network node is at the lth t The updating process on a time scale is as follows:
[0111]
[0112] Where LN(·) represents the normalization operation, Indicates the first t -1 time scale division result, MHA(·) is the multi-head self-attention mechanism, are the parameters of the temporal multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder, is the result of time scale division, For the first t The intermediate representation of the time scale, FFN(·) represents the feed-forward layer;
[0113] Next, the final output of the temporal Transformer unit is represented as Input to the spatial Transformer unit; the spatial Transformer unit has a total of l s Layer, the spatial Transformer unit is divided into spatial scales, and the final output is expressed as The network node is at the lth s The updating process on each spatial scale is as follows:
[0114]
[0115] Where LN(·) represents the normalization operation, Indicates the first s -1 spatial scale division result, MHA(·) is the multi-head self-attention mechanism, are the parameters of the spatial multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder, is the result of spatial scale division, For the first s The intermediate representation of spatial scales, FFN(·) represents the feed-forward layer;
[0116] Finally, the output of the spatial Transformer unit Input into a fully connected layer to get the predicted value of the road network node speed in the future The calculation is as follows:
[0117]
[0118] Among them, FC(·) is a fully connected layer;
[0119] In step 7, the training loss function and evaluation function of the spatiotemporal dual-cross-scale Transformer prediction module are designed as follows:
[0120] The training loss function design process is:
[0121] The mean absolute error (MAE) is selected as the training loss function to reduce the mean absolute error between the predicted value and the true value; the training loss function calculation formula is:
[0122]
[0123] Among them, Y i represents the true value of the node speed of the path network at the i-th time. represents the predicted value of the network node speed at the i-th time step, and Q represents the length of the i-th time step;
[0124] The evaluation function design process is as follows:
[0125] The mean square error (MSE) and root mean square error (RMSE) are selected as the evaluation loss function to evaluate the performance of the model on the test set; the evaluation loss function calculation formula is:
[0126]
[0127] Among them, Y i represents the true value of the node speed of the path network at the i-th time. represents the predicted value of the network node speed at the i-th time step, and Q represents the length of the i-th time step;
[0128] Beneficial effect: Due to the adoption of the above technical solution, the present invention adopts data block technology to improve computational efficiency, designs a method for generating spatiotemporal masked blocks, enhances the generalization ability of the model, and improves the accuracy of distinguishing spatiotemporal heterogeneity. The spatiotemporal dual-scale Transformer prediction module effectively captures topologically irrelevant features by dividing different time and space scales, thereby improving the prediction accuracy. The present invention first pre-processes the GPS trajectory data of the open-pit mine truck, generates a road network adjacency matrix using a road network automatic generation algorithm, and generates a speed sequence sample set based on the road network nodes; then, long time series data is randomly selected from the sample set, and after being processed by a high-dimensional spatiotemporal data embedding module and a spatiotemporal position encoding module together with the adjacency matrix, it is input into a spatiotemporal coupling masked pre-training module to obtain the parameters of the spatiotemporal encoder in the spatiotemporal masked autoencoder, which are used to configure the spatiotemporal dual-scale Transformer prediction module; finally, short time series data is randomly selected from the speed sequence sample set, and after being processed by a high-dimensional spatiotemporal data embedding module together with the adjacency matrix, it is input into a spatiotemporal dual-scale Transformer prediction module to realize the speed prediction of future road network nodes. The main advantages compared with existing technologies are:
[0129] (1) The present invention designs a spatiotemporal masking pre-training framework, and uses a spatiotemporal masking autoencoder to process the long time series data and the road network adjacency matrix processed by the high-dimensional spatiotemporal data embedding module and the spatiotemporal position encoding module, so as to obtain the parameters of the spatiotemporal encoder in the spatiotemporal masking autoencoder; that is, the long time series data is cut and divided to realize block processing of the data, thereby improving the model calculation efficiency and ensuring the smooth operation of the model under limited resources; in addition, a spatiotemporal masking block generation method is designed to enhance the generalization ability of the model and improve the accuracy of distinguishing spatiotemporal heterogeneity;
[0130] (2) The present invention uses the parameters of the spatiotemporal encoder in the spatiotemporal masked autoencoder to configure the spatiotemporal dual-scale Transformer prediction module, which receives the short time series data and the road network adjacency matrix processed by the high-dimensional spatiotemporal data embedding module as input, and then efficiently captures topology-independent features in the spatial and temporal dimensions, thereby realizing the prediction of speed data of road network nodes in the future. BRIEF DESCRIPTION OF THE DRAWINGS
[0131] Figure 1 The present invention is a flow chart of the cross-scale velocity prediction method of open-pit mine trucks based on spatiotemporal masking pre-training. DETAILED DESCRIPTION
[0132] An embodiment of the present invention is further described below in conjunction with the accompanying drawings:
[0133] The cross-scale speed prediction method of an open-pit mine truck based on spatiotemporal masking pre-training of the present invention designs a speed prediction model based on the speed data of the open-pit mine truck, pre-processes the GPS trajectory data of the open-pit mine truck, obtains a speed sequence sample set and a road network adjacency matrix; selects a long time series from the speed sequence sample set, and processes it together with the adjacency matrix through a high-dimensional spatiotemporal data embedding module and a spatiotemporal position encoding module, and then inputs it into a spatiotemporal coupling masking pre-training module to obtain the parameters of the spatiotemporal encoder in the spatiotemporal masking autoencoder, and the parameters are used to configure the spatiotemporal dual cross-scale Transformer prediction module. Select a short time series from the speed sequence sample set, and process it together with the adjacency matrix through a high-dimensional spatiotemporal data embedding module, and then input it into a spatiotemporal dual cross-scale Transformer prediction module to predict the speed data of the road network node for a period of time in the future; the specific steps are as follows:
[0134] Step 1: pre-process the GPS trajectory data of the open-pit mine truck, use the existing road network automatic generation algorithm, and generate a road network adjacency matrix based on the pre-processed GPS trajectory data; represent each node of the road network as a speed sequence sample, and use the speed sequence samples of all nodes to construct a speed sequence sample set X;
[0135] The preprocessing of the GPS trajectory data of the open-pit mine truck includes generating a road network adjacency matrix and constructing a speed sequence sample set. The processing process is as follows:
[0136] The road network adjacency matrix generation process is:
[0137] First, collect GPS track data of open-pit mining trucks. These data record the driving information of mining trucks in the mining area, including mining truck ID, latitude and longitude coordinates, speed and timestamp;
[0138] Next, the GPS trajectory data is preprocessed, including removing noise, outliers, and duplicate data;
[0139] Then, the existing automatic road network generation algorithm is used to analyze the preprocessed GPS trajectory data to identify key locations in the trajectory, such as road intersections, turning points, etc. These locations will be set as nodes of the road network V = {v1, v2, ... v N}; After determining the road network nodes, the algorithm forms a road network adjacency matrix A∈R according to the connectivity of the trajectory data N×N , forming a road network G;
[0140] The constructed speed sequence sample set:
[0141] First, for each node v in the road network i , extract the speed data of all mining trucks passing through the node, including the speed value and the corresponding timestamp;
[0142] Then, based on the extracted speed values and timestamps, a speed sequence sample x is constructed for each node in the road network at a time interval of 5 minutes. i ∈R T×C ;
[0143] Finally, the speed sequence samples of all road network nodes are aggregated to construct the speed sequence sample set X = {x1, x2, …x N}∈R T×N×C ; Where T represents time, N represents the number of road network nodes, C represents node speed, and C=1 represents a single speed dimension.
[0144] Step 2: Use random sampling method to select long time series data X from the speed series sample set X obtained in step 1 long , the long time series data X long Together with the adjacency matrix A, it is input into the high-dimensional spatiotemporal data embedding module to realize the mapping of data from the original feature space to the high-dimensional feature space, and obtain the high-dimensional spatiotemporal representation of the sequence data;
[0145] The long time series data X long Together with the adjacency matrix generated in step 1, it is input into the high-dimensional spatiotemporal data embedding module for processing.
[0146] The processing process is as follows:
[0147] The process of randomly extracting long time series data:
[0148] First, set the long time series extraction ratio to
[0149] Then, determine the length of the long time series
[0150] Finally, a long time series data with a time length of
[0151] The high-dimensional spatiotemporal data embedding module includes time dimension embedding, space dimension embedding, and period embedding;
[0152] The time dimension embedding process is:
[0153] First, the long time series data X long Divide into L time periods of length T P =T long / L time blocks, and splice along the feature dimension to obtain the spliced data
[0154] Then, the concatenated data X p Through a linear layer, we get the time embedding vector
[0155] E t =W p X p +b p
[0156] Among them, W p is the weight matrix, b p is the bias vector;
[0157] The spatial dimension embedding process is:
[0158] First, the normalized Laplacian matrix L∈R is calculated from the adjacency matrix A N×N :
[0159] L=I n -D -1 / 2 AD -1 / 2 =U T ΛU
[0160] Among them, I n is the identity matrix, D is the degree matrix, and U is the eigenvector matrix;
[0161] Then, in the eigenvector matrix U, select the K smallest non-zero eigenvectors and project the K smallest non-zero eigenvectors to the D-dimensional features to generate the spatial embedding vector
[0162] The cycle embedding includes daily cycle embedding and weekly cycle embedding, and the embedding process is:
[0163] First, the timestamp corresponding to the speed sequence sample is converted into the time proportion T of the sample in a day d and the proportion of time in a week T w, each of which passes through the daily embedding linear layer and the weekly embedding linear layer to obtain the daily period embedding vector and the periodic embedding vector The processing process is:
[0164] E d =W d T d +b d
[0165] E w =W w T w +b w
[0166] Among them, W d and b d is the weight matrix and bias vector of the embedding linear layer, W w and b w The weight matrix and bias vector for the week embedding linear layer;
[0167] Finally, the above embedding vectors are summed to obtain the output E of the high-dimensional spatiotemporal data embedding module e :
[0168] E e =E t +E s +E d +E w
[0169] Among them, E t is the time embedding vector, E s is the spatial embedding vector, E d is the daily period embedding vector, E w is the periodic embedding vector, It is a high-dimensional spatiotemporal representation of sequence data;
[0170] Step 3: Represent the obtained sequence data in high-dimensional space-time e Input includes time position code E tp , spatial position encoding E sp The spatiotemporal position encoding module is used to obtain the final embedding representation E of the spatiotemporal coupling masking pre-training module;
[0171] The high-dimensional spatiotemporal representation of sequence data E e Input includes time position code E tp , spatial position encoding E sp The spatiotemporal position encoding module of ; the processing process is as follows:
[0172] First, the high-dimensional space-time representation E ePerform time position encoding E tp , the formula is as follows:
[0173] E tp [t,n,2i]=sin(t / 10000 2i / D )
[0174] E tp [t,n,2i+1]=cos(t / 10000 2i / D )
[0175] Where t represents the time step, n represents the length of the time series, i represents the index in the time embedding dimension D, sin(·) represents the sine function, and cos(·) represents the cosine function. Indicates time position coding;
[0176] Then, the high-dimensional space-time representation E e Spatial position encoding sp , the formula is as follows:
[0177] E sp [t,n,2i]=sin(n / 10000 2i / D )
[0178] E sp [t,n,2i+1]=cos(n / 10000 2i / D )
[0179] Where t represents the time step, n represents the length of the time series, i represents the index in the spatial embedding dimension D, sin(·) represents the sine function, and cos(·) represents the cosine function. Represents spatial position encoding;
[0180] Finally, the high-dimensional space-time is represented by E e , time position coding E tp and spatial position encoding E sp Add them together to obtain the output E of the spatiotemporal position encoding module, and the formula is as follows:
[0181] E=E e +E tp +E sp
[0182] in, The final embedding representation of the pre-trained module is masked for spatiotemporal coupling.
[0183] Step 4: Input the final embedding representation E of the spatiotemporal coupled masking pre-training module into the spatiotemporal data masking unit of the spatiotemporal coupled masking pre-training module, and perform masking using a randomly sampled spatiotemporal masking block generation method to obtain masked data X;
[0184] The processing process of the spatiotemporal data masking unit of the spatiotemporal coupled masking pre-training module is as follows:
[0185] First, the final embedding representation E is input into the spatiotemporal data masking unit, and the masking ratio is determined to be α;
[0186] Then, the final embedding representation E is split into patches as follows:
[0187] N patch =T p ND / l m
[0188] Among them, N patch is the total number of patches, T p is the time step of the final embedding representation E, N is the number of nodes, D is the feature dimension of the final embedding representation E, l m The length of each patch block is based on the total number of patch blocks N patch Set the patch block index set N = {1, 2, ..., N patch};
[0189] Next, according to the masking ratio α and the total number of patches N patch Calculate the number of masked patches N m =αN patch ; Randomly select N from the patch block index set N m index, set the mask patch index set
[0190] Next, create a masking matrix with all initial values 1 Using the masked patch index set P, set the values of the patches corresponding to these indexes in the masking matrix M to 0 to generate a new masking matrix
[0191] Finally, the masked data is calculated Masking data As input to the spatiotemporal masked autoencoder;
[0192]
[0193] Where E is the final embedding representation of the spatiotemporal coupled masked pre-training module, M′ is the masking matrix, To mask data;
[0194] Step 5, design a pre-training loss function, input the masked data X obtained in step 4 into the spatiotemporal masked autoencoder of the spatiotemporal coupled masked pre-training module for pre-training, and obtain the parameters of the temporal multi-head attention layer and the spatial multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder; the spatiotemporal masked autoencoder of the spatiotemporal coupled masked pre-training module includes three core parts: a spatiotemporal encoder, a masked token replacement module, and a spatiotemporal reconstruction decoder;
[0195] The masked data X is input into the spatiotemporal masked autoencoder of the spatiotemporal coupled masked pre-training module for pre-training. During the pre-training process, a pre-training loss function is designed to evaluate the performance of the model in the reconstruction task and guide the model to continuously optimize its parameters during the training process; the spatiotemporal encoder includes a temporal multi-head attention layer and a spatial multi-head attention layer, the masked token replacement module replaces the masked part of the spatiotemporal encoder output data with a specific masked token, and the spatiotemporal reconstruction decoder includes an attention layer and a fully connected layer. The processing process is as follows:
[0196] The pre-training process of spatiotemporal masked autoencoder is:
[0197] First, mask the data Input to the spatiotemporal encoder, masking data After passing through the time multi-head attention layer and the space multi-head attention layer respectively, the outputs of the two attention layers are merged to obtain the output X of the spatiotemporal encoder enc-out ; The calculation process of the space-time encoder is as follows:
[0198]
[0199] X temp-att =Attention(Q1,K1,V1)
[0200]
[0201] X spatial-att =Attention(Q2,K2,V2)
[0202] X enc-out =X temp-att +X spatial-att
[0203] in, To mask the data, Q1 is the query vector of the temporal multi-head attention layer, K1 is the key vector of the temporal multi-head attention layer, and V1 is the value vector of the temporal multi-head attention layer. is the weight matrix of Q1, is the weight matrix of K1, is the weight matrix of V1, X temp-attis the output of the temporal multi-head attention layer; Q2 is the query vector of the spatial multi-head attention layer, K2 is the key vector of the spatial multi-head attention layer, and V2 is the value vector of the temporal multi-head attention layer. is the weight matrix of Q2, is the weight matrix of K2, is the weight matrix of V2, X spatial-att is the output of the spatial multi-head attention layer; Attention represents the attention mechanism calculation, X enc-out is the output of the space-time encoder;
[0204] Second, the output X of the space-time encoder enc-out Passed to the masked token replacement module;
[0205] The process of masked token replacement is as follows: First, construct an output X enc-out Masked token vector T of the same shape m ; Next, construct an output X that is consistent with the spatiotemporal encoder enc-out A mask tensor M of the same shape p , the mask tensor M p The position value specified by the masked patch index set P is 1, and the other position values are 0; finally, apply M p T m With X enc-out Merge to generate masked token replacement result X dec-in ; The merging process is as follows:
[0206] X dec-in =M p ⊙T m +(1-M p )⊙X enc-out
[0207] Among them, M p is the mask tensor, T m is the masked token vector, X enc-out is the output of the spatiotemporal encoder, ⊙ represents element-by-element multiplication;
[0208] Third, replace the masked token with the result X dec-in Input into the space-time reconstruction decoder, decode and reconstruct the data, and obtain the output representation X of the space-time reconstruction decoder dec-out ; The reconstruction decoder process is as follows:
[0209] X dec-out =FC(Attention(X dec-in ))
[0210] Among them, Attention represents the attention layer, FC(·) represents the fully connected layer;
[0211] Fourth, calculate the masked reconstruction part after training and mask the true value Q∈R T×N×D , the calculation process is as follows:
[0212]
[0213] Among them, X dec-in is the masked token replacement result, M′ is the masking matrix, To mask the data, ⊙ represents element-by-element multiplication;
[0214] Fifth, based on the masked true value Q and the masked reconstruction part Design a pre-training loss function and adjust the model parameters through optimization strategy; the calculation formula of the pre-training loss function is as follows:
[0215]
[0216] Among them, L is the loss value, T is the time step, N is the number of nodes, and D is the hidden dimension. is the masked reconstruction part after pre-training, Q is the masked true value, t is the time index, n is the node index, and d is the hidden dimension index;
[0217] Sixth, after pre-training, obtain the weight matrix of the temporal multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder Weight matrix of the spatial multi-head attention layer
[0218] Step 6: Select short time series data X from the speed series sample set obtained in step 1 by random sampling. short , the short time series data X short The high-dimensional spatiotemporal data embedding module is input together with the adjacency matrix to obtain a high-dimensional spatiotemporal representation. The high-dimensional spatiotemporal representation and the parameters of the temporal multi-head attention layer and the spatial multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder obtained in step 5 are used in the spatiotemporal dual-cross-scale Transformer prediction module to predict the speed of road network nodes in the future.
[0219] The spatiotemporal dual-scale Transformer prediction module involves the division of spatial scale and temporal scale, and the division process is as follows:
[0220] The short time series data and the road network adjacency matrix are processed by the high-dimensional spatiotemporal data embedding module, and the obtained high-dimensional spatiotemporal representation H t-T:t ∈R T×N×D It is expressed as follows:
[0221]
[0222] in, Represents the value of the Nth node in the time period tT to t;
[0223] The spatial scale division refers to dividing the geographic space area into multiple grid units according to the pre-set length and width standards; the size of these grid units will vary due to the use of different length and width standards, thus producing different spatial scale divisions; then, the nodes in the road network will be assigned to the corresponding grid units according to their geographical locations, and the nodes in the same grid unit will be aggregated as the representation of the grid unit; the first s Aggregate representation of the mth grid at each spatial scale as follows:
[0224]
[0225] Where LN(·) represents the normalization operation, Indicates that the first s All values of the mth grid at the spatial scale are aggregated, represents the value of the i-th node in the time period tT to t, represents the weight matrix, represents the bias vector;
[0226] The time scale division refers to dividing a time period with a fixed length into several small time segments according to a preset unit time length. When different unit times are used, different time scale divisions will be generated. Assuming that the first t In the time scale, the unit time length is Then the input data will be split into time segments, where T is the length of the original time segment; the i-th node is in the l-th t The jth time segment of the time scale represents as follows:
[0227]
[0228] in, represents the value of the i-th node at time t, j represents the time segment index, Indicates the unit time length;
[0229] Will Input to the fully connected layer to get the lth t Aggregate representation of the jth time segment of a time scale
[0230]
[0231] Where LN(·) represents the normalization operation, Indicates the first t The j-th time segment of a time scale is represented by represents the weight matrix, represents the bias vector;
[0232] The temporal and spatial dual-scale Transformer prediction module includes a temporal Transformer unit and a spatial Transformer unit; the prediction process is as follows:
[0233] First, randomly select a speed sequence sample set X with a length of T short =θT long Short time series data The short time series data X short It is input into the high-dimensional spatiotemporal data embedding module together with the road network adjacency matrix A to obtain the high-dimensional spatiotemporal representation of the sequence data.
[0234] H t-T:t =FC(X short +X d +X w )+X s
[0235] Among them, FC(·) represents the fully connected layer, X d and X w represents the daily and weekly period embedding, X s represents the spatial embedding of the adjacency matrix after Laplace eigentransformation;
[0236] Then, the high-dimensional spatiotemporal data is embedded into the high-dimensional spatiotemporal representation H of the module. t-T:t Input to the temporal Transformer unit; the temporal Transformer unit has a total of l t Layer, divide the temporal Transformer unit into time scales, and the final output is expressed as The network node is at the lth t The updating process on a time scale is as follows:
[0237]
[0238] Where LN(·) represents the normalization operation, Indicates the first t -1 time scale division result, MHA(·) is the multi-head self-attention mechanism, are the parameters of the temporal multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder, is the result of time scale division, For the first t The intermediate representation of the time scale, FFN(·) represents the feed-forward layer;
[0239] Next, the final output of the temporal Transformer unit is represented as Input to the spatial Transformer unit; the spatial Transformer unit has a total of l s Layer, the spatial Transformer unit is divided into spatial scales, and the final output is expressed as The network node is at the lth s The updating process on each spatial scale is as follows:
[0240]
[0241] Where LN(·) represents the normalization operation, Indicates the first s -1 spatial scale division result, MHA(·) is the multi-head self-attention mechanism, are the parameters of the spatial multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder, is the result of spatial scale division, For the first s The intermediate representation of spatial scales, FFN(·) represents the feed-forward layer;
[0242] Finally, the output of the spatial Transformer unit Input into a fully connected layer to get the predicted value of the road network node speed in the future The calculation is as follows:
[0243]
[0244] Among them, FC(·) is a fully connected layer;
[0245] Step 7: Design the loss function of the spatiotemporal dual-cross-scale Transformer prediction module, perform iterative training, optimize the learning parameters of the prediction module, and use a variety of evaluation functions to evaluate the prediction performance of the prediction module;
[0246] The training loss function and evaluation function of the spatiotemporal dual-cross-scale Transformer prediction module are designed as follows:
[0247] The training loss function design process is:
[0248] The mean absolute error (MAE) is selected as the training loss function to reduce the mean absolute error between the predicted value and the true value; the training loss function calculation formula is:
[0249]
[0250] Among them, Y i represents the true value of the node speed of the path network at the i-th time. represents the predicted value of the network node speed at the i-th time step, and Q represents the length of the i-th time step;
[0251] The evaluation function design process is as follows:
[0252] The mean square error (MSE) and root mean square error (RMSE) are selected as the evaluation loss function to evaluate the performance of the model on the test set; the evaluation loss function calculation formula is:
[0253]
[0254] Among them, Y i represents the true value of the node velocity of the step network at the Ith time. represents the predicted value of the network node speed at the i-th time step, and Q represents the length of the i-th time step.
Claims
1. A method for predicting cross-scale velocity of open-pit mine trucks based on spatiotemporal masking pre-training, characterized by: A speed prediction model is designed based on the speed data of open-pit mine trucks, and the GPS trajectory data of open-pit mine trucks is preprocessed to obtain the speed sequence sample set and the road network adjacency matrix; A long time series is selected from the speed sequence sample set, and after being processed by the high-dimensional spatiotemporal data embedding module and the spatiotemporal position encoding module together with the adjacency matrix, it is input into the spatiotemporal coupling masked pre-training module to obtain the parameters of the spatiotemporal encoder in the spatiotemporal masked autoencoder, and the parameters are used to configure the spatiotemporal dual-cross-scale Transformer prediction module; a short time series is selected from the speed sequence sample set, and after being processed by the high-dimensional spatiotemporal data embedding module together with the adjacency matrix, it is input into the spatiotemporal dual-cross-scale Transformer prediction module to predict the speed data of the road network nodes for a period of time in the future; The specific steps are as follows: Step 1: pre-process the GPS trajectory data of the open-pit mine truck, use the existing road network automatic generation algorithm, and generate a road network adjacency matrix based on the pre-processed GPS trajectory data; represent each node of the road network as a speed sequence sample, and use the speed sequence samples of all nodes to construct a speed sequence sample set X; Step 2: Use random sampling method to select long time series data X from the speed series sample set X obtained in step 1 long , the long time series data X long Together with the adjacency matrix A, it is input into the high-dimensional spatiotemporal data embedding module to realize the mapping of data from the original feature space to the high-dimensional feature space, and obtain the high-dimensional spatiotemporal representation of the sequence data; Step 3: Represent the obtained sequence data in high-dimensional space-time e Input includes time position code E tp , spatial position encoding E sp The spatiotemporal position encoding module is used to obtain the final embedding representation E of the spatiotemporal coupling masking pre-training module; Step 4: Input the final embedding representation E of the spatiotemporal coupled masking pre-training module into the spatiotemporal data masking unit of the spatiotemporal coupled masking pre-training module, and perform masking using a randomly sampled spatiotemporal masking block generation method to obtain masked data X; Step 5, design a pre-training loss function, input the masked data X obtained in step 4 into the spatiotemporal masked autoencoder of the spatiotemporal coupled masked pre-training module for pre-training, and obtain the parameters of the temporal multi-head attention layer and the spatial multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder; the spatiotemporal masked autoencoder of the spatiotemporal coupled masked pre-training module includes three core parts: a spatiotemporal encoder, a masked token replacement module, and a spatiotemporal reconstruction decoder; Step 6: Select short time series data X from the speed series sample set obtained in step 1 by random sampling. short , the short time series data X short The high-dimensional spatiotemporal data embedding module is input together with the adjacency matrix to obtain a high-dimensional spatiotemporal representation. The high-dimensional spatiotemporal representation and the parameters of the temporal multi-head attention layer and the spatial multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder obtained in step 5 are used in the spatiotemporal dual-cross-scale Transformer prediction module to predict the speed of road network nodes in the future. Step 7: Design the loss function of the spatiotemporal dual-cross-scale Transformer prediction module, perform iterative training, optimize the learning parameters of the prediction module, and use a variety of evaluation functions to evaluate the prediction performance of the prediction module.
2. The method for predicting cross-scale velocity of open-pit mine trucks based on spatiotemporal masking pre-training according to claim 1 is characterized in that: In step 1, the GPS trajectory data of the open-pit mine truck is preprocessed, including the generation of the road network adjacency matrix and the construction of the speed sequence sample set; The road network adjacency matrix generation process is: First, collect GPS track data of open-pit mining trucks. These data record the driving information of mining trucks in the mining area, including mining truck ID, latitude and longitude coordinates, speed and timestamp; Next, the GPS trajectory data is preprocessed, including removing noise, outliers, and duplicate data; Then, the existing automatic road network generation algorithm is used to analyze the preprocessed GPS trajectory data to identify key locations in the trajectory, such as road intersections, turning points, etc. These locations will be set as nodes of the road network V = {v1, v2, ... v N }; After determining the road network nodes, the algorithm forms a road network adjacency matrix A∈R according to the connectivity of the trajectory data N×N , forming a road network G; The constructed speed sequence sample set: First, for each node v in the road network i , extract the speed data of all mining trucks passing through the node, including the speed value and the corresponding timestamp; Then, based on the extracted speed values and timestamps, a speed sequence sample x is constructed for each node in the road network at a time interval of 5 minutes. i ∈R T×C ; Finally, the speed sequence samples of all road network nodes are aggregated to construct the speed sequence sample set X = {x1, x2, …x N }∈R T ×N×C ; Where T represents time, N represents the number of road network nodes, C represents node speed, and C=1 represents a single speed dimension.
3. The method for predicting cross-scale velocity of open-pit mine trucks based on spatiotemporal masking pre-training according to claim 1 is characterized in that: In step 2, the long time series data X long Together with the adjacency matrix generated in step 1, it is input into the high-dimensional spatiotemporal data embedding module for processing: The process of randomly extracting long time series data: First, set the long time series extraction ratio to Then, determine the length of the long time series Finally, a long time series data with a time length of The high-dimensional spatiotemporal data embedding module includes time dimension embedding, space dimension embedding, and period embedding; The time dimension embedding process is: First, the long time series data X long Divide into L time periods of length T P =T long / L time blocks, and splice along the feature dimension to obtain the spliced data Then, the concatenated data X p Through a linear layer, we get the time embedding vector E t =W p X p +b p Among them, W p is the weight matrix, b p is the bias vector; The spatial dimension embedding process is: First, the normalized Laplacian matrix L∈R is calculated from the adjacency matrix A N×N : L=I n -D -1 / 2 AD -1 / 2 =U T ΛU Among them, I n is the identity matrix, D is the degree matrix, and U is the eigenvector matrix; Then, in the eigenvector matrix U, select the K smallest non-zero eigenvectors and project the K smallest non-zero eigenvectors to the D-dimensional features to generate the spatial embedding vector The cycle embedding includes daily cycle embedding and weekly cycle embedding, and the embedding process is: First, the timestamp corresponding to the speed sequence sample is converted into the time proportion T of the sample in a day d and the proportion of time in a week T w , each of which passes through the daily embedding linear layer and the weekly embedding linear layer to obtain the daily period embedding vector and the periodic embedding vector The processing process is: E d =W d T d +b d E w =W w T w +b w Among them, W d and b d is the weight matrix and bias vector of the embedding linear layer, W w and b w The weight matrix and bias vector for the week embedding linear layer; Finally, the above embedding vectors are summed to obtain the output E of the high-dimensional spatiotemporal data embedding module e : AND e =And t +E s +E d +E w Among them, E t is the time embedding vector, E s is the spatial embedding vector, E d is the daily period embedding vector, E w is the periodic embedding vector, It is a high-dimensional spatiotemporal representation of sequence data.
4. The method for predicting cross-scale velocity of open-pit mine trucks based on spatiotemporal masking pre-training according to claim 1 is characterized in that: In step 3, the high-dimensional spatiotemporal representation E of the sequence data e Input includes time position code E tp , spatial position encoding E sp The spatiotemporal position encoding module process is as follows: First, the high-dimensional spacetime is represented by E e Perform time position encoding E tp , the formula is as follows: Yes tp [t,n,2i]=sin(t / 10000 2i / D ) E tp [t,n,2i+1]=cos(t / 10000 2i / D ) Where t represents the time step, n represents the length of the time series, i represents the index in the time embedding dimension D, sin(·) represents the sine function, and cos(·) represents the cosine function. Indicates time position coding; Then, the high-dimensional space-time representation E e Spatial position encoding sp , the formula is as follows: Yes sp [t,n,2i]=sin(n / 10000 2i / D ) It is sp [t,n,2i+1]=cos(n / 10000 2i / D ) Where t represents the time step, n represents the length of the time series, i represents the index in the spatial embedding dimension D, sin(·) represents the sine function, and cos(·) represents the cosine function. Represents spatial position encoding; Finally, the high-dimensional spacetime is represented by E e , time position coding E tp and spatial position encoding E sp Add them together to obtain the output E of the spatiotemporal position encoding module, and the formula is as follows: E=E e +E tp +E sp in, The final embedding representation of the pre-trained module is masked for spatiotemporal coupling.
5. The method for predicting cross-scale velocity of open-pit mine trucks based on spatiotemporal masking pre-training according to claim 1 is characterized in that: In step 4, the processing process of the spatiotemporal data masking unit of the spatiotemporal coupled masking pre-training module is as follows: First, the final embedding representation E is input into the spatiotemporal data masking unit, and the masking ratio is determined to be α; Then, the final embedding representation E is split into patches as follows: N patch =T p ND / l m Among them, N patch is the total number of patches, T p is the time step of the final embedding representation E, N is the number of nodes, D is the feature dimension of the final embedding representation E, l m The length of each patch block is based on the total number of patches N patch Set the patch block index set N = {1, 2, ..., N patch }; Next, according to the masking ratio α and the total number of patches N patch Calculate the number of masked patches N m =αN patch ; Randomly select N from the patch block index set N m index, set the mask patch index set Next, create a masking matrix with all initial values 1 Using the masked patch index set P, set the values of the patches corresponding to these indexes in the masking matrix M to 0 to generate a new masking matrix Finally, the masked data is calculated Masking data As input to the spatiotemporal masked autoencoder; Where E is the final embedding representation of the spatiotemporal coupled masked pre-training module, M' is the masking matrix, To mask the data.
6. The method for predicting cross-scale velocity of open-pit mine trucks based on spatiotemporal masking pre-training according to claim 1 is characterized in that: In step 5, the masked data X is input into the spatiotemporal masked autoencoder of the spatiotemporal coupled masked pre-training module for pre-training. During the pre-training process, a pre-training loss function is designed to evaluate the performance of the model in the reconstruction task and guide the model to continuously optimize its parameters during the training process; the spatiotemporal encoder includes a temporal multi-head attention layer and a spatial multi-head attention layer, the masked token replacement module replaces the masked part of the spatiotemporal encoder output data with a specific masked token, and the spatiotemporal reconstruction decoder includes an attention layer and a fully connected layer; The pre-training process of spatiotemporal masked autoencoder is: First, mask the data Input to the spatiotemporal encoder, masking data After passing through the time multi-head attention layer and the space multi-head attention layer respectively, the outputs of the two attention layers are merged to obtain the output X of the spatiotemporal encoder enc-out ; The calculation process of the space-time encoder is as follows: X temp-att =Attention(Q1,K1,V1) X spatial-att =Attention(Q2,K2,V2) X enc-out =X temp-att +X spatial-att in, To mask the data, Q1 is the query vector of the temporal multi-head attention layer, K1 is the key vector of the temporal multi-head attention layer, and V1 is the value vector of the temporal multi-head attention layer. is the weight matrix of Q1, is the weight matrix of K1, is the weight matrix of V1, X temp-att is the output of the temporal multi-head attention layer; Q2 is the query vector of the spatial multi-head attention layer, K2 is the key vector of the spatial multi-head attention layer, and V2 is the value vector of the temporal multi-head attention layer. is the weight matrix of Q2, is the weight matrix of K2, is the weight matrix of V2, X spatial-att is the output of the spatial multi-head attention layer; Attention represents the attention mechanism calculation, X enc-out is the output of the space-time encoder; Second, the output X of the space-time encoder enc-out Passed to the masked token replacement module; The process of masked token replacement is as follows: First, construct an output X enc-out Masked token vector T of the same shape m ; Next, construct an output X that is consistent with the spatiotemporal encoder enc-out A mask tensor M of the same shape p , the mask tensor M p The position value specified by the masked patch index set P is 1, and the other position values are 0; finally, apply M p T m With X enc-out Merge to generate masked token replacement result X dec-in ; The merging process is as follows: X dec-in =M p ⊙T m +(1-M p )⊙X enc-out Among them, M p is the mask tensor, T m is the masked token vector, X enc-out is the output of the spatiotemporal encoder, ⊙ represents element-by-element multiplication; Third, replace the masked token with the result X dec-in Input into the space-time reconstruction decoder, decode and reconstruct the data, and obtain the output representation X of the space-time reconstruction decoder dec-out ; The reconstruction decoder process is as follows: X dec-out =FC(Attention(X dec-in )) Among them, Attention represents the attention layer, FC(·) represents the fully connected layer; Fourth, calculate the masked reconstruction part after training and mask the true value Q∈R T×N×D , the calculation process is as follows: Among them, X dec-in is the masked token replacement result, M′ is the masking matrix, To mask the data, ⊙ represents element-by-element multiplication; Fifth, based on the masked true value Q and the masked reconstruction part Design a pre-training loss function and adjust the model parameters through optimization strategy; the calculation formula of the pre-training loss function is as follows: Among them, L is the loss value, T is the time step, N is the number of nodes, and D is the hidden dimension. is the masked reconstruction part after pre-training, Q is the masked true value, t is the time index, n is the node index, and d is the hidden dimension index; Sixth, after the pre-training is completed, the weight matrix of the temporal multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder is obtained Weight matrix of the spatial multi-head attention layer 7. The method for predicting cross-scale velocity of open-pit mine trucks based on spatiotemporal masking pre-training according to claim 1 is characterized in that: In step 6, the spatiotemporal dual-scale Transformer prediction module involves the division of spatial scale and temporal scale. The division process is: The short time series data and the road network adjacency matrix are processed by the high-dimensional spatiotemporal data embedding module, and the obtained high-dimensional spatiotemporal representation H t-T:t ∈R T×N×D It is expressed as follows: in, Represents the value of the Nth node in the time period tT to t; The spatial scale division refers to dividing the geographic space area into multiple grid units according to the pre-set length and width standards; the size of these grid units will vary due to the use of different length and width standards, thus producing different spatial scale divisions; then, the nodes in the road network will be assigned to the corresponding grid units according to their geographical locations, and the nodes in the same grid unit will be aggregated as the representation of the grid unit; the first s Aggregate representation of the mth grid at each spatial scale as follows: Where LN(·) represents the normalization operation, Indicates that the first s All values of the mth grid at the spatial scale are aggregated, represents the value of the i-th node in the time period tT to t, represents the weight matrix, represents the bias vector; The time scale division refers to dividing a time period with a fixed length into several small time segments according to a preset unit time length. When different unit times are used, different time scale divisions will be generated. Assuming that the first t In the time scale, the unit time length is Then the input data will be split into time segments, where T is the length of the original time segment; the i-th node is in the l-th t The jth time segment of the time scale represents as follows: in, represents the value of the i-th node at time t, j represents the time segment index, Indicates the unit time length; Will Input to the fully connected layer to get the lth t Aggregate representation of the jth time segment of a time scale Where LN(·) represents the normalization operation, Indicates the first t The j-th time segment of a time scale is represented by represents the weight matrix, Represents the bias vector.
8. The method for predicting cross-scale velocity of open-pit mine trucks based on spatiotemporal masking pre-training according to claim 1 is characterized in that: In step 6, the spatiotemporal dual-cross-scale Transformer prediction module includes a temporal Transformer unit and a spatial Transformer unit; the prediction process is as follows: First, randomly select a speed sequence sample set X with a length of T short =θT long Short time series data The short time series data X short It is input into the high-dimensional spatiotemporal data embedding module together with the road network adjacency matrix A to obtain the high-dimensional spatiotemporal representation of the sequence data. H t-T:t =FC(X short +X d +X w )+X s Among them, FC(·) represents the fully connected layer, X d and X w represents the daily and weekly period embedding, X s represents the spatial embedding of the adjacency matrix after Laplace eigentransformation; Then, the high-dimensional spatiotemporal data is embedded into the high-dimensional spatiotemporal representation H of the module. t-T:t Input to the temporal Transformer unit; the temporal Transformer unit has a total of l t Layer, divide the temporal Transformer unit into time scales, and the final output is expressed as The network node is at the lth t The updating process on a time scale is as follows: Where LN(·) represents the normalization operation, Indicates the first t -1 time scale division result, MHA(·) is the multi-head self-attention mechanism, are the parameters of the temporal multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder, is the result of time scale division, For the first t The intermediate representation of the time scale, FFN(·) represents the feed-forward layer; Next, the final output of the temporal Transformer unit is represented as Input to the spatial Transformer unit; the spatial Transformer unit has a total of l s Layer, the spatial Transformer unit is divided into spatial scales, and the final output is expressed as The network node is at the lth s The updating process on each spatial scale is as follows: Where LN(·) represents the normalization operation, represents the division result of the ls-1th spatial scale, MHA(·) is the multi-head self-attention mechanism, are the parameters of the spatial multi-head attention layer of the spatiotemporal encoder in the spatiotemporal masked autoencoder, is the result of spatial scale division, For the first s The intermediate representation of spatial scales, FFN(·) represents the feed-forward layer; Finally, the output of the spatial Transformer unit Input into a fully connected layer to get the predicted value of the road network node speed in the future The calculation is as follows: Among them, FC(·) is a fully connected layer.
9. The method for predicting cross-scale velocity of open-pit mine trucks based on spatiotemporal masking pre-training according to claim 1 is characterized by: In step 7, the training loss function and evaluation function of the spatiotemporal dual-cross-scale Transformer prediction module are designed as follows: The training loss function design process is: The mean absolute error (MAE) is selected as the training loss function to reduce the mean absolute error between the predicted value and the true value; the training loss function calculation formula is: Among them, Y i represents the true value of the node speed of the path network at the i-th time. represents the predicted value of the network node speed at the i-th time step, and Q represents the length of the i-th time step; The evaluation function design process is as follows: The mean square error (MSE) and root mean square error (RMSE) are selected as the evaluation loss function to evaluate the performance of the model on the test set; the evaluation loss function calculation formula is: Among them, Y i represents the true value of the node speed of the path network at the i-th time. represents the predicted value of the network node speed at the i-th time step, and Q represents the length of the i-th time step.