Space-time data prediction method based on space grid coding and large language model

By using spatial grid coding and large language model methods in spatiotemporal data prediction, the problems of insufficient capture of space-time dependency relationships, low real-time, computational efficiency, and insufficient interpretation in the prior art are solved, and higher generalization ability, dynamic modeling ability and interpretation ability are achieved.

CN120045633AActive Publication Date: 2025-05-27BEIJING BIG DATA CENT +2

Patent Information

Application Number
CN202510525900.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Existing spatiotemporal data prediction methods are difficult to effectively capture complex spatiotemporal dependencies, lack real-time and computational efficiency, and are insufficiently interpretable, making it difficult to meet the needs of practical applications.

Method used

The spatial and temporal data prediction method based on spatial grid coding and large language model is adopted. By dividing the geographical area into grid units, sequential encoding or two-dimensional encoding is used to generate unique position identifiers, and a Transformer model based on attention mechanism is constructed. The spatial and temporal instruction optimization prompt template is designed in combination with the large language model to generate the next position prediction and provide explanation.

Benefits of technology

Explicitly capturing the proximity and regionalized semantics of geospatial space improves the generalization ability and spatiotemporal dynamic modeling ability of the model, reduces the dependence on labeled data, and enhances the interpretability and reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045633A_ABST
    Figure CN120045633A_ABST
Patent Text Reader

Abstract

The invention relates to a spatio-temporal data prediction method based on space grid coding and a large language model, and belongs to the field of artificial intelligence and spatio-temporal data analysis. The method comprises the following steps: according to a spatial range of a geographic area, dividing the area into grid units, processing the grid units by adopting a coding method, and identifying each unit through a unique identifier; constructing a space-time dependency model based on the stay record of the user, representing each stay as a tuple, and predicting the next position of the user by analyzing a historical stay sequence and a context stay sequence; and designing a space-time instruction optimization prompt template in combination with a large language model, guiding the model to analyze historical data and context data by introducing grid coding and context perception reasoning, generating next position prediction, and providing explanation for each prediction. According to the invention, through grid coding and hierarchical modeling, the generalization ability of the model is enhanced, and accurate spatio-temporal behavior prediction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and spatio-temporal data analysis, and specifically relates to a spatio-temporal data prediction method based on spatial grid coding and large language models. Background Art

[0002] Spatio-temporal data prediction is a fundamental task in the study of human mobility, aiming to predict a user's next destination by analyzing their movement trajectories. With the popularization of mobile devices and the rapid development of positioning technologies, a large amount of spatio-temporal data (such as GPS trajectories, check-in data, etc.) has been recorded, providing a rich data basis for studying human movement behavior. These data not only contain the geographical location information of users, but also cover multi-dimensional features such as time, stay duration, activity type, etc., providing important support for understanding and predicting human movement patterns. However, existing spatio-temporal data prediction methods still face some challenges: one is data complexity: spatio-temporal data has the characteristics of high-dimensionality, non-linearity, and dynamic changes, and traditional prediction models are difficult to capture complex spatio-temporal dependencies; the second is spatio-temporal dependence: human movement behavior is not only affected by time factors (such as time points, days of the week, seasons), but also restricted by spatial factors (such as geographical location, regional functions), and it is necessary to model spatio-temporal dependencies simultaneously. The third is the requirement for real-time performance: many application scenarios (such as real-time traffic prediction) require the model to be able to quickly process dynamically changing spatio-temporal data, and existing methods have deficiencies in real-time performance and computational efficiency. Fourth is the lack of interpretability: existing prediction models are mostly "black box" models, lacking interpretability of prediction results, and it is difficult to meet the requirements for transparency and credibility in practical applications. Summary of the Invention

[0003] In order to solve the above problems, the present invention provides a spatio-temporal data prediction method based on spatial grid coding and large language models.

[0004] To achieve the above object, the present invention is realized through the following technical solutions: The present invention provides a spatio-temporal data prediction method based on spatial grid coding and large language models, including the following steps: S1. According to the spatial range of the geographical area, divide the area into grid cells, and calculate the number of rows and columns of the grid according to the height and width of the geographical area and the size of the cells. Process the grid cells using an encoding method, and each cell is identified by a unique identifier; the encoding method includes sequential encoding and two-dimensional encoding; Further, according to the height of the geographical area and width as well as the size of the cells , calculate the number of rows and columns of the grid, and the formula is expressed as follows: , Among them, represents the height of the grid, represents the width of the grid.

[0005] Furthermore, the encoding method includes: Sequential encoding: The grid cells are sequentially numbered in row-major order, numbered from left to right and top to bottom. The calculation formula for numbering the grid cell in the i-th row and j-th column is: , where, represents the cell number of sequential encoding, C is the total number of columns of the grid; in the two-dimensional encoding method, each cell is represented by a row index and a column index; Two-dimensional encoding: The grid cells are represented as two-dimensional coordinates (x, y), where x represents the column index and y represents the row index. The calculation formula for numbering the grid cell in the i-th row and j-th column is: , where, represents the cell number of two-dimensional encoding; Each grid cell generates a unique location identifier through the above encoding method. The unique identifier is a sequential number or two-dimensional coordinates, which is used to represent the location identifier in the tuple.

[0006] S2. Based on the user's stay records, construct a spatio-temporal dependence model, represent each stay as a tuple, and predict the user's next location by analyzing the historical stay sequence and the context stay sequence; the tuple includes start time, day of the week, stay duration, location identifier, and location identifier; Furthermore, the user's stay records are generated through the following steps: S211. Extract the timestamped latitude and longitude coordinate points from the GPS trajectory data, and then identify the stay points based on distance and time thresholds; S212. Use the DBSCAN clustering algorithm to cluster the stay points and generate a unique location ID, which is used to represent the location identifier in the tuple; S213. Associate time features and spatial features with each stay point.

[0007] Furthermore, the spatio-temporal dependence model is a Transformer model based on the attention mechanism. Input the historical stay sequence and the context stay sequence into the spatio-temporal dependence model to predict the user's next location, including the following steps: S221. Data preprocessing: Each stop point is represented as a tuple S=(st, dow, dur, pid, tid), where the time features are: st represents the start time, dow represents the day of the week, and dur represents the stay duration; the spatial features are: pid represents the location identifier, and tid represents the position identifier; S222. Sequence generation and filling mechanism: Select all stop point tuples of the user in the past T days , a total of stop points, truncate or pad the stop points to a fixed length after sorting them by time to obtain the historical stay sequence M; select the stop points within the most recent Δt hours , a total of stop points, truncate or pad the stop points to a fixed length after sorting them by time to obtain the context stay sequence N; merge the historical stay sequence M and the context stay sequence N to obtain the input sequence , , represents the feature representation dimension of each stop point; S223. Perform joint spatio-temporal feature encoding on the tuples in the input sequence : Encode the start time st through sine position encoding to obtain the start time vector , , and the formula is as follows: , , where, represents the dimension index, ; represents the dimension of time encoding, represents the cosine function, represents the sine function, represents the start time at the encoding value of the dimension in the even dimension, represents the start time at the encoding value of the dimension in the odd dimension; through the learnable embedding table , convert the day of the week dow into the week embedding vector , represents the original dimension of the embedding vector; Based on grid division to construct the embedding matrix , and then obtain the position identifier vector , Denotes the dimension of the embedding matrix; the operation steps are as follows: each grid cell generates a unique position identifier tid through sequential encoding or two-dimensional encoding, and the embedding matrix is initialized to a Gaussian distribution , and is optimized through backpropagation during training to enhance the cosine similarity of the embedding vectors of adjacent grids and cluster them in the embedding space. Finally, the position identifier vector is obtained from the embedding matrix through a look-up table operation ; The start time vector , the week embedding vector , and the position identifier vector are concatenated and then linearly mapped to obtain a fused comprehensive feature vector, which is expressed by the formula as follows: , where denotes the fused comprehensive feature vector, denotes the weight matrix of the linear projection, denotes the bias term of the linear projection, denotes the concatenated vector, denotes matrix multiplication; The fused comprehensive feature vectors corresponding to L stop points in the input sequence are stacked to form a comprehensive feature vector matrix: , , where denotes the th fused comprehensive feature vector, denotes the comprehensive feature vector matrix; Furthermore, the learnable embedding table is randomly initialized with a uniform distribution at the initial stage of model training and contains 7 row vectors from Monday to Sunday; the parameters are optimized through end-to-end training so that the embedding vectors from Monday to Friday are clustered in the latent space, and Saturday and Sunday form independent semantic clusters, and the week embedding vector is obtained through linear transformation.

[0008] Furthermore, the look-up table operation is specifically as follows: the unique position identifier tid of the input grid cell is used as the row index to locate the corresponding row of the embedding matrix , and the output directly returns that row vector.

[0009] S224. Input the comprehensive feature vector matrix into the spatio-temporal dependence model to predict the probability distribution of the user's next location.

[0010] Furthermore, the spatio-temporal dependence model includes 6 layers of encoders and a prediction layer. Each encoder includes a multi-head attention sub-layer, a residual connection and layer normalization layer, a feed-forward network sub-layer, and a secondary residual and normalization layer. The operation process is as follows: In the first layer encoder, the comprehensive feature vector matrix is input into the multi-head attention sub-layer and split into 8 independent sub-spaces according to the number of heads; each head generates a query and a key and a value through a learnable weight matrix; the scaled dot-product attention is calculated for each head to obtain the output of each attention head ; the outputs of the 8 attention heads are concatenated and linearly transformed through a weight matrix to obtain the output of the multi-head attention sub-layer , and the formula is expressed as: ; the output of the multi-head attention sub-layer passes through the residual connection and layer normalization layer and is residually connected with the comprehensive feature vector matrix and normalized to obtain the output of the residual connection and layer normalization layer, and the formula is expressed as follows: , where represents the layer normalization operation; the output of the residual connection and layer normalization layer passes through two fully-connected networks of the feed-forward network sub-layer to obtain the output of the feed-forward network sub-layer, and the formula is expressed as follows: , where represents the non-linear activation function, and respectively represent the weight matrices of the first fully-connected layer and the second fully-connected layer; and respectively represent the bias vectors of the first fully-connected layer and the second fully-connected layer; the output of the feed-forward network sub-layer passes through the secondary residual and normalization layer and is residually connected with the output of the residual connection and layer normalization layer and normalized to finally obtain the output of the first layer encoder, and the formula is expressed as: ; The output of the first layer encoder is used as the input of the second layer encoder, and the above operations are repeated until the output of the sixth layer encoder is obtained; In the prediction layer, the sequence end vector is extracted from the output of the sixth-layer encoder. , and through the weight matrix the sequence end vector is mapped to the set of candidate positions to obtain the probability distribution of the user's next position, which is expressed by the formula as follows: , where, represents the total number of candidate positions, represents learning the association weights between different positions and , represents the probability distribution that the user's next position is , represents the bias vector of the prediction layer.

[0011] S3. Combining with the large language model, design a spatio-temporal instruction optimization prompt template. By introducing grid encoding and context-aware reasoning, guide the model to analyze historical data and context data, generate the next position prediction, and provide an explanation for each prediction.

[0012] Further, the spatio-temporal instruction optimization prompt template includes a probability mapping module, an instruction generation module, an explanation generation module, and a result verification module; Input the probability distribution of the user's next position into the probability mapping module. The probability distribution is sorted in descending order of values, and the top K positions with the highest probability values are selected. Then, they are filtered according to the probability threshold, and the positions with probabilities greater than the threshold are retained; if the number of candidates is less than K, all positions that meet the threshold are retained; finally, a sorted list of candidate positions is obtained; The historical stay sequence M, the context stay sequence N, and the sorted list of candidate positions are input into the instruction generation module. The structured template engine converts the original input data into a canonical JSON instruction; first, the structured data is converted into serialized text data; then, the serialized text data is injected into a predefined JSON template to obtain a parsable complete JSON instruction; The complete JSON instruction is input into the explanation generation module to generate a structured explanation. The explanation generation module uses a large language model; The structured explanation is input into the result verification module for quality control and standardization processing.

[0013] The advantages of the present invention are: Through grid coding, the model can explicitly capture the proximity and regional semantics of the geographical space, avoiding the prediction bias caused by isolated location IDs in traditional methods; the unified grid division rule enables the model to maintain consistency in different geographical regions (such as across cities), further enhancing the generalization ability; enhancing spatio-temporal dynamic modeling ability: the combination of the Transformer model and the hierarchical modeling strategy realizes the accurate prediction of users' spatio-temporal behavior. The hierarchical modeling strategy adopts a phased and multi-granularity architecture design: at the feature fusion level, primary linear fusion is first performed on time coding, week embedding, and grid embedding, and then the cross-enhancement of long-term and short-term features is achieved through the Transformer high layer. This hierarchical processing enables the model to explicitly distinguish long-term periodic patterns from short-term sudden behaviors; the zero-shot inference ability based on natural language instruction templates greatly reduces the dependence on labeled data. For example, when deploying in a new city, only the grid division parameters (such as grid size) need to be adjusted, and the model can generate predictions without retraining. In contrast, traditional deep learning methods (such as Transformer) need to retrain the embedding layer for each region, consuming a large amount of computing resources. In addition, the inference process driven by the large language model supports few-shot learning, and only a small number of examples are needed to adapt to new user behavior patterns, further expanding the application scope; by directly outputting the prediction reasons through the large language model, users can intuitively understand the model's decision-making logic and improve the decision-making credibility. Description of the Drawings

[0014] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention.

[0015] Figure 1 It is a flowchart of the steps of the method of the present invention. Detailed Embodiments

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0017] Embodiment 1 In this embodiment, as Figure 1 shown, the present invention provides a spatio-temporal data prediction method based on spatial grid coding and a large language model. The specific steps include: S1. Divide the region into grid cells according to the spatial scope of the geographical region, calculate the number of rows and columns of the grid based on the height and width of the geographical region and the size of the cells, and process the grid cells using a coding method to ensure the preservation of spatial continuity and adjacency relationships. Each cell is identified by a unique identifier; the coding method includes sequential coding and two-dimensional coding; Specifically, according to the height of the geographical region and width as well as the size of the cells , calculate the number of rows and columns of the grid, and the formula is as follows: wherein, represents the height of the grid, represents the width of the grid.

[0018] Specifically, the coding method includes: Sequential coding: Number the grid cells in a row-major order, numbering them sequentially from left to right and from top to bottom. The calculation formula for numbering the grid cell in the i-th row and j-th column is: , wherein, represents the cell number of sequential coding, and C is the total number of columns of the grid; in the two-dimensional coding method, each cell is represented by a row index and a column index; Two-dimensional coding: Represent the grid cells as two-dimensional coordinates (x, y), where x represents the column index and y represents the row index. The calculation formula for numbering the grid cell in the i-th row and j-th column is: , wherein, represents the cell number of two-dimensional coding; Each grid cell generates a unique location identifier through the above coding method, and the unique identifier is a sequential number or two-dimensional coordinates, which is used to represent the location identifier in the tuple.

[0019] S2. Based on the user's stay records, construct a spatio-temporal dependence model, represent each stay as a tuple, and predict the user's next location by analyzing the historical stay sequence and the context stay sequence; the tuple includes the start time, day of the week, stay duration, location identifier, and location identifier; Specifically, the user's stay records are generated through the following steps: S211. Extract the longitude and latitude coordinate points with timestamps from the GPS trajectory data, and then identify the stay points based on distance and time thresholds; if the user stays within a radius of 200 meters for more than 30 minutes, it is regarded as a stay point; S212. Use the DBSCAN clustering algorithm to cluster the stay points and generate a unique location ID to represent the location identifier in the tuple; this location ID is dynamically generated by DBSCAN and depends on the actual distribution density of the user's stay points; even if two stay points belong to the same grid cell (the same , if the physical distance exceeds the clustering threshold, they may be divided into different clusters (different location IDs); conversely, the stay points across grid cells may be clustered into the same location ID; where the parameters of DBSCAN are set as the neighborhood radius of 200 meters, and the minimum number of samples is adjusted according to the data distribution; the isolated points that do not form an effective cluster (such as short stays or GPS drift points) are filtered, and only the points with significant stay significance are retained; S213. Associate the time features and space features with each stay point.

[0020] Specifically, the spatio-temporal dependence model is a Transformer model based on the attention mechanism. Input the historical stay sequence and the context stay sequence into the spatio-temporal dependence model to predict the user's next location, including the following steps: S221. Data preprocessing: Represent each stay point as a tuple S=(st, dow, dur, pid, tid), where, time features: st represents the start time, dow represents the day of the week, dur represents the stay duration; space features: pid represents the location identifier, tid represents the position identifier; S222. Sequence generation and filling mechanism: Select all the stay point tuples of the user in the past T days , a total of stay points, sort the stay points by time and then truncate or fill them to a fixed length to obtain the historical stay sequence M; Select the stay points within the most recent Δt hours , a total of stay points, sort the stay points by time and then truncate or fill them to a fixed length to obtain the context stay sequence N; Merge the historical stay sequence M and the context stay sequence N to obtain the input sequence , , represents the feature representation dimension of each stay point; Preferably, 1. Historical sequence M: Time range: Stay points within the past T = 14 days, truncated / filled to length = 100; After sorting by time from early to late, if the number of stop points exceeds 100, only the last 100 are retained (i.e., the old data in the front is truncated); if the number of stop points is less than 100, zero vectors are padded at the beginning of the sequence. 2. Context sequence N: Time range: Stop points within the most recent Δt = 24 hours, truncated / filled to length = 100 (the truncation / filling method is the same as the historical sequence).

[0021] S223. Joint spatio-temporal feature encoding for the tuples in the input sequence : Encode the start time st through sine position encoding to obtain the start time vector , , the formula is as follows: , , Among them, represents the dimension index, , represents the dimension of time encoding, which determines the dimension of the expression ability of time features, dimension ; represents the cosine function, represents the sine function, represents the start time The encoding value at the dimension in the even dimension, represents the start time The encoding value at the dimension in the odd dimension; Through the learnable embedding table , convert the day of the week dow into the week embedding vector , represents the original dimension of the embedding vector; Preferably, the learnable embedding table is randomly initialized by a uniform distribution at the initial stage of model training, and contains 7 row vectors from Monday to Sunday, is defaulted to 64; Optimize the parameters through end-to-end training, so that the embedding vectors from Monday to Friday cluster in the latent space, and Saturday and Sunday form independent semantic clusters. If the output dimension is required to be 64, the week embedding vector , is obtained through linear transformation.

[0022] Based on grid division to construct the embedding matrix , and then obtain the location identifier vector , represents the dimension of the embedding matrix, = 256; The operation steps are as follows: Each grid cell generates a unique location identifier tid through sequential encoding or two-dimensional encoding, and the embedding matrix is initialized to a Gaussian distribution , and is optimized through backpropagation during training to increase the cosine similarity of the embedding vectors of adjacent grids and cluster them in the embedding space. Finally, a 256-dimensional location identifier vector is obtained from the embedding matrix through a look-up table operation ; Preferably, the look-up table operation is specifically as follows: Input the unique location identifier tid of the grid cell, use tid as the row index to locate the corresponding row of the embedding matrix , and directly return the 256-dimensional vector of this row as the output.

[0023] The start time vector , the week embedding vector , and the location identifier vector are concatenated and then linearly mapped to obtain a fused comprehensive feature vector, which is expressed by the following formula: , where represents the fused comprehensive feature vector, ; represents the weight matrix of the linear projection, ; represents the bias term of the linear projection, represents the concatenated vector, represents matrix multiplication; Stack the fused comprehensive feature vectors corresponding to L stop points in the input sequence to form a comprehensive feature vector matrix: , , where represents the -th fused comprehensive feature vector, represents the comprehensive feature vector matrix; S224. Input the comprehensive feature vector matrix into the spatio-temporal dependence model to predict the probability distribution of the user's next location.

[0024] Specifically, the spatio-temporal dependence model includes 6 encoder layers and a prediction layer. Each encoder includes a multi-head attention sub-layer, a residual connection and layer normalization layer, a feed-forward network sub-layer, and a secondary residual and normalization layer. The operation process is as follows: In the first encoder layer, the comprehensive feature vector matrix is input into the multi-head attention sub-layer and split into 8 independent sub-spaces according to the number of heads; the dimension of each attention head is , each head generates a query through a learnable weight matrix , the key and the value , which is expressed by the formula as follows: , , , , where, , , represent the query , the key and the value of the learnable weight matrix respectively; calculate the scaled dot-product attention for each head to obtain the output of each attention head , which is expressed by the formula as follows: , where, represents the Softmax function; concatenate the outputs of the 8 attention heads into a 512-dimensional vector and perform a linear transformation through the weight matrix to obtain the output of the multi-head attention sublayer , which is expressed as: ; the output of the multi-head attention sublayer passes through a residual connection and a layer normalization layer and a comprehensive feature vector matrix for residual connection and normalization processing to obtain the output of the residual connection and the layer normalization layer , which is expressed by the formula as follows: , where, represents the layer normalization processing operation; the output of the residual connection and the layer normalization layer passes through two fully-connected networks of the feed-forward network sublayer for processing to obtain the output of the feed-forward network sublayer , which is expressed by the formula as follows: , where, represents the non-linear activation function, and represent the weight matrices of the first fully-connected layer and the second fully-connected layer respectively, ; and represent the bias vectors of the first fully-connected layer and the second fully-connected layer respectively; the output of the feed-forward network sublayer passes through a secondary residual and normalization layer and the output of the residual connection and the layer normalization layer Perform residual connection and normalization, and finally obtain the output of the first-layer encoder , which is expressed by the formula: ; Take the output of the first-layer encoder as the input of the second-layer encoder, and repeat the above operations until the output of the sixth-layer encoder is obtained ; In the prediction layer, extract the sequence end vector from the output of the sixth-layer encoder . Further, after the user stay sequence is processed by the six-layer encoder, the output of the sixth layer contains the high-order spatio-temporal feature representations at each time step. Since the input sequence is strictly arranged in ascending order of time (the earliest behavior is in the front and the latest behavior is at the end), the system directly extracts the last row vector of this matrix as the spatio-temporal state encoding of the user's most recent moment, where L is the sequence length (L = ); map the sequence end vector to the set of candidate positions through the weight matrix to obtain the probability distribution of the user's next position. The formula is expressed as follows: , where represents the total number of candidate positions (i.e., the number of grids ), represents learning the association weights between different positions and , represents the probability distribution that the user's next position is , represents the bias vector of the prediction layer.

[0025] S3. Combine with the large language model, design a spatio-temporal instruction optimization prompt template, and through introducing grid encoding and context-aware reasoning, guide the model to analyze historical data and context data, generate the next position prediction, and provide an explanation for each prediction to enhance the interpretability and reliability of the model.

[0026] Specifically, the spatio-temporal instruction optimization prompt template includes a probability mapping module, an instruction generation module, an explanation generation module, and a result verification module; Take the probability distribution of the user's next position Input to the probability mapping module, the probability distribution is sorted in descending order of values, and the top K positions with the highest probability values are selected. Then, filtering is performed according to the probability threshold, and the positions with probabilities greater than the threshold are retained; if the number of candidates is less than K, all positions that meet the threshold are retained; finally, a sorted list of candidate positions is obtained; the steps are as follows: The latitude and longitude coordinates in the user trajectory data are mapped to a unified location identifier tid through geographic grid division, and then DBSCAN spatial clustering analysis is performed on all coordinate points corresponding to each location identifier tid. By setting parameters such as the neighborhood radius eps and the minimum number of points min_samples, regions with different semantics within the same grid are distinguished into independent clusters, and a semantic location with a hierarchical structure is assigned to each cluster: location identifier pid. Noise points that are not assigned to any cluster during the clustering process are directly removed. Finally, a mapping dictionary with the location identifier tid as the key and the location identifier pid as the value is constructed. In the preprocessing stage, the system first maps the latitude and longitude coordinates in the user trajectory data to a unified grid cell code (tid) through geographic grid division. For example, when using a 10×10 grid division, each coordinate point will be assigned a unique tid value according to its row and column positions. Next, to endow these grid positions with richer semantic information, the system performs DBSCAN spatial clustering analysis on all coordinate points corresponding to each tid. By setting parameters such as the neighborhood radius (eps) and the minimum number of points (min_samples), regions with different semantics within the same grid (such as the main entrance of a shopping mall and a parking lot) are distinguished into independent clusters, and a semantic location ID (pid) with a hierarchical structure is assigned to each cluster. Its format is usually a combination of "tid_clusterIndex" to ensure global uniqueness. Noise points that are not assigned to any cluster during the clustering process are directly removed. Finally, the system constructs a mapping dictionary with tid as the key and the corresponding pid list as the value. In the candidate location screening stage, for each location to be queried, its tid is first calculated based on the coordinates, and then the mapping table is queried to obtain its semantic pid. If the tid does not exist in the mapping table or the corresponding pid list is empty, the location is marked as an "unknown location" and removed from the candidate list, thereby ensuring that the location data processed by the model not only retains the grid-based spatial structure characteristics but also incorporates the semantic hierarchical information brought by clustering, while effectively filtering out noise points in the data. This processing method enables subsequent spatio-temporal prediction models to be trained and inferred based on richer semantic location information.

[0027] The historical stay sequence M, the context stay sequence N, and the sorted candidate position list are input into the instruction generation module, and the structured template engine converts the original input data into a strictly standardized JSON instruction; first, the structured data is converted into serialized text data; then the serialized text data is injected into a predefined JSON template to obtain a parsable complete JSON instruction; The structured template engine first receives the historical stay sequence M ([{pid: "1230", tid: 35, st: "2023-10-01 09:00:00", dow: "Monday", dur: 60}]), the context stay sequence N (in the same format as M), and the filtered Top-K candidate position list. The candidate position list is as follows: ([{pid: "7890", tid: 55, prob: 0.72}, {pid: "5671", tid: 46, prob: 0.15}]). These structured data are converted into serialized text data: The historical sequence is organized as "User's stay records in the past 14 days: Location 1230 (grid 35) stayed for 60 minutes at 09:00 on October 1st (Monday)...", the context sequence is converted into "Recent 24-hour activities: Location 5671 (grid 46) stayed for 30 minutes...", and the candidate positions are sorted by probability as "1. Location 7890 (grid 55, probability 72%) 2. Location 5671 (grid 46, probability 15%)". These serialized text data are then injected into a predefined JSON template, which has three key data blocks (history, context, candidates) and strict response requirements built-in. For example, it is mandatory to include the predicted_pid field and stipulate that the explanation must be presented in three dimensions: historical mode (the frequency of the target position in the historical data needs to be calculated), spatial relationship (the Manhattan distance needs to be calculated based on the grid encoding), and time dependence (the proportion of accesses in the same time period needs to be counted). When processing the candidate position where pid = "7890" and prob = 0.72, the template engine will automatically emphasize in the instruction the need to focus on analyzing the historical access pattern of this position in the early morning on weekdays (such as the occurrence frequency reaching 60%), the grid distance from the current position (such as 2 units apart), and the access proportion in the 09:00 - 10:00 time period (such as 42%). Finally, a machine-parsable complete JSON instruction is generated, and its structure fully follows the preset template to ensure that the downstream large language model can accurately understand the task requirements.

[0028] The complete JSON instruction is input into the interpretation generation module to generate a structured interpretation. The interpretation generation module uses a large language model. After receiving the JSON structured instruction, the GPT-3.5 model generates a structured interpretation through the following process: First, the model receives the serialized text input from the instruction generation module, which includes the user's historical stay records, context location information, and candidate location list. The model calculates the token weights through the self-attention mechanism of the decoder and uses beam search (beam width k = 3) to generate coherent text. During the generation process, the model organizes the content based on a predefined logical chain (historical pattern → spatial proximity → temporal dependence). The output is an itemized list of interpretation content (historical pattern, spatial relationship, temporal dependence). Historical pattern: The access frequency is statistically calculated from the user's historical records (e.g., pid1 appears 5 times every Monday), and is summarized into natural language by the large language model. Spatial relationship: Calculate the grid distance (such as Manhattan distance) between the predicted location and the context location: , if ≤ 1, it is marked as adjacent. Temporal dependence: Statistically calculate the access proportion of this location in the same time period (e.g., 17:00 - 18:00) in the total data, and the model converts it into a peak period description.

[0029] The structured interpretation is input into the result verification module for quality control and standardization processing. This module conducts comprehensive quality control and standardization processing on the prediction results generated by GPT-3.5. When receiving the natural language text output by the model (such as JSON data containing the predicted location ID and three interpretations), the verification engine will first strictly check whether the output structure conforms to the preset specifications, including confirming that the predicted_pid field exists and is a valid location identifier, the probability value is within the range of 0 - 1, and the three interpretation dimensions of historical pattern / spatial relationship / temporal dependence are complete without omission. Then the system will conduct in-depth logical consistency verification: verify the accuracy of the spatial relationship description by recalculating the Manhattan distance of the grid coordinates; confirm whether the access frequency matches the original data by querying the historical database; and at the same time check whether the time distribution statistical values are reasonably calculated. For the results that pass the verification, the module will perform a standardized conversion and extract key indicators to generate a structured output. When field missing, data contradiction, or calculation deviation is found, the system will mark the verification as failed and trigger the result regeneration process to ensure that the final delivered prediction results not only maintain the readability advantage of natural language but also have the strict structured characteristics that can be processed by machines.

[0030] Example 2 In this factual example, based on the actual application scenario, combined with data and processes, the prediction process of the method of the present invention is described in detail. This example takes the user location prediction in an urban area as an example, covering the complete process from geographical area division to final location prediction. In an urban area, the geographical range is a height = 10 km and a width = 15 km, the goal is to predict the user's next location within this area. The following are the specific steps: Step S1: According to the spatial scope of the geographical area, divide the area into grid cells. Calculate the number of rows and columns of the grid based on the height and width of the geographical area and the size of the cells; then process the spatial data using sequential encoding or two-dimensional encoding methods to ensure the preservation of spatial continuity and adjacency relationships; finally, each cell is identified by a unique identifier; Step S2: Based on the user's stay records, construct a spatio-temporal dependence model. Represent each stay as a tuple (start time, day of the week, stay duration, location identifier, position identifier). Predict the user's next location by analyzing the historical stay sequence and the context stay sequence; Step S3: Combine with a large language model, design a spatio-temporal instruction optimization prompt template. By introducing grid encoding and context-aware reasoning, guide the model to analyze historical data and context data, generate the next location prediction, and provide an explanation for each prediction to enhance the interpretability and reliability of the model.

[0031] Furthermore, in step S1, the steps of grid division and encoding include: Grid division: The size of the grid cell is set to a height = 1 km and a width = 1 km. Calculate the number of rows and columns of the grid: , , Therefore, the geographical area is divided into 10×15 grid cells.

[0032] Encode the grid: Sequential encoding: Each grid cell is encoded in a row-major order. For example, the encoding of the grid cell in the i-th row and j-th column is: . For example, the encoding of the grid cell in the 3rd row and 5th column is: 3 15 5 = 50; Two-dimensional encoding: Each grid cell is represented as a two-dimensional encoding (x, y), where x represents the column index and y represents the row index. For example, the grid cell in the 3rd row and 5th column is represented as: (5 - 1, 3 - 1) = (4, 2); Each grid cell generates a unique identifier (sequential encoding or two-dimensional encoding) through the above encoding method for input to the spatio-temporal dependence model.

[0033] Furthermore, in step S2, the steps of constructing the spatio-temporal dependence model include: Data Preparation: The user's stay records are represented as a tuple S = (st, dow, dur, pid, tid), where st: start time, the starting time of the stay (e.g., "2023-10-01 09:00:00"); dow: day of the week, the day of the week when the stay occurs (e.g., "Monday"); dur: stay duration, the length of the stay (e.g., "60" minutes); pid: location identifier, the unique identifier of the stay location (e.g., "1234"); tid: position identifier, the unique identifier of the grid cell where the stay location is located (the encoding result of encoding the grid, e.g., "50" or "(4, 2)"); Example Data: = (2023-10-01 09:00:00, Monday, 60, 1234, 35); = (2023-10-01 10:00:00, Monday, 30, 5678, 46); = (2023-10-01 12:00:00, Monday, 90, 9012, 55); Historical Stay Sequence and Contextual Stay Sequence: Assume the user's stay sequence is , where Q is the sequence length of 3. The historical stay sequence M includes the user's stay records in the past month, which is used to capture long-term movement patterns. The contextual stay sequence N includes the user's stay records within the past 24 hours, which is used to capture recent movement patterns.

[0034] Constructing a Spatiotemporal Dependence Model: The model input is the historical stay sequence M and the contextual stay sequence N, which is represented as: . The model uses a Transformer model based on the attention mechanism to fuse temporal and spatial features. The temporal features st and dow are encoded as temporal vectors. The spatial feature tid is encoded as a grid embedding vector (based on the grid division of R and C). The model output is the predicted probability distribution of the user's next location, for example: Output description: tid: Target grid position identifier (sequential encoding or two-dimensional encoding); prob: Probability that the user's next position belongs to this grid cell; The output is sorted in descending order of probability, and the Top3 results are retained to enhance interpretability.

[0035] Furthermore, in step S3, the design method of the spatio-temporal instruction optimization prompt template includes: Probability mapping module: Input: Probability distribution output by the spatio-temporal dependence model: ; Construction method and processing flow 1. Candidate position screening: Top-K selection (K = 3): Select the top 3 items with the highest probability; Probability threshold filtering (threshold = 0.05): Remove items with probability < 0.05 (keep all in the example); 2. tid→pid mapping: Convert the semantic position ID based on the mapping table generated by preprocessing (such as {55:1234,46:5678,35:9012}). If a certain tid is not in the mapping table (such as a new grid cell), mark it as "unknown" and remove it; Output: Sorted list of candidate positions: [{"pid":1234,"tid":55,"prob":0.72}, {"pid":5678,"tid":46,"prob":0.15}, {"pid":9012,"tid":35,"prob":0.08}]。

[0036] Instruction generation module: After receiving the structured input data, this module converts it into a JSON instruction that can be parsed by the large language model through a template engine; The input includes the historical stay sequence M: [{"pid":1234,"st":"2023-09-01 08:30:00","tid":35,"dur":45}, ...] (showing the user's behavior pattern in the past month), the context stay sequence N: [{"pid":1234,"st":"2023-10-01 09:00:00","tid":35,"dur":60}, ...] (reflecting recent activities), and a list of candidate locations (sorted by probability). First, the historical data is converted into natural language descriptions, such as "The user's stay records in the past 14 days show frequent visits to grid 35 (location 1234) on Monday mornings", the context data is converted into "Stayed in grid 35 for 60 minutes in the last 24 hours", and the candidate locations are formatted as "1. Grid 55 (probability 72%) 2. Grid 46 (15%)". These pieces of information are injected into a strictly defined JSON template to generate a structured instruction containing three major data blocks (historical records, recent activities, candidate locations) and clear response requirements, which specifically stipulates that predicted_pid and explanations including three dimensions of historical patterns (such as periodic access statistics), spatial relationships (grid distance calculation), and time dependencies (analysis of time period occupancy ratio) must be output.

[0037] Explanation generation module: In the explanation generation module, when GPT-3.5 receives the instruction from the instruction generation module, the system processes it into the final output: the predicted location pid = 1234, with detailed explanations including historical patterns (visited 8 times at 9:00 am every Monday in the past month, with significant periodicity), spatial relationships (although the predicted grid 55 is actually 9 units away from the current grid 46, it is along the subway line and conforms to the commuting habit), and time dependencies (the occupancy ratio of visits to this location in the 9:00 - 10:00 time period is 78%, and 09:30 on the current Monday perfectly matches the user's behavior pattern).

[0038] Result verification module Input: The text generated by the model (example): {predicted_pid:1234, explanation:{...}}; Output: Format verification result (passed / failed) and the standardized final output; Example output: {status: passed, predicted_pid:1234, explanation:,...}.

[0039] Example 3 In this example, experimental data comparison between the method of the present invention and the existing method is carried out: 1. Experimental setup and dataset The experiment is based on the Geolife dataset from Microsoft Research Asia. This dataset contains GPS trajectory data of 182 users, covering more than 3 years of movement records, with a total of 17,621 trajectories. Data preprocessing includes identifying stay points (stay time ≥ 30 minutes, radius ≤ 200 meters) and using the DBSCAN algorithm for clustering to generate unique location identifiers. The experiment divides the data into a training set (70%), a validation set (15%), and a test set (15%). The evaluation metrics include accuracy (Acc@k), weighted F1-score, and normalized discounted cumulative gain (nDCG@k).

[0040] 2. Comparative Methods LSTM: A prediction method based on long short-term memory networks; LSTM-SA: An LSTM model combined with a hierarchical self-attention mechanism; DeepMove: A multi-modal embedded recurrent neural network; MobTcast: A Transformer-based mobile feature extractor; MHSA: A multi-head self-attention network; LLM-Mob: A mobile prediction method based on LLMs; LLM-ZS: A zero-shot learning method of LLMs; T-LENs (the present invention): A prediction framework combining grid coding and LLMs.

[0041] 3. Comparison of Experimental Results 3.1 Prediction Result Data Table 1 Performance comparison of different models using various metrics on the Geolife dataset The results shown in the table demonstrate the performance comparison of different models using various metrics on the Geolife dataset. Among them, "without time information" means that the target time information is not used in the prediction model, which is used to measure the performance of the model when ignoring time factors; while "with time information" means that the target time information, such as time points, day of the week, etc., is used in the prediction model to help the model better align its predictions with the user's movement behavior. The proposed T-LENs framework in the present invention, which combines two different coding schemes (the bold part is sequential coding and the underlined part is two-dimensional coding), significantly outperforms the existing baseline models in terms of accuracy and ranking-based metrics.

[0042] 3.1.1 Improvement in Accuracy The most significant improvement was observed in the Acc@1 metric, where both sequential encoding (without temporal information: 49.6%, with temporal information: 55.6%) and two-dimensional encoding (without temporal information: 49.6%, with temporal information: 53.7%) showed significant gains compared to the best learning-based baseline (MHSA, 31.4%). This improvement can be attributed to the fact that spatial encoding captures spatial relationships more effectively than traditional location ID-based methods. Specifically, by encoding locations based on their spatial relationships, the model can better understand and predict frequently visited or adjacent locations, especially for top-ranked predictions.

[0043] 3.1.2 Enhancement of Weighted F1 and nDCG@10 The T-LENs framework also performs well on weighted F1 and nDCG@10 ranking-based metrics, which consider not only the correctness of predictions but also their relevance order. The sequential encoding using target time information in the prediction model achieved a weighted F1 score of 0.473 and nDCG@10 of 0.659, while two-dimensional encoding achieved 0.458 and 0.656 respectively. These improvements reflect the model's ability to capture long-term and short-term spatio-temporal dependencies. The inclusion of the tiling scheme helps maintain geographical proximity and context patterns, enabling more reasonable ranking of possible next locations.

[0044] 3.1.3 Impact of Target Time Information Another factor contributing to the performance improvement is the integration of target time information. For both encoding methods, adding time information led to consistent improvements in all metrics. This indicates that incorporating time patterns, such as time of day or day of the week, helps the model better align its predictions with user movement behavior.

[0045] 3.2 Case Study Table 2 Case Study Results The following is a concise analysis of the case study results presented in Table 2, where the next location in the real scenario is location code 13; the target time information (15:13, Friday) is a key input for the model prediction, which helps the model determine which historical data is the most relevant and is used to speculate on the most likely destination of the user at this specific time point. Different models handle time differently, so the prediction results also vary. The present invention compares the prediction results of LLM-Mob (k = 10) with two variants of T-LENs - one using sequential encoding and the other using two-dimensional encoding. LLM-Mob places location 13 in the third position in its prediction list, while both T-LENs (sequential encoding) and T-LENs (two-dimensional encoding) identify location 13 as the top candidate. This ranking difference is particularly important for metrics such as Acc@1 (accuracy of the first rank) and nDCG@10, which reward higher correct prediction ranks. In the prediction of LLM-Mob (k = 10), it emphasizes the user's weekend activities at location codes 54 and 55, indicating that location code 13 is also frequently visited. However, the model only ranks 13 in the third position, meaning that although LLM-Mob recognizes the relevance of 13, based on recent or temporal patterns, it considers 54 and 55 slightly more likely. In the sequential encoding prediction of the present invention, it attributes the higher ranking of 13 to the patterns of frequent visits on weekends and Mondays, as well as the context that recently points to continued activities at 13. By utilizing the sequential encoding scheme, T-LENs (sequential encoding) more precisely captures the user's preference for 13 and its surrounding locations, resulting in it promoting 13 to the top of the list. T-LENs (two-dimensional encoding) highlights spatial proximity, citing the encodings near (16,13) and (13,18), which surround location code 13. The two-dimensional encoding leaves geographical coordinates, enabling the model to identify the strong association of 13 with these closely adjacent points, further confirming that 13 is the most likely next location. Therefore, T-LENs demonstrates stronger prediction accuracy and better consistency with the actual movement patterns of users.

[0046] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A spatiotemporal data prediction method based on spatial grid coding and large language model, characterized in that: The following steps are involved: S1. Divide the region into grid cells according to the spatial range of the geographic region, calculate the number of rows and columns of the grid according to the height and width of the geographic region and the size of the cell, and process the grid cells using a coding method, wherein each cell is identified by a unique identifier; the coding method includes sequential coding and two-dimensional coding; S2. Based on the user's stay records, a spatiotemporal dependency model is constructed, each stay is represented as a tuple, and the user's next location is predicted by analyzing the historical stay sequence and the context stay sequence; the tuple includes the start time, day of the week, stay duration, place identifier, and location identifier; S3. Combined with a large language model, a spatiotemporal instruction optimization prompt template is designed. By introducing grid coding and context-aware reasoning, the model is guided to analyze historical data and context data, generate the next position prediction, and provide explanations for each prediction.

2. The method for spatiotemporal data prediction based on spatial grid coding and large language model according to claim 1, characterized in that: In step S1, according to the height of the geographical area and width and the size of the unit , calculate the number of rows in the grid and number of columns , the formula is as follows: , in, Indicates the height of the grid, Indicates the width of the grid.

3. The method for spatiotemporal data prediction based on spatial grid coding and large language model according to claim 2, characterized in that: The encoding method in step S1 includes: Sequential coding: The grid cells are numbered sequentially in a row-first manner, from left to right and from top to bottom. The calculation formula for numbering the grid cell in the i-th row and j-th column is: , in, represents the cell number of the sequential coding, and C is the total number of columns of the grid; in the two-dimensional coding method, each cell is represented by a row index and a column index; The two-dimensional encoding is: the grid unit is represented as a two-dimensional coordinate (x, y), where x represents the column index and y represents the row index. The calculation formula for numbering the grid unit in the i-th row and j-th column is: , in, Indicates the unit number of the two-dimensional code; Each grid unit generates a unique location identifier through the above encoding method. The unique identifier is a sequential number or a two-dimensional coordinate, which is used to represent the location identifier in the tuple.

4. The method for spatiotemporal data prediction based on spatial grid coding and large language model according to claim 3, characterized in that: The user's stay record in step S2 is generated by the following steps: S211. Extracting the latitude and longitude coordinates with timestamps from the GPS trajectory data, and then identifying the stop points based on the distance and time thresholds; S212. Cluster the stay points using the DBSCAN clustering algorithm to generate a unique location ID for representing the location identifier in the tuple; S213. Each stay point is associated with a time feature and a space feature.

5. The method for spatiotemporal data prediction based on spatial grid coding and large language model according to claim 4, characterized in that: In step S2, the spatiotemporal dependency model is a Transformer model based on the attention mechanism. The historical stay sequence and the context stay sequence are input into the spatiotemporal dependency model to predict the next location of the user, including the following steps: S221. Data preprocessing: Each stop point is represented as a tuple S = (st, dow, dur, pid, tid), where the temporal feature is: st represents the start time, dow represents the day of the week, and dur represents the duration of stay; the spatial feature is: pid represents the location identifier, and tid represents the location identifier; S222. Sequence generation and filling mechanism: select all the user's stay points in the past T days ,common A stopover point Stops are sorted by time and then truncated or padded to a fixed length , get the historical stay sequence M; select the stay points within the last Δt hours ,common A stopover point Stops are sorted by time and then truncated or padded to a fixed length , get the context stay sequence N; merge the historical stay sequence M and the context stay sequence N to get the input sequence , , The feature representation dimension representing each stop point; S223. Input sequence The tuples in are jointly encoded with spatiotemporal features: Encode the start time st through sinusoidal position coding to obtain the start time vector , , the formula is as follows: , , in, represents the dimension index, ; represents the dimension of time encoding, represents the cosine function, represents the sine function, Indicates the start time In even dimensions The encoded value of the dimension, Indicates the start time In odd dimensions The encoded value of the dimension; through the learnable embedding table , convert the day of the week dow into a week embedding vector , represents the original dimension of the embedding vector; based on Meshing to construct embedding matrix , and then obtain the position identifier vector , Represents the dimension of the embedding matrix; the operation steps are: each grid cell generates a unique location identifier tid through sequential encoding or two-dimensional encoding, and the embedding matrix is ​​initialized to a Gaussian distribution In the training, the cosine similarity of the embedding vectors of adjacent grids is improved through back propagation optimization and clustered in the embedding space. Finally, the embedding matrix is ​​obtained through table lookup operation. Get the location identifier vector from ; The start time vector , weekday embedding vector and a vector of position identifiers After splicing, linear mapping is performed to obtain the fused comprehensive feature vector, which is expressed as follows: , in, represents the integrated feature vector after fusion, represents the weight matrix of the linear projection, represents the bias term of the linear projection, represents the concatenation vector, Represents matrix multiplication; The input sequence The fused comprehensive feature vectors corresponding to the L stop points are stacked to form a comprehensive feature vector matrix: , , in, Indicates The fused comprehensive feature vector, represents the comprehensive eigenvector matrix; S224. Input the comprehensive feature vector matrix into the spatiotemporal dependency model to predict the probability distribution of the user's next location.

6. The method for spatiotemporal data prediction based on spatial grid coding and large language model according to claim 5, characterized in that: The learnable embedding table in step S223 At the beginning of model training, the uniformly distributed random initialization is used to contain 7 row vectors from Monday to Sunday. The parameters are optimized through end-to-end training, so that the embedding vectors from Monday to Friday are clustered in the latent space, and Saturday and Sunday form independent semantic clusters. The embedding vector of the week is obtained through linear transformation. .

7. The method for spatiotemporal data prediction based on spatial grid coding and large language model according to claim 6, characterized in that: The table lookup operation in step S223 is specifically to input the unique location identifier tid of the grid unit, use tid as the row index, and locate the embedded matrix The corresponding row of , the output directly returns the row vector.

8. The method for spatiotemporal data prediction based on spatial grid coding and large language model according to claim 7, characterized in that: In step S224, the spatiotemporal dependency model includes 6 layers of encoders and prediction layers, each encoder includes a multi-head attention sublayer, a residual connection and layer normalization layer, a feedforward network sublayer, and a secondary residual and normalization layer, and the operation process is as follows: In the first encoder layer, the comprehensive feature vector matrix The input to the multi-head attention sublayer is split into 8 independent subspaces according to the number of heads; each head generates a query through a learnable weight matrix ,key Sum ; Calculate the scaled dot product attention for each head and get the output of each attention head ; Concatenate the outputs of the 8 attention heads and pass them through the weight matrix Perform a linear transformation to obtain the output of the multi-head attention sub-layer , the formula is: ; The output of the multi-head attention sublayer After residual connection and layer normalization layer and comprehensive feature vector matrix Perform residual connection and normalization to obtain the output of residual connection and layer normalization layer , the formula is as follows: , in, Represents a layer normalization operation; the residual connection is connected to the output of the layer normalization layer After being processed by two layers of fully connected networks in the feedforward network sublayer, the output of the feedforward network sublayer is obtained , the formula is as follows: , in, represents a nonlinear activation function, and Represent the weight matrices of the first and second fully connected layers respectively; and Represents the bias vectors of the first and second fully connected layers respectively; the output of the feedforward network sublayer After the quadratic residual and normalization layer and the residual connection and layer normalization layer output Perform residual connection and normalization to finally get the output of the first layer encoder , the formula is: ; The output of the first encoder As the input of the second layer encoder, repeat the above operation until the output of the sixth layer encoder is obtained ; In the prediction layer, the sequence end vector is extracted from the output of the sixth layer encoder , through the weight matrix The end vector of the sequence Mapped to the candidate location set, the probability distribution of the user's next location is obtained, and the formula is as follows: , in, represents the total number of candidate positions, Represents learning different positions and The associated weight of Indicates the user's next location for The probability distribution of Represents the bias vector of the prediction layer.

9. The method for spatiotemporal data prediction based on spatial grid coding and large language model according to claim 8, characterized in that: Step S3 specifically includes: The spatiotemporal instruction optimization prompt template includes a probability mapping module, an instruction generation module, an explanation generation module and a result verification module; The probability distribution of the user's next location Input to the probability mapping module, sort the probability distribution in descending order, select the first K positions with the highest probability value, and then filter according to the probability threshold, retaining the positions with probability greater than the threshold; if the number of candidates is less than K, retain all positions that meet the threshold; finally, get the sorted candidate position list; The historical stay sequence M, the context stay sequence N and the sorted candidate position list are input into the instruction generation module, and the structured template engine converts the original input data into standard JSON instructions; first, the structured data is converted into serialized text data; then the serialized text data is injected into the predefined JSON template to obtain a complete JSON instruction that can be parsed; The complete JSON instruction is input into the explanation generation module to generate a structured explanation, and the explanation generation module adopts a large language model; The structured interpretation is input into the result verification module for quality control and standardization.

Citation Information

Patent Citations

  • Private car parking position prediction method and system based on enhanced recurrent neural network

    CN113780665A

  • Ship path classification method based on grid coding and TextRCNN

    CN117390506A

  • Dynamic long-distance flight delay prediction method and system based on delay time delay perception

    CN117876184A

  • Ship trajectory prediction method based on grid coding and natural language generation model

    CN117933514A

  • Method and system for comparative data analysis

    WO2015164910A1

Cited By

  • Natural language space-time retrieval method and system based on large model

    CN120653659A

  • Video generation method based on visual identifier

    CN120935377A

  • Multi-modal data prediction and analysis method and system

    CN121302064A

  • Hierarchical zero sample trajectory prediction method based on large language model

    CN121365120A