Traffic spatio-temporal data storage method and device, computer equipment and storage medium

By pre-constructing an R*-tree index structure through grid partitioning and predictive network models, the problems of high latency and frequent adjustments in index construction in traffic spatiotemporal data storage are solved, achieving efficient data writing and index management, which is suitable for edge computing scenarios.

CN121029818APending Publication Date: 2025-11-28SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511114476.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing traffic spatiotemporal data storage solutions suffer from high index building latency and frequent adjustments under the requirements of high-frequency writing and real-time storage, resulting in wasted system resources and increased performance overhead.

Method used

Vehicle data is discretized using a grid partitioning method, and an R*-tree index structure is pre-built through a prediction network model. Lightweight encoding is performed using the Index2Vec encoding scheme, and a prediction network model is designed to generate the index structure for future time moments, reducing the index update frequency and computational burden.

Benefits of technology

It effectively reduces the frequency of index updates and computational overhead, improves data writing efficiency, and is particularly suitable for high-frequency writing requirements in edge computing scenarios, thereby enhancing system stability and processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029818A_ABST
    Figure CN121029818A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic spatio-temporal data storage method and device, computer equipment and a storage medium, and is applied to the technical field of traffic spatio-temporal data storage, and the method comprises the steps: obtaining the traffic flow data of a vehicle in a road, and carrying out the preprocessing of the traffic flow data; receiving a data storage request, and judging whether an optional index structure exists in the index structure set or not; if yes, selecting a target index structure; if not, constructing a new index structure according to the preprocessed traffic flow data; based on a pre-established prediction network model, generating an index structure of a next moment through the index structure of the historical moment; according to the method, data storage in an intelligent traffic scene is optimized, and the data writing efficiency is improved while effective data query is ensured, so that the management requirements of high-dynamic and high-concurrency intelligent traffic system spatio-temporal data are met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of traffic spatiotemporal data storage, and particularly relates to a traffic spatiotemporal data storage method and device, computer equipment and a storage medium. BACKGROUND

[0002] With the continuous development of intelligent transportation systems, there are more and more Internet of Things sensing devices, mainly including two types of fixed source devices installed on the side of the road and mobile source devices deployed on mobile traffic entities. These devices continuously collect time series data, video, images and point clouds, etc., providing basic data support for the upper-layer applications of the intelligent transportation system. However, due to the huge amount of data and real-time nature of these spatiotemporal data, higher requirements are put forward for the effective organization and management of data.

[0003] Since the traffic spatiotemporal data has time series nature in the time dimension, the previous scheme usually establishes an index based on the timestamp, thereby organizing the data in the time dimension. However, traditional relational databases such as MySQL and PostgreSQL cannot meet the real-time storage requirements of massive spatiotemporal data. These databases usually use B-trees or improved structures as index structures, which are suitable for organizing ordered data, but their performance is not good under the requirements of high-frequency writing and real-time storage.

[0004] In order to overcome these problems, the present application provides a traffic spatiotemporal data storage method, device, computer equipment and storage medium. SUMMARY

[0005] The purpose of the present application is to provide a traffic spatiotemporal data storage method, device, computer equipment and storage medium, aiming to solve the problem of inefficient storage of massive real-time spatiotemporal data in the intelligent transportation scenario.

[0006] To achieve the above purpose, the present application provides the following technical scheme:

[0007] In a first aspect, the present application provides a traffic spatiotemporal data storage method, comprising the steps of:

[0008] Obtaining traffic flow data of a vehicle in a road, and preprocessing the traffic flow data;

[0009] Receiving a data storage request, and judging whether there is a selectable index structure in the index structure set; if there is, selecting a target index structure; if there is not, constructing a new index structure according to the preprocessed traffic flow data;

[0010] Based on a pre-established prediction network model, generating an index structure of the next time based on the index structure of the historical time.

[0011] In a second aspect, the application provides a traffic spatiotemporal data storage device, specifically comprising:

[0012] a data collection and preprocessing module, configured to acquire traffic flow data of a vehicle in a road, and to preprocess the traffic flow data;

[0013] a dynamic index management module, configured to receive a data storage request, and to determine whether there is a selectable index structure in the index structure set; if yes, select a target index structure; if not, construct a new index structure according to the preprocessed traffic flow data;

[0014] an index prediction and generation module, configured to generate an index structure at a next time point based on a prediction network model established in advance and an index structure at a historical time point.

[0015] In a third aspect, the application provides a computer device, comprising a processor and a memory coupled to the processor, wherein the memory stores program instructions for implementing a traffic spatiotemporal data storage method; and the processor is configured to execute the program instructions stored in the memory to implement a traffic spatiotemporal data storage method.

[0016] In a fourth aspect, the application provides a storage medium storing program instructions executable by a processor, wherein the program instructions are used to execute a traffic spatiotemporal data storage method.

[0017] The application provides a traffic spatiotemporal data storage method, device, computer device and storage medium, which has the following beneficial effects:

[0018] (1) In the intelligent traffic scenario, intelligent traffic system spatiotemporal data is generated and collected in real time. If the vehicle continuous trajectory data is directly stored and indexed, the index will be constantly updated due to the frequent changes of the accurate position of the vehicle, resulting in a substantial increase in query cost. The application uses a grid division method to divide the visible area, and discretizes the vehicle data, effectively reducing the index update frequency, and thereby reducing the index update overhead, making the system more efficient and stable when processing a large amount of real-time vehicle data.

[0019] (2) The application realizes the pre-construction of the R*-tree structure through a deep learning algorithm, marks each node in the R*-tree as a unique index structure node mark according to the node type and the size of the R*-tree, and then constructs an R*-tree sequence using a pre-order traversal algorithm. A prediction network model is designed based on the Transformer architecture, which can predict the R*-tree at a future time point according to the R*-tree sequence structure at a historical time point, so as to construct the spatial index structure in advance. This avoids real-time index construction during data writing, reduces the computational burden during data writing, and improves the data writing efficiency.

[0020] (3) In the prediction network model, an Index2Vec encoding scheme is designed as an embedding code, which is lightweight and can efficiently encode the node information in the R*-tree, converting the complex tree structure into a vector form that is easy for the model to process. Not only does it enhance the model's perception of the R*-tree structure, enabling it to more accurately understand and process spatial index information, but also due to its lightweight nature, it is suitable for deployment on edge devices, allowing efficient inference even in resource-constrained environments, providing strong data processing support for edge devices in intelligent transportation scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 A flowchart of a traffic spatiotemporal data storage method according to Embodiment 1 of the present application;

[0022] Figure 2 A flowchart of a traffic spatiotemporal data storage method according to Embodiment 1 of the present application;

[0023] Figure 3 A process diagram for preprocessing traffic flow data according to Embodiment 1 of the present application;

[0024] Figure 4 A process diagram for constructing a minimum rectangular frame and an R*-tree according to Embodiment 1 of the present application;

[0025] Figure 5 A process diagram for converting an R*-tree structure into a node feature unit according to Embodiment 1 of the present application;

[0026] Figure 6 A flowchart of an embedding encoding based on an index structure according to Embodiment 1 of the present application;

[0027] Figure 7 A training process diagram for word vector embedding encoding of an R*-tree according to Embodiment 1 of the present application;

[0028] Figure 8 A diagram of sequential sequence encoding and tree structure encoding according to Embodiment 1 of the present application;

[0029] Figure 9 A diagram of path-based location embedding encoding according to Embodiment 1 of the present application;

[0030] Figure 10 A prediction network architecture diagram according to Embodiment 1 of the present application;

[0031] Figure 11 A diagram comparing the index read-write performance of different traffic densities according to Embodiment 1 of the present application;

[0032] Figure 12FIG. 1 is a schematic diagram of index read-write performance comparison under mixed traffic density for Embodiment 1 of the present application;

[0033] Figure 13 FIG. 4 is a schematic diagram of data read-write throughput of different data volumes for Embodiment 1 of the present application;

[0034] Figure 14 FIG. 5 is a heat map of attention weight distribution based on a Transformer model for Embodiment 1 of the present application;

[0035] Figure 15 FIG. 6 is a structural schematic diagram of a traffic spatio-temporal data storage device for Embodiment 2 of the present application;

[0036] Figure 16 FIG. 7 is a structural schematic diagram of a computer device for Embodiment 3 of the present application;

[0037] Figure 17 FIG. 8 is a structural schematic diagram of a storage medium for Embodiment 4 of the present application. DETAILED DESCRIPTION

[0038] It should be understood that the specific embodiments described herein are merely intended to explain the present application and are not intended to limit the present application.

[0039] The following analyzes the prior art in combination with related technologies.

[0040] In the prior art, in order to cope with the write-intensive scenarios of intelligent transportation systems, time series databases such as Cassandra, FluteDB, OpenTSDB, and InFluxDB adopt LSM-tree structure as their underlying storage structure, and LSM-tree converts random write into sequential write, which has a significant advantage in terms of time series data write. At the same time, traffic data also shows obvious spatial distribution characteristics, and effective spatial index support is needed when querying data. Spatial databases such as PostGIS, Oracle Spatial, and GeoMesa usually use quadtree, grid index, KD-tree, R-tree and its improved structure as spatial index, and R-tree index is the most commonly used spatial index.

[0041] In addition, in order to support the retrieval of spatio-temporal data in both time and space dimensions at the same time, various database combination storage has become the mainstream scheme of spatio-temporal data storage, for example, adding GeoMesa with spatial index on Cassandra, TimescaleDB based on PostGIS extension, etc. The spatio-temporal index is usually composed of another dimension index in a certain specific dimension index, and the spatio-temporal data is usually organized by means of dimension reduction or multi-level index, such as HSTI, TPR-tree, PO-tree, PARINET, LSM R-tree and its variants, etc. The commonly used spatio-temporal index LSM R*-tree improves the LSM R-tree in the spatial dimension and has good spatial retrieval performance.

[0042] In the face of massive real-time multi-modal spatio-temporal data storage generated by multiple sources and multiple devices, the existing LSM R*-tree data storage scheme still has a large optimization space, mainly manifested as high data write delay. From the data storage process, first, data needs to be collected, then based on the collected data, a data write request is sent to the database, then the database writes data and "writes data and constructs index at the same time", and finally the data storage process is completed. In this process, the database only performs real-time index construction when data is written. At this time, the system not only needs to transmit data but also needs to construct index, increasing the index construction overhead and wasting system resources. In addition, when the R*-tree structure is constructed, the position information of the spatial data needs to be frequently adjusted to accurately construct the index structure, further increasing the index construction overhead. Since the R*-tree structure needs to be traversed once when data is written, and multiple node comparisons and splitting may also be involved, the total index construction time complexity can be represented as O(N·H·C), where N represents the amount of data collected in the current time slice, H is the tree height of the R*-tree, and C is the average cost of each comparison, splitting or re-insertion operation. In the intelligent transportation system scenario of high-frequency data continuous writing, as the amount of data increases, the dynamic maintenance of the R*-tree will frequently trigger structure updates, and the average cost of node splitting or re-insertion will increase significantly, the time complexity will increase rapidly, and great performance overhead will be brought.

[0043] Therefore, in view of the challenges of high real-time index construction delay and frequent index adjustment, the application innovatively proposes a data storage strategy of pre-construction of LSM R*-tree index structure, solving the problem of high index construction delay. In addition, an end-to-end multi-task prediction network is designed for the pre-constructed index structure, which generates the index of the future time from the index structure of the current time, realizes the prediction of the R*-tree index, and solves the problem of frequent index adjustment.

[0044] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work are within the scope of protection of the present application.

[0045] Embodiment 1

[0046] Please refer to Figure 1 and 2 , a flowchart and a flowchart of a traffic spatiotemporal data storage method in Embodiment 1 of the present application.

[0047] Compared with the conventional scheme, the technical solution proposed in the present application can construct an index in advance and generate an index structure at a future time by using a prediction network model, thereby accelerating the data writing process. The specific advantages are lower time complexity and lower storage delay. The traditional real-time index data storage step of “collecting data and constructing an index at the same time” is as follows: first, collecting real-time spatiotemporal data from various Internet of Things devices and performing preliminary preprocessing; second, constructing an R*-tree index according to the position information of the data; and finally, performing a disk operation on the collected data and the corresponding index structure. As shown in Figure 2 , the present application aims to optimize the data writing performance by predicting and constructing an R*-tree index in advance. When constructing the index structure for the first time, since there is no R*-tree structure, the R*-tree prediction network cannot be used, and the data storage process is consistent with the traditional data storage process. However, from the second time slice, there is a historical R*-tree structure, and the R*-tree structure is pre-constructed by using the prediction network RTPN4ITS. The main steps include: first, collecting data; second, selecting the tree structure corresponding to the timestamp according to the model prediction result as the index; third, performing a disk operation on the collected data and the corresponding index structure; and finally, instead of waiting for the data of the next timestamp to construct the index structure, the R*-tree index structure of the future several time slices is directly predicted according to the R*-tree structure information at the current time.

[0048] Therefore, from the time complexity, the time delay of data storage is the overhead from data arrival to data writing. The traditional scheme needs to dynamically maintain the R*-tree index at each data point arrival, and needs to traverse the tree structure and compare each time data is inserted; resulting in an overall write complexity of O(N·H·C); where N is the amount of data collected in each time slice, H represents the tree height of the R*-tree, and C represents the cost of each comparison or split. In contrast, the pre-construction scheme proposed in the present application generates an index structure in advance through a prediction network model for prediction, and only needs to map the data to the constructed index, reducing the write complexity to O(N·H+P), where P represents the time delay of selecting the R*-tree index. The calculation time of this part is independent of the data amount, and is usually considered as a constant. When the data amount is large, the present application significantly reduces the index construction overhead in the write process, improves the overall data write throughput efficiency, and is especially suitable for high-frequency write requirements in edge computing scenarios.

[0049] The specific steps of a traffic spatio-temporal data storage method include:

[0050] S1: Obtain traffic flow data of vehicles in a road, and pre-process the traffic flow data;

[0051] In an embodiment, the collection of traffic flow data is based on the cooperative work of intelligent vehicles, road test units and edge servers, and specifically collects traffic flow data in the same section and same direction of the road. The traffic flow data includes but is not limited to vehicle unique identifier ID, timestamp information t, latitude and longitude coordinates Image and audio data carried by the vehicle The traffic flow data of vehicle i at time t is represented as

[0052] Please refer to Figure 3 , which is a process diagram for pre-processing traffic flow data according to Embodiment 1 of the present application. Specifically, it includes:

[0053] UTM projection is adopted to convert the latitude and longitude coordinates into plane rectangular coordinates The visible area is divided into an M×N grid array, and the plane rectangular coordinates are used to distribute the vehicles into grid elements G(m,n) to obtain the discretized traffic flow data; where m∈[1,M], n∈[1,N]. The specific formula is represented as:

[0054]

[0055] It should be noted that, since the space-time data of the intelligent transportation system is generated and collected in real time, in order to ensure the spatial retrieval characteristics of the data, it is usually necessary to retrieve according to the continuous trajectory data of the vehicle; if directly stored and indexed, the index will be frequently updated with the accurate position of the vehicle, and the query cost will be high. Therefore, the application divides the visible area into a grid array before storing the data, discretizes the vehicle data, stores the vehicle data within a reasonable accuracy, and reduces the index update overhead.

[0056] S2: receiving a data storage request, judging whether there is a selectable index structure in the index structure set; if there is, selecting a target index structure; if there is not, constructing a new index structure according to the preprocessed traffic flow data.

[0057] In an embodiment, the index structure is established according to the preprocessed traffic flow data, and the index structure is an R*-tree structure. The R*-tree belongs to the multi-way self-balancing tree type, and in the tree structure, each node is associated with a minimum rectangular box (MBR). Specifically, the leaf node corresponds to the minimum rectangular box of the spatial entity, and the non-leaf node generates the minimum rectangular box of the current level by aggregating the minimum rectangular boxes of its child nodes.

[0058] Please refer to Figure 4 , which is a schematic diagram of the construction process of the minimum rectangular box and the R*-tree of embodiment 1 of the application. As shown in the left graph a) of Figure 4 , the spatial range of the minimum rectangular box is defined by using the lower left coordinate (x left ,y left ) and the upper right coordinate (x right ,y rigjt ), and the formula is expressed as:

[0059] MBR = (x left ,y left ,x right ,y right ), where x * ∈{A,B,C,D,…}, y * ∈{0,1,2,…} (2)

[0060] where MBR represents the minimum rectangular box, x * represents the horizontal coordinate of the minimum rectangular box, and y * represents the vertical coordinate of the minimum rectangular box, which corresponds to the grid coordinate.

[0061] Since the number of entries of the R*-tree node is not fixed, but there are explicit minimum and maximum capacity constraints. Its construction process uses an iterative approach to gradually advance, as Figure 4As shown in the right graph a) of FIG. 1, the initial stage of construction will generate minimum rectangular frames for each entity in priority, and these minimum rectangular frames will be included in the tree structure as leaf nodes. For non-leaf nodes, the generation of their minimum rectangular frames is determined according to the minimum area increment principle recursively integrated from the lower layer nodes, that is, when a new rectangle is inserted into a parent node, the insertion position that makes the area increment of the parent node's circumscribed rectangle the smallest is selected in priority, so as to ensure the compactness of the tree structure and the query efficiency.

[0062] For example, assuming that a leaf node C needs to be inserted into the tree structure, and two candidate non-leaf nodes A and B are available. After node C is inserted into non-leaf node A, the area of the circumscribed rectangle increases from 40 to 60 (an increase of 20); and after node C is inserted into node B, the area of the circumscribed rectangle increases from 50 to 55 (an increase of 5). At this time, according to the minimum area increment principle, node B should be selected for insertion. Because its area increment is smaller, the overlapping area between nodes is reduced, and the retrieval efficiency is improved.

[0063] During the construction process, if a node reaches the maximum capacity limit, that is, the full load condition occurs, a split operation will be triggered. This operation not only causes the structure adjustment of the current node, but also may trigger a series of chain reactions, such as the reconstruction of related minimum rectangular frames and the reinsertion of sub-tree nodes, so as to ensure that the tree structure always maintains balance and high efficiency. Therefore, since the R*-tree needs to perform full traversal of the tree structure every time data is written, and multiple node comparisons and possible split operations are involved, the total time complexity of the index construction can be represented as O(N·H·C). Where N represents the total amount of data collected in the current time slice, H represents the tree height of the R*-tree, and C represents the average cost of each comparison, split or reinsertion operation.

[0064] It should be noted that in the scenario of intelligent transportation systems (ITS) with high-frequency data continuous writing, the R*-tree needs to be dynamically maintained frequently to adapt to data changes, which inevitably triggers frequent tree structure updates. Such frequent update operations will bring significant performance overhead, and thus affect the overall running efficiency of the system. Therefore, how to effectively reduce the performance overhead of the R*-tree during dynamic maintenance has become a key problem for index optimization. In the present application, the data storage method of "constructing index first and then writing data" optimizes the index construction timing, improves the data storage efficiency, and solves the problem of high index construction delay. In addition, the prediction network model established generates the index of the future time based on the index structure of the historical time, realizes the prediction of the R*-tree index, and solves the problem of frequent index adjustment.

[0065] Further, the vocabulary of the R*-tree structure is regarded as a mapping between the nodes and the index structure node tokens. Since the R*-tree itself is a tree structure and each node of the R*-tree corresponds to a minimum rectangular box, each index structure node token is composed of two parts: the node type information and the coordinate information of the minimum rectangular box.

[0066] In representing the R*-tree node types, a specific label is used to distinguish the different types of nodes. Specifically, the label <nl>to mark non-leaf nodes, using tags <l>For the leaf nodes, the marking scheme defined by equation (2) is used. In addition, in order to clearly distinguish different R*-tree structures and different branches of an R*-tree, some special symbols are added to the vocabulary. These special symbols include <beg>, a marker sequence end position <end>and as a placeholder for <s>In summary, the vocabulary can be represented as:

[0067] Vocab = { <beg> , <end> , <l> <s>, <l>A0B1,... <nl> <s>, <nl>A0B1,...} (3)

[0068] Based on the above constructed vocabulary, the R*-tree structure is serialized by traversal algorithm. The serialization algorithm of R*-tree structure is as follows:

[0069]

[0070]

[0071] The specific operation process is as follows: first, insert the start tag at the starting position of R*-tree sequence <beg>. Next, each node is traversed in a recursive manner starting from the root node until a leaf node is reached. During the traversal, according to the type of the current node, the corresponding <nl>or <l>the label, and a corresponding minimum rectangular frame representation is added according to the minimum rectangular frame information of the node. If the entry of the current node is empty, the label is used <s>To identify its minimum rectangular frame. For example: add a label to non-leaf nodes <nl>and the coordinates of its minimum rectangular frame; add a label to the leaf node <l>and its minimum bounding box coordinate representation; add placeholder for null entry <s>and its minimum-rectangular-box coordinate representation. Finally, when backtracking to the root node, add the label <end>, which indicates that the construction of the whole sequence is completed.

[0072] Referring to Figure 5 , the process of converting the R*-tree structure of Embodiment 1 of the present application into a node feature unit is shown. The specific steps include:

[0073] According to the above vocabulary and the traversal algorithm designed for the R*-tree structure, the spatial feature token (MBR Token, denoted as T MBR ) and the type feature token (Type Token, denoted as T type ) are obtained. Among them, the spatial feature token corresponds to the geometric parameters of the minimum rectangular box of the node, and the type feature token corresponds to the node type. After obtaining T MBR and T type , the spatial feature token and the type feature token are combined to construct the node feature unit (Node Token, denoted as T Node ). The node feature unit completely covers the spatial information and type information of the R*-tree node, and provides convenience for subsequent efficient processing and analysis of the R*-tree structure.

[0074] S3: Based on the pre-established prediction network model, the index structure of the next time is generated based on the index structure of the historical time.

[0075] In an embodiment, for resource-constrained edge devices, a lightweight encoding scheme is proposed, i.e. the encoder of the prediction network model adopts Index2Vec based embedding encoding, which realizes fast perception and efficient representation of the R*-tree structure. Specifically, it includes generating word embedding encoding (WordEmbedding) and position encoding (Positional Encoding) for the minimum rectangular box of the R*-tree node.

[0076] Referring to Figure 6 , the process of embedding encoding based on the index structure of Embodiment 1 of the present application is shown. Specifically, taking node R6 as an example.

[0077] In natural language sequences, the context information of a word mainly revolves around its linear relationship with the words before and after it in the text, semantic association, etc. However, the R*-tree structure is completely different, which belongs to a spatial tree structure with hierarchical characteristics. In this structure, the nodes are not simply linearly arranged, but there are complex and ordered relationships between them, including parent-child relationship, sibling relationship, etc. When focusing on the minimum rectangular frame, the relationship between these nodes has a direct geometric manifestation. The parent-child relationship of the nodes is manifested as the containing relationship of the minimum rectangular frames, that is, the minimum rectangular frame of the parent node will completely cover the minimum rectangular frame of the child node; the sibling relationship of the nodes corresponds to the adjacency between the minimum rectangular frames, which means that the minimum rectangular frames of the sibling nodes are close to each other and adjacent to each other in space. Based on the above characteristics, for any node n in the R*-tree i , its context is defined as a set composed of the parent node n parent , the sibling node n sibling and the child node n child ; represented as:

[0078]

[0079] In natural language processing (NLP) tasks, the common methods of word vector embedding encoding are continuous bag-of-words model (CBOW) and skip-gram model. In the training process of the continuous bag-of-words model, multiple context information is aggregated at a time, and the aggregated context information is used to predict the center word. The skip-gram model adopts a modeling method one by one, describes the relationship between the center word and the context, and uses the center word to predict the word in the context. When these two methods are applied to the R*-tree structure, it can be found that they have advantages and disadvantages. The continuous bag-of-words model has relatively insufficient modeling capability when dealing with fine-grained information. If only the continuous bag-of-words model is used, the strong local dependency relationship between the R*-tree nodes may be ignored, such as the close relationship between the parent-child relationship, which may lead to the loss of feature information of the center node. On the contrary, the skip-gram model is more suitable for processing scenes with strong dependency. However, it also has some problems, such as increasing the computational overhead when dealing with weak context relationships, for example, predicting the parent node through the leaf node, which reduces the processing efficiency. In view of the above situation, the continuous bag-of-words model and the skip-gram model are combined innovatively to jointly complete the encoding of word embedding.

[0080] Please refer to Figure 7 This diagram illustrates the training process of word vector embedding encoding for an R*-tree according to Embodiment 1 of this application. The steps include: first, extracting global context features of non-leaf nodes in the tree structure using a skip-word model to generate an embedding matrix. After the skip-word model is trained, its parameters are fixed; the resulting embedding matrix is ​​used as input to a continuous bag-of-words model, and the vector representation of the center node is predicted by aggregating the feature information of the context nodes in the R*-tree, thus establishing a mapping relationship from the context nodes to the center node of the R*-tree.

[0081] Specifically, according to step S2, the serialization result obtained after serializing the R*-tree structure through the traversal algorithm is the R*-tree node sequence; specifically, it is expressed as:

[0082]

[0083] Among them, T MBR Spatial feature labeling; T type This is used to label type features.

[0084] For each node sequence T Node One-Hot encoding node x Node :

[0085] x Node =OneHot(Concat(T) Type ,T MBR ))(6)

[0086] Where, x Node ∈R |V| , where |V| is the size of the vocabulary.

[0087] The embedding matrix W∈R of the jump character model |V|×d Through uniform distribution Perform random initialization, where W ij Let represent the value of the element in the i-th row and j-th column of the embedding matrix W. Let d represent a uniformly distributed random variable on the interval [a, b], where d is the dimension of the embedding vector.

[0088] The matrix parameters are progressively optimized through backpropagation during training, where each row corresponds to a node; that is, the word vector embedding encoding e of node x. node ∈R d By embedding the transpose of matrix W and the One-Hot encoding of node x... Node Multiplying them together gives:

[0089] e node =W T ·x Node (7)

[0090] The skip-gram model and the continuous bag-of-words model are described in detail as follows:

[0091] 1. In the skip-gram model, the context nodes of a center node are predicted by the center node, thereby capturing the node features of the R*-tree structure relationship. The input of the skip-gram model is a d-dimensional embedding vector e of the center node v v , and the output is a linear transformation of the center node v through an output embedding matrix W' ∈ R |V|×d , generating a context score vector o ∈ R |V| . The formula is as follows:

[0092] o = W'·e v (8)

[0093] Each element of the score vector o represents the relevance intensity of the center node v to all nodes in the vocabulary.

[0094] In theory, the score vector o can be converted into a conditional probability distribution P(u|v) by a Softmax function, that is, the probability of the occurrence of the context node u of the center node v; the formula is as follows:

[0095]

[0096] Where v' is all nodes in the vocabulary V except the center node v. However, since the vocabulary V of the R*-tree is usually large in scale, the calculation of the normalization term in the denominator needs to traverse the entire vocabulary, resulting in extremely high computational complexity, which is difficult to realize in practical applications. Therefore, in order to reduce the computational overhead, the skip-gram model uses a negative sampling strategy to optimize the calculation of the conditional probability. The objective function is decomposed into two parts:

[0097] Positive sample optimization: maximize the similarity between the center node v and the context node u. That is: logσ(e u ·e v ), where σ is a sigmoid function, and e u ·e v is the inner product of the node embedding vectors, which is used to measure the relevance between nodes. Negative sample optimization: minimize the similarity between the center node v and the negative sample node v' k . That is:

[0098] Where the negative sample node v' k ∈ V represents any node irrelevant to the context node u, and the negative sign indicates that the negative sample should be as irrelevant to the center node as possible. Combining the above two parts, the loss function of the skip-gram model is as follows:

[0099]

[0100] ​

[0101] where Context(v) is the context set of the center node v, and K is the number of negative samples.

[0102] 2. In the continuous bag-of-words model, after the skip-gram model training is completed and its parameters are frozen, the continuous bag-of-words model predicts the center node through the context node, realizing reasoning from the context node to the center node of the R*-tree. Specifically, it includes:

[0103] The input is the embedding vector e u of the context node obtained through the skip-gram model Context ; the formula is:

[0104]

[0105] where u∈Context(v) is the context node u set of the center node v; e Context ∈R d is the average embedding vector of the context node.

[0106] Based on the embedding matrix W''∈R |V|×d output by the continuous bag-of-words model, the context aggregation embedding vector e Context is calculated, and the center word score o∈R |V| is calculated; the formula is:

[0107] o=W''·e Context (12)

[0108] where each dimension in o represents the prediction score of the context corresponding to a node in the vocabulary as the center node; the score value reflects the matching degree of the node and the current context, that is, the higher the score, the higher the probability of the node appearing in the current context.

[0109] In theory, the score vector o can be converted into the conditional probability P(v|u) through Softmax, but due to the limitation of computational complexity, the continuous bag-of-words model adopts a similar strategy to the skip-gram model to optimize the objective function. Specifically, it includes two parts:

[0110] Positive sample optimization: maximize the prediction probability of the real center node v, i.e., logσ(e Context ·e v ).

[0111] Negative sample optimization: minimize the prediction probability of the negative sample node v' k , i.e.

[0112] Therefore, the loss function of the continuous bag-of-words model is represented as,

[0113]

[0114] The loss function realizes reverse modeling of the local-global relationship between nodes in the R*-tree structure by optimizing the similarity between the context aggregation vector and the center node vector.

[0115] Referring to Figure 8 , a schematic diagram of the sequential sequence encoding and tree structure encoding of embodiment 1 of the present application is shown.

[0116] In natural language processing, the traditional Transformer model generates position encoding through the method of positive sine position encoding, capturing the relative position relationship of elements in the sequence. However, there are great differences between the serialized tree structure sequence and the tree structure, mainly in two aspects: nodes far apart in the sequence have strong correlation in the tree, such as Figure 8 In (a), R6 and its sibling nodes R7 and R8 are far apart in the sequential sequence, but the three nodes are sibling nodes and are closely related. Nodes adjacent in the sequence may have no direct correlation in the tree, such as Figure 8 In (b), R7 and the child nodes (R1, R2, R3) of R6 are close in distance, but have no correlation in the tree structure. Therefore, if the encoding scheme of positive sine is used for encoding, the serialized tree structure can only show the relationship between nodes through recursion, and the characteristics cannot be directly shown.

[0117] To solve the above problems, the present application proposes a path-based position embedding encoding, which specifically includes: recording the access path from the root node to the current node, combining the path sequence information and the preset dimension parameter, and using the sine-cosine function to calculate and generate the position encoding of the current node. Combine the word vector embedding encoding and the position encoding of the current node to obtain the encoding result output by the prediction network model.

[0118] Referring to Figure 9 , a schematic diagram of the path-based position embedding encoding of embodiment 1 of the present application is shown. For the adaptation problem of hierarchical path information and serialized representation in the R*-tree spatial index structure, the present application proposes a hybrid position encoding method of 8421 path encoding and positive sine modulation, and the specific steps include:

[0119] Let each node in the R*-tree have at most n child entries. For any node x in the tree, the path from the root node to the leaf node is represented as:

[0120] PE Path (x)=[c1,c2,…,c L ],c j ∈{1,2,…,n} (14)

[0121] Where c j wherein represents the path number of the node from the (j-1)th layer to the jth layer; and L is the maximum height of the tree. Since the R*-tree root node contains only one entry, its path number is marked as n+1, which is different from the maximum number of entries n of other nodes.

[0122] From the foregoing R*-tree construction and adjustment, it can be seen that the R*-tree backbone node changes slowly, while the branch and leaf nodes change quickly. Therefore, the path encoding of the node PE Path (x) represents the position vector of the node x in the global sequence.

[0123]

[0124] wherein d is the dimension of the position encoding; i is the ith dimension in the vector; k is the dimension sequence number corresponding to the current calculation, and when k=2i, the sine function value is taken, and when k=2i+1, the cosine function value is taken; PE Path (x) is the node path encoding.

[0125] The node path encoding of the node x is: PE seq (x)=[PE1, PE2, …, PE d ] (16)

[0126] The word vector embedding encoding e Node of the node x and the position encoding PE seq (x) are added to obtain the result X Index2Vec ∈R B*L*d after the encoding of the prediction network model.

[0127] X Index2Vec (x) = e Node (x) + PE seq (x) (17)

[0128] For an input with a batch size of B, the output tensor shape is: X Index2Vec ∈R B*L*d , L is the sequence length, and d is the embedding vector dimension.

[0129] Please refer to Figure 10 Fig. 1 is a schematic diagram of a prediction network architecture of Embodiment 1 of the present application. Based on the Index2Vec index structure embedding encoding, the present application innovatively constructs an R*-tree prediction network for ITS (R*-tree Prediction Network for ITS, RTPN4ITS), and in the training strategy, it draws on the multi-label prediction paradigm of DeepSeekV3. The prediction network model is designed based on the Transformer architecture, and a number of sequentially stacked multi-label prediction modules are used to recursively predict the index structure node labels to be predicted. Specifically, the multi-label prediction module is based on the Transformer basic block for feature processing; the features at each level are converted into a probability distribution on the vocabulary table through a linear projection head, where the linear projection head generates a probability distribution through a softmax function; and finally the Euclidean distance loss values of a number of multi-label prediction modules are arithmetically averaged to obtain the total loss function of the prediction network.

[0130] Further, the model uses D sequentially stacked multi-label prediction modules to predict the future D index structure node labels (Tokens). Each multi-label prediction module recursively predicts the index structure node labels to be predicted based on shared Transformer basic blocks and output heads. Specifically, the input of each multi-label prediction module is the index structure node label sequence after slicing processing, denoted as:

[0131]

[0132] Among all the multi-label prediction modules, the first multi-label prediction module is set as the main module, and its initial input is converted into an embedding representation based on the index structure embedding encoding The formula is expressed as:

[0133]

[0134] And for the remaining D-1 multi-label prediction modules, they act as shared parameter optimizers in the training phase to make the main module achieve the optimal parameter state.

[0135] For the input of the kth layer of the Transformer basic block (transformer_block), the input vector is constructed in the following way: the output representation of the previous layer is concatenated with the Index2Vec embedding vector representation X Tree2Vec (t i+k of the current input index structure node label, and the concatenated result is used as the input vector of the layer. The formula is expressed as:

[0136]

[0137] wherein M k ∈R d×2d is a linear transformation matrix, and RMSNorm(·) is a normalization operation.

[0138] Each layer in the prediction network adopts a shared Transformer basic block to process the input, so as to capture the context dependency in the sequence. The formula is expressed as:

[0139]

[0140] wherein T is the length of the input sequence; the subscript 1:T-k is from the first element to the T-k element (including the left and right boundaries); and k is the Transformer basic block of the kth layer. The output of each layer is converted into a probability distribution on the vocabulary table through a shared linear projection head OutHead(·), which adopts a softmax activation function and is expressed as:

[0141]

[0142] wherein i∈T-k, that is, represents the probability distribution of the i+1+kth index structure node marker of the kth layer, that is, the prediction of the i+1+kth index structure node marker using the first i index structure node markers in the sequence; and |V| is the size of the vocabulary table.

[0143] Since each multi-label prediction module is used for recursive prediction of the index structure node marker to be predicted, and the R*-tree structure essentially belongs to the minimum rectangular frame hierarchy and is closely related to the spatial position relationship, the commonly used cross entropy in the Transformer model is replaced by the Euclidean distance loss function. The formula is expressed as:

[0144]

[0145] wherein is a loss function, the superscript k represents the kth multi-label prediction module, the subscript MTP represents the multi-label prediction strategy; T is the length of the input sequence; t i is the true value node feature unit of the i th position after serialization of the R*-tree; is the predicted value of the kth multi-label prediction module.

[0146] Finally, the loss function of the prediction network is obtained by calculating the average value of the Euclidean distance loss value of the D-layer multi-label prediction module . The formula is expressed as:

[0147]

[0148] In summary, the present application pre-constructs the R*-tree index data storage process, puts the index construction time before the data writing, and only needs to establish the mapping with the index during data writing, which reduces the system overhead and improves the writing efficiency. Secondly, a prediction network model based on the Transformer architecture for intelligent transportation systems is proposed to realize R*-tree structure prediction. In addition, considering the characteristics of the edge deployment of Internet of Things devices in the intelligent transportation system scenario, an Index2Vec lightweight coding structure is designed in the prediction network model, which covers word vector embedding coding and position coding, enhances the model's perception ability of R*-tree structure, and supports real-time inference under resource constraints.

[0149] In another embodiment, the technical solutions of the present application are verified by experiments, which are described in detail below.

[0150] (1) Data set and experimental environment

[0151] The US101 and i-80 data sets in NGSIM are used as experimental data sets. The data set provides detailed vehicle trajectory information on the US101 highway in Los Angeles, California, and the I-80 interstate highway, including timestamp, vehicle position, vehicle length and width, etc. The important information meets the experimental data requirements. The data storage environment uses ScyllaDB v6.1 as the benchmark database. The database is a distributed data storage system, which uses LSM-tree as the underlying data organization method, and has very low data write latency, which meets the current data storage requirements of the intelligent transportation field. The scheme discussed in the present application is to construct a two-level index R*-tree in memory and optimize the data query method; ScyllaDB stores the vehicle information table and the R*-tree structure index table.

[0152] (2) Experimental design and evaluation index

[0153] Based on the key technology of the present application, three groups of experiments are designed here. First, verify the influence of pre-built R*-tree on data storage. Since the experimental database is ScyllaDB, its underlying index is LSM-tree, therefore, this experiment compares the data write latency, read latency, etc. of three kinds of space index construction: no space index construction (LSM-tree Only), real-time space index construction (LSM Real R*-tree), and pre-built space index construction (LSM Prebuilt R*-tree). Second, the prediction ability of R*-tree structure of different types of models. Since the current task refers to the natural language processing task for implementation, and the prediction accuracy of R*-tree also affects data storage, therefore, this place uses perplexity (PPL) to measure the overall modeling ability, and uses accuracy (Acc) and the intersection over union (IoU) of the smallest rectangular frame to measure the local modeling performance. Finally, verify the Index2Vec encoding scheme based on the R*-tree structure proposed in the present application. This experiment qualitatively analyzes the effect of Index2Vec by visualizing the attention weight inside the model.

[0154] (3) Data index strategy evaluation

[0155] This experiment analyzes the three index strategies of LSM-tree Only, LSM Real R*-tree, and LSM PreBuilt R*-tree. Referring to the method of the most widely used time series database evaluation benchmark "Time Series Benchmark Suite", the data write-read load ratio of the data storage in the traffic scene is set to 10:1, that is, write data according to time slice, and query once every 10 time slice data. At the same time, traffic data is always associated with space, so the data query is set to query vehicle information in a random area.

[0156] Since the vehicle density in the traffic scene has a time sequence, the data set is first divided according to the vehicle density to simulate the data write-read efficiency of LSM-tree, LSM Real R*-tree, and LSM PreBuilt R*-tree under three types of traffic data density: sparse, medium, and dense.

[0157] Please refer to Figure 11 for the index read-write performance comparison diagram of different traffic densities of Embodiment 1 of the present application. Overall, LSM-tree has better performance when the data volume is small (such as when the vehicle density is sparse and the data volume is less than 35,000), but as the data volume gradually increases, the effect of LSM-tree becomes worse and shows an exponential growth trend, and LSM PreBuilt R*-tree has better performance.

[0158] Please refer to Figure 12 Fig. 1 is a schematic diagram of index read-write performance comparison under mixed traffic density for Embodiment 1 of the present application. In order to be closer to the traffic fluctuation characteristics in the real traffic environment, further simulation of different traffic states from sparse to dense traffic flow is carried out to verify the data write-read delay of different data indexing strategies under mixed traffic density. The experimental results show that under mixed traffic density, the write-read effect of LSM PreBuilt R*-tree is better, which is consistent with the foregoing conclusion. When the data amount is small, the effect of LSM Real-time R*-tree index is the worst, but as the data amount increases, the effect of LSM-Tree Only is the worst, and it shows an exponential growth trend.

[0159] Referring to Figure 13 Fig. 2 is a schematic diagram of data read-write throughput of different data amounts for Embodiment 1 of the present application. In order to further explore the data write efficiency, here the data write and query throughput are used as evaluation indexes according to the data amount and data write-read delay results, which intuitively show the difference; the throughput represents the number of operations per millisecond, and the higher the value, the faster the data write / read rate. Overall, the LSM-Tree Only scheme has high data write efficiency but very low query efficiency, LSM Real R*-tree shows the opposite trend, and LSM PreBuilt R*-tree scheme has balanced performance and the best effect. Its data write efficiency is improved by nearly 50% compared with the LSM Real R*-tree scheme under the current accuracy. From the data retrieval point of view, the LSM PreBuilt R*-tree scheme improves the retrieval performance of LSM-Tree by about 90%.

[0160] (4) Performance evaluation of R*-tree structure prediction model

[0161] The experiment evaluates the model performance by comparing the PPL, IoU and ACC of RTPN4ITS (Ours) and the individual comparative benchmarks under sparse, medium and dense traffic flow states; the prediction results of different models are shown in the table below:

[0162]

[0163] The experimental results show that RTPN4ITS (Ours) is superior to the comparative methods in various indicators, verifying its effectiveness in the task of modeling the time-series R*-tree structure. Further analysis of the model architecture, in the current task, Reformer, Transformer-XL and RTPN4ITS (Ours) based on the Transformer architecture all achieved good results under three traffic flow conditions. Reformer architecture in the traffic flow in the medium and dense scenes, its IoU is 0.65 and 0.73, ACC is 0.69 and 0.77, after introducing the Index2Vec encoding mechanism, the IoU in RTPN4ITS (Ours) is 0.71 and 0.78, and the ACC is 0.78 and 0.85. The prediction accuracy is improved by nearly 10%. This shows that the use of Index2Vec encoding scheme can improve the model's attention to the features of R*-tree index structure, making it easier for the model to learn the features of the tree structure.

[0164] Comparing the prediction performance of the model under three traffic flow conditions, all models have better prediction results in the dense traffic flow situation, because when there are fewer vehicles, their movement is random and it is difficult to form a stable structure evolution pattern, so the prediction accuracy is low. On the contrary, in the dense traffic flow scenario, the movement of vehicles is significantly affected by surrounding vehicles, so the overall traffic state tends to be stable over time and presents obvious group movement rules, which is reflected in the R*-tree structure as its continuity and regularity are enhanced, and the model can more easily learn the structure features, thus having higher structure prediction accuracy. In addition, when the vehicle density is medium, Transformer-XL performs best in the perplexity index, but its accuracy and intersection over union are slightly worse, which shows that the model has good language modeling ability and better fits the overall R*-tree vocabulary distribution, but pays less attention to local modeling of the tree structure.

[0165] (5) Verification of structure modeling ability based on Index2Vec

[0166] In order to further analyze the influence of Index2Vec on the model's structure perception ability, this experiment compares the attention weight distribution of different model architectures and whether they use Index2Vec. The selected models include the original Reformer, Transformer-XL, and the improved models RTPN4ITS (Ours) and Transformer-XL with Index2Vec.

[0167] Please refer to Figure 14 Fig. 1 is a diagram illustrating the attention weight distribution of the Transformer model based on the embodiment 1 of the present application. The diagram shows the attention distribution of four models, Fig. a) is the attention weight of Reformer, Fig. b) is the attention weight distribution after introducing Index2Vec, Fig. c) is the attention weight distribution of Transformer-XL, and Fig. d) is the model improved using Index2Vec. Through the comparison of the two groups of models, the Index2Vec coding method can enable the model to capture the features of the tree structure at the first layer, significantly enhancing the model's perception ability of the tree structure. The boxes in the diagram are two sub-trees in the R*-tree, and the attention values between the internal nodes are relatively concentrated. In addition, the yellow boxes represent two non-leaf nodes <nl>A2E7, <nl>B7D13) the attention weight of each respective child node, e.g. <nl>The attention weights of A2E7 are distributed among its child nodes, without focusing on <nl>B7D13. This shows that Index2Vec can effectively distinguish different branches of R*-tree structure and capture the structural dependencies within the same sub-tree, demonstrating the model's ability to understand the R*-tree structure.

[0168] Embodiment 2

[0169] Referring to Figure 15 , it is a structural schematic diagram of a traffic spatio-temporal data storage device according to Embodiment 2 of the present application. The specific content includes:

[0170] The data acquisition and preprocessing module is configured to acquire traffic flow data of vehicles in a road and to preprocess the traffic flow data.

[0171] The dynamic index management module is configured to receive a data storage request, to determine whether there is a selectable index structure in the index structure set, to select a target index structure if there is one, and to construct a new index structure according to the preprocessed traffic flow data if there is not one.

[0172] The index prediction and generation module is configured to generate an index structure at a next time based on a prediction network model established in advance and a historical index structure.

[0173] Embodiment 3

[0174] Referring to Figure 16 , it is a structural schematic diagram of a computer device according to Embodiment 3 of the present application. The computer device 50 includes a processor 51 and a memory 52 coupled to the processor 51.

[0175] The memory 52 stores program instructions for implementing the above traffic spatio-temporal data storage method.

[0176] The processor 51 is configured to execute the program instructions stored in the memory 52 to implement a traffic spatio-temporal data storage.

[0177] The processor 51 can also be referred to as a CPU (Central Processing Unit).

[0178] The processor 51 can be an integrated circuit chip with signal processing capability. The processor 51 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a ready programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0179] Embodiment 4

[0180] Referring to Figure 17 Fig. 4 is a structural schematic diagram of a storage medium of Embodiment 4 of the present application. The storage medium of the present application stores a program file 61 capable of implementing all the methods described above, wherein the program file 61 can be stored in the storage medium in the form of a software product, and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various other media capable of storing program codes, or a computer, a server, a mobile phone, a tablet, etc.

[0181] It should be noted that in this document, the terms "comprising", "containing" or any other variant thereof are intended to cover non-exclusive inclusions, such that a process, device, article or method that includes a list of elements not only includes those elements, but also includes other elements not explicitly listed, or inherent to such process, device, article or method. Without more limitations, the element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, device, article or method that includes the element.

[0182] The above description is only the preferred embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.

[0183] Although the embodiments of the present application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made thereto without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.

[0184] Of course, the present application can have other various embodiments, and based on the present embodiments, other embodiments obtained by those skilled in the art without any creative labor are within the scope of protection of the present application.< / nl> < / nl> < / nl> < / nl> < / end> < / s> < / l> < / nl> < / s> < / l> < / nl> < / beg> < / nl> < / s> < / nl> < / l> < / s> < / l> < / end> < / beg> < / s> < / end> < / beg> < / l> < / nl>

Claims

1. A method for storing traffic spatiotemporal data, characterized in that, include: Acquire traffic flow data of vehicles on the road, and preprocess the traffic flow data; Receive data storage requests and determine whether there are any available index structures in the index structure set; If the target index structure exists, select it; otherwise, construct a new index structure based on the preprocessed traffic flow data. Based on a pre-established prediction network model, the index structure for the next time step is generated using the index structure for historical time steps.

2. The traffic spatiotemporal data storage method according to claim 1, characterized in that, The index structure is an R*-tree structure, where each node in the tree is associated with a minimum rectangle. Among them, leaf nodes correspond to the smallest rectangle of the spatial entity, and non-leaf nodes generate the smallest rectangle of their own level by aggregating the smallest rectangles of their child nodes. The spatial range of the minimum rectangle is determined by using the lower left coordinate (x) left ,y left ) and the upper right coordinate (x right ,y right Define it.

3. The traffic spatiotemporal data storage method according to claim 2, characterized in that, The construction process of the R*-tree structure includes: Initialize the smallest rectangle of the spatial entity as a leaf node, and recursively construct the smallest rectangle of the non-leaf nodes according to the principle of minimum area increment. When the node capacity reaches the preset threshold, a split operation is performed, triggering the associated minimum rectangle reconstruction and tree structure maintenance operations. The vocabulary of the R*-tree structure is constructed by mapping nodes and index structure node labels; wherein the index structure node labels consist of node type and the coordinates of the minimum rectangle.

4. The traffic spatiotemporal data storage method according to claim 3, characterized in that, Based on the vocabulary, the R*-tree structure is serialized using a traversal algorithm; specifically including: Insert a start label at the beginning of the R*-tree sequence. <beg> ;< / beg> Recursively traverse from the root node to the leaf nodes; add labels to non-leaf nodes. <nl>The coordinates of its minimum bounding box; adding labels to leaf nodes. <l>Its minimum rectangle coordinates; add placeholders for empty entries. <s> Its minimum rectangular frame coordinate representation;< / s> < / l> < / nl> <s> When the traversal ends, backtrack to the root node and add an ending tag. <end> ;< / end> Based on the vocabulary and the traversal algorithm, spatial feature labels and type feature labels are obtained; wherein, the spatial feature label corresponds to the geometric parameters of the minimum bounding box of the node, and the type feature label corresponds to the node type; Spatial feature labels and type feature labels are combined to construct node feature units. The node feature units are then arranged in a tree structure order using a traversal algorithm to generate the final sequence.

5. The traffic spatiotemporal data storage method according to claim 4, characterized in that, The encoder of the prediction network model employs an index-based embedding encoding method to generate word vector embedding and positional encodings for the minimum bounding box of the R*-tree node; specifically including: The word vector embedding encoding is trained using the continuous bag-of-words model and the skip-word model; The global context features of non-leaf nodes in the tree structure are extracted using the jump character model to generate an embedding matrix; After the jump character model is trained, the parameters of the jump character model are fixed. The embedding matrix is ​​used as input to the continuous bag-of-words model. The vector representation of the center node is predicted by aggregating the feature information of the context nodes in the R*-tree, so as to establish the mapping relationship from the context nodes to the center node of the R*-tree.

6. The traffic spatiotemporal data storage method according to claim 5, characterized in that, The location encoding employs a path-based location embedding encoding method; specifically including: Record the access path from the root node to the current node, and combine the path sequence information and preset dimension parameters to calculate and generate the position code of the current node using a sine-cosine function; By combining the word vector embedding encoding and the position encoding of the current node, the encoding result output by the prediction network model is obtained.

7. The traffic spatiotemporal data storage method according to claim 6, characterized in that, The prediction network model is based on the Transformer architecture and uses a series of sequentially stacked multi-label prediction modules to recursively predict the labels of the index structure nodes to be predicted. The multi-label prediction module performs feature processing based on the Transformer base block; Each level of feature is transformed into a probability distribution on the vocabulary through a linear projection head; wherein, the linear projection head generates the probability distribution through a softmax function; The total loss function of the prediction network is obtained by arithmetically averaging the Euclidean distance loss values ​​of several multi-label prediction modules.

8. A traffic spatiotemporal data storage device, characterized in that, For performing the traffic spatiotemporal data storage method according to any one of claims 1 to 7, the traffic spatiotemporal data storage device comprises: The data acquisition and preprocessing module is used to acquire traffic flow data of vehicles on the road and to preprocess the traffic flow data. The dynamic index management module receives data storage requests and determines whether there is an available index structure in the set of index structures. If it exists, the target index structure is selected; if it does not exist, a new index structure is constructed based on the preprocessed traffic flow data. The index prediction and generation module generates the index structure for the next time step based on the index structure of the previous time step using the index structure of the previous time step.

9. A computer device, characterized in that, The computer device includes a processor and a memory coupled to the processor, wherein the memory stores program instructions for implementing the traffic spatiotemporal data storage method according to any one of claims 1-7; the processor is used to execute the program instructions stored in the memory to implement traffic spatiotemporal data storage.

10. A storage medium, characterized in that, The system stores processor-executable program instructions for performing the traffic spatiotemporal data storage method according to any one of claims 1-7. < / s>

Citation Information

Cited By

  • Partition tree decomposition-based time-dependent label constraint efficient query method

    CN121561014A