A road network spatiotemporal data visualization system and method based on kernel density estimation
By decomposing the kernel function through spatiotemporal indexing and second-order difference method, the problems of high complexity and repeated calculation of kernel density estimation algorithm in the existing technology are solved, and efficient online query and visualization of spatiotemporal data are achieved.
Patent Information
- Application Number
- CN202411683761.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-11-22
AI Technical Summary
Existing kernel density estimation algorithms have high computational complexity when processing high resolution or large amounts of data, cannot solve spatiotemporal data online, suffer from serious repeated calculations, lack nonlinear kernel function analysis, and cannot meet real-time visualization requirements.
A road network spatiotemporal data visualization system based on kernel density estimation is adopted. Through spatiotemporal indexing and second-order difference method, the kernel function is decomposed into query vector and aggregation vector. The spatiotemporal index is used to quickly obtain data point aggregates, and nonlinear kernel function is used for calculation to reduce redundant calculations.
It realizes fast online query and visualization of spatiotemporal data, improves computing efficiency, reduces repeated calculations, supports nonlinear kernel function analysis, and meets real-time visualization needs.
Smart Images

Figure CN119597990B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road network graph spatiotemporal data visualization, and in particular to a road network graph spatiotemporal data visualization system and method based on kernel density estimation. Background Art
[0002] Kernel density estimation is a common visualization algorithm that can fit discrete data points into a continuous, smooth data distribution for further visualization analysis. By varying the kernel function and related parameters, different results can be presented. The kernel density estimation method for spatiotemporal data in road network graphs shifts the data carrier from a two-dimensional plane to the road network, and the associated distance calculations also shift from Euclidean distance to the shortest path distance on the road network. This approach is more practical for road network-related data points (such as traffic flow data and traffic accident data).
[0003] The biggest bottleneck of kernel density estimation is its high computational complexity. It requires calculating weights for all pixels and data points pairwise, making it unsuitable for high-resolution or large-scale data. Currently, popular kernel density estimation methods often use a pre-indexing approach, merging points with close spatial locations into a cluster and pre-calculating some values. Compared to baseline methods, this index- and cluster-based algorithm significantly reduces the complexity of online queries. However, some flaws and shortcomings still exist:
[0004] (1) Existing kernel density estimation algorithms only consider spatial data and do not take into account the relationship in the time dimension. Therefore, queries for a specific time interval require screening all data and recalculating them, which cannot be solved online.
[0005] (2) The computational process of the existing kernel density estimation algorithm is still not efficient enough. Although the pre-established index can merge local similarity calculations to a certain extent, from a global perspective, a large number of weights are still repeatedly calculated.
[0006] (3) Existing kernel density estimation algorithms only use linear kernel functions, but there is little analysis of nonlinear kernel functions, and they are mainly fitted with multiple linear kernel functions, lacking precise numerical calculations.
[0007] Therefore, it is necessary to provide a visualization system and method that can utilize the index of spatiotemporal data to accelerate the kernel density estimation algorithm on the road network and perform real-time online visualization analysis. Summary of the Invention
[0008] The purpose of the present invention is to provide a road network graph spatiotemporal data visualization system and method based on kernel density estimation, which can query, calculate and visualize data points in space and time, and improve query and calculation efficiency.
[0009] To achieve the above objectives, the present invention provides a road network graph spatiotemporal data visualization system and method based on kernel density estimation, comprising the following steps:
[0010] S1. Obtain the road network graph structure data consisting of points and edges, collect the data point coordinates of relevant events, and map them to the edges of the road network graph based on spatial distance;
[0011] S2. Determine the kernel function and decompose it into two parts: the query vector and the aggregation vector;
[0012] S3. Establish a spatiotemporal index of the data points on each edge of the road network graph based on the time sequence and space sequence;
[0013] S4. Based on the query time period, obtain the index of the data points within the time range, and obtain the distance range of the current edge pixel according to the upper limit of the distance of the kernel function. Aggregate to obtain the cluster vector containing all target data points, and calculate the kernel density estimation of the current edge pixel using the kernel function with the edge as the basic unit;
[0014] S5. During the query process, when all data points on any edge are within the query range and are within the distance range of the current edge pixel, the kernel density estimate of the current edge pixel is obtained by interpolation using the second-order difference method;
[0015] S6. Obtain kernel density estimates for all pixels, normalize them, and map them onto a color band. Color each edge and draw a heat map.
[0016] Preferably, in step S2, the kernel function selected includes a linear function, a trigonometric function and an exponential function.
[0017] Preferably, in step S3, a spatiotemporal index of the data points on each edge of the road network graph is established, including sorting the data points in spatial order and encoding them in sequence, and then sorting them in chronological order and inserting them into the index in sequence;
[0018] The spatial order refers to the distance between a data point and any endpoint of the edge on which it is located, and the temporal order refers to the time label of the data point.
[0019] Preferably, in step S4, the expression of the kernel function value is as follows:
[0020]
[0021] f(q, oi)=Ks(d(q, pi))·Kt(|t-ti|);
[0022] In the formula, o i represents a data point, p i , t i Represents data point o iThe position and time of f(q , o i ) represents the data point o i The kernel density generated for pixel q, d(q, p i ) represents the spatial distance between the data point and the pixel point, K s , K t They represent kernel functions in time and space dimensions respectively, t represents query time, and F(q) represents the kernel density of pixel q.
[0023] A road network spatiotemporal data visualization system based on kernel density estimation, comprising:
[0024] Data acquisition module, used to obtain road network spatiotemporal data and data points of related events, and build the road network graph structure;
[0025] Establish an indexing module to construct a spatiotemporal index of data points;
[0026] Query optimization module, used to query data points that affect pixels;
[0027] Kernel function value calculation module, used to calculate the kernel function value of pixel points;
[0028] Visualization module, used to color edges and draw heat maps.
[0029] Therefore, the present invention adopts the above-mentioned road network graph spatiotemporal data visualization system and method based on kernel density estimation, which has the following technical effects:
[0030] (1) By adding analysis of time information, the original spatial index is expanded to the spatiotemporal index, which enables online queries in different time periods, realizes the indexing of spatiotemporal data points, and realizes the rapid acquisition of data point clusters and corresponding clustering vectors.
[0031] (2) Accurately locate duplicate information during the query process, design an efficient computing process, and eliminate redundant calculations.
[0032] (3) More nonlinear kernel functions are introduced to meet the requirements of splitting data into constants and variables, merging data points into aggregates, and establishing indexes and query calculations by replacing the corresponding operation types and operation symbols without reconstructing the main code.
[0033] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A spatiotemporal index schematic diagram established in an embodiment of a road network graph spatiotemporal data visualization system and method based on kernel density estimation;
[0035] Figure 2 A schematic diagram of spatiotemporal index links in an embodiment of a spatiotemporal data visualization system and method for a road network graph based on kernel density estimation;
[0036] Figure 3 The present invention is a schematic diagram of an index query process in an embodiment of a road network graph spatiotemporal data visualization system and method based on kernel density estimation;
[0037] Figure 4 The present invention is a schematic diagram of the core structure of a kernel density estimation algorithm of an existing road network graph in an embodiment of a road network graph spatiotemporal data visualization system and method based on kernel density estimation;
[0038] Figure 5 A schematic diagram of query process merging calculation in an embodiment of a road network graph spatiotemporal data visualization system and method based on kernel density estimation;
[0039] Figure 6 The invention discloses a second-order difference diagram in an embodiment of a road network graph spatiotemporal data visualization system and method based on kernel density estimation. DETAILED DESCRIPTION
[0040] The present invention can be explained in more detail by the following examples. The purpose of disclosing the present invention is to protect all changes and improvements within the scope of the present invention. The present invention is not limited to the following examples.
[0041] The present invention provides a road network graph spatiotemporal data visualization system and method based on kernel density estimation, comprising the following steps:
[0042] S1. Collect the structural data of the road network graph and abstract it into a data structure consisting of points and edges. Collect the coordinates of the data points of the relevant events. For each data point, find the nearest edge in the road network graph and map the data point to this edge. Maintain a fixed distance ratio between the data point's original coordinates and the mapped coordinates to the two endpoints. The time stamp of the data point remains unchanged.
[0043] S2. Select a kernel function and confirm whether it can achieve accelerated operation. The kernel function is usually a function with a domain within [0, 1] and monotonically decreasing. It can be split into two parts: the query vector and the aggregation vector. It can be a linear function, such as It can also be a nonlinear function, such as the exponential function K(x)=e x , cosine function K(x)=cos(x), etc.
[0044] Kernel function acceleration primarily involves splitting operations into constant and variable components, namely the query vector and the aggregate vector. The constant component is fixed for a given query, while the variable components across multiple data points can be combined into a single aggregate. The idea behind kernel function decomposition in space and time is as follows:
[0045] If the kernel function in the spatial domain can be decomposed into Q s (q)·A s (Γ), where Q s (q) is the query vector related to q in the spatial domain, A s (Γ) is the aggregation vector related to Γ in the spatial domain, and the kernel function in the time domain can be split into Q t (q)·A t (Γ), then their combination can be expressed as:
[0046]
[0047] where Q ij (q) = Q i Q j , A ij (Γ)=A i ·A j , similarly organized as a query vector and an aggregation vector, whose size is the product of the original spatial and temporal query or aggregation vector sizes. More generally, any kernel function that satisfies the decomposition form can be directly replaced by simply overloading the underlying operators, without affecting the overall algorithm framework.
[0048] Specifically, assuming that the kernel function values of all data points in a data point set Γ need to be calculated, the following decomposition operation can be performed:
[0049]
[0050] Where t represents the query time, p i , t i Represents data point o i The position and time, d(q,v c ) represents the pixel q and the vertex v c The spatial distance, d(q, p i ) represents the pixel q and the data point o i The spatial distance, d(v c , p i ) represents the data point o i With vertex v c The spatial distance, b s 、b t Respectively represent the upper limit of the statistical range of the kernel function in time and space dimensions, F Γ(q) represents the kernel density of all data points in the cluster Γ for the query point (pixel q). The first vector is called the query vector and the second vector is called the cluster vector. The key to the decomposition is that only q is a variable in the query vector and only p is a variable in the cluster vector. i is a variable that satisfies the commutative and associative laws.
[0051] This transformation simplifies the original process of calculating each data point separately into a combined calculation of a constant and an aggregate, greatly improving computational efficiency. Moreover, this aggregate can be freely merged and split to handle multiple queries.
[0052] S3. Since traditional work only considers the spatial range, the position of the aggregate can be simply located by the binary algorithm, and the aggregate can be obtained by the prefix sum algorithm. However, after adding the time range restriction, the estimation of kernel density needs to consider both time and space dimensions, and the two-dimensional prefix sum algorithm will make the query very inefficient. Therefore, in this embodiment, a tree-like spatiotemporal index is established for all data points on each edge, all data points are inserted into the index in chronological order, and the variable part of the data point in the index is dynamically updated for subsequent kernel density calculations. The specific process is as follows:
[0053] The data points are sorted and renumbered according to spatial order (i.e., the distance from a certain endpoint), and then sorted and inserted into the index in sequence according to time order (i.e., time label). The index presents a binary tree structure, and each tree node is an aggregate of several data points. First, the data point is merged into the aggregate of the root node, and then divided according to the number of the data point. If the number is smaller, it is merged into the aggregate of the left child node, otherwise it is merged into the aggregate of the right child node; it continues to recursively merge into the aggregate of the corresponding tree node until the last leaf node has only one data point. In order to record all intermediate results in a persistent way, when a data point is inserted into a tree node, the index will generate a new tree node to record the merged aggregate instead of directly merging it into the original tree node. These two tree nodes correspond to the states before and after the insertion of the new data point, respectively. Therefore, the index can record and restore the index state after each insertion operation.
[0054] like Figure 1As shown, the spatiotemporal index has a tree structure, and each node stores the clustered vectors of several data points. For the four data points (o1, o2, o3, o4) on the current edge, they are inserted into the tree in chronological order from early to late. The insertion position is determined by the spatial position label, where the root node stores all the data points, and then the left and right child nodes store the data points labeled with the first half or the second half, respectively, and continue recursively until the leaf node. T1, T2, T3, and T4 respectively represent the results after the four data points are inserted, and the insertion positions are represented by gray nodes. In addition, since the data of the white nodes remain unchanged, there is no need to copy them each time. Just link to the previous position, as shown in Figure 2 shown.
[0055] When querying an index, first determine the corresponding index range based on the query time range, calculate the difference between the two indexes at each position, and obtain the index consisting of all data points in the time range, that is, extract all data points updated in the query time period, such as Figure 3 As shown in the figure. Next, based on the calculated spatial range, the cluster vectors on these nodes are converted to several nodes on the index. The cluster vectors on these nodes are further aggregated to obtain a cluster vector containing all target data points. This means that the tree nodes covered by a certain distance range are extracted and merged into a single cluster containing the queried data point. This cluster vector is then multiplied by the previously determined query vector to obtain the kernel function value generated for all data points within the query time and spatial range.
[0056] S4, such as Figure 4 As shown in Figure 2, the core structure of the existing kernel density estimation method for road network graphs is as follows:
[0057] For a a , v b ) on the pixel q, suppose we need to calculate the other edge (v c , v d ) The kernel function value of all data points on . First, the shortest path algorithm is used to calculate the distance from q to v c and v d The distance between the two endpoints is determined to determine which endpoint is closer; secondly, the boundary b of the region is calculated. s -d(q,v_c) and b s -d(q,v d ); Finally, we use the binary method to find out which data points will affect the pixel. It can be seen that v c There is one data point on the side, v d There are two data points on the side, and it is necessary to calculate the kernel function value generated by these data points.
[0058] Since the computation overhead for each data point is too high, this embodiment expands the spatial distance of the existing road network graph to the spatiotemporal distance. The corresponding data point clusters can be obtained by querying the time period. As long as the cluster vectors corresponding to these data points can be quickly obtained, the kernel function value can be quickly calculated in constant time using the decomposed kernel function, that is, there is no need to calculate each data point separately, as shown below:
[0059] Given a query time period, a kernel density estimate is calculated for each pixel, i.e., the kernel function value of pixel q and all data points. This process is performed on an edge-by-edge basis. First, a naive shortest path algorithm (e.g., Dijkstra's algorithm) is used to determine the distance from pixel q to all graph vertices. The upper limit of the kernel function distance is then used to determine the distance range on that edge. Next, the corresponding data point cluster is retrieved from the pre-established data point index, without having to re-traverse all data points each time. Finally, the kernel density estimate for that edge is calculated by combining the constant corresponding to the query with the variable of the cluster. This process is repeated for each edge to obtain the final result.
[0060] Among them, for a pixel point q, its kernel density value is calculated as follows:
[0061]
[0062] f(q, oi)=Ks(d(q, pi))·Kt(|t-ti|);
[0063] In the formula, o i represents the data point, f(·) is used to measure each data point o i The effect on pixel q is obtained by multiplying two kernel functions, which are measured in the spatial dimension and the temporal dimension respectively; K s , K t Represent the kernel functions in time and space dimensions, f(q, o i ) represents the data point o i The kernel density generated by pixel q, F(q) represents the kernel density of pixel q. In the road network graph, the spatial distance d(q, p i ) uses the shortest path distance calculation, the closer the data point is to the pixel point (d(q, p i ) is smaller) and closer to the center of the query time range (|tt i |The smaller it is), the greater the impact.
[0064] S5. Since the above method only applies to data points on a single edge, the indexes on multiple edges still need to be calculated separately and sequentially, and these results cannot be shared across multiple indexes. Therefore, this embodiment uses a query optimization algorithm to perform unified calculations on the indexes on multiple edges, thereby reducing the amount of calculation and improving calculation efficiency, including:
[0065] During the query process, there is a special case where all data points on a certain edge are within the distance range of all pixel points on the other edge. In this case, these data points will contribute to the kernel function value of these pixel points, and these contributions have a second-order linear relationship. If this special data is found during the query process, the linear variation characteristics can be directly used to directly determine the correlation coefficient in the kernel density calculation formula through interpolation, and then the kernel density value can be directly calculated without the need for specific query on the index, further improving efficiency. The specific calculation process is as follows: Figure 5 shown.
[0066] Assume that the edge (v c , v d ) are within the range of the query point and are at a distance v from any pixel q1 to q5. c Closer, that is, these five data points are always in the same cluster, then the kernel function values generated by this cluster for the five pixels can be combined. Further assume that q1, q2, q3 are at a distance v a Closer, while q4,q5 are closer to v b If the distance between the two clusters is closer, the kernel function value of the cluster decreases on q1 to q3 (the spatial distance is getting farther), and increases on q4 to q5 (the spatial distance is getting closer).
[0067] In this implementation, since the kernel function expression is about d(q,v c ), so the difference between the kernel function values on q1 and q2 is a constant multiple of the difference between the distances of the two pixels. Similarly, the difference between the kernel functions of q2 and q3, q4 and q5 is the same. However, the difference between q3 and q4 is uncertain because the shortest path distances of these two pixels are calculated from two different directions. In summary, (v a ,v b ) is expressed as two arithmetic progressions.
[0068] Considering that the difference of the sum of a series is equivalent to the sum of the differences of each series, we can only store the second-order differences of the arithmetic series, such as Figure 6 As shown. The first-order difference of an arithmetic sequence is a constant function, and the difference of a constant function is a unique function (i.e., it has only one non-zero value). Figure 5In the case shown, the second-order difference only needs to be saved as two values. For other edges, if this condition is met, the second-order differences can be further superimposed. Finally, the second-order differences can be restored to the original sequence. Compared to the prior art method of calculating the kernel function value of each cluster for all pixels, the method of this embodiment only needs to calculate a constant kernel function value each time, and then perform a unified calculation, which can reduce the amount of calculation and improve processing efficiency.
[0069] S6. Visualization. Collect and normalize kernel density estimates for all pixels, map them onto a color band using common software, color each edge, and create a heat map.
[0070] A system for visualizing spatiotemporal data of a road network graph based on kernel density estimation includes: a data acquisition module for acquiring spatiotemporal data of a road network, abstracting it into a structure of points and edges, collecting data points of events, and constructing a road network graph structure; an indexing module for constructing a spatiotemporal index of data points; a query optimization module for querying data points that affect pixel points; and a visualization module for coloring edges and drawing heat maps.
[0071] Therefore, the present invention adopts the above-mentioned road network graph spatiotemporal data visualization system and method based on kernel density estimation, expands the spatial index to the spatiotemporal index, realizes online query for different time periods, and accurately locates repeated information in the query process to improve computing efficiency.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for visualizing spatiotemporal data of road network graphs based on kernel density estimation, characterized in that: The following steps are involved: S1. Obtain the road network graph structure data consisting of points and edges, collect the data point coordinates of relevant events, and map them to the edges of the road network graph based on spatial distance; S2. Determine the kernel function and decompose it into two parts: the query vector and the aggregation vector; S3. Establishing a spatiotemporal index of the data points on each edge of the road network graph according to the temporal and spatial order, including sorting the data points in spatial order and encoding them in sequence, and then sorting them in temporal order and inserting them into the index in sequence; Among them, spatial order refers to the distance between a data point and any endpoint of the edge, and temporal order refers to the time label of the data point; S4. Based on the query time period, obtain the index of the data points within the time range, and obtain the distance range of the current edge pixel according to the upper limit of the distance of the kernel function. Aggregate to obtain the cluster vector containing all target data points, and calculate the kernel density estimation of the current edge pixel using the kernel function with the edge as the basic unit; Among them, the expression of kernel density estimation is: ; ; Where, represents a data point, 、 Represents data points location and time, Represents a data point Pixel The resulting kernel density, Represents the spatial distance between the data point and the pixel point, 、 denote the kernel functions in time and space dimensions respectively, Indicates the query time, Represents pixel points The kernel density of S5. During the query process, when all data points on any edge are within the query range and are within the distance range of the current edge pixel, the kernel density estimate of the current edge pixel is obtained by interpolation using the second-order difference method; S6. Obtain kernel density estimates for all pixels, normalize them, and map them to a color band. Color each edge and draw a heat map.
2. The method for visualizing spatiotemporal data of a road network graph based on kernel density estimation according to claim 1, characterized in that: The choices of kernel functions include linear functions, trigonometric functions and exponential functions.
3. A road network spatiotemporal data visualization system based on kernel density estimation, used to implement the road network spatiotemporal data visualization method based on kernel density estimation according to any one of claims 1-2, characterized in that: include: Data acquisition module, used to obtain road network spatiotemporal data and data points of related events, and build the road network graph structure; Establish an indexing module to construct a spatiotemporal index of data points; Query optimization module, used to query data points that affect pixels; Kernel function value calculation module, used to calculate the kernel function value of pixel points; Visualization module, used to color edges and draw heat maps.