Spatial position sensitive spatio-temporal data lake data management method
Through the localized computing strategy of unified spatiotemporal encoding and optimization, the problem of low data management and analysis processing efficiency of data lakes in complex cloud-edge distribution scenarios is solved, and the efficiency of data localization processing and the balance of computing response speed is achieved.
Patent Information
- Application Number
- CN202411989843.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology is difficult to effectively improve the efficiency of data lakes in data management and analysis processing, especially when facing complex cloud-edge distribution scenarios and high-performance data analysis needs.
Unified spatiotemporal encoding is used to organize storage, computing and data acquisition resources, and use spatiotemporal information as the main indexing mechanism for defining data localization, and uniformly handle computing scenarios such as stream computing, batch computing and spatiotemporal query through optimization-oriented localized computing strategies.
The efficiency of data localization processing is improved, and the performance of data lakes in data management and analysis processing is improved by reducing data transmission delay and improving computing response speed.
Smart Images

Figure CN120067228A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and in particular relates to a data management method for a spatio-temporal data lake sensitive to spatial location, mainly used to improve the efficiency of the data lake in data management and analysis processing. Background Art
[0002] Digital twins pose higher requirements for high-performance data analysis and processing. For various intelligent applications, it is necessary to process and analyze massive data in real time. The deployment scenarios of digital twins are becoming increasingly complex, involving a large number of cloud-edge distribution requirements. Therefore, enhancing data localization and sinking data processing to the edge has become an important means to improve data processing efficiency and user response capabilities. As a platform that supports digital twins and manages various complex data including spatio-temporal information, the data lake has high requirements for processing performance, throughput, and response capabilities. Therefore, it is necessary to improve the data-computation high-performance platform capabilities required for the support platform of digital twin applications by enhancing computing power and reducing data transmission latency. In the computing scenarios required to support digital twins, the most important ones include: batch processing, stream processing, and spatio-temporal information processing modes. Therefore, in terms of improving computing efficiency, it is necessary to provide data localization for batch processing, stream processing data localization, and data query processing for various types of data.
[0003] This project proposes to achieve computing localization based on spatio-temporal attributes, thereby improving the efficiency of the data lake in data management and analysis processing. This technology respectively corresponds to several typical computing scenarios of the data lake: stream processing computing, batch processing computing, and high-performance data query and update. Summary of the Invention
[0004] Data localization is mainly achieved through reasonable data distribution and load balancing to achieve on-demand computing and timely data extraction. In the scenario facing digital twins, the spatio-temporal relationship is an important indexing mechanism for organizing data and computing resources. Therefore, by limiting the allocation of data and computing resources based on the spatio-temporal correlation relationship, it is an important means to improve data localization processing. At the same time, in digital twin applications, the data lake needs to support stream computing, batch processing computing, and data query (especially spatio-temporal query). Therefore, spatio-temporal-based data localization needs to take into account the scenarios of stream computing, batch processing computing, and spatio-temporal query. For this reason, the present invention proposes to use unified spatio-temporal coding to organize storage, computing, and data acquisition resources, and use spatio-temporal information as the main indexing mechanism for defining data localization to locate the computing nodes or storage media where data transmission and computing are located, and use unified spatio-temporal-based localized computing to process computing scenarios such as stream computing, batch processing computing, and spatio-temporal query.
[0005] The present invention provides the following technical solution: a method for managing spatio-temporal data lake data sensitive to spatial location, which uses a unified spatio-temporal encoding to organize storage, computing, and data acquisition resources. The spatio-temporal encoding is represented in a four-dimensional space and includes a spatial cube and a time range. Using spatio-temporal information as the main indexing mechanism for defining data localization, it is used to locate the computing nodes or storage media where data transmission and computing occur. Through an optimization-oriented local computing strategy, it uniformly processes computing scenarios such as stream computing, batch processing computing, and spatio-temporal queries, achieving an effective balance between local computing and response speed.
[0006] Preferably, it includes establishing resource management based on spatio-temporal indexing. The resource management includes the following contents: Spatio-temporal index; Coding indexes of computing, storage, and sensor resources; Application-based data flow graph.
[0007] Preferably, the spatio-temporal index is represented in a four-dimensional space. The spatial cube is demarcated by two spatial points from {x i , y i , z i} to {x j , y j , z j}, and the time range is demarcated from t i to t j .
[0008] Preferably, each computing device, storage device, and sensor has a corresponding spatial index, and the spatial indexes can overlap. The spatial index is based on the corresponding spatial range in the digital twin space and also considers the physical device deployment topology.
[0009] Preferably, the node coordinates of the data flow graph G D are represented as {V D , E D}, where V D represents the node type and E D represents the connection type.
[0010] Preferably, it also includes spatio-temporal data source management and computing location positioning. 7. According to a method for managing spatio-temporal data lake data sensitive to spatial location as described in claim 6, it is characterized in that: by locating the data source v ds and the consumer end v cs , and collecting the corresponding data flow graph nodes and spatio-temporal index sets to achieve computing location positioning.
[0011] Preferably, it also includes an optimization-oriented local computing strategy.
[0012] Preferably, the localization calculation strategy is based on heuristic search, and the input includes data source v ds , consumer side v cs , data flow graph nodes and their corresponding spatio-temporal index sets, and application calculation amount requirement C vol .
[0013] Preferably, the heuristic search is a cluster of computing nodes that meets the calculation amount requirement, and nodes in the same cluster are preferentially selected to improve the spatio-temporal localization effect. During the search process of the cluster of computing nodes C CLUSTER , nodes in the same cluster are preferentially selected to improve the spatio-temporal localization effect, and the communication delay between nodes is used as one of the judgment conditions for meeting the calculation amount.
[0014] Compared with the prior art, the present invention provides a spatio-temporal data lake data management method sensitive to spatial position, having the following beneficial effects: Adopt a unified spatio-temporal coding to organize storage, computing, and data acquisition resources, and use spatio-temporal information as the main index mechanism for defining data localization to locate the computing nodes or storage media where data transmission and computing are located. Use unified spatio-temporal-based localization computing to process computing scenarios such as stream computing, batch computing, and spatio-temporal queries. Through the search of an optimization-oriented localization computing strategy, effectively balance the localization computing and response speed and improve the efficiency of the data lake in data management and analysis processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is the spatio-temporal index of the present invention; Figure 2 is the spatio-temporal clustering of the present invention; Figure 3 is the data flow graph of the present invention; Figure 4 is the flow chart of the heuristic search of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] The implementation manner of the present invention includes organizing storage, computing, and sensor data acquisition through a spatio-temporal index; managing spatio-temporal data sources (data producers) and data consumers (including related computing), and based on this, locating the computing location to achieve data localization processing; through the search of an optimization-oriented localization computing strategy, effectively balance the localization computing and response speed. The detailed implementation manner is described as follows: Resource management based on spatio-temporal index Establishing resource management based on spatio-temporal index includes three parts: 1. Spatio-temporal index; 2. Coding index of computing, storage, and sensor resources; 3. Data flow graph based on applications.
[0017] As Figure 1 shown, the spatio-temporal index is represented in a four-dimensional space as I ST = ({x i , y i , z i , t i}, {x j , y j , z j ,t j}), where the index represents the spatio-temporal range, including the spatial cube demarcated by two spatial points from {x i , y i , z i} to {x j , y j , z j}, and the time range demarcated from t i to t j . Each computing device, storage device, and sensor has its corresponding spatial index, and there can be an overlap between the spatial indices I STi and I STj . The spatio-temporal coordinates are all based on the spatio-temporal in the digital space.
[0018] Determine that the coordinates related to space are based on their corresponding spatial ranges in the digital twin space, and at the same time consider the deployment topologies of physical computing devices, storage devices, and sensors to avoid data localization failures caused by mismatches between digital spatio-temporal and physical spaces.
[0019] To achieve digital localization analysis, it is necessary to determine the data flow graph starting from the application. The nodes in the graph represent sensors, storage, and computing nodes, and the edges represent the data channels between the nodes. The purpose of establishing the data flow graph is to identify the data production end, consumption end, and participating nodes in application computing (including storage and sensors) based on the spatio-temporal index, facilitating the positioning of the computing resources required for a given production end (or multiple production ends) and consumption end (or multiple consumption ends) and application computing to achieve computing localization.
[0020] The definition of the data flow graph is as follows: G D = {V D , E D}, V D = {v d0 , v d1 , … v dm}, E D = {e d0 , e d1 , … e dr}.
[0021] v di represents 3 types of node types: data storage nodes, computing nodes, and sensor nodes; e dj represents 5 types of connection types 1) Standard TCP connection; 2) RDMA connection; 3) RoCE connection; 4) IoT protocol connection; 5) PCIe connection.
[0022] Computing Location Based on Spatiotemporal Resources After the data source and the consumer end of a given data application (stream processing, batch processing, or data query) are given, locate the computing resources through the following steps: First, locate the data source v ds and the consumer end v cs , and collect the data flow graph nodes reaching both ends to obtain v as = { v di , v di+1 , … v dr}, collect the corresponding spatiotemporal index set of v as ; Based on the spatiotemporal index set, locate the local computing cluster (heuristic condition: close to the data production end) C CLUSTRER . The specific process is as follows: Step 1: Locate the data source v ds and the consumer end v cs ; Step 2: Collect the data flow graph nodes reaching both ends to obtain v as = { v di , v di+1 , … v dr}; Step 3: Collect the corresponding spatiotemporal index set I as of v as ; Step 4: Locate the local computing cluster C CLUSTRER .
[0023] For the specific process details, please refer to the appendix Figure 3 .
[0024] Localized Computing Strategy for Optimization The localized computing strategy for optimization is based on heuristic search. The search inputs include: v ds , v cs , v as = { v di , v di+1 , … v dr}, v as The corresponding spatio-temporal index set I as , and the application computing workload requirement C vol (The computing power requirement can be estimated). The search process starts from the data source v ds side and traverses v as in v ds to v cs computing nodes. The specific process is as follows: 1. First, the search input includes: v ds , v cs , v as = { v di , v di+1 , … v dr}, v as The corresponding spatio-temporal index set I as , and the application computing workload requirement C vol ; 2. Obtain the node set ne; 3. Cluster according to the spatio-temporal index in the adjacent node set; 4. Search in the clustering of the adjacent node set to see if there is a computing node cluster C vol that meets the computing workload C CLUSTER ; 5. If it meets the requirement, try to lock it. If it does not meet the requirement, further expand the adjacent node set and search again.
[0025] For the details of the process, see the appendix Figure 4 .
[0026] As Figure 2 shown, get_neighbour(ne, v as ) obtains the adjacent node set of the node set ne in the set v as . st_cluster clusters according to the spatio-temporal index in the adjacent node set. search_comp searches in the clustering of the adjacent node set to see if there is a computing node cluster C vol that meets the computing workload C CLUSTER . The computing nodes prefer to select the nodes belonging to the same cluster to improve the spatio-temporal localization effect. If it meets the requirement, try to lock it. If it does not meet the requirement, further expand the adjacent node set and search again. If the lock of C CLUSTER fails, also expand the adjacent node set and search again.
[0027] During the search_comp process, while considering the computing power of the nodes, the communication delay between the nodes is also used as a condition to meet the computing workload. Specifically, the user can determine the degree of communication delay according to the type of node connection.
[0028] Through the search of an optimization-oriented local computing strategy, an effective balance between local computing and response speed is achieved, and the efficiency of the data lake in data management and analysis processing is improved.
Claims
1. A spatial location-sensitive spatiotemporal data lake data management method, characterized by: Using a unified space-time code to organize storage, computing and data acquisition resources, the space-time code is represented in four-dimensional space, including a space cube and a time range; Using spatiotemporal information as the main indexing mechanism to define data locality, used to locate the computing nodes or storage media where data transmission and computation are located; Through optimization-oriented localized computing strategies, computing scenarios such as stream computing, batch computing, and spatiotemporal query are uniformly processed to achieve an effective balance between localized computing and response speed.
2. A spatial location-sensitive spatiotemporal data lake data management method according to claim 1, characterized in that: It includes establishing resource management based on spatiotemporal index, and the resource management includes the following contents: Spatiotemporal index; Coding indexes of computing, storage, and sensor resources; Application-based data flow diagram.
3. A spatial location-sensitive spatiotemporal data lake data management method according to claim 2, characterized in that: The spatiotemporal index is represented in four-dimensional space, and the space cube is represented by {x i , y i , z i } to {x j , y j , z j } two spatial points, the time range is from t i to j The marked.
4. A spatial location-sensitive spatiotemporal data lake data management method according to claim 3, characterized in that: Each computing device, storage device, and sensor has a corresponding spatial index, and the spatial indexes may overlap. The spatial index is based on the corresponding spatial range in the digital twin space, while taking into account the physical device deployment topology.
5. The spatial location-sensitive spatiotemporal data lake data management method according to claim 2, characterized in that: The data flow graph G D The node coordinates are expressed as {V D , E D }, where V D Indicates the node type, E D Indicates the connection type.
6. The spatial location-sensitive spatiotemporal data lake data management method according to claim 1, characterized in that: It also includes time-space based data source management and location calculation.
7. A spatial location-sensitive spatiotemporal data lake data management method according to claim 6, characterized in that: By locating the data source v ds and consumer side cs , and collect the corresponding data flow graph nodes and time-space index sets to realize the calculation position positioning.
8. The spatial location-sensitive spatiotemporal data lake data management method according to claim 1, characterized in that: It also includes localized computation strategies for optimization.
9. A spatial location-sensitive spatiotemporal data lake data management method according to claim 8, characterized in that: The localization calculation strategy is based on heuristic search, and the input includes the data source v ds 、Consumer side cs , data flow graph nodes and their corresponding spatiotemporal index sets and application computing requirements C vol .
10. A spatial location-sensitive spatiotemporal data lake data management method according to claim 9, characterized in that: The heuristic search is for a computing node cluster that meets the computing demand, and nodes in the same cluster are preferentially selected to improve the spatiotemporal localization effect. The computing node cluster C CLUSTER In the search process, nodes in the same cluster are given priority to improve the spatiotemporal localization effect, and the communication delay between nodes is used as one of the judgment conditions for satisfying the calculation amount.