Space-time retrieval method, device and equipment based on multi-element data fusion and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-27
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本申请的目的在于提供一种基于多元数据融合的时空检索方法、装置、设备和介质,解决了需要对数据进行预先设置索引层级,以及在高级别层级索引量巨大而导致构建索引的效率较低和存储困难的问题
通过提取多元异构数据的时空范围信息,可以将多元异构数据按照时空范围信息进行空间索引,提升包含时空范围信息的空间查询效率,无需预先设置索引层级,索引编码与地理网格进行关联构建目标索引结构,在高级别层索引量巨大时,保证了构建索引的效率。
Smart Images

Figure CN116775722B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a spatiotemporal retrieval method, apparatus, device, and medium based on multi-source data fusion. Background Technology
[0002] With the development of Earth observation, the Internet of Things, 5G, and mobile internet, massive amounts of spatiotemporal data are exploding. This data exhibits characteristics such as spatiotemporal multidimensionality, large volume, and scarce value. Effectively utilizing this massive amount of data presents two severe challenges: how to store the data and how to quickly retrieve the data needed.
[0003] Among related technologies, such massive amounts of data can be stored based on distributed technologies (such as HDFS, HBase, etc.). However, solutions based on the Hadoop ecosystem are mostly used for offline or streaming computing and storage, which are insufficient for real-time query and multi-dimensional spatiotemporal query capabilities, and do not support more complex vector data. Using pre-computation technology (such as Apache Kylin) can enhance multi-dimensional query capabilities, but the pre-computation period also introduces latency, which cannot adapt to scenarios with high real-time requirements due to real-time changes.
[0004] Therefore, there is currently a problem that various types of massive amounts of data, such as surveying and mapping, remote sensing, meteorology and oceanography, cannot be searched and queried in a unified manner. Summary of the Invention
[0005] The purpose of this application is to provide a spatiotemporal retrieval method, apparatus, device, and medium based on multi-source data fusion, which solves the problems of needing to pre-set the index hierarchy of data, and the low efficiency and storage difficulties caused by the huge amount of index at high-level levels.
[0006] In a first aspect, the present invention provides a spatiotemporal retrieval method based on multi-source data fusion. The method includes: acquiring multi-source heterogeneous data and extracting spatiotemporal range information from the multi-source heterogeneous data; determining effective index boundaries based on the spatial positional relationship between the spatiotemporal range information of the multi-source heterogeneous data and grid encoding; performing index encoding processing within the effective index boundaries based on a preset grid level, and associating the index encoding with a geographic grid to construct a target index structure for spatiotemporal retrieval; the target index structure contains one or more data types within different grid ranges at different levels.
[0007] In optional implementations, the data source types of diverse heterogeneous data include at least remote sensing image data, navigation video data, surveying and mapping geographic data, meteorological and oceanographic data, geological data, and nuclear, biological and chemical data; the data structure types of diverse heterogeneous data include at least document data, video data, image data, vector data, and raster data.
[0008] In an optional implementation, the effective index boundary is determined based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid codes. This includes: acquiring the data to be queried; calculating the corresponding grid codes of the data to be queried according to the grid level order based on preset coding constraints; wherein the coding constraints include at least preset segmentation rules, grid code values, grid coding modes, effective coding range, and coding model types; and determining the effective index boundary based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid codes.
[0009] In an optional implementation, the grid code corresponding to the data to be queried is calculated according to the grid level order based on preset coding constraints. This includes: determining whether there is an intersection between the grid code and the geographic range corresponding to the data to be queried; the grid code is used to represent the range of the corresponding geographic grid at a fixed subdivision level; if so, determining whether the grid level exceeds a preset limit level; if it exceeds, determining the grid code corresponding to the current level as the grid code contained within the geographic range corresponding to the data to be queried; if it does not exceed the preset limit level, determining whether the spatial range completely contains the range of the geographic grid; if they intersect and contain each other, determining the grid code corresponding to the current level as the grid code contained within the geographic range corresponding to the data to be queried; if they intersect but do not contain each other, calculating the grid code of the next grid level, until all grid codes at all levels have been calculated.
[0010] In an optional implementation, the index codes are associated with geographic grids to construct a target index structure, including: calculating the number of associated grids at different levels according to the data range and in a specified order; if the number of associated grids meets the preset range, performing grid association indexing between the geographic grids and index codes at the current level; if the number of associated grids does not meet the preset range, expanding the current level, and performing grid association processing between the geographic grids and index codes at the expanded level; and constructing the target index structure after all index codes are associated with geographic grids.
[0011] In an optional implementation, the target index structure is stored in array format, and the data objects and their codes have a one-to-many relationship.
[0012] In an optional implementation, the data type contained in each grid in the target index structure is represented by color and / or numbers.
[0013] Secondly, the present invention provides a spatiotemporal retrieval device based on multi-source data fusion. The device includes: a spatiotemporal extraction module for acquiring multi-source heterogeneous data and extracting the spatiotemporal range information of the multi-source heterogeneous data; an index boundary determination module for determining the effective index boundary based on the spatial positional relationship between the spatiotemporal range information of the multi-source heterogeneous data and the grid encoding; and an index structure construction module for performing index encoding processing based on a preset grid level within the effective index boundary, and associating the index encoding with a geographic grid to construct a target index structure for spatiotemporal retrieval through the target index structure; the target index structure contains one or more data types within different grid ranges at different levels.
[0014] Thirdly, the present invention provides an electronic device including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the spatiotemporal retrieval method based on multi-source data fusion according to any of the foregoing embodiments.
[0015] Fourthly, the present invention provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are invoked and executed by a processor, the computer-executable instructions cause the processor to implement the spatiotemporal retrieval method based on multi-source data fusion according to any of the foregoing embodiments.
[0016] The beneficial effects of the spatiotemporal retrieval method, apparatus, device, and medium based on multivariate data fusion provided in this application are as follows: By extracting the spatiotemporal range information of diverse heterogeneous data, spatial indexing can be performed on the diverse heterogeneous data according to the spatiotemporal range information, thereby improving the efficiency of spatial querying containing spatiotemporal range information. There is no need to pre-set the index level. The index code is associated with the geographic grid to build the target index structure, which ensures the efficiency of index building when the amount of high-level index is huge. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 A schematic diagram of the system architecture corresponding to a spatiotemporal retrieval method based on multi-source data fusion provided in an embodiment of this application; Figure 2 A flowchart illustrating a spatiotemporal retrieval method based on multivariate data fusion, provided for embodiments of this application; Figure 3(a) is a flowchart of a grid encoding to two-dimensional and three-dimensional spatial range provided in an embodiment of this application; Figure 3(b) is a flowchart of a two-dimensional and three-dimensional spatial range to grid encoding provided in an embodiment of this application; Figure 4 A schematic diagram of an effective boundary provided for an embodiment of this application; Figure 5 A schematic diagram of two-dimensional and three-dimensional encoding subdivision provided for embodiments of this application; Figure 6 A data caching process flowchart provided in this application embodiment; Figure 7(a) is a schematic diagram of the association between geographic entities and grids provided in the embodiments of this application; Figure 7(b) is a schematic diagram of the association between geographic entities and high-precision grids provided in the embodiments of this application; Figure 8 This is a schematic diagram of a mesh generation tool interface provided in an embodiment of this application; Figure 9 A 3D display effect diagram of mesh subdivision provided in an embodiment of this application; Figure 10 An illustration of a spacetime cube construction effect provided in an embodiment of this application; Figure 11 A structural diagram of a spatiotemporal retrieval device based on multi-source data fusion provided in an embodiment of this application; Figure 12 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0020] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0021] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0022] For ease of understanding, the system architecture corresponding to the spatiotemporal retrieval method based on multi-source data fusion provided in the embodiments of this application will be described first. See [link to relevant documentation]. Figure 1 As shown, the overall architecture adopts a layered structure, consisting of a visualization layer, an API layer, a service layer, and a storage layer from top to bottom. Functionally, it can be divided into four modules: data storage, index building and querying, index cache building, and encoding / decoding algorithms.
[0023] First, the data storage module provides the function of parsing the data source, extracting spatiotemporal information, and parsing the data source into data objects. After the data source is processed by the storage module, spatiotemporal range information is extracted, and structured and unstructured information are separated. Second, encoding and decoding algorithms are the core algorithms for constructing multidimensional index grids, and can accelerate spatial range retrieval efficiency. Encoding and decoding algorithms are used to encode or decode the spatial range of data objects according to encoding and decoding rules.
[0024] Third, index caching is used to accelerate index query efficiency. Frequently used indexes or complex aggregation results are temporarily stored using caching techniques, and then read from the cache during queries.
[0025] Fourth, indexed queries utilize spatial range, time range, and other query conditions for rapid retrieval of multidimensional data. Through encoding and decoding algorithms, query conditions are converted into encoded query conditions. The system first checks the cache; if the cache is not found, it then queries the index storage.
[0026] This application provides a spatiotemporal retrieval method based on multi-source data fusion, see [link to relevant documentation]. Figure 2 As shown, the method mainly includes the following steps: Step S210: Obtain multivariate heterogeneous data and extract the spatiotemporal range information of the multivariate heterogeneous data.
[0027] In one implementation, the multi-dimensional heterogeneous data are data that all contain spatiotemporal range information. For example, remote sensing image data, navigation video data, surveying and mapping geographic data, meteorological and oceanographic data, geological data, etc., all contain latitude and longitude information, so the corresponding spatiotemporal range information can be extracted from different data.
[0028] The data source types of multi-source heterogeneous data should include at least remote sensing image data, navigation video data, surveying and mapping geographic data, meteorological and oceanographic data, geological data, and nuclear, biological and chemical data; the data structure types of multi-source heterogeneous data should include at least document data, video data, image data, vector data, and raster data.
[0029] In addition, according to the data production level, multi-source heterogeneous data can be divided into: raw data and product data at all levels processed based on raw data.
[0030] Step S220: Determine the effective index boundary based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid encoding.
[0031] The above grid coding was obtained by global grid subdivision according to the GB / T 40087-2021 "Rules for Geospatial Grid Coding". The subdivision level of the geospatial grid is divided into 4 categories and 33 levels: degree grid includes 10 levels from 0 to 9; minute grid includes 6 levels from 10 to 15; second grid includes 6 levels from 16 to 21; and sub-second grid includes 11 levels from 22 to 32.
[0032] The spatial relationship between the spatiotemporal range information and grid encoding of multivariate heterogeneous data can include at least overlap, complete coverage, partial coverage, and non-intersection. To maximize retrieval efficiency, index boundaries can be determined based on the spatial relationship between the spatiotemporal range information and grid encoding of multivariate heterogeneous data. Retrieval can then be performed according to appropriate index boundaries, and data outside these boundaries can be ignored, thereby improving the retrieval efficiency using the target retrieval structure.
[0033] Step S230: Within the effective index boundary, perform index encoding processing based on the preset grid level, and associate the index encoding with the geographic grid to construct the target index structure for spatiotemporal retrieval; the target index structure contains one or more data types within different grid ranges at different levels.
[0034] In one implementation, the target index structure is stored in array format, with a one-to-many relationship between data objects and their codes. Furthermore, to provide a clear view of the index interface, the data type contained in each grid within the target index structure is represented using color and / or numbers.
[0035] For ease of understanding, the following provides a detailed explanation of each implementation method.
[0036] In an optional implementation, the effective index boundary is determined based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid encoding. In specific implementation, this may include the following steps 1 and 2: Step 1: Obtain the data to be queried. Calculate the corresponding grid code for the data to be queried according to the grid level order based on preset coding constraints. The coding constraints include at least the preset segmentation rules, grid code values, grid coding modes, effective coding range, and coding model type. Step 2: Determine the effective index boundaries based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid encoding.
[0037] In one implementation, step 1 above, which calculates the corresponding grid code for the data to be queried according to the grid level order based on preset coding constraints, may include the following steps 1.1 to 1.6 in a specific implementation: Step 1.1: Determine whether there is an intersection between the grid code and the geographic area corresponding to the data to be queried; the grid code is used to represent the range of the corresponding geographic grid at a fixed subdivision level; Step 1.2: If yes, determine whether the grid level exceeds the preset limit level; If step 1.3 exceeds the limit, the grid code corresponding to the current level will be determined as the grid code contained within the geographical range corresponding to the data to be queried; Step 1.4 If the preset limit level is not exceeded, determine whether the spatial range completely includes the geographic grid range; Step 1.5 If they intersect and contain each other, then the grid code corresponding to the current level is changed to the grid code contained within the geographic range corresponding to the data to be queried; Step 1.6 If the grids intersect but do not contain each other, calculate the grid code for the next grid level, until all grid codes for all levels have been calculated.
[0038] In one example, this application embodiment provides a method for encoding and decoding between spatial range and grid encoding: First, by using regular expressions, the input parameters are constrained according to the encoding standards to restrict the input of erroneous data. This allows for the early detection of erroneous data and the provision of warnings, thus ensuring the quality of processing.
[0039] Second, we abstract the segmentation rules, encoding methods, and encoding values to achieve the scalability of tree structure encoding models such as quadtrees and octrees. Third, the setting of the effective range restricts the effective range of the encoding, realizing the concept of virtual space area in the national standard; Fourth, spatial relationship (intersecting but not containing, intersecting and containing, non-intersecting) judgment is used to avoid excessive data splitting.
[0040] Figure 3(a) illustrates a flowchart for converting grid encoding into two-dimensional and three-dimensional spatial ranges. First, the input encoding code is subjected to regular expression matching for rule checking to determine if it conforms to the encoding rules. If it does, the encoding level is calculated, and data is read and parsed; otherwise, parsing fails. During data reading and parsing, constraints may include grid segmentation rules at different levels in the x, y, and z directions, grid encoding values, grid encoding pattern (Z-code), setting an effective range, and setting the dimension (2D or 3D). Further, the encoded spatial range is calculated based on the calculated encoding level and the parsed information. It is then determined whether this range exceeds the defined boundaries. If it does not exceed the boundaries, the effective spatial range is determined; if it exceeds the boundaries, the effective range is truncated, and non-empty filtering is performed. By summarizing the filtering results, the final effective spatial range is obtained.
[0041] Figure 3(b) illustrates a flowchart for converting two-dimensional and three-dimensional spatial ranges into grid codes. First, the grid level and spatial range are acquired and read / parsed. During data reading and parsing, constraints may include grid segmentation rules at different levels in the x, y, and z directions, grid code values, grid coding patterns (Z-code), setting the effective range, and setting the dimension (2D or 3D). Further calculations are initiated, setting the initial grid level to 1. It is determined whether the grid and the input range intersect. If there is no intersection, completely non-intersecting grids are discarded. If there is an intersection, it is determined whether the grid level exceeds the limit. If it does, the grid code result corresponding to the spatial range is obtained. If it does not exceed the limit, it is further determined whether the spatial range contains the grid. For cases where there is intersection but no overlap, the next level of coding is calculated, and the intersection between the grid and the input range is re-evaluated. For cases where there is both intersection and overlap, the result is added to the result set, determining the grid code result corresponding to the spatial range.
[0042] In one implementation, after performing the above encoding and decoding operations, this application provides an efficient storage method for data object indexes that combine encoding and decoding: First, to save storage space and computing performance, if the grid code is completely contained within the space range of the query input, then no further segmentation is performed, and it is preserved as a whole.
[0043] Second, when calculating the grid range, the effective boundaries of the constraints are used to eliminate location data that is not of interest. For details, see [link to relevant documentation]. Figure 4As shown, if the geographic object completely contains grid A, meaning grid A is within the coverage area of the geographic object, then no further subdivision is needed. If the grid is exactly on the boundary, with part of the grid within the coverage area of the geographic object and another part outside, then a judgment is made based on the subdivision level to determine whether the highest subdivision level has been reached (this is specified according to the actual application; setting it too high will lead to wasted storage and computing performance, while setting it too low will result in grid subdivision not meeting the actual application requirements; the highest subdivision level can be set according to different applications). If the level of grid B (B and C are at the same level, with B being one level higher than A) is the highest subdivision level, even if part of grid B is outside the coverage area of the geographic object, no further subdivision is needed, because this level is sufficient to meet the application's needs according to the user's judgment.
[0044] In addition to setting a threshold based on the highest subdivision level, you can also set a threshold constraint on the number of subdivisible grids, and set a minimum and maximum number of subdivision grids to limit the grid into which a geographic object is subdivided to a controllable range. This is crucial for planning the storage space and computing resources required for the operating environment in advance.
[0045] Whether grid A is further subdivided determines the number of grids created. In a quadtree algorithm (two-dimensional space), subdivision causes the number of grids to grow exponentially by 4. In an octree algorithm (three-dimensional space), subdivision causes the number of grids to grow exponentially by 8. Controlling the amount of data during indexing can greatly improve indexing efficiency, storage efficiency, and save computing resources.
[0046] The grid encoding used in the embodiments of this application may include three-dimensional grids and two-dimensional grids. The three-dimensional grid is compatible with the two-dimensional grid. The encoding of the ground surface of the two-dimensional grid is the same as that of the three-dimensional grid. The three-dimensional grid also includes a height domain dimension relative to the two-dimensional grid.
[0047] In one implementation, this application provides a processing method for unifying two-dimensional and three-dimensional data indexing encoding. Specifically, the two-dimensional grid encoding rule follows a quadtree encoding method, consisting of an uppercase letter G followed by the numbers 0, 1, 2, and 3. The three-dimensional grid encoding rule follows an octree encoding method, consisting of an uppercase letter G followed by the numbers 0, 1, 2, 3, 4, 5, 6, and 7. Under the same grid partitioning level, taking partitioning level 1 as an example, such as... Figure 5 As shown, according to the two-dimensional coding partitioning rules, the Earth can be divided into G0, G1, G2, and G3. The specific partitioning is as follows: G0 represents the longitude space of (-256°, 0°) and the latitude space of (0°, 256°); G1 represents the longitude space of (0°, 256°) and the latitude space of (0°, 256°); G2 represents the longitude space of (-256°, 0°) and the latitude space of (0°, -256°); G3 represents the longitude space of (0°, 256°) and the latitude space of (0°, -256°).
[0048] Considering that in the widely used geographic coordinate systems or projected coordinate systems on Earth, longitude is divided into -180° to 180° and latitude into -90° to 90°, longitude and latitude need to be constrained by the actual usage environment. The range represented by each code must also be realistically valid. Therefore, in the two-dimensional coding partitioning rules, G0 represents the longitude space of (-180°, 0°) and the latitude space of (0°, 90°). In the three-dimensional grid partitioning rules, taking G0 and G4 as examples, in addition to latitude and longitude, the concept of altitude is added. The partitioning is as follows: G0 represents the longitude space of (-256°, 0°), the latitude space of (0°, 256°), and the altitude space of (0°, 256°). G4 represents the longitude space (-256°, 0°), the latitude space (0°, 256°), and the altitude space (0°, -256°).
[0049] While the above method of grid partitioning introduces the concept of height in 3D grid partitioning, the encoding of the first layer above the ground is actually the same as that in 2D space; the first layer above the ground will not be represented by numbers 4, 5, 6, and 7. Considering that the Earth's surface is the focus of most applications in actual use and production, this method can, to a certain extent, achieve compatibility with free transformation of multi-dimensional scenes when expanding dimensions.
[0050] Furthermore, in order to improve retrieval efficiency, this application embodiment also provides a design and implementation method for an efficient index structure. First, most applications of spatiotemporal cubes query data object information through spatiotemporal attributes. When constructing the index, the data object ID is used as the primary key, which can be better integrated with the application. In subsequent applications, it is more convenient to aggregate and count the number of statistical data objects (the same as the number of index records), and it also avoids the problem of duplicate data storage in data objects (selection of index primary key).
[0051] Secondly, the grid levels can be divided according to the application, which not only improves query efficiency but also enhances the scalability of the index structure, making it suitable for applications with different data volumes (index structure scalability design).
[0052] In one implementation, all the codes can be stored together, with the data object ID used as the index ID. As the application grows, the amount of data also increases, with each object corresponding to one index record.
[0053] Then, when segmenting, you can segment according to different needs, or segment at each level, or segment using objects and codes.
[0054] When segmenting according to different needs, the code can be segmented. Depending on the application scenario, the code can be specifically segmented for querying. Although this requires more storage space, the query efficiency is relatively fast.
[0055] When segmenting at each level, each level can be segmented into one segment, which is similar to query scenarios aimed at a single level, where data objects are queried at a fixed level.
[0056] When segmenting by object and code, it is suitable for queries with fixed objects and codes.
[0057] In practical applications, the above segmentation method can be adapted to meet actual business needs.
[0058] Third, the index is stored in array format to realize a one-to-many relationship between data objects and their codes (query design).
[0059] Fourth, aggregated information is stored in the cache to improve query efficiency. See [link to relevant documentation]. Figure 6 As shown.
[0060] Furthermore, the above-mentioned association of index encoding with geographic grids to construct the target index structure may include the following steps A to D in its implementation: Step A: Calculate the number of associated grids at different levels according to the data range, based on the specified order of the geographic grids; Step B: If the number of associated grids meets the preset range, perform grid association indexing between the geographic grid and index code corresponding to the current level; Step C: If the number of associated grids does not meet the preset range, the current level is expanded, and the geographic grids and index codes corresponding to the expanded level are associated. Step D: After all the index codes are associated with the geographic grid, construct the target index structure.
[0061] To address the issue of how to index and associate global grid partitioning rules with geographic entities, and which specific grid level to use for association, this application proposes a solution that limits the range of grids associated with geographic entities. This ensures that each type of geographic entity can be associated with an appropriate grid level, thereby maximizing grid accuracy while maintaining query efficiency for each type of geographic entity.
[0062] The specific implementation method is to calculate the number of grid associations for geographic entities at different levels according to the data range, from small to large. If the number falls within the set range, the grid association index is only performed at this level. When querying and counting a specified grid, it is only necessary to set the grid code start and inclusion conditions according to the current grid code.
[0063] Figures 7(a) and 7(b) show schematic diagrams of the association between two geographic entities and grids, respectively. Figure 7(a) shows the association between a geographic entity and a lower precision grid, and Figure 7(b) shows the association between a geographic entity and a higher precision grid.
[0064] Furthermore, embodiments of this application provide an indexing system for diverse heterogeneous data, utilizing... Figure 8 The grid partitioning tool shown can draw a range or perform a rough preview of grid partitioning according to administrative divisions, thus allowing you to set different grid partitioning levels for different data.
[0065] See the 3D display effect of mesh partitioning. Figure 9 As shown, the amount of data contained in a grid is represented by color. You can also index and query according to multiple dimensions such as spatial range (administrative division or self-drawn), data type, grid scale (grid level), data collection time, height range limit, grid number, and keywords. You can also query the data information contained in a specific grid by clicking on it.
[0066] Furthermore, the target retrieval structure described above is a spatiotemporal cube. Figure 10 This demonstrates a spatiotemporal cube construction effect, where the amount of data contained in each grid can also be displayed numerically, providing a more intuitive view of the amount of available data at each location and laying the foundation for further data utilization.
[0067] In summary, this application embodiment, by constructing an index structure for a spatiotemporal cube, enables rapid querying, retrieval, extraction, and service assurance of diverse and heterogeneous business data across disciplines and formats based on a unified spatiotemporal cube.
[0068] Based on the above method embodiments, this application also provides a spatiotemporal retrieval device based on multi-source data fusion, see [link to relevant documentation]. Figure 11 As shown, the device mainly includes the following parts: The spatiotemporal extraction module 112 is used to acquire multivariate heterogeneous data and extract the spatiotemporal range information of the multivariate heterogeneous data; The index boundary determination module 114 is used to determine the effective index boundary based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid encoding; The index structure construction module 116 performs index encoding processing based on a preset grid level within the effective index boundary, and associates the index encoding with the geographic grid to construct a target index structure for spatiotemporal retrieval. The target index structure contains one or more data types within different grid ranges at different levels.
[0069] In one feasible implementation, the data source types of the diverse heterogeneous data include at least remote sensing image data, navigation video data, surveying and mapping geographic data, meteorological and oceanographic data, geological data, and nuclear, biological and chemical data; the data structure types of the diverse heterogeneous data include at least document data, video data, image data, vector data, and raster data.
[0070] In one feasible implementation, the index boundary determination module 114 is further configured to: The system retrieves the data to be queried and calculates the corresponding grid code for the data according to the grid level order based on preset coding constraints. The coding constraints include at least the preset segmentation rules, grid code values, grid coding modes, effective coding range, and coding model type. The effective index boundaries are determined based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid encoding.
[0071] In one feasible implementation, the index boundary determination module 114 is further configured to: Determine if there is an intersection between the grid code and the geographic area corresponding to the data to be queried; the grid code is used to represent the range of the corresponding geographic grid at a fixed subdivision level; If so, determine whether the grid level exceeds the preset limit level; If it exceeds the limit, the grid code corresponding to the current level will be determined as the grid code contained within the geographical range corresponding to the data to be queried; If the preset limit level is not exceeded, determine whether the spatial range completely includes the geographic grid range; If they intersect and contain each other, the grid code corresponding to the current level will be changed to the grid code contained within the geographic range corresponding to the data to be queried. If they intersect but do not contain each other, calculate the grid code for the next grid level, until all grid codes for all levels have been calculated.
[0072] In one feasible implementation, the index structure construction module 116 is further configured to: Calculate the number of associated grids at different levels according to the data range, in a specified order for the geographic grids; If the number of associated grids meets the preset range, perform grid association indexing with the corresponding geographic grid and index code at the current level; If the number of associated grids does not meet the preset range, the current level is expanded, and the geographic grids and index codes corresponding to the expanded level are associated. Once all index codes are associated with the geographic grid, the target index structure is constructed.
[0073] In one feasible implementation, the target index structure is stored in an array format, and the data objects and their codes have a one-to-many relationship.
[0074] In one feasible implementation, the data type contained in each grid in the target index structure is represented by color and / or numbers.
[0075] The spatiotemporal retrieval device based on multi-data fusion provided in this application has the same implementation principle and technical effects as the aforementioned method embodiments. For the sake of brevity, any parts of the spatiotemporal retrieval device based on multi-data fusion that are not mentioned in the embodiments can be referred to the corresponding content in the aforementioned spatiotemporal retrieval method embodiments based on multi-data fusion.
[0076] This application also provides an electronic device, such as... Figure 12 The diagram shows the structure of the electronic device 100, which includes a processor 121 and a memory 120. The memory 120 stores computer-executable instructions that can be executed by the processor 121. The processor 121 executes the computer-executable instructions to implement any of the above-mentioned spatiotemporal retrieval methods based on multi-source data fusion.
[0077] exist Figure 12 In the illustrated embodiment, the electronic device further includes a bus 122 and a communication interface 123, wherein the processor 121, the communication interface 123 and the memory 120 are connected via the bus 122.
[0078] The memory 120 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 123 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 122 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 122 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0079] Processor 121 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 121 or by instructions in software form. Processor 121 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in the memory. The processor 121 reads the information in the memory and, in conjunction with its hardware, completes the steps of the spatiotemporal retrieval method based on multi-source data fusion in the aforementioned embodiment.
[0080] This application also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are called and executed by a processor, they cause the processor to implement the above-described spatiotemporal retrieval method based on multi-source data fusion. For specific implementation details, please refer to the foregoing method embodiments, which will not be repeated here.
[0081] The computer program product of the spatiotemporal retrieval method, apparatus, device and medium based on multi-data fusion provided in the embodiments of this application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0082] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of this application.
[0083] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0084] In the description of this application, it should also be noted that, unless otherwise expressly specified and limited, the terms "set up," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A spatiotemporal retrieval method based on multi-source data fusion, characterized in that, The method includes: Acquire diverse heterogeneous data and extract the spatiotemporal range information of the diverse heterogeneous data; the data source types of the diverse heterogeneous data include at least remote sensing image data, navigation video data, surveying and mapping geographic data, meteorological and oceanographic data, geological data, and nuclear, biological and chemical data; the data structure types of the diverse heterogeneous data include at least document data, video data, image data, vector data, and raster data. Based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid codes, the effective index boundaries are determined. Specifically, this includes: acquiring the data to be queried; determining whether there is an intersection between the grid code and the geographical range corresponding to the data to be queried; if so, determining whether the grid level exceeds a preset limit; if it does, determining the grid code corresponding to the current level as the grid code contained within the geographical range corresponding to the data to be queried; if it does not exceed the preset limit, determining whether the spatial range completely includes the geographical grid range; if they intersect and contain each other, determining the grid code corresponding to the current level as the grid code contained within the geographical range corresponding to the data to be queried; if they intersect but do not contain each other, calculating the grid code of the next grid level, until all grid codes at all levels have been calculated; and determining the effective index boundaries based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid codes. Within the effective index boundary, index encoding is performed based on a preset grid level. The geographic grids are then linked at different levels according to a specified order and data range. This includes: calculating the number of linked grids for geographic entities from smallest to largest at different levels according to data range; if the number of linked grids meets a preset range, the geographic grid at the current level is linked to the index code; if the number of linked grids does not meet the preset range, the current level is expanded, and the geographic grid at the expanded level is linked to the index code. After all index codes are linked to geographic grids, a target index structure is constructed for spatiotemporal retrieval. The target index structure contains one or more data types within different grid ranges at different levels. The encoding storage format of the target index structure is an array format, with a one-to-many relationship between data objects and codes. The data type contained in each grid in the target index structure is represented by color and / or numbers.
2. A spatiotemporal retrieval device based on multi-source data fusion, characterized in that, The device includes: The spatiotemporal extraction module is used to acquire multi-source heterogeneous data and extract the spatiotemporal range information of the multi-source heterogeneous data; the data source types of the multi-source heterogeneous data include at least remote sensing image data, navigation video data, surveying and mapping geographic data, meteorological and oceanographic data, geological data, and nuclear, biological and chemical data; the data structure types of the multi-source heterogeneous data include at least document data, video data, image data, vector data, and raster data. The index boundary determination module is used to determine the effective index boundary based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid codes. Specifically, it includes: acquiring the data to be queried; determining whether there is an intersection between the grid code and the geographical range corresponding to the data to be queried; if so, determining whether the grid level exceeds a preset limit; if it does, determining the grid code corresponding to the current level as the grid code contained within the geographical range corresponding to the data to be queried; if it does not exceed the preset limit, determining whether the spatial range completely encompasses the geographical grid range; if they intersect and contain each other, determining the grid code corresponding to the current level as the grid code contained within the geographical range corresponding to the data to be queried; if they intersect but do not contain each other, calculating the grid code for the next grid level, until all grid codes for all levels have been calculated; and determining the effective index boundary based on the spatiotemporal range information of the multivariate heterogeneous data and the spatial positional relationship between the grid codes. The index structure construction module performs index encoding processing based on a preset grid level within the effective index boundary, and calculates the number of associated grids at different levels according to the data range, in a specified order. This includes: calculating the number of associated grids at different levels according to the data range, from smallest to largest, for geographic entities; if the number of associated grids meets the preset range, then the geographic grids at the current level are associated with the index codes; if the number of associated grids does not meet the preset range, then the current level is expanded, and the geographic grids at the expanded level are associated with the index codes; after all index codes are associated with geographic grids, a target index structure is constructed for spatiotemporal retrieval; the target index structure contains one or more data types within different grid ranges at different levels, and the encoding storage format of the target index structure is an array format, with a one-to-many relationship between data objects and codes; the data type contained in each grid in the target index structure is represented by color and / or numbers.
3. An electronic device, characterized in that, It includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the spatiotemporal retrieval method based on multi-source data fusion as described in claim 1.
4. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the spatiotemporal retrieval method based on multi-source data fusion as described in claim 1.
Citation Information
Patent Citations
Space-time coding method, space-time index and query method and device
CN109992636A