Ecological space multi-source information association query and traceability system based on graph database
By using a graph database-based ecological space multi-source information association query and tracing system, the problem of low query efficiency between different databases is solved, achieving efficient and stable query tracing and improving the targeting and efficiency of query strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-02
- Publication Date
- 2026-03-27
AI Technical Summary
How to achieve efficient and stable querying of ecological space data across different types of databases, especially when receiving query requests, and addressing the reduced query efficiency caused by cross-database queries and in-depth source tracing analysis.
An ecological space multi-source information association query and tracing system based on graph database is adopted. The system receives query and tracing requests through the input module, determines the cost set associated with the query request through the cost determination module, determines the query and tracing strategy based on the cost set, and performs a query in the target database through the query and tracing module to output the target data.
It improves the query and tracing efficiency of the multi-source information association query and tracing system for ecological space, shortens the query path and time, and enhances the pertinence and efficiency of the query strategy.
Smart Images

Figure CN121743548A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data management of ecological space correlation, in particular to, but not limited to, a multi-source information correlation query and traceability system for ecological space based on a graph database. BACKGROUND
[0002] With the continuous promotion of ecological civilization construction, ecological space data presents the significant characteristics of multi-source heterogeneity, multi-temporal and spatial scale, and multi-management subject. Correspondingly, storing multi-source, multi-type, and diversified ecological space data in different types of databases becomes an effective technical means for ecological space data storage.
[0003] However, when receiving a query request for ecological space data, how to realize efficient and stable query for ecological space data between different types of databases becomes a technical problem to be solved. SUMMARY
[0004] Based on the above technical problems, the embodiment of the present application provides a multi-source information correlation query and traceability system for ecological space based on a graph database, which can stably improve the query and traceability efficiency of the multi-source information correlation query and traceability system for ecological space.
[0005] The technical scheme provided by the embodiment of the present application is as follows: The embodiment of the present application provides a multi-source information correlation query and traceability system for ecological space based on a graph database, comprising: An input module configured to receive a query and traceability request for a target database; wherein the target database comprises the graph database and a relational database; the relational database is configured to store at least part of multi-source information correlated with the ecological space; and the graph database is configured to store graph data corresponding to at least part of the multi-source information; A cost determination module configured to determine a cost set corresponding to the query and traceability request; wherein the cost set is associated with the storage state of data of a target type in the target database; and the target type is associated with the query and traceability request; A strategy determination module configured to determine a query and traceability strategy based on the cost set; A query and traceability module configured to query the data stored in the target database based on the query and traceability strategy, obtain, and output target data associated with the target type.
[0006] The multi-source information correlation query and traceability system for ecological space based on a graph database provided by the embodiment of the present application has at least the following beneficial effects: In the ecological space multi-source information association query and tracing system based on a graph database provided by the embodiments of the present application, an input module receives a query and tracing request for a target database, the target database includes a graph database and a relational database, the relational database is used to store at least part of the multi-source information associated with the ecological space, and the graph database is used to store graph data corresponding to at least part of the multi-source information. In this way, the query and tracing request for the target database can be captured through the input module. Moreover, a cost determination module determines a cost set corresponding to the query and tracing request, the cost set is associated with the storage state of the data of a target type in the target database, and the target type is associated with the query and tracing request. In this way, the association between the cost set and the storage state of the data of the target type corresponding to the query and tracing request in the target database is improved, so that the complexity of querying and tracing the data of the target type in the target database can be represented through the cost set. On this basis, the query and tracing strategy is determined based on the cost set, which can improve the pertinence of the query and tracing strategy. At the same time, in the case that the costs in the cost set are different, the query and tracing strategy is determined based on the cost set, which can improve the efficiency of the query and tracing strategy. On the other hand, a query and tracing module queries the multi-source information stored in the target database based on the query and tracing strategy, obtains and outputs the target data associated with the target type, which can shorten the path and time of querying and tracing the target data in the target database, thereby stably improving the query and tracing efficiency of the ecological space multi-source information association query and tracing system. BRIEF DESCRIPTION OF DRAWINGS
[0007] Figure 1 A structural schematic diagram of the ecological space multi-source information association query and tracing system provided by the embodiments of the present application is shown. Figure 2 A process schematic diagram of the ecological space multi-source information association query and tracing provided by the embodiments of the present application is shown. DETAILED DESCRIPTION
[0008] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application.
[0009] It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0010] With the continuous advancement of ecological civilization construction, ecological space data presents the characteristics of multi-source heterogeneity, multi-temporal and spatial scale, and multi-management subject. The multi-source heterogeneous ecological space data includes structured survey data, semi-structured business data, and unstructured spatial data. The multi-temporal and spatial scale ecological space data includes remote sensing images from meters to kilometers, and monitoring data from hours to years. The multi-management subject includes at least forestry management subject, grassland management subject, wetland management subject, and desert management subject, etc.
[0011] Given the above situation, storing diverse, multi-source, and multi-type ecological spatial data in different types of databases has become an effective technical means for the scientific management of ecological spatial data. However, when receiving queries for ecological spatial data, cross-database queries, complex queries, and in-depth source tracing analysis are involved between different types of databases, thereby reducing the query efficiency of ecological spatial data. Therefore, how to achieve efficient and stable querying of ecological spatial data through different types of databases has become an urgent technical problem to be solved.
[0012] To address the above technical issues, this application provides an ecological space multi-source information association query and tracing system based on graph database. Figure 1 This is a schematic diagram of the structure of the ecological space multi-source information association query and tracing system provided in the embodiments of this application, as shown below. Figure 1 As shown, the ecological space multi-source information association query and tracing system 100 may include: The input module 101 is used to receive query and tracing requests for the target database; wherein the target database includes a graph database and a relational database; the relational database is used to store at least some multi-source information related to the ecological space; and the graph database is used to store graph data corresponding to at least some multi-source information.
[0013] The cost determination module 102 is used to determine the cost set corresponding to the query tracing request; wherein, the cost set is associated with the storage status of the target type data in the target database; the target type is associated with the query tracing request.
[0014] The strategy determination module 103 is used to determine the query tracing strategy based on the cost set.
[0015] The query tracing module 104 is used to query the data stored in the target database based on the query tracing strategy, and obtain and output the target data associated with the target type.
[0016] In some embodiments, multi-source information may include ecological space data of various types, from multiple sources, with multiple data structures, associated with multiple different management entities, associated with multiple entities, and possessing different approval or privacy levels; for example, multiple entities may include type entities, engineering entities, mechanism entities, regulatory entities, administrative entities, and functional entities associated with multi-source information, etc., as shown in Table 1: ; Table 1 Among them, ecological functional zones include water conservation areas and biodiversity protection areas.
[0017] For example, the relationships between data and entities in multi-source information can be shown in Table 2.
[0018] Among them, the Land Administration Law, the Regulations on Ecological Space Management, and the Regulations on Ecological Compensation can be legal entities related to administrative management rules.
[0019] In some embodiments, a relational database may store entities, features, and attributes associated with at least some of the multi-source information in the form of tables, and may also use foreign keys (FK) to represent the dependencies between the multi-source information stored in different tables.
[0020] In some embodiments, the ecological space multi-source information association query and tracing system can perform a first collection operation based on relational statistical information through extended PostgreSQL to obtain a first result and update the data stored in the relational database. Exemplarily, the above collection operation may include incremental operation or full operation. Exemplarily, the relational statistical information may include table cardinality, column histogram, and index selectivity, etc. Exemplarily, the first collection operation may be executed according to a preset first time period, which is not limited in this embodiment.
[0021] ; Table 2 In some embodiments, a graph database can represent the dependencies between different entities in the form of graph data; for example, graph data may include nodes and edges connecting nodes, wherein nodes may represent entities, and edges connecting nodes may represent the strength of the dependencies between entities.
[0022] In some embodiments, the ecological space multi-source information association query and tracing system can perform a second collection operation of graph topology features through a graph sampler, and update the structure of the graph in the graph database in a targeted manner based on the second result obtained from the second collection operation of graph topology features; for example, graph topology features may include node degree distribution, relation type frequency, and average path length, etc.; for example, the second collection operation may be performed periodically with a second time period.
[0023] In some embodiments, at least some of the multi-source information stored in a relational database and the graph data corresponding to at least some of the multi-source information stored in a graph database can be associated with each other; for example, the multi-source information stored in the relational database and the graph data stored in the graph database can have at least one of the following associations: The first graph data associated with the first data stored in the relational database and / or the second graph data associated with the attributes of the first data can be stored in the graph database.
[0024] The additional attributes of graph nodes in the third graph data stored in the graph database can be the second data stored in the relational database.
[0025] The information contained in the data stored in relational databases overlaps with that contained in the graph data stored in graph databases.
[0026] Version numbers enable the interrelationship between graph databases and relational databases. For example, when a graph node in the graph database is updated, a globally incrementing version number can be assigned to the node, and this version number can be tracked and recorded as an attribute of the node. At the same time, the updated graph node and its version number can be synchronously recorded in the relational database. In this way, the consistency between the data stored in the graph database and the relational database can be verified through the graph node and its version number.
[0027] Create a source event graph related to the data change process between a relational database and a graph database. For example, change logs of multi-source information can be obtained through Change Data Capture (CDC), and an Extract-Transform-Load (ETL) process can be performed on the change logs to obtain a third result. Then, the source event graph can be constructed based on the third result. For example, in the source event graph, each change of multi-source information can be modeled as a change event node. The multi-source information associated with the change event node can be updated and stored in the relational database and / or graph database. Furthermore, in the specific query source tracing process, the actual query source tracing request can be transformed into a traversal query of the change event nodes.
[0028] In some embodiments, when the relationship between multiple sources of information in the target database changes, a temporal graph database design pattern can be adopted to retain the historical relationship between multiple sources of information but mark the timeliness of the historical relationship. For example, the historical relationship can be set as an effective time period instead of being deleted, and the historical relationship can be updated to the current relationship.
[0029] In some embodiments, a query tracing request may include a request to query and trace at least one type of multi-source information stored in a target database; for example, the query tracing request may be sent by a user or by a requesting device.
[0030] In some embodiments, the cost set may include a set of time costs and / or computational costs required to query and trace target data from the target database; for example, the time cost may include the length of time required to query the target data; for example, the computational cost may include the amount of data processing resources used to query the target data; for example, the data processing resources may include memory usage, central processing unit (CPU) utilization, and network bandwidth usage; for example, the target data may include multi-source information requested by the query and tracing request from the data stored in the target database.
[0031] In some embodiments, the query sourcing cost in the cost set may include a set of time costs and / or computational costs corresponding to all possible data query paths for obtaining target data; for example, a query path may include the number of tables in a relational database, the number of rows or columns in the tables, and may also include the number of graph nodes in a graph database, the number of edges between different graph nodes, etc.
[0032] In some embodiments, the target type may include the data type of the target data described in the query tracing request; for example, the target type may include the data type of the data carried by the rows and / or columns in a table stored in a relational database, and may also include the attribute type of graph nodes or edges connecting graph nodes in a graph database.
[0033] In some embodiments, the storage status of target type data in the target database may include the storage location of the target type data in the target database; for example, the storage location may be represented by the identifier of a table in a relational database, or by the identifier of graph data in a graph database.
[0034] In some embodiments, the cost set is associated with the storage state and may include a query cost that increases as the storage state becomes more diverse. For example, if the storage state represents data of the target type that is interleaved in relational databases and graph databases, then the query tracing cost in the cost set may be a first cost, and the first cost may be greater than the second cost. The storage state corresponding to the second cost may include: the database of the target type is stored only in relational databases or graph databases.
[0035] In some embodiments, the query sourcing cost in the cost set can be determined in the following ways: If the storage state is a cross-storage state, then the query tracing cost can be determined as the first cost; if the storage state is a single storage state, then the query tracing cost can be determined as the second cost. For example, the single cost may include the storage state representing the target type of the database being stored only in a relational database or a graph database.
[0036] In some embodiments, a query tracing strategy may include steps, stages, conditions, and whether to perform cross-queries based on the query conditions and query requirements contained in the query tracing request; for example, cross-queries may include paths for cross-querying data stored in relational databases and graph databases.
[0037] In some embodiments, the query tracing strategy can be determined in the following ways: The request content contained in the query tracing request is analyzed to determine the data storage range of the target type in the target database. Based on the data correlation between the data stored within the data storage range in the target database, an abstract syntax tree corresponding to the data storage range is constructed. Based on the query tracing cost corresponding to different query paths in the abstract syntax tree in the cost set, the query paths in the abstract syntax tree are sorted, and the query path with the lowest query tracing cost is determined as the query tracing strategy.
[0038] In some embodiments, the query tracing module is used to query data stored in a relational database and / or a graph database based on the query path represented by the query tracing strategy, thereby obtaining and outputting target data.
[0039] As can be seen from the above, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the input module receives query and tracing requests for a target database. The target database includes a graph database and a relational database. The relational database is used to store at least some of the multi-source information associated with the ecological space, and the graph database is used to store the graph data corresponding to at least some of the multi-source information. Thus, the input module can capture query and tracing requests for the target database. Furthermore, the cost determination module determines the cost set corresponding to the query and tracing request, and the cost set is associated with the storage state of the target type data in the target database. The target type is also associated with the query and tracing request. This improves the accuracy of the relationship between the cost set and the target type data corresponding to the query and tracing request. The correlation between the storage states in the target database allows the cost set to characterize the complexity of querying and tracing target type data in the target database. Based on this, determining the query and tracing strategy based on the cost set can improve the targeting of the query and tracing strategy. At the same time, when the costs in the cost set are different, determining the query and tracing strategy based on the cost set can improve the efficiency of the query and tracing strategy. On the other hand, the query and tracing module queries the multi-source information stored in the target database based on the query and tracing strategy, obtains and outputs the target data associated with the target type, which can shorten the path and time of querying and tracing target data in the target database, thereby steadily improving the query and tracing efficiency of the ecological space multi-source information association query and tracing system.
[0040] Based on the foregoing embodiments, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to determine the cost of the k-th query path in the abstract syntax tree and the k-th relation corresponding to the relational database, determine the cost of the k-th query path and the k-th graph corresponding to the graph database, and determine the cost of the k-th query tracing in the cost set based on the cost of the k-th relation and the cost of the k-th graph.
[0041] In this context, the abstract syntax tree is associated with the query tracing request; k is an integer greater than or equal to 1; the k-th query path includes the path from the root node of the abstract syntax tree to the target node where the target data is located.
[0042] In some embodiments, the cost of the k-th relation can characterize the number of nodes in the k-th query path; correspondingly, the cost of the k-th relation can be determined by counting the number of nodes in the k-th path.
[0043] In some embodiments, the cost of the k-th graph may include the number of graph nodes contained in the k-th query path; correspondingly, the cost of the k-th graph may be determined by counting the number of graph nodes contained in the k-th query path.
[0044] In some embodiments, the source tracing cost of the k-th query can be determined in the following ways: The cost of the k-th relation and the cost of the k-th graph are statistically averaged to obtain the first statistical result, which is then determined as the source cost of the k-th query.
[0045] In some embodiments, by traversing all query paths in the abstract syntax tree using the above method, the query origination cost corresponding to all query paths can be obtained. At this time, the set of query origination costs corresponding to all query paths can be determined as the cost set.
[0046] Figure 2 This is a schematic diagram of the ecological space multi-source information association query and tracing process provided in the embodiments of this application, such as... Figure 2 As shown in the diagram 200, the process of multi-source information association query and tracing in ecological space may include the following steps: Execute query sourcing strategy: For example, after determining the query sourcing strategy, the target query path can be determined from the multiple query paths contained in the abstract syntax tree based on the cost set, and the target query path can be analyzed to determine the first query branch for relational databases and the second query branch for graph databases contained therein; wherein, the target query path may include the query path with the minimum query sourcing cost in the cost set.
[0047] For example, after executing the query tracing strategy, the first branch and the second branch can be executed in parallel, or their order can be flexibly adjusted. This application embodiment does not limit this. The first branch includes querying through a relational database and obtaining relational query results, and the second branch includes querying through a graph database and obtaining graph query results.
[0048] Querying a relational database: For example, a query can be performed on a relational database based on the first query branch to obtain relational query results.
[0049] Obtain relational query results.
[0050] Querying via a graph database: For example, the data contained in the graph database can be queried according to the second query branch to obtain graph query results.
[0051] Obtain the graph query results.
[0052] After both the first and second branches have been queried, the following steps can be performed: Result deduplication and sorting: For example, duplicate data in relational query results and graph query results can be deduplicated to obtain deduplicated results. Then, the data in the deduplicated results can be sorted according to the query conditions described in the query source request to obtain the query results.
[0053] Output the query results.
[0054] Through the above process, we can achieve efficient, independent, and parallel queries on data contained in relational databases and graph databases, thereby improving the accuracy and efficiency of query results.
[0055] As can be seen from the above, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the k-th query path in the abstract syntax tree and the k-th relation cost corresponding to the relational database are determined. The k-th query path includes the path from the root node of the abstract syntax tree to the target node where the target data is located. Thus, the cost corresponding to querying the target data in the relational database can be accurately reflected through the k-th relation cost. Furthermore, the k-th query path and the k-th graph cost corresponding to the graph database are determined. Thus, the cost corresponding to querying the target data in the graph database can be accurately reflected through the k-th graph cost. Based on this, the k-th query tracing cost in the cost set is determined based on the k-th relation cost and the k-th graph cost, so that the k-th query tracing cost can be associated with both the k-th relation cost and the k-th graph cost, thereby improving the accuracy and comprehensiveness of the k-th query tracing cost.
[0056] Based on the foregoing embodiments, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to determine the cost of the k-th relationship based on the k-th scan parameter associated with the k-th query path.
[0057] The k-th scan parameters include the number of k-th scan rows and the computational complexity of the k-th predicate.
[0058] In some embodiments, the number of rows scanned at the kth level may include the number of rows scanned in the relational database that are included in the kth query path; for example, the number of rows scanned at the kth level can be determined in the following manner: Based on the query conditions contained in the query tracing request, the number of rows contained in the table containing the node of the k-th query path is counted to obtain the number of rows scanned at the k-th level.
[0059] In some embodiments, the computational complexity of the k-th predicate may include: the k-th cost incurred in querying and tracing a single node or record in the k-th query path; exemplarily, the k-th cost incurred may include: the k-th time cost and / or the k-th computational cost included in querying and tracing a single node or record; accordingly, the computational complexity of the k-th predicate can be obtained in the following ways: Determine the unit cost value of each predicate in the k-th query path, then multiply the unit cost value of the predicate by the number of searches performed on the predicate to determine the computational complexity of the predicate, and sum the computational complexities of all predicates in the k-th query path to determine the computational complexity of the k-th predicate.
[0060] In some embodiments, the cost of the k-th relation can be determined in the following way: The throughput cost of the k-th relation is determined based on the number of rows scanned at k, and the computation cost of the k-th relation is determined based on the computational complexity of the k-th predicate and the number of rows scanned at k. Then, the sum of the throughput cost of the k-th relation and the computation cost of the k-th relation is determined as the cost of the k-th relation; specifically, it can be shown in equations (1) to (3): (1); (2); (3); in, The cost of the k-th relation, Calculate the cost for the k-th relation. Let the throughput cost be the cost of the k-th relation. Let k be the number of the scanned rows. The computational complexity for the k-th predicate is: Disk throughput is used to characterize the amount of data that the disk of the target device configured with the target database can read or write per unit of time.
[0061] As can be seen from the above, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to determine the cost of the k-th relation based on the k-th scan parameters associated with the k-th query path. The k-th scan parameters include the number of k-th scan rows and the computational complexity of the k-th predicate. Thus, by associating the number of k-th scan rows and the computational complexity of the k-th predicate with the cost of the k-th relation, the cost of the k-th relation can accurately and comprehensively reflect the query and tracing costs corresponding to the k-th query path, thereby improving the accuracy and comprehensiveness of the cost of the k-th relation.
[0062] Based on the foregoing embodiments, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to determine the length of the k-th path and the number of hops of the k-th path of the k-th query path, and to determine the cost of the k-th graph based on the length of the k-th path and the number of hops of the k-th path.
[0063] In some embodiments, the number of hops on the k-th path may include the number of edges in the k-th query path that connect different graph nodes in the graph data; correspondingly, the number of hops on the k-th path can be obtained by counting the number of edges in the k-th query path from the root node to the end leaf node; wherein, the end leaf node may be associated with the target data, or the end leaf node may carry the target data.
[0064] In some embodiments, the length of the k-th path can be related to the number of hops in the k-th path and the weights of the edges contained in the number of hops in the k-th path; for example, the weight of an edge can include the strength of the association between the graph nodes it connects to; accordingly, the length of the k-th path can be obtained by counting the number of edges contained in the k-th query path and the weights corresponding to each edge.
[0065] In some embodiments, the cost of the k-th graph can be determined in the following way: The product of the length of the k-th path and the number of hops on the k-th path is determined as the cost of the k-th graph; for example, the cost of the k-th graph... It can be calculated using equation (4): (4); Where n is an integer greater than or equal to 1 and less than or equal to N, where N is greater than 1 and is used to characterize the number of graph nodes contained in the k-th query path. Let be the degree of the nth graph node. Let be the initial weight of the nth edge. Let be the number of hops on the k-th path.
[0066] For example, the initial weight of the edge used to connect graph nodes can be predefined according to the association between graph nodes. For instance, when the association between graph nodes is "belongs to", the weight of the edge can be 1, while when the association between graph nodes is "constrained by", the weight of the edge can be 2.
[0067] As can be seen from the above, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module determines the length and hop count of the k-th query path, and determines the cost of the k-th graph based on the length and hop count of the k-th path. Thus, by associating the cost of the k-th graph with the specific length and hop count of the k-th query path, the accuracy of the cost of the k-th graph can be improved.
[0068] Based on the foregoing embodiments, in the ecological space multi-source information association query and tracing system based on graph database provided in this application, the cost determination module is used to determine the initial graph cost of the kth graph based on the length of the kth path and the number of hops of the kth path contained in the kth query path, obtain the kth correction coefficient associated with the kth query path, and correct the initial graph cost of the kth graph based on the kth correction coefficient to obtain the kth graph cost.
[0069] Among them, the k-th correction coefficient is associated with administrative rules.
[0070] In some embodiments, when the k-th query path is not associated with administrative rules related to multi-source information in the ecological space, the k-th graph cost can be determined based on the length and hop count of the k-th path using the method provided in the foregoing embodiments. Correspondingly, if the graph nodes contained in the k-th query path are related to administrative rules, it is necessary to obtain the k-th correction coefficient associated with the k-th query path and correct the k-th initial cost based on the k-th correction coefficient to obtain the k-th graph cost.
[0071] In some embodiments, the k-th correction coefficient may include at least one correction coefficient for correcting the k-th initial graph cost; specifically, the k-th correction coefficient may be calculated as follows: Obtain the quantized weight coefficients corresponding to all administrative rules associated with the graph nodes and / or edges in the k-th query path, obtain a coefficient set, and perform integrated calculation on the coefficient set to obtain the k-th correction coefficient; for example, the integrated calculation on the coefficient set may include summing the M coefficients contained in the coefficient set to obtain the k-th summation result, and summing the k-th summation result with 1 to determine the k-th correction coefficient; where M is an integer greater than 1.
[0072] It should be noted that the values of the coefficients in the coefficient set are related to specific administrative rules, and may include: The values of the coefficients in the coefficient set can increase as the strictness of the administrative rules increases. For example, if the administrative rule is "strictly prohibited", the corresponding coefficient value can be the first value; if the administrative rule is "restricted development", the corresponding coefficient value can be the second value; if the administrative rule is "guided exit", the corresponding coefficient value can be the third value. Furthermore, the first value can be greater than the second and third values, and the second value can be greater than the third value.
[0073] For example, when there are multiple types and / or numbers of administrative rules, the mapping relationship between administrative rules and coefficient values can be preset. When it is necessary to calculate graph cost, the coefficient values corresponding to the administrative rules associated with the k-th query path can be determined according to the administrative rules associated with the k-th query path and the above mapping relationship, thereby obtaining the coefficient set.
[0074] Specifically, the cost of the k-th graph can be calculated using equations (5) to (6): (5); (6); in, The cost of the k-th initial graph. For the cost of the k-th graph, The m-th coefficient is the one associated with the m-th administrative rule contained in the k-th correction coefficient. Let m be the k-th correction coefficient, where m is an integer greater than or equal to 1 and less than or equal to M.
[0075] As can be seen from the above, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to determine the initial graph cost of the k-th graph based on the length of the k-th path and the number of hops in the k-th query path, obtain the k-th correction coefficient associated with the k-th query path, and correct the initial graph cost of the k-th graph based on the k-th correction coefficient to obtain the graph cost of the k-th graph. Thus, by associating the calculation process of the k-th graph cost with the path length and the number of hops in the k-th query path, the graph structure characteristics of the graph data covered by the k-th query path can be reflected. Furthermore, by associating the calculation process of the k-th graph cost with the k-th correction coefficient, and the k-th correction coefficient with administrative rules, the k-th graph cost can be directly associated with the administrative rules related to multi-source information in the ecological space, thereby improving the constraint and correction degree of administrative rules on the k-th graph cost, and thus improving the accuracy and comprehensiveness of the k-th graph cost.
[0076] Based on the foregoing embodiments, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to determine the cost of relational database and corresponding k-th relation, and the cost of graph database corresponding k-th graph, and to determine the k-th query tracing cost in the cost set based on the k-th relation cost, k-th graph cost, k-th load cost and k-th network cost.
[0077] Among them, the k-th load cost is associated with the usage status of the data processing resources of the target device; the k-th network cost is associated with the network status of the target network; the target network includes at least the network between the target device and the ecological space multi-source information association query and tracing system; the target device is equipped with a relational database and a graph database; the k-th relation cost, the k-th graph cost, the k-th network cost, and the k-th load cost are associated with the k-th query path in the abstract syntax tree; the abstract syntax tree is associated with the query tracing request; k is an integer greater than or equal to 1.
[0078] In some embodiments, the target device may include a physical machine device or a virtual machine device.
[0079] In some embodiments, the target device may include an edge computing device, a local device, or a cloud device; for example, the target device may include a server device or a terminal device.
[0080] In some embodiments, the target network may include a wired network and / or a wireless network.
[0081] In some embodiments, when the number of target devices is at least two, the target network may further include a data transmission network between different target devices; for example, the number of relational databases and graph databases may also be at least two, and at least some relational databases and / or at least some graph databases may be deployed in any target device, that is, the relational databases and graph databases may be distributed among at least two target devices.
[0082] In some embodiments, the k-th load cost may include: the amount of data processing resources and time required by the target device in determining the k-th relation cost and the k-th graph cost; for example, data processing resources may include memory capacity and CPU, etc.; accordingly, the k-th load cost can be obtained in the following ways: Based on the coupling degree between the nodes included in the k-th query path and the data stored in the graph database and relational database respectively, determine the number of searches or indexes for the data in the graph database and relational database, and predict the k-th load cost based on the above search or index counts.
[0083] In some embodiments, the k-th network cost may include the bandwidth of the target network used during data exchange between the target device and the ecological space multi-source information association query and tracing system, between target devices, and between graph databases and relational databases; correspondingly, the k-th network cost can be determined in the following ways: If the relational database and the graph database are deployed on different target devices, the number of searches or indexes for data in the graph database and the relational database included in the k-th query path can be predicted. Based on the above search or index counts and the network status between different target devices, the k-th network cost can be predicted.
[0084] In some embodiments, the source tracing cost of the k-th query can be determined in the following ways: The cost of the k-th relation, the cost of the k-th graph, the cost of the k-th load, and the cost of the k-th network are statistically averaged to obtain the second statistical result, which is then determined as the cost of the k-th query tracing.
[0085] In some embodiments, the network status of the target network may include the remaining available bandwidth, packet loss rate, network latency, and network type of the target network.
[0086] As can be seen from the above, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module determines the k-th query tracing cost based on the k-th relation cost, k-th graph cost, k-th load cost, and k-th network cost. The k-th load cost is associated with the usage status of the data processing resources of the target device, and the k-th network cost is associated with the network status of the target network. This not only improves the comprehensiveness of the k-th query tracing cost but also enables a comprehensive and accurate evaluation of the query tracing cost of the k-th query path from the perspectives of relational databases, graph databases, network dimensions, and target device dimensions. This allows for a precise and intuitive quantification of the actual query tracing cost of the k-th query path.
[0087] Based on the foregoing embodiments, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to process the k-th memory occupancy rate and the k-th CPU utilization rate of the target device based on the load weight set to obtain the k-th load cost.
[0088] Among them, the load weights in the load weight set are associated with the priority of the administrative entity corresponding to the k-th query path.
[0089] In some embodiments, the administrative entity corresponding to the k-th query path may include: a set of administrative units or organizations for approving and / or managing the data corresponding to the nodes included in the k-th query path, and may also include the administrative unit or organization that sends the query tracing request.
[0090] In some embodiments, the load weights in the load weight set and the administrative entities corresponding to the k-th query path may include administrative entities whose load weights increase as their priority level increases.
[0091] For example, the cost of the kth load can be calculated using equation (7): (7); in, For the cost of the k-th load, For the k-th CPU utilization, Let k be the memory usage rate. and This is a set of load weights.
[0092] As can be seen from the above, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to process the k-th memory occupancy rate and k-th CPU utilization rate of the target device based on the load weight set to obtain the k-th load cost. Furthermore, the load weights in the load weight set are associated with the priority of the administrative entity corresponding to the k-th query path. Thus, through the above steps, the k-th load cost is not only associated with the memory occupancy rate and CPU utilization rate of the target device, but also with the priority of the administrative entity corresponding to the k-th query path. This achieves a balance between the priority of the administrative entity and the memory occupancy rate and CPU utilization rate of the target device, and also ensures that query paths corresponding to high-priority administrative entities receive high-priority processing at the target device.
[0093] Based on the foregoing embodiments, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to process the kth data transmission volume and the kth network delay of the target network based on the network weight set to obtain the kth network cost.
[0094] Among them, the network weights in the network weight set are associated with the privacy level of the target data corresponding to the query tracing request.
[0095] In some embodiments, the network weight is associated with the privacy level of the data corresponding to the query tracing request, which may include: the weight in the network weight set corresponding to the k-th network delay increases with the increase of the privacy level; for example, if the privacy level of the target data corresponding to the query tracing request is public, then its corresponding weight value can be a fourth value; if the privacy level of the target data corresponding to the query tracing request is internal, then its corresponding weight value can be a fifth value; if the privacy level of the data corresponding to the query tracing request is classified, then its corresponding weight value can be a sixth value. Here, the public level can be less than the internal level, and the internal level can be less than the classified level. Correspondingly, the fourth value can be less than the fifth and sixth values, and the fifth value can be less than the sixth value.
[0096] In some embodiments, the amount of data transmitted for the kth query path may include the amount of data transmitted in the target network during the construction of the kth query path.
[0097] In some implementation sets, the k-th network latency may include the predicted transmission latency of the target network when constructing the k-th query path.
[0098] For example, the cost of the k-th network It can be calculated using equation (8): (8); in, For the k-th data transmission volume, For the network bandwidth of the target network, For the network delay of the kth digit, 1 and This constitutes the network weight set, where, The value is associated with the privacy level of the data corresponding to the query tracing request.
[0099] As can be seen from the above, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to process the k-th data transmission volume and k-th network latency of the target network based on the network weight set to obtain the k-th network cost. Furthermore, the network weights in the network weight set are associated with the privacy level of the data corresponding to the query tracing request. Thus, the k-th network cost is not only related to the data transmission volume and network latency of the target network, but also to the privacy level of the data requested in the query tracing request. This achieves a balance between the network transmission status of the target network and the privacy level of the target data requested in the query tracing request, thereby improving the comprehensiveness and accuracy of the k-th network cost.
[0100] Based on the foregoing embodiments, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the cost determination module is used to process the k-th relation cost, the k-th graph cost, the k-th load cost and the k-th network cost based on the cost weight set to obtain the k-th query tracing cost.
[0101] The cost weight set includes the relation weight corresponding to the k-th relation cost, the graph weight corresponding to the k-th graph cost, the load weight corresponding to the k-th load cost, and the network weight corresponding to the k-th network cost; the relation weight is greater than the load weight and the network weight; the graph weight is greater than the load weight and the network weight.
[0102] In some embodiments, the cost of tracing the source of the k-th query It can be calculated using equation (9): (9); in, For relation weights, For graph weights, For load weight, This refers to network weights.
[0103] As can be seen from the above, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the k-th relation cost, k-th graph cost, k-th load cost, and k-th network cost are processed based on the weights in the weight set to obtain the k-th query tracing cost. Furthermore, the graph weight is greater than the load weight and network weight, and the relation weight is greater than the load weight and network weight. In this way, the influence of the k-th relation cost and the k-th graph cost in the k-th query tracing cost is increased, thereby reducing the probability of drastic fluctuations in the k-th query tracing cost caused by fluctuations in the target network. This further stabilizes the core position of the data storage status of the target database in the k-th query tracing cost and improves the stability and accuracy of the k-th query tracing cost.
[0104] Based on the foregoing embodiments, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the network weight changes with the fluctuation of the target network, and the load weight changes with the change of network weight.
[0105] In some embodiments, the ecological space multi-source information association query and tracing system can continuously monitor the fluctuation status of the target network and dynamically adjust the network weight and load weight according to the fluctuation status. Specifically, if the network latency of the target network is greater than or equal to the latency threshold, the network weight can be increased proportionally to the latency, and the load weight can be decreased proportionally to the latency, thereby achieving a balance between the network weight and the load weight.
[0106] For example, the latency ratio may include the ratio between the latency difference and the latency threshold; wherein the latency difference may include the difference between the network latency and the latency threshold.
[0107] As can be seen from the above, in the ecological space multi-source information association query and tracing system based on graph database provided in this application embodiment, the network weight changes with the fluctuation of the target network, and the load weight changes with the change of the network weight. In this way, targeted adjustment of the network weight is realized, as well as dynamic adjustment of the load weight and network weight, so that the adjustment of the network weight and load weight can reflect the change process of the target network.
[0108] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0109] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict.
[0110] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0111] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0112] It should be noted that the aforementioned computer-readable storage media can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various electronic devices including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0113] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0114] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0115] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware nodes. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0116] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0117] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0118] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0119] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A system for multi-source information association query and tracing in ecological space based on graph databases, characterized in that, include: An input module is used to receive query and tracing requests for a target database; wherein the target database includes the graph database and a relational database; the relational database is used to store at least some of the multi-source information associated with the ecological space; and the graph database is used to store graph data corresponding to at least some of the multi-source information. A cost determination module is used to determine a cost set corresponding to the query tracing request; wherein, the cost set is associated with the storage status of data of the target type in the target database; and the target type is associated with the query tracing request. The strategy determination module is used to determine the query tracing strategy based on the cost set; The query and tracing module is used to query the data stored in the target database based on the query and tracing strategy, and to obtain and output the target data associated with the target type.
2. The system according to claim 1, characterized in that, The cost determination module is used to determine the cost of the k-th query path in the abstract syntax tree and the k-th relation corresponding to the relational database, to determine the cost of the k-th query path and the k-th graph corresponding to the graph database, and to determine the cost of the k-th query tracing in the cost set based on the cost of the k-th relation and the cost of the k-th graph; wherein the abstract syntax tree is associated with the query tracing request; k is an integer greater than or equal to 1; the k-th query path includes the k-th path from the root node of the abstract syntax tree to the target node where the target data is located.
3. The system according to claim 2, characterized in that, The cost determination module is used to determine the cost of the k-th relation based on the k-th scan parameters associated with the k-th query path; wherein the k-th scan parameters include the number of k-th scan rows and the computational complexity of the k-th predicate.
4. The system according to claim 2, characterized in that, The cost determination module is used to determine the length of the k-th path and the number of hops of the k-th path in the k-th query path, and to determine the cost of the k-th graph based on the length of the k-th path and the number of hops of the k-th path.
5. The system according to claim 2, characterized in that, The cost determination module is used to determine the initial cost of the kth graph based on the length of the kth path and the number of hops of the kth path contained in the kth query path, obtain the kth correction coefficient associated with the kth query path, and correct the initial cost of the kth graph based on the kth correction coefficient to obtain the cost of the kth graph; wherein the kth correction coefficient is associated with administrative rules.
6. The system according to claim 1, characterized in that, The cost determination module is used to determine the k-th query tracing cost in the cost set based on the k-th relation cost corresponding to the relational database and the k-th graph cost corresponding to the graph database, and based on the k-th relation cost, the k-th graph cost, the k-th load cost, and the k-th network cost; wherein, the k-th load cost is associated with the usage status of the data processing resources of the target device; the k-th network cost is associated with the network status of the target network; the target network includes at least the network between the target device and the ecological space multi-source information association query and tracing system; the target device deploys the relational database and the graph database; the k-th relation cost, the k-th graph cost, the k-th network cost, and the k-th load cost are associated with the k-th query path in the abstract syntax tree; the abstract syntax tree is associated with the query tracing request; k is an integer greater than or equal to 1.
7. The system according to claim 6, characterized in that, The cost determination module is used to process the k-th memory occupancy rate and the k-th CPU utilization rate of the target device based on the load weight set to obtain the k-th load cost; wherein, the load weights in the load weight set are associated with the priority of the administrative entity corresponding to the k-th query path.
8. The system according to claim 6, characterized in that, The cost determination module is used to process the kth data transmission volume and the kth network delay of the target network based on the network weight set to obtain the kth network cost; wherein, the network weights in the network weight set are associated with the privacy level of the target data corresponding to the query tracing request.
9. The system according to claim 6, characterized in that, The cost determination module is used to process the k-th relation cost, the k-th graph cost, the k-th load cost, and the k-th network cost based on a cost weight set to obtain the k-th query origination cost; wherein, the cost weight set includes the relation weight corresponding to the k-th relation cost, the graph weight corresponding to the k-th graph cost, the load weight corresponding to the k-th load cost, and the network weight corresponding to the k-th network cost, wherein the relation weight is greater than the load weight and the network weight, and the graph weight is greater than the load weight and the network weight.
10. The system according to claim 9, characterized in that, The network weights change with fluctuations in the target network; the load weights change with changes in the network weights.
Citation Information
Patent Citations
Query request processing method and data selection model training method and device
CN120523829A
Graph database and relational database mapping
US20200201909A1