Query optimization method based on edge calculation

By analyzing the SPARQL query patterns in an edge computing environment, determining the induced subgraphs of frequent patterns, and building a resource allocation model, the problem of high SPARQL query latency when storing RDF graphs in cloud storage is solved, and more efficient query processing and resource utilization is achieved.

CN120162367APending Publication Date: 2025-06-17NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510233182.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art faces performance challenges when using cloud storage RDF graphs and supporting SPARQL queries, especially because queries need to be submitted to the cloud via the Internet, resulting in high latency and reducing the efficiency of SPARQL query processing.

Method used

Using edge computing-based query optimization method, the induction subgraph of frequent patterns is determined by analyzing the pattern of workloads, and storing it into the edge server of the edge network, and a resource allocation model and a minimum query cost model are built to optimize the allocation and processing of SPARQL queries.

Benefits of technology

It effectively reduces the latency of SPARQL query, improves the efficiency of SPARQL query processing, reduces data transmission requirements, reduces the computing pressure of cloud servers, and effectively utilizes the computing and storage resources of edge servers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162367A_ABST
    Figure CN120162367A_ABST
Patent Text Reader

Abstract

The invention provides an edge computing-based query optimization method, which comprises the following steps of: acquiring a resource description framework (RDF) data graph and a workload, and analyzing a mode of the workload to obtain a mode set; determining a frequent pattern in the pattern set, obtaining an inducer sub-graph corresponding to the frequent pattern from the RDF data graph, and storing the inducer sub-graph in an edge server of the edge network; constructing a resource allocation model according to the query information corresponding to each query and the resource information of the edge network, and obtaining an optimal network resource allocation scheme corresponding to the workload according to the resource allocation model; obtaining an optimal query unloading scheme corresponding to the workload according to the optimal network resource allocation scheme and a pre-constructed minimum query cost model; and querying according to the optimal network resource allocation scheme and the optimal query unloading scheme to obtain a query result of the workload. According to the method, the SPARQL query delay can be reduced, and the SPARQL query processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and particularly relates to a query optimization method based on edge computing. Background Art

[0002] The Resource Description Framework (RDF) is a standard for describing and representing information about Web resources in a structured format. RDF uses triples to represent resource attributes or relationships between resources, and each triple consists of a subject, a predicate, and an object element. A set of RDF triples is typically regarded as a graph, where the subjects and objects act as vertices and the triples act as edges. In contrast, SPARQL is a query language tailored for RDF. When using SPARQL, users can extract triple patterns from an RDF dataset. Notably, a SPARQL query can also be regarded as a query graph. In this regard, answering a SPARQL query in an RDF dataset corresponds to identifying subgraph matches in the RDF graph according to the query graph.

[0003] With the surging use of RDF graphs in many applications, it is a current trend to develop cloud-based RDF data management solutions by using the cloud to store RDF graphs and support SPARQL queries. To overcome latency and bandwidth challenges, integrating edge computing technology to transfer RDF graph data storage and processing to the edge environment can significantly improve SPARQL query performance. For example, many public SPARQL endpoints are provided on the Internet, such as the DBpedia SPARQL endpoint. In addition, Amazon Web Services (AWS) has also launched Neptune, a cloud-based graph database service that supports RDF graphs and SPARQL queries. However, storing RDF graphs in the cloud to support SPARQL queries usually faces performance challenges because queries need to be submitted to the cloud through the Internet, which introduces significant latency due to its inherent latency and reduces the efficiency of SPARQL query processing. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a query optimization method based on edge computing to reduce SPARQL query latency and improve SPARQL query processing efficiency.

[0005] In a first aspect, the present invention provides a query optimization method based on edge computing, and the method includes the following steps:

[0006] Obtain a Resource Description Framework (RDF) data graph and a workload, and parse the patterns of the workload to obtain a pattern set; wherein, the workload includes at least one query, the pattern set includes at least one pattern, and at least one query corresponds to at least one pattern one by one;

[0007] Determine the frequent patterns in the pattern set, obtain the induced subgraphs corresponding to the frequent patterns from the RDF data graph, and store the induced subgraphs in the edge servers of the edge network;

[0008] Construct a resource allocation model according to the query information corresponding to each query and the resource information of the edge network, and obtain the optimal network resource allocation scheme corresponding to the workload according to the resource allocation model; wherein, the network resource allocation scheme represents the computing resource allocation of each edge server in the edge network to each user terminal;

[0009] Obtain the optimal query offloading scheme corresponding to the workload according to the optimal network resource allocation scheme and the pre-constructed minimum query cost model; wherein, the query offloading scheme represents the query allocation of each edge server;

[0010] Perform queries according to the optimal network resource allocation scheme and the optimal query offloading scheme to obtain the query results of the workload.

[0011] Optionally, the query information includes the total number of CPU cycles required for each query and the query result size corresponding to each query;

[0012] The resource information includes the computing resources allocated by the edge server to the user terminal, the bandwidth allocated by the edge server to the user terminal, the transmission power from the edge server to the user terminal, and the channel gain.

[0013] Optionally, the expression of the resource allocation model is as follows:

[0014]

[0015] s.t.C2:0≤f n,k ≤F k

[0016] C3:

[0017] wherein, O total represents the total execution cost of all queries, n represents the nth user terminal, N represents the total number of user terminals, k represents the kth edge server, K represents the total number of edge servers, D n,k represents the decision vector for the query Q n submitted by the nth user terminal to be offloaded to the kth edge server, e n,k represents that the query submitted by the nth user terminal is processed in the kth edge server, c n represents Q n 's query information, f n,k represents the resource information allocated by the kth edge server to the nth user terminal, w n represents the query result size of the query submitted by the nth user terminal, rn,k denotes the transmission rate between the nth user terminal and the kth edge server, where B represents the bandwidth, and tp i is the transmission rate, and h i,o is the channel gain, and σ 2 is the background noise, and r n,c denotes the transmission rate between the nth user terminal and the cloud server, and F k denotes the computing resources of the kth edge server, and F n ={f n,1 ,..., f n,k}, and f n,k denotes the allocated computing resource vector provided by the kth edge server to the nth user terminal.

[0018] Optionally, the expression of the optimal network resource allocation scheme is:

[0019]

[0020] where denotes the optimal network resource allocation scheme.

[0021] Optionally, the expression of the minimum query cost model is:

[0022]

[0023] s.t.C1:D n,k ∈[0,1]

[0024] C2:

[0025] where denotes the query offloading decision vector, and D n,k denotes the decision to offload the query submitted by the nth user terminal to the kth edge server, and D n,k takes the value of 1 indicating offloading, and D n,k takes the value of 0 indicating no offloading.

[0026] Optionally, the query optimization method further includes performing transaction processing on the cloud server and the edge server respectively.

[0027] Optionally, the query optimization method further includes implementing serializable isolation of transactions by adopting a two-phase locking strategy.

[0028] In a second aspect, the present invention provides a query optimization system based on edge computing, including:

[0029] A data acquisition module, configured to acquire a Resource Description Framework (RDF) data graph and a workload, and parse the schema of the workload to obtain a schema set; wherein, the workload includes at least one query, the schema set includes at least one schema, and the at least one query corresponds one-to-one to the at least one schema;

[0030] An induced subgraph module, configured to determine frequent schemas in the schema set, obtain induced subgraphs corresponding to the frequent schemas from the RDF data graph, and store the induced subgraphs in edge servers of the edge network;

[0031] A resource allocation module, configured to construct a resource allocation model according to the query information corresponding to each query and the resource information of the edge network, and obtain an optimal network resource allocation scheme corresponding to the workload according to the resource allocation model; wherein, the network resource allocation scheme characterizes the computing resource allocation of each edge server in the edge network to each user terminal;

[0032] A query offloading module, configured to obtain an optimal query offloading scheme corresponding to the workload according to the optimal network resource allocation scheme and a pre-constructed minimum query cost model; wherein, the query offloading scheme characterizes the query allocation of each edge server;

[0033] A query module, configured to perform queries according to the optimal network resource allocation scheme and the optimal query offloading scheme to obtain query results of the workload.

[0034] In a third aspect, the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the above method is implemented.

[0035] In a fourth aspect, the present invention provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0036] The beneficial effects of the present invention are:

[0037] The query optimization method based on edge computing provided by the present invention can effectively distribute the storage and query processing of graph data to edge servers by parsing the patterns of workloads to obtain a pattern set, determining the frequent patterns in the pattern set, obtaining the induced subgraphs corresponding to the frequent patterns from the RDF data graph, and storing the induced subgraphs in the edge servers of the edge network, thereby realizing localized queries. This can not only reduce the demand for data transmission, reduce the computing pressure on cloud servers, improve query performance, but also effectively utilize the computing and storage resources of edge servers; the constructed resource allocation model and minimum query cost model can optimize the allocation of SPARQL queries, reduce query costs, optimize network resource configuration, and allocate SPARQL queries to appropriate edge servers or cloud servers, thereby improving overall performance, reducing SPARQL query latency, and improving SPARQL query processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a structural diagram of an edge computing network in one embodiment of the present application;

[0039] Figure 2 It is a flowchart of a query optimization method based on edge computing in one embodiment of the present application;

[0040] Figure 3 It is a Resource Description Framework (RDF) data graph in one embodiment of the present application;

[0041] Figure 4 It is a schematic diagram of a frequent pattern in one embodiment of the present application;

[0042] Figure 5a It is an induced subgraph corresponding to a frequent pattern in one embodiment of the present application;

[0043] Figure 5b It is another induced subgraph corresponding to a frequent pattern in one embodiment of the present application;

[0044] Figure 6 It is a schematic diagram of query optimization in one embodiment of the present application;

[0045] Figure 7 It is a schematic diagram of a decision tree constructed by a backtracking algorithm in one embodiment of the present application;

[0046] Figure 8 It is a structural diagram of a query optimization system in one embodiment of the present application;

[0047] Figure 9 It is a structural diagram of a terminal device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Aiming at the problems of high latency and low query processing efficiency in traditional query methods, the present invention discloses a query optimization method based on edge computing. This method parses the patterns of the workload to obtain a pattern set, determines the frequent patterns in the pattern set, obtains the induced subgraphs corresponding to the frequent patterns from the RDF data graph, and stores the induced subgraphs in the edge servers of the edge network, which can effectively distribute the graph data storage and query processing to the edge servers, thereby realizing localized queries. This can not only reduce the need for data transmission, reduce the computing pressure on the cloud server, improve query performance, but also effectively utilize the computing and storage resources of the edge servers; the constructed resource allocation model and minimum query cost model can optimize the allocation of SPARQL queries, reduce the query cost, optimize the network resource configuration, and allocate SPARQL queries to appropriate edge servers or cloud servers, thereby improving the overall performance, reducing the SPARQL query latency, and improving the SPARQL query processing efficiency.

[0049] For ease of understanding, edge computing is first described.

[0050] It should be understood that edge computing is a hierarchical network topology. For example, in an embodiment of the present invention, as Figure 1 shown, the edge computing network has 2 edge servers (Edge Sever1, Edge Sever2) and 5 user terminals (EU1, EU2, EU3, EU4, EU5). Multiple user terminals can be associated with an edge server at the same time. In the edge computing network, the cloud server (Cloud) owns the entire RDF graph G, while the edge server k contains a subgraph (G1, G1) of G. In most existing use cases, the EU is an Internet of Things device that generates SPARQL queries (Query) and stores them in the edge server. The questions submitted by the user terminal EU will be converted into SPARQL queries according to some templates, and each edge server stores the subgraphs derived from these templates. When the user terminal submits a query, some edge servers can process the query, the cloud server can process the query, and the user terminal cannot process the query.

[0051] Based on the above generalization of the edge computing network, the query optimization method based on edge computing provided by the present invention is described below.

[0052] As Figure 2 shown, the query optimization method based on edge computing provided by the present invention includes the following steps:

[0053] Step 21, obtain the Resource Description Framework (RDF) data graph and the workload, and parse the patterns of the workload to obtain a pattern set.

[0054] In an embodiment of the present invention, the Resource Description Framework (RDF) data graph is asFigure 3 as shown

[0055] In an embodiment of the present invention, the above workload includes at least one query, the pattern set includes at least one pattern, and at least one query corresponds one-to-one with at least one pattern. Among them, a pattern refers to replacing the subject in the form of a constant and the object in the form of a constant in the query with variables. Exemplarily, in a certain scenario, the workload includes <Student A plays basketball>, <Student B plays table tennis>, <Student C plays badminton>. In this scenario, by parsing the workload, the constant subject (Student A, Student B, Student C) can be replaced with the variable "a certain student", and the constant object (basketball, table tennis, badminton) can be replaced with the variable "ball", and the patterns of the obtained workload are P1: , P2: , P3: . Each query in the workload can be parsed into a pattern, and the patterns corresponding to different queries can be the same or different, depending on the query content.

[0056] Step 22: Determine the frequent patterns in the pattern set, obtain the induced subgraphs corresponding to the frequent patterns from the RDF data graph, and store the induced subgraphs in the edge servers of the edge network.

[0057] After step 21, the patterns corresponding to each query in the workload can be obtained. In an embodiment of the present invention, by counting the number of each pattern, when the number of any pattern is greater than a preset threshold (for example, 3), it is determined as a frequent pattern, indicating that the query corresponding to this pattern is relatively many and is suitable for local queries to reduce the network communication pressure caused by multiple repeated queries. As Figure 3 shown in the RDF data graph case, the frequent patterns are determined as Q1, Q2, and Q3, specifically as Figure 4 shown.

[0058] From the RDF data graph as Figure 3 shown, the induced subgraphs corresponding to the frequent patterns are respectively as Figure 5a , Figure 5b shown, and they can be stored in the edge servers respectively to ensure that the results of the three frequent queries can be matched in a certain edge server.

[0059] Step 23: Construct a resource allocation model according to the query information corresponding to each query and the resource information of the edge network, and obtain the optimal network resource allocation scheme corresponding to the workload according to the resource allocation model.

[0060] Among them, the network resource allocation scheme characterizes the computing resource allocation of each edge server to each user terminal in the edge network. In an embodiment of the present invention, the above query information includes the total number of CPU cycles required for each query and the query result size corresponding to each query; the resource information includes the computing resources allocated by the edge server to the user terminal, the bandwidth allocated by the edge server to the user terminal, the transmission power from the edge server to the user terminal, and the channel gain.

[0061] It should be noted that each user terminal EU processes only one SPARQL query (denoted as Q n ) at a time. It is atomic and cannot be further divided. The characteristics of the query are two-parameter tuples, namely J n =(c n , w n ). Among them, c n represents the amount of computation to complete Q n , that is, the total number of CPU cycles required to process Q n , and w n represents the query result size of Q n .

[0062] In the edge computing network, the computing power of each edge server is limited. For edge server k, F k represents all computing resources, expressed in CPU cycles per second. The computing power of the cloud server is infinite, and the computing delay can be ignored compared with the transmission delay. Let F n ={f n,1 ,…,f n,k} represent the K-dimensional computing resource (in units of central processing unit (CPU) cycles per second) vector allocated by edge server k to EU n (so 0≤f n,k ≤F k ).

[0063] Edge servers can communicate with cloud servers but not with each other. Therefore, the queries submitted by user terminals can be executed by edge servers or assigned to cloud servers for execution, and the results are forwarded by the edge servers that submitted the queries to the user terminals.

[0064] Assume that the transmission rate of the Qn result from the cloud to the EUs is fixed, denoted by r n,c . Without loss of generality, assume that the MEC architecture of the present invention is based on orthogonal frequency division multiple access (OFDMA) technology, which means that user terminals do not interfere with each other when transmitting data, and each user terminal EU is allocated the same bandwidth B. Denote the transmission power and channel gain of EU as tp i , h i,o. The transmission rate of the result of Qn from the edge server to the EU is expressed as:

[0065]

[0066] where σ 2 is the background noise.

[0067] For ease of understanding, in the embodiments of the present invention, the cost of the SPARQL query is divided into two parts. One part of the cost is the cost of executing the SPARQL query on the cloud server or on the edge server, which is respectively expressed as

[0068]

[0069] and The other part of the cost is the cost for each user to request the query result, which specifically includes query upload, query execution, and query download.

[0070] The expression of the resource allocation model constructed by the present invention is as follows:

[0071]

[0072] s.t.C2:0≤f n,k ≤F k

[0073] C3:

[0074] where, O total represents the total execution cost of all queries, n represents the nth user terminal, N represents the total number of user terminals, k represents the kth edge server, K represents the total number of edge servers, D n,k represents the decision vector for the query Q n submitted by the nth user terminal to be offloaded to the kth edge server, e n,k represents that the query submitted by the nth user terminal is processed in the kth edge server, c n represents the query information of Q n , f n,k represents the resource information allocated by the kth edge server to the nth user terminal, w n represents the size of the query result submitted by the nth user terminal, r n,k represents the transmission rate between the nth user terminal and the kth edge server, r n,c represents the transmission rate between the nth user terminal and the cloud server, F k represents the computing resource of the kth edge server, F n ={f n,1 ,...,f n,k}, fn,k Denote the allocated computing resource vector provided by the k-th edge server to the n-th user terminal

[0075] Solving the computing resource allocation is a convex optimization problem, which can be solved by the Karush-Kuhn-Tucker (KKT) conditions. The expression of the obtained optimal network resource allocation scheme is as follows:

[0076]

[0077] where Denote the optimal network resource allocation scheme

[0078] Step 24: According to the optimal network resource allocation scheme and the pre-constructed minimum query cost model, obtain the optimal query offloading scheme corresponding to the workload

[0079] where the query offloading scheme characterizes the query allocation situation of each edge server

[0080] Specifically, the expression of the minimum query cost model is as follows:

[0081]

[0082] s.t.C1:D n,k ∈[0,1]

[0083] C2:

[0084] where Denote the query allocation decision vector. D n,k Denote the decision to offload the query submitted by the n-th user terminal to the k-th edge server. D n,k Taking the value of 1 means offloading, and D n,k Taking the value of 0 means not offloading

[0085] The solution of the above minimum query cost model involves solving a non - linear programming problem. Exemplarily, the query offloading vector D can be solved based on the backtracking method: The backtracking algorithm systematically explores the decision tree, tries all query offloading strategies and calculates the cost to identify the optimal cost - benefit offloading strategy for query tasks in edge computing. It achieves this goal by carefully considering the functions of the server and a series of constraints. During the process of searching the decision tree, it evaluates the feasibility of offloading each query task. If a given server proves insufficient for offloading, the algorithm backtracks and selects an alternative exploration path. This iterative process continues until all query tasks are considered, at which point the algorithm records the cost of the selected strategy. In this way, the backtracking algorithm skillfully balances the trade - offs and finds the query offloading strategy with the lowest cost while adhering to the constraints. Based on the backtracking algorithm, methods such as branch - and - bound can also be used for optimization, reducing some unnecessary decision tree searches through pruning to improve the algorithm execution efficiency.

[0086] Step 25, according to the optimal network resource allocation scheme and the optimal query offloading scheme, perform a query to obtain the query result of the workload.

[0087] In an embodiment of the present invention, an example diagram of performing a query according to the optimal network resource allocation scheme and the optimal query offloading scheme is as Figure 6 shown. In this embodiment, the computing resources of edge servers 1 and 2 (ES1 and ES2) are 40 and 100 respectively. Query1 is sent to the cloud server, while Query2 is processed by edge server 1. On the other hand, Query3 and Query5 are processed by edge server 2, and Query4 is sent to the cloud server. Since Query1 and Query4 are sent to the cloud, all values in the first and fourth rows of D n,k and f n,k are 0. Since only one query is sent to edge server 1, all computing resources are allocated to Query2. Therefore, D 2,1 = 1, F 2,1 = 40. At the same time, Query3 and Query5 are processed by edge server 2. Assume that the computing resources scheduled by edge server 2 for Query3 and Query5 are 70 and 30 respectively. Therefore, D 3,2 = 1, F 3,2 = 70, D 5,2 = 1, F 5,2= 30. In the above embodiments, the edge server 1 allocates 40 computing resources to the user terminal EU2 to execute the query Query2. The edge server 2 allocates 70 to EU3 to execute Query3 and 30 to EU5 to execute Query5 respectively. The queries allocated to the edge server resources can be executed on the edge server or on the cloud server. The queries Query1 and Query4 that do not receive computing resources allocated from the edge server can only be executed on the cloud server.

[0088] Furthermore, the backtracking algorithm is used to solve the optimal query offloading strategy, and a decision tree as shown in Figure 7 is established. Since the edge server 1 is not capable of processing the query Query1, the query task in EU1 can only be processed by the cloud server. Therefore, there is only one child node (00) in the second layer of the decision tree. EU2 can be executed by ES1 or the cloud server, so there are two child nodes in the second layer based on the first layer (00 and 10 in the second row respectively). However, the query task in EU3 can be processed by the edge server 1, the edge server 2, or the cloud server. Therefore, there are 3 child nodes in the fourth layer of the decision tree. And so on, the entire decision tree is established and traversed. The cost is calculated for each node of the decision tree, and the node with the minimum cost is taken as the optimal offloading strategy.

[0089] Finally, according to the optimal network resource allocation scheme and the optimal query offloading scheme, the query Q1 requested by user 1 is executed on the cloud, the Q2 requested by user 2 is executed on the edge server 1, the Q3 requested by user 3 is executed on the edge server 2, the Q4 requested by user 4 is executed on the cloud server, and the Q5 requested by user 5 is executed on the edge server 2.

[0090] For the traditional scheme, without using the optimal network resource allocation scheme and the optimal SPARQL query offloading scheme of the present invention, the SPARQL queries will be randomly sent to the terminal or the cloud for execution. Since users 1, 2, and 3 are in the same local area network, there is a situation where queries are simultaneously allocated to the edge server 1 or the cloud, causing congestion. In contrast, for the present invention, the queries submitted by users 1, 2, and 3 are executed on the cloud, the edge server 1, and the edge server 2 respectively. The present invention can allocate SPARQL queries to appropriate edge servers or cloud servers, reduce the SPARQL query latency, improve the SPARQL query processing efficiency, and thus improve the overall performance.

[0091] As can be seen from the above, the query optimization method based on edge computing provided by the present invention can effectively distribute graph data storage and query processing to edge servers by parsing the patterns of workloads, obtaining a pattern set, determining the frequent patterns in the pattern set, obtaining the induced subgraphs corresponding to the frequent patterns from the RDF data graph, and storing the induced subgraphs in the edge servers of the edge network, thereby realizing localized queries. This can not only reduce the need for data transmission, reduce the computing pressure on cloud servers, improve query performance, but also effectively utilize the computing and storage resources of edge servers; the constructed resource allocation model and minimum query cost model can optimize the distribution of SPARQL queries, reduce query costs, optimize network resource configuration, and distribute SPARQL queries to appropriate edge servers or cloud servers, thereby improving overall performance, reducing SPARQL query latency, and improving SPARQL query processing efficiency.

[0092] The query optimization system based on edge computing provided by the present invention will be described below. As Figure 8 shown, the query optimization system 800 includes:

[0093] A data acquisition module 801, configured to acquire a Resource Description Framework (RDF) data graph and a workload, and parse the patterns of the workload to obtain a pattern set; wherein, the workload includes at least one query, the pattern set includes at least one pattern, and at least one query corresponds one-to-one with at least one pattern;

[0094] An induced subgraph module 802, configured to determine the frequent patterns in the pattern set, obtain the induced subgraphs corresponding to the frequent patterns from the RDF data graph, and store the induced subgraphs in the edge servers of the edge network;

[0095] A resource allocation module 803, configured to construct a resource allocation model according to the query information corresponding to each query and the resource information of the edge network, and obtain an optimal network resource allocation scheme corresponding to the workload according to the resource allocation model; wherein, the network resource allocation scheme represents the computing resource allocation situation of each edge server in the edge network to each user terminal;

[0096] A query offloading module 804, configured to obtain an optimal query offloading scheme corresponding to the workload according to the optimal network resource allocation scheme and a pre-constructed minimum query cost model; wherein, the query offloading scheme represents the query allocation situation of each edge server;

[0097] A query module 805, configured to perform queries according to the optimal network resource allocation scheme and the optimal query offloading scheme to obtain the query results of the workload.

[0098] It should be noted that, for the content such as information interaction and execution process between the above modules, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be repeated here. Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be repeated here.

[0099] As Figure 9 shown, an embodiment of the present invention provides a terminal device. As Figure 9 shown, the terminal device D10 in this embodiment includes: at least one processor D100 ( Figure 9 only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, the steps in any of the above method embodiments are implemented.

[0100] Specifically, when the processor D100 executes the computer program D102, it obtains a Resource Description Framework (RDF) data graph and a workload, parses the schema of the workload to obtain a schema set, determines the frequent schemas in the schema set, obtains the induced subgraphs corresponding to the frequent schemas from the RDF data graph, and stores the induced subgraphs in the edge servers of the edge network. It constructs a resource allocation model based on the query information corresponding to each query and the resource information of the edge network, and obtains the optimal network resource allocation scheme corresponding to the workload according to the resource allocation model. It obtains the optimal query offloading scheme corresponding to the workload according to the optimal network resource allocation scheme and the pre-constructed minimum query cost model. It performs queries according to the optimal network resource allocation scheme and the optimal query offloading scheme to obtain the query results of the workload. Among them, by parsing the schema of the workload to obtain a schema set, determining the frequent schemas in the schema set, obtaining the induced subgraphs corresponding to the frequent schemas from the RDF data graph, and storing the induced subgraphs in the edge servers of the edge network, the storage and query processing of graph data can be effectively distributed to the edge servers, thereby realizing localized queries. This can not only reduce the need for data transmission, reduce the computing pressure on the cloud server, improve query performance, but also effectively utilize the computing and storage resources of the edge servers. The constructed resource allocation model and minimum query cost model can optimize the allocation of SPARQL queries, reduce query costs, optimize network resource configuration, and distribute SPARQL queries to appropriate edge servers or cloud servers, thereby improving overall performance, reducing SPARQL query latency, and improving SPARQL query processing efficiency.

[0101] The so-called processor D100 may be a central processing unit (CPU, Central Processing Unit), and this processor D100 may also be other general-purpose processors, digital signal processors (DSP, Digital Signal Processor), application specific integrated circuits (ASIC, Application Specific Integrated Circuit), field-programmable gate arrays (FPGA, Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0102] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In some other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk equipped on the terminal device D10, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory D101 may also be used to temporarily store data that has been output or will be output.

[0103] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor can implement the steps in each of the above method embodiments.

[0104] An embodiment of the present application provides a computer program product, which when running on a terminal device enables the terminal device to execute the steps in each of the above method embodiments.

[0105] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the protection scope of the present application is limited to these examples; under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of the present application as described above, which are not provided in detail for the sake of brevity.

[0106] One or more embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the present application. Therefore, any omission, modification, equivalent substitution, improvement, etc. made within the spirit and principle of one or more embodiments of the present application shall be included in the protection scope of the present application.

Claims

1. A query optimization method based on edge computing, characterized in that: include: Obtain a resource description framework RDF data graph and a workload, and parse a pattern of the workload to obtain a pattern set; wherein the workload includes at least one query, the pattern set includes at least one pattern, and the at least one query corresponds to the at least one pattern in a one-to-one manner; Determine a frequent pattern in the pattern set, obtain an induced subgraph corresponding to the frequent pattern from the RDF data graph, and store the induced subgraph in an edge server of an edge network; A resource allocation model is constructed according to the query information corresponding to each query and the resource information of the edge network, and an optimal network resource allocation scheme corresponding to the workload is obtained according to the resource allocation model; wherein the network resource allocation scheme represents the computing resource allocation of each edge server in the edge network to each user terminal; According to the optimal network resource allocation scheme and the pre-built minimum query cost model, an optimal query offloading scheme corresponding to the workload is obtained; wherein the query offloading scheme represents the query allocation of each edge server; According to the optimal network resource allocation scheme and the optimal query offloading scheme, a query is performed to obtain a query result of the workload.

2. The query optimization method according to claim 1, characterized in that: The query information includes the total number of CPU cycles required for each query and the size of the query result corresponding to each query; The resource information includes computing resources allocated by the edge server to the user terminal, bandwidth allocated by the edge server to the user terminal, transmission power from the edge server to the user terminal, and channel gain.

3. The query optimization method according to claim 2, characterized in that: The resource allocation model is expressed as follows: Among them, O total represents the total execution cost of all queries, n represents the nth user terminal, N represents the total number of user terminals, k represents the kth edge server, K represents the total number of edge servers, and D n,k represents the query Q submitted by the nth user terminal n The decision vector offloaded to the kth edge server, e n,k indicates that the query submitted by the nth user terminal is processed in the kth edge server, c n Indicates Q n Query information, f n,k represents the resource information allocated by the kth edge server to the nth user terminal, w n Indicates the size of the query result submitted by the nth user terminal, r n,k represents the transmission rate between the nth user terminal and the kth edge server, B represents bandwidth, tp i Indicates the transmission rate, h i,o represents the channel gain, σ 2 represents the background noise, r n,c represents the transmission rate between the nth user terminal and the cloud server, F k represents the computing resources of the kth edge server, F n ={f n,1 ,...,f n,k }, f n,k It represents the allocated computing resource vector provided by the k-th edge server to the n-th user terminal.

4. The query optimization method according to claim 3, characterized in that: The expression of the optimal network resource allocation scheme is: in, represents the optimal network resource allocation solution.

5. The query optimization method according to claim 4, characterized in that: The expression of the minimum query cost model is: in, represents the query offloading decision vector, D n,k represents the decision to offload the query submitted by the nth user terminal to the kth edge server, D n,k A value of 1 indicates uninstallation. n,k A value of 0 means not to uninstall.

6. The query optimization method according to claim 1, characterized in that: The query optimization method also includes performing transaction processing on the cloud server and the edge server respectively.

7. The query optimization method according to claim 6, characterized in that: The query optimization method also includes adopting a two-phase lock strategy to achieve serializable isolation of transactions.

8. A query optimization system based on edge computing, characterized in that: include: A data acquisition module, used to acquire a resource description framework RDF data graph and a workload, and parse a pattern of the workload to obtain a pattern set; wherein the workload includes at least one query, the pattern set includes at least one pattern, and the at least one query corresponds to the at least one pattern in a one-to-one manner; An induced subgraph module, used to determine a frequent pattern in the pattern set, obtain an induced subgraph corresponding to the frequent pattern from the RDF data graph, and store the induced subgraph in an edge server of an edge network; A resource allocation module, configured to construct a resource allocation model according to the query information corresponding to each query and the resource information of the edge network, and obtain an optimal network resource allocation scheme corresponding to the workload according to the resource allocation model; wherein the network resource allocation scheme represents the computing resource allocation of each edge server in the edge network to each user terminal; A query offloading module, used to obtain an optimal query offloading scheme corresponding to the workload according to the optimal network resource allocation scheme and a pre-built minimum query cost model; wherein the query offloading scheme represents the query allocation of each edge server; The query module is used to perform a query according to the optimal network resource allocation scheme and the optimal query offloading scheme to obtain a query result of the workload.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.