Data structured processing method and system based on structured query tree
Through dynamic path optimization technology based on structured query trees, query paths and execution order are adjusted in real time, the problems of low query efficiency and high system load in large-scale data environments are solved, and the efficiency and compatibility of cross-system data processing are achieved, and the query performance and scalability of distributed systems are improved.
Patent Information
- Application Number
- CN202510741520.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing technology lacks dynamic optimization capabilities in large-scale data environments and cannot adjust the query path and execution order in real time, resulting in low query efficiency, unintelligent resource scheduling, insufficient cross-system data processing capabilities, affecting query performance and system load.
Through dynamic query path optimization technology based on structured query tree, the execution performance and load of query nodes are evaluated in real time, combined with intelligent scheduling algorithms, the query path and execution sequence are dynamically adjusted, node configuration and data transmission path are optimized, and cross-system data format conversion and compatibility are achieved.
It significantly improves query efficiency, reduces system load, ensures the efficiency and compatibility of query paths among different systems, and improves the scalability and query response speed of large-scale distributed systems.
Smart Images

Figure CN120256436A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a data structuring processing method and system based on a structured query tree. Background Art
[0002] With the rapid development of big data technology, cross-system data processing and query optimization have become key challenges in distributed computing systems. Especially in a large-scale data environment, how to efficiently process and query massive data, improve query response speed, and reduce system load has become an urgent problem to be solved in the field of data processing. Traditional query optimization methods usually rely on static query paths and fixed query strategies, lacking dynamic adaptability and targeted optimization. When facing constantly changing data loads, network topologies, and query tasks, these methods show the following obvious deficiencies:
[0003] 1. Lack of dynamic optimization ability: Existing query optimization technologies are mostly based on static models and cannot dynamically adjust query paths and execution orders in real time according to factors such as node loads and transmission delays, resulting in low query efficiency, especially more significant in complex distributed environments.
[0004] 2. Poor query path optimization effect: Traditional methods usually determine query paths before query execution, making it difficult to handle bottleneck nodes or resource competition during query execution and unable to adjust query paths in a timely manner, affecting query performance.
[0005] 3. Unintelligent resource scheduling: Existing query optimization technologies fail to fully consider the dependency relationships and load conditions between nodes and lack an intelligent scheduling mechanism, resulting in sub-optimal resource allocation and data transmission paths and increasing the system burden.
[0006] 4. Insufficient cross-system data processing ability: In cross-platform or cross-system query scenarios, existing technologies usually cannot effectively solve data format conversion and data flow problems between different systems, affecting data interaction and query processing between different platforms.
[0007] Therefore, how to provide a data structuring processing method and system based on a structured query tree is an urgent problem for those skilled in the art to solve. Summary of the Invention
[0008] An object of the present invention is to propose a data structuring method and system based on a structured query tree. By real-time evaluating multi-dimensional information such as the execution performance, load, and data traffic of query nodes, and combining an intelligent scheduling algorithm, the present invention dynamically adjusts the query path and execution order, and at the same time optimizes the node configuration and data transmission path, which can effectively improve the query efficiency, reduce the system load, solve the problem of cross-system data format conversion, and has the advantages of high-efficient data processing capabilities, intelligent query optimization, and cross-system compatibility.
[0009] The data structuring method based on a structured query tree according to an embodiment of the present invention includes the following steps:
[0010] S1. Based on the structured features and metadata description information of the input data, automatically generate a preliminary query tree framework, and adaptively adjust the topological structure, node configuration, and query path of the query tree according to the storage mode, distribution characteristics, and access mode of the data source;
[0011] S2. Introduce a query tree dynamic optimization algorithm based on data dimension and query load analysis, and use the access frequency, storage location, and processing requirement information of the data source to real-time adjust the connection method and query path between query tree nodes;
[0012] S3. Combine the adaptive conversion mechanism of data format and data structure, and dynamically design and adjust the data format conversion rules by analyzing the format features and hierarchical relationships of nodes to achieve the parsing and formatting processing of nodes in the query tree under different task requirements;
[0013] S4. During the execution process of the query tree, real-time monitor and analyze the execution performance of the query path, and through the dynamic evaluation of data flow and query load, intelligently adjust the query path and execution strategy;
[0014] S5. Adopt a cross-level data interaction mechanism based on the hierarchical structure of the query tree, automatically identify and manage the data transmission requirements between different systems, and achieve cross-platform data transmission and collaborative processing through intelligent data mapping and format conversion strategies;
[0015] S6. Use the structured information extraction in the query tree and the intelligent reasoning mechanism based on the Bayesian inference algorithm to automatically identify the internal associations between data, infer the conditional dependence relationships between multi-dimensional data by constructing a probability model, and dynamically adjust the data processing order and execution strategy of relevant nodes in the query tree according to the inference results;
[0016] S7. Based on the dynamic query path optimization technology of the query tree, through the intelligent scheduling and optimization of the query execution path, automatically identify and adjust the execution path of the query tree, and achieve query performance in a large-scale distributed system.
[0017] Optionally, the S1 specifically includes:
[0018] S11, extract the structural features of the input data, analyze the hierarchical relationship, field type, and inter-field dependency information of the data, construct preliminary nodes, and assign a unique identifier to each node;
[0019] S12, extracting metadata description information of input data, obtaining storage mode, distribution characteristics and access mode of data source, including physical storage location, data update frequency, data access frequency and corresponding query load distribution of data source, and associating with nodes;
[0020] S13. Based on the structured features and metadata description information of the input data, a preliminary query tree framework is constructed. The topological structure of the preliminary query tree is determined according to the relevance and storage location between the nodes, and the preliminary node configuration and query path of the query tree are generated according to the distribution characteristics of the data source.
[0021] S14. Adjust the preliminary query tree framework and dynamically optimize the topology of the query tree according to the storage mode and access mode of the data source: ;
[0022] in, Representation Node and nodes The query path priority between Is a node The query frequency, Is a node The access delay of Describes the impact of storage mode and access mode on query path optimization;
[0023] S15, by configuring the nodes of the query tree, adjusting the connection mode between the nodes, determining the execution order of the nodes, and making adaptive adjustments based on the access mode and load prediction to achieve flexible configuration of the query path;
[0024] S16. Minimize the total cost of the query path and optimize the order of the query path by the following objective function: ;
[0025] in, Representative Node The query frequency, Is a node To Node The data transmission load value is and Node and nodes The access delay, is the total number of nodes in the query tree;
[0026] S17. Dynamically adjust the path formation rule of the query tree according to the mutual relationship between the metadata description information and the input data to adapt to different data query scenarios;
[0027] S18. Based on the dynamic adjustment, generate an adaptive query tree framework that conforms to the data source characteristics and query requirements, and optimize the topological structure, node configuration, and query path of the query tree.
[0028] Optionally, the S2 specifically includes:
[0029] S21. Construct a multi-dimensional feature model of the query tree nodes based on the access frequency, storage location, and processing requirement information of the data source, including the access frequency, storage location, data size, and data source attributes of the processing requirements of each node;
[0030] S22. Calculate the load value of the data source node according to the multi-dimensional feature model: ;
[0031] Among them, represents the load value of the th node, is the access frequency of node , is the storage location attribute of node , is the data processing requirement of node , is the calculation function of the node load value;
[0032] S23. Based on the load value of the node, apply the dynamic optimization algorithm of the query tree to calculate and adjust the connection weight between each node in the query tree: ;
[0033] Among them, represents the connection weight between node and node , and are the weight adjustment factors, represents the load value of the th node, represents the load value of the th node, is the data transmission requirement between node and node , and are the data processing requirements of node and node ;
[0034] S24. Optimize the query path in real time by adjusting the connection weights between nodes, dynamically adjust the connection mode of nodes in the query tree, and optimize the data transmission delay and the execution efficiency of the query tree;
[0035] S25. Recalculate the query path of the query tree according to the load value and the connection weight, adjust the data transmission path between nodes and the execution order of nodes, and optimize the topological structure of the query path;
[0036] S26. Minimize the total delay and data transmission load of the query path through the following objective function: ;
[0037] where, is the connection weight between node and node , is the data transmission delay between nodes, and are the load values of node and node , is the total number of nodes in the query tree;
[0038] S27. Optimize the execution order and query path of the query tree by adjusting the query path and node configuration in real time, so that the execution delay and calculation load of the query path reach the optimum.
[0039] Optionally, the specific steps of S3 include:
[0040] S31. Extract the format features of each node in the query tree, including data type, field characteristics, hierarchical relationship, and dependency between fields. Build a node formatting model through the hierarchical relationship and correlation between nodes, and assign a unique formatting identifier to each node;
[0041] S32. Establish a data format conversion rule set according to the formatting requirements of nodes. Each rule is dynamically generated according to the structural characteristics, data volume, and task requirements of nodes to meet the processing requirements of different query tasks;
[0042] S33. Define the hierarchical structure of node format conversion. Combine the storage mode, distribution characteristics, and access mode of the data source, analyze and determine the format conversion path between nodes, and automatically select the optimal conversion path to minimize the resource consumption of data format conversion during the execution of the query tree;
[0043] S34. Calculate the format conversion complexity of each node through the following formula: ;
[0044] Among them, represents the format conversion complexity of the th node, is the data volume of the node , is the field hierarchy feature of the node , is the storage mode of the node , is a function representing the format conversion complexity of the node;
[0045] S35. Dynamically optimize the data format conversion rules according to the format conversion complexity of the nodes, adjust the format conversion order and path between the nodes, and optimize the data conversion overhead and execution efficiency of each node during the execution of the query tree;
[0046] S36. By dynamically adjusting the format conversion rules between the nodes, optimize the parsing and formatting processing methods of the nodes in the query tree in real time, automatically select a suitable formatting method according to the task requirements, and process the data formats of the nodes under different query tasks;
[0047] S37. By analyzing the data flow and load distribution during the execution of the query tree, based on the hierarchical relationship and formatting requirements of the nodes, adjust the format conversion strategy in real time, and adjust the data format conversion rules according to different query requirements during the execution.
[0048] Optionally, the specific steps of S4 include:
[0049] S41. During the execution of the query tree, monitor the execution performance of the query path in real time. By monitoring the data flow volume, processing time, access frequency, and latency of each query node, construct a query path performance evaluation model;
[0050] S42. By analyzing the load situation and data flow situation of each node during the execution of the query path, identify the bottleneck nodes in the query path, determine the impact of the bottleneck nodes on the overall query performance, and calculate the performance metrics of the bottleneck nodes: ;
[0051] Among them, represents the bottleneck performance metric of the th node, is the processing time of the node , is the load value of the node , is the data flow frequency of the node , is the bottleneck node performance evaluation function;
[0052] S43. Dynamically adjust the connection mode between nodes in the query path and reconfigure the execution order of the query path based on the performance metrics of the bottleneck nodes, and adjust the query path by calculating the optimization weight of each node: ;
[0053] Among them, represents the optimization weight of node . and are optimization adjustment factors, is the maximum processing time of all nodes in the query tree, is the total number of nodes in the query tree, represents the bottleneck performance metric of the th node;
[0054] S44. Optimize the execution order of the query tree through an intelligent scheduling algorithm according to the optimization weight of the nodes, dynamically adjust the data transmission path and processing order in the query path, and optimize the latency and data transmission load in the query path;
[0055] S45. Based on the node performance data and load information collected in real time during the execution of the query tree, adjust the execution strategy of the query tree adaptively, and adjust the data processing strategy and query path of each query node to optimize the execution efficiency of the query tree: ;
[0056] Among them, is the processing load of node , is the optimization weight of node , and are the processing time and load value of node respectively, is the total number of nodes in the query tree;
[0057] S46. Make the latency, load, and execution order of the query path reach the optimal by adjusting the path and execution strategy of the query tree in real time.
[0058] Optionally, the specific content of S5 includes:
[0059] S51. Based on the hierarchical structure of the query tree, identify the data transmission requirements between different systems by analyzing the hierarchical relationships of the nodes in the query tree, and construct a cross-system data transmission requirements model. The cross-system data transmission requirements model considers the storage modes, access methods, and data types of different systems where the nodes are located, and automatically identifies the data transmission paths and data transmission volumes between different levels;
[0060] S52. Dynamically generate a data mapping strategy according to the cross-system data transmission requirement model, map the formats and structures of nodes at different levels in the query tree, determine the mapping rules between each node, and ensure the consistency of the structure and format of data during cross-system transmission;
[0061] S53. Dynamically establish cross-platform data format conversion rules according to the data transmission path and the mapping relationship of nodes. The format conversion rules combine the data types, data formats and query task requirements of the query tree nodes, adjust the format conversion methods between each node in real time, and automatically adjust the encoding method and storage form of data according to the data processing requirements of the target system;
[0062] S54. Calculate the transmission load of cross-system data transmission requirements through the following formula: ;
[0063] where, represents the transmission load of the cross-system data transmission requirement between node and node , is the data transmission delay between node and node , and are the load values of node and node respectively, and are the storage modes of node and node respectively, is the calculation function of the cross-system transmission load;
[0064] S55. Based on the calculation result of the cross-system data transmission load, adjust the connection mode between query tree nodes and the data transmission path in real time, and optimize the data flow and load distribution in the cross-platform data interaction process in the query tree;
[0065] S56. Through an intelligent data mapping and format conversion strategy, during the execution of the query tree, adjust the cross-system data format conversion rules according to the dynamic changes of the query task, so that data on different platforms can be collaboratively processed in the query tree, and cross-system data interaction and transmission can be realized;
[0066] S57. By analyzing the cross-system data interaction performance data during the execution of the query tree, adjust the data transmission path, format conversion strategy and data mapping rules in real time to optimize the efficiency of cross-platform data transmission.
[0067] Optionally, the specific content of S6 includes:
[0068] S61. Extract the features of each node based on the structured information in the query tree, including data type, hierarchical relationship, field dependency, and access pattern, construct a relationship graph between nodes, and identify the inherent associations existing between nodes;
[0069] S62. Use the Bayesian inference algorithm to construct a conditional probability model based on the structured information between nodes: ;
[0070] where, represents the conditional probability of node when the parent node is given, is the prior probability of node , is the conditional probability of the parent node when node is given, is the prior probability of the parent node ;
[0071] S63. Based on the conditional probability model, infer the dependency relationships between the nodes in the query tree, calculate the conditional dependency probabilities between multi-dimensional nodes, and identify the strength of the correlations under different query tasks to model the dependency relationships between nodes in an automated manner;
[0072] S64. According to the inference results, dynamically adjust the processing order and execution strategy of the relevant nodes in the query tree, and adjust the execution order of the nodes to optimize the query path by calculating the inference weights of the nodes: ;
[0073] where, represents the inference weight of node , and are weight adjustment factors, is the load of node , is the conditional probability between node and its parent node;
[0074] S65. According to the inference weights, use the greedy algorithm to optimize the scheduling of the query tree nodes. By calculating the inference weight and load of each node, sort the nodes by priority and give priority to processing those with an inference weight higher than other nodes and a load lower than other nodes: ;
[0075] where, represents node The scheduling priority is the inference weight of the node, and is the load of the node;
[0076] S66. Adjust the execution order of the nodes in the query path according to the node order optimized by the scheduling algorithm, and optimize the execution strategy of the query tree in real time.
[0077] Optionally, the S7 specifically includes:
[0078] S71. Based on the dynamic query path optimization technology of the query tree, evaluate each node in the query tree and its corresponding data processing requirements, construct the execution performance model of each query node, and generate the preliminary query path topology structure;
[0079] S72. By collecting the execution performance data of each query node in real time, including the load, data transmission delay, processing time and access frequency of the node, use the intelligent scheduling algorithm to dynamically adjust the execution path of the query tree, automatically identify the bottleneck nodes and high-load nodes in the query path, and optimize the data transmission and processing order in the path;
[0080] S73. During the execution path optimization process, calculate the execution efficiency of each query path based on the execution performance, data traffic and network bandwidth of each node, and optimize the topology structure of the query path: ;
[0081] wherein, represents the execution efficiency of the query path , is the data traffic of the query path , is the query path , is the data traffic, and are the load values of the node and the node , is the node , and is the node , is the load value, is the node , and is the node , is the data transmission delay between the nodes;
[0082] S74. According to the calculation result of the execution efficiency, sort the query paths by priority, give priority to processing the one with the largest execution efficiency value compared with other paths, and optimize the priority of the one with the smallest execution efficiency value compared with other paths;
[0083] S75. During the execution of the query tree, continuously monitor the running state of the query path, adjust the node order in the query path in real time, and dynamically adjust the connection method between the nodes in the query path according to the load of the nodes, so as to optimize the data flow and the allocation of computing resources;
[0084] S76. Combining the characteristics of large-scale distributed systems, schedule and optimize the execution of query paths through distributed load balancing technology, and dynamically adjust the node configuration and data transmission path of the query path according to the load conditions of nodes, data traffic, and the network topology of the system.
[0085] Optionally, a data structuring system based on a structured query tree includes the following modules:
[0086] Query tree construction module: Automatically generate a query tree framework based on the input data source and its structuring characteristics, and dynamically adjust the topology of the query tree according to the type of nodes, query load, and system requirements;
[0087] Node analysis module: By introducing a data dependency analysis algorithm, analyze the field dependencies, hierarchical relationships, and data formats of each node in the query tree, construct a relationship graph between nodes, and model its structuring characteristics;
[0088] Data processing module: Automatically construct a data processing flow based on the relationship graph between nodes, and optimize the data processing order according to the characteristics and dependency relationships of each node in the query tree to achieve cross-node data structuring;
[0089] Query path optimization module: Based on the load conditions of nodes and the multi-dimensional requirements of query tasks, collect the execution performance data of query nodes in real time, dynamically adjust the query path through an intelligent scheduling algorithm, automatically identify bottleneck nodes, and optimize the data transmission order and processing order in the path;
[0090] Execution efficiency calculation module: According to the execution efficiency calculation formula of the query path, optimize the topology of the query path through the load of nodes, data transmission delay, and traffic data;
[0091] Query path priority sorting module: Through the greedy algorithm, sort the query paths according to the execution efficiency calculation results of the query paths, and give priority to processing the one with the largest execution efficiency value compared to other paths;
[0092] Data format conversion module: Adjust the conversion rules of data formats and structures according to the characteristics of query tree nodes and cross-platform data processing requirements to achieve conversion between different data formats and cross-system collaborative processing;
[0093] System load balancing module: Combining the characteristics of large-scale distributed systems, monitor the load conditions of nodes, data traffic, and network topology in the query path through load balancing technology, and dynamically adjust the node configuration and data transmission path in the query path.
[0094] The beneficial effects of the present invention are:
[0095] (1) Through the dynamic query path optimization technology based on the structured query tree, combined with real-time evaluation of the execution performance data of query nodes, and using an intelligent scheduling algorithm to dynamically adjust the query path, the present invention solves the problems of low processing efficiency and high response latency of traditional methods in a large-scale data environment. The present invention can real-time identify bottleneck nodes and high-load nodes in the query path, and by optimizing the topological structure and execution order of the query path, greatly improves the execution efficiency of the query tree and the response speed of the system, meeting the query requirements in a high-concurrency and high-load data environment.
[0096] (2) By constructing a cross-system data processing model, the present invention solves the problems of data format differences and mismatched processing capabilities between different systems. By dynamically analyzing the format characteristics, data transmission requirements, and data types of each node in the query tree, designing and adjusting data format conversion rules, the data can be seamlessly transmitted and processed between different systems. The present invention also automatically adjusts the cross-system data transmission path according to the execution requirements of the query tree, optimizes data flow and load distribution, ensuring the efficiency, compatibility, and consistency of the query path between different systems, and greatly reducing the performance bottleneck caused by data format incompatibility or transmission delay.
[0097] (3) The present invention supports the dynamic expansion and flexible adjustment of the query tree structure, can perform intelligent adjustment in real-time according to the performance data of the query execution path, and by continuously monitoring the node order and data flow in the query path, timely adjusts the connection method and processing order between nodes to achieve optimal data transmission and computing resource allocation. The present invention can automatically adjust the query path node configuration according to the load changes and network topology structure of a large-scale distributed system, greatly improving the scalability and flexibility of the query tree in a distributed environment, effectively reducing query latency, and ensuring the efficient operation of the system in a large-scale data processing environment. Description of the Drawings
[0098] The drawings are used to provide further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0099] Figure 1 is the overall framework diagram of the data structuring processing method and system based on the structured query tree proposed by the present invention;
[0100] Figure 2 is the working schematic diagram of the cross-system data processing model proposed by the present invention. Detailed Embodiments
[0101] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0102] Reference Figure 1-2 , a data structuring processing method and system based on a structured query tree, comprising the following steps:
[0103] S1. Based on the structured features and metadata description information of the input data, automatically generate a preliminary query tree framework, and adaptively adjust the topological structure, node configuration, and query path of the query tree according to the storage mode, distribution characteristics, and access mode of the data source;
[0104] S2. Introduce a query tree dynamic optimization algorithm based on data dimension and query load analysis, and use the access frequency, storage location, and processing requirement information of the data source to adjust the connection method and query path between query tree nodes in real time;
[0105] S3. Combine an adaptive conversion mechanism for data format and data structure, and dynamically design and adjust data format conversion rules by analyzing the format features and hierarchical relationships of nodes to achieve the parsing and formatting processing of nodes in the query tree under different task requirements;
[0106] S4. During the execution of the query tree, monitor and analyze the execution performance of the query path in real time, and intelligently adjust the query path and execution strategy through the dynamic evaluation of data flow and query load;
[0107] S5. Adopt a cross-level data interaction mechanism based on the hierarchical structure of the query tree, automatically identify and manage the data transmission requirements between different systems, and achieve cross-platform data transmission and collaborative processing through intelligent data mapping and format conversion strategies;
[0108] S6. Use the structured information extraction in the query tree and an intelligent reasoning mechanism based on the Bayesian inference algorithm to automatically identify the internal relationships between data, infer the conditional dependence relationships between multi-dimensional data by constructing a probability model, and adjust the data processing order and execution strategy of relevant nodes in the query tree dynamically;
[0109] S7. Based on the dynamic query path optimization technology of the query tree, through the intelligent scheduling and optimization of the query execution path, automatically identify and adjust the execution path of the query tree, and achieve query performance in a large-scale distributed system.
[0110] In this embodiment, the specific content of S1 includes:
[0111] S11. Extract the structured features of the input data, analyze the hierarchical relationship, field types, and interdependency information between fields of the data, construct preliminary nodes, and assign unique identifiers to each node;
[0112] S12. Extract the metadata description information of the input data, obtain the storage mode, distribution characteristics, and access mode of the data source, including the physical storage location of the data source, data update frequency, data access frequency, and the corresponding query load distribution, and associate them with the nodes;
[0113] S13. Based on the structured features and metadata description information of the input data, construct a preliminary query tree framework. The topological structure of the preliminary query tree is determined according to the association between nodes and the storage location, and the preliminary node configuration and query path of the query tree are generated through the distribution characteristics of the data source;
[0114] S14. Adjust the preliminary query tree framework, and dynamically optimize the topological structure of the query tree according to the storage mode and access mode of the data source: ;
[0115] Among them, represents the query path priority between node and node , is the query frequency of node , is the access latency of node , and the function describes the impact of the storage mode and access mode on the query path optimization;
[0116] S15. Through the node configuration of the query tree, adjust the connection method between nodes, determine the execution order of nodes, and perform adaptive adjustment according to the access mode and load prediction to achieve flexible configuration of the query path;
[0117] S16. Minimize the total cost of the query path through the following objective function to optimize the order of the query path: ;
[0118] Among them, represents the query frequency of node , is the data transmission load value from node to node , and are the access latencies of nodes and respectively, and is the total number of nodes in the query tree;
[0119] S17. Dynamically adjust the path formation rule of the query tree according to the mutual relationship between the metadata description information and the input data to adapt to different data query scenarios;
[0120] S18. Based on the dynamic adjustment, generate an adaptive query tree framework that conforms to the data source characteristics and query requirements, and optimize the topological structure, node configuration, and query path of the query tree.
[0121] In this embodiment, by constructing an adaptive query tree framework and combining the structured features of the input data with the metadata description information, the dynamic optimization and adaptive adjustment of the query path are realized. By analyzing the hierarchical relationship, field dependence, and access pattern of the data, this method can effectively optimize the topological structure of the query tree, reduce the query cost, improve the query efficiency, and ensure the efficient response to different query requirements and data source characteristics in a large-scale data environment.
[0122] In this embodiment, the specific steps of S2 are as follows:
[0123] S21. Based on the access frequency, storage location, and processing requirement information of the data source, construct a multi-dimensional feature model of the query tree nodes, including the access frequency, storage location, data size, and data source attributes of processing requirements for each node;
[0124] S22. According to the multi-dimensional feature model, calculate the load value of the data source node: ;
[0125] Wherein, represents the load value of the th node, is the access frequency of node , is the storage location attribute of node , is the data processing requirement of node , is the calculation function of the node load value;
[0126] S23. Based on the load value of the node, apply the dynamic optimization algorithm of the query tree to calculate and adjust the connection weight between each node in the query tree: ;
[0127] Wherein, represents the connection weight between node and node , and are the weight adjustment factors, represents the th node load value, represents the The load value of a node, for the node and the node the data transmission requirement between them, and for the node and the node the data processing requirement;
[0128] S24. By adjusting the connection weights between nodes, optimize the query path in real time, dynamically adjust the connection mode of nodes in the query tree, and optimize the data transmission delay and the execution efficiency of the query tree;
[0129] S25. According to the load value and the connection weight, recalculate the query path of the query tree, adjust the data transmission path between nodes and the execution order of nodes, and optimize the topological structure of the query path;
[0130] S26. Minimize the total delay and data transmission load of the query path through the following objective function: ;
[0131] wherein, is the connection weight between the node and the node ; is the data transmission delay between nodes, and are the load values of the node and the node ; is the total number of nodes in the query tree;
[0132] S27. By adjusting the query path and node configuration in real time, optimize the execution order and query path of the query tree, so that the execution delay and calculation load of the query path reach the optimum.
[0133] In this embodiment, by constructing a multi-dimensional feature model of query tree nodes, based on access frequency, storage location and processing requirement information, dynamically optimize the node configuration and query path of the query tree. By accurately calculating the load value of nodes and the connection weight between nodes, the optimization adjustment of the query path is realized, the data transmission delay is reduced, and the query efficiency is improved.
[0134] In this embodiment, the specific steps of S3 are as follows:
[0135] S31. Extract the format features of each node in the query tree, including data type, field characteristics, hierarchical relationship and dependence between fields, construct a node formatting model through the hierarchical relationship and correlation between nodes, and assign a unique formatting identifier to each node;
[0136] S32. Establish a data format conversion rule set according to the formatting requirements of the nodes. Each rule is dynamically generated based on the structural characteristics, data volume, and task requirements of the nodes to handle the processing requirements of different query tasks.
[0137] S33. Define the hierarchical structure of node format conversion. Combine the storage mode, distribution characteristics, and access mode of the data source, analyze and determine the format conversion paths between nodes, and automatically select the optimal conversion path to minimize resource consumption during the execution of the query tree for data format conversion.
[0138] S34. Calculate the format conversion complexity of each node through the following formula: ;
[0139] where, represents the format conversion complexity of the th node, is the data volume of node , is the field-level feature of node , is the storage mode of node , is a function representing the format conversion complexity of the node;
[0140] S35. Dynamically optimize the data format conversion rules based on the format conversion complexity of the nodes, adjust the format conversion order and paths between nodes, and optimize the data conversion overhead and execution efficiency of each node during the execution of the query tree.
[0141] S36. Dynamically adjust the format conversion rules between nodes, optimize the parsing and formatting processing methods of the nodes in the query tree in real time, automatically select a suitable formatting method according to the task requirements, and process the data formats of the nodes under different query tasks.
[0142] S37. By analyzing the data flow and load distribution during the execution of the query tree, based on the hierarchical relationship and formatting requirements of the nodes, adjust the format conversion strategy in real time, and adjust the data format conversion rules according to different query requirements during the execution process.
[0143] In this embodiment, by constructing a formatting model of the nodes, for each node in the query tree, extract its format feature information, and dynamically generate data format conversion rules adapted to different query tasks. By analyzing the structural characteristics, data volume, and task requirements of the nodes, optimize the data format conversion paths and order, and minimize resource consumption and improve the execution efficiency of the query tree to the greatest extent.
[0144] In this embodiment, the specific content of S4 includes:
[0145] S41. During the execution of the query tree, monitor the execution performance of the query path in real time. By monitoring the data flow volume, processing time, access frequency, and latency of each query node, construct a query path performance evaluation model;
[0146] S42. By analyzing the load situation and data flow situation of each node during the execution of the query path, identify the bottleneck nodes in the query path, determine the impact of the bottleneck nodes on the overall query performance, and calculate the performance metrics of the bottleneck nodes: ;
[0147] Among them, represents the bottleneck performance metric of the th node, is the processing time of node , is the load value of node , is the data flow frequency of node , is the bottleneck node performance evaluation function;
[0148] S43. Based on the performance metrics of the bottleneck nodes, dynamically adjust the connection method between nodes in the query path, reconfigure the execution order of the query path, and adjust the query path by calculating the optimization weight of each node: ;
[0149] Among them, represents the optimization weight of node , and are the optimization adjustment factors, is the maximum processing time of all nodes in the query tree, is the total number of nodes in the query tree, represents the bottleneck performance metric of the th node;
[0150] S44. According to the optimization weight of the nodes, optimize the query tree execution order through an intelligent scheduling algorithm, dynamically adjust the data transmission path and processing order in the query path, and optimize the latency and data transmission load in the query path;
[0151] S45. Based on the node performance data and load information collected in real time during the execution of the query tree, adjust the execution strategy of the query tree adaptively, adjust the data processing strategy and query path of each query node to optimize the execution efficiency of the query tree: ;
[0152] Among them, is the node processing load is the optimization weight of node ; and are the processing time and load value of node respectively, and
[0153] S46. By adjusting the path and execution strategy of the query tree in real time, the latency, load, and execution order of the query path are optimized to the best.
[0154] In this embodiment, by monitoring the node performance during the execution of the query tree in real time, a performance evaluation model of the query path is constructed, bottleneck nodes are identified and their performance metrics are calculated, and then the execution order of the query tree is dynamically optimized. By adjusting the optimization weights and execution paths of the nodes, the data flow and processing order are optimized, and the latency and data transmission load are reduced.
[0155] In this embodiment, step S5 specifically includes:
[0156] S51. Based on the hierarchical structure of the query tree, by analyzing the hierarchical relationships of the nodes in the query tree, the data transmission requirements between different systems are identified, and a cross-system data transmission requirements model is constructed. The cross-system data transmission requirements model takes into account the storage modes, access methods, and data types of different systems where the nodes are located, and automatically identifies the data transmission paths and data transmission volumes between different levels.
[0157] S52. According to the cross-system data transmission requirements model, a data mapping strategy is dynamically generated to map the formats and structures of the nodes at different levels in the query tree, and the mapping rules between each node are determined to ensure the consistency of the structure and format of the data during cross-system transmission.
[0158] S53. According to the data transmission path and the mapping relationship of the nodes, cross-platform data format conversion rules are dynamically established. The format conversion rules combine the data types, data formats, and query task requirements of the query tree nodes, adjust the format conversion methods between the nodes in real time, and automatically adjust the encoding method and storage form of the data according to the data processing requirements of the target system.
[0159] S54. Calculate the transmission load of the cross-system data transmission requirements through the following formula: ;
[0160] where represents the transmission load of the cross-system data transmission requirements between node and node ; is the optimization weight of node and node The data transmission delay between and are the load values of node and node respectively, and are the storage modes of node and node respectively; is a calculation function for cross-system transmission load;
[0161] S55. Based on the calculation result of the cross-system data transmission load, adjust the connection method between query tree nodes and the data transmission path in real time, and optimize the data flow and load distribution in the cross-platform data interaction process in the query tree;
[0162] S56. Through an intelligent data mapping and format conversion strategy, during the execution of the query tree, adjust the format conversion rules of cross-system data according to the dynamic changes of the query task, so that data on different platforms can be collaboratively processed in the query tree, and cross-system data interaction and transmission can be realized;
[0163] S57. By analyzing the cross-system data interaction performance data during the execution of the query tree, adjust the data transmission path, format conversion strategy and data mapping rules in real time, and optimize the efficiency of cross-platform data transmission.
[0164] In this embodiment, by constructing a cross-system data transmission requirement model, analyzing the hierarchical relationship and data transmission requirements of nodes in the query tree, automatically identifying the data transmission path and data volume between different systems, ensuring the consistency of data during cross-system transmission, combined with an intelligent data mapping and format conversion strategy, it can dynamically adjust the data format, encoding method and storage form, optimize the resource consumption during data transmission, and further optimize the efficiency of cross-system data transmission by real-time analyzing data interaction performance, thereby improving the data interaction and transmission efficiency during the execution of the query tree.
[0165] In this embodiment, the specific steps of S6 are as follows:
[0166] S61. Based on the structured information in the query tree, extract the features of each node, including data type, hierarchical relationship, field dependency and access mode, construct a relationship graph between nodes, and identify the internal associations existing between nodes;
[0167] S62. Use the Bayesian inference algorithm to construct a conditional probability model based on the structured information between nodes: ;
[0168] Among them, represents the probability of node given the parent node The conditional probability is the prior probability of node ; is the conditional probability of the parent nodes given node ; is the prior probability of the parent nodes ; ;
[0169] S63. Based on the conditional probability model, infer the dependencies between the nodes in the query tree, calculate the conditional dependence probabilities between the multi-dimensional nodes, and identify the strength of the correlations under different query tasks, and model the dependencies between the nodes in an automated manner;
[0170] S64. According to the inference results, dynamically adjust the processing order and execution strategy of the relevant nodes in the query tree, and adjust the execution order of the nodes to optimize the query path, which is achieved by calculating the inference weights of the nodes: ;
[0171] wherein represents the inference weight of node ; and are weight adjustment factors; is the load of node ; is the conditional probability between node and its parent nodes;
[0172] S65. According to the inference weights, use the greedy algorithm to optimize the scheduling of the query tree nodes. By calculating the inference weights and loads of each node, sort the nodes by priority, and give priority to processing those with higher inference weights and lower loads than other nodes: ;
[0173] wherein represents the scheduling priority of node ; is the inference weight of the node; is the load of the node;
[0174] S66. According to the node order optimized by the scheduling algorithm, adjust the execution order of the nodes in the query path, and optimize the execution strategy of the query tree in real time.
[0175] This embodiment combines the Bayesian inference algorithm and the greedy algorithm. By dynamically inferring and optimizing the execution order of nodes in the query tree, it automatically identifies the dependency relationships and conditional probabilities between nodes, optimizes the processing order and execution strategy of the query path, and can significantly improve the execution efficiency of the query tree, reduce query latency, optimize system load under different query tasks by efficiently adjusting the scheduling priorities of nodes, and provides an efficient query optimization method.
[0176] In this embodiment, the S7 specifically includes:
[0177] S71. The dynamic query path optimization technology based on the query tree evaluates each node in the query tree and its corresponding data processing requirements, constructs an execution performance model for each query node, and generates a preliminary query path topological structure;
[0178] S72. By collecting the execution performance data of each query node in real time, including the load of the node, data transmission latency, processing time, and access frequency, use an intelligent scheduling algorithm to dynamically adjust the execution path of the query tree, automatically identify the bottleneck nodes and high-load nodes in the query path, and optimize the data transmission and processing order in the path;
[0179] S73. During the execution path optimization process, based on the execution performance, data traffic, and network bandwidth of each node, calculate the execution efficiency of each query path, and optimize the topological structure of the query path: ;
[0180] Among them, represents the execution efficiency of the query path , is the data traffic of the query path , and are the load values of node and node , is the data transmission latency between node and node ;
[0181] S74. According to the calculation results of the execution efficiency, sort the priorities of the query paths, give priority to processing the one with the largest execution efficiency value compared with other paths, and optimize the priority of the one with the smallest execution efficiency value compared with other paths;
[0182] S75. During the execution of the query tree, by continuously monitoring the running state of the query path, adjust the node order in the query path in real time, and dynamically adjust the connection method between nodes in the query path according to the load of the nodes, and optimize the allocation of data flow and computing resources;
[0183] S76. Combining the characteristics of large-scale distributed systems, the execution of query paths is scheduled and optimized through distributed load balancing technology, and the node configuration of query paths and data transmission paths are dynamically adjusted according to the load conditions of nodes, data traffic, and the network topology of the system.
[0184] Through the dynamic query path optimization technology, this embodiment combines the execution performance model and intelligent scheduling algorithm to automatically identify and optimize bottleneck nodes and high-load nodes in the query path, real-time adjust the node order and data transmission path. Based on the distributed load balancing technology, the execution efficiency of the query path is significantly improved, data transmission latency and computing load are reduced, ensuring the efficiency and scalability of query operations in large-scale distributed systems.
[0185] In this embodiment, the data structuring processing system based on the structured query tree includes the following modules:
[0186] Query tree construction module: Automatically generate the query tree framework based on the input data source and its structured characteristics, and dynamically adjust the topology of the query tree according to the type of nodes, query load, and system requirements;
[0187] Node analysis module: By introducing the data dependency analysis algorithm, analyze the field dependencies, hierarchical relationships, and data formats of each node in the query tree, construct the relationship graph between nodes, and model its structured characteristics;
[0188] Data processing module: Automatically construct the data processing flow based on the relationship graph between nodes, and optimize the data processing order according to the characteristics and dependency relationships of each node in the query tree to achieve cross-node data structuring processing;
[0189] Query path optimization module: According to the load conditions of nodes and multi-dimensional requirements of query tasks, collect the execution performance data of query nodes in real-time, dynamically adjust the query path through the intelligent scheduling algorithm, automatically identify bottleneck nodes, and optimize the data transmission order and processing order in the path;
[0190] Execution efficiency calculation module: According to the execution efficiency calculation formula of the query path, optimize the topology of the query path through the load of nodes, data transmission latency, and traffic data;
[0191] Query path priority sorting module: Through the greedy algorithm, sort the priorities of query paths according to the execution efficiency calculation results of query paths, and preferentially process the one with the largest execution efficiency value compared to other paths;
[0192] Data format conversion module: According to the characteristics of query tree nodes and cross-platform data processing requirements, adjust the conversion rules of data formats and structures to achieve conversion between different data formats and cross-system collaborative processing;
[0193] System load balancing module: Combining the characteristics of large-scale distributed systems, it monitors the load conditions, data traffic, and network topology of nodes in the query path through load balancing technology, and dynamically adjusts the node configuration and data transmission path in the query path.
[0194] Example 1:
[0195] To verify the actual application effect of the present invention in query tree optimization and dynamic adjustment of execution paths, the present invention is applied to the query system of a large e-commerce platform. The platform processes millions of query requests per day on average, and the system involves multi-level data storage, complex query conditions, and cross-system data interaction. Against the background of a sharp increase in data volume, there is a significant delay in the query response time of the platform, seriously affecting the user experience and the service quality of the platform. The traditional static optimization method of query paths can no longer adapt to the rapidly changing system load and data requirements. To solve the above problems, the platform decides to introduce the dynamic query path optimization technology provided by the present invention to improve the performance and response efficiency of the query system.
[0196] In the query system of this platform, query requests are processed in the form of a query tree. Each node in the query tree represents a query operation or data processing task, and the nodes are connected through query paths to form complex dependency relationships. Due to the diverse business types of the e-commerce platform, the execution efficiency of the query path directly determines the length of the query response time. However, due to the variety of query requests, complex data types, and frequent load fluctuations, the original query optimization method is difficult to effectively handle. To improve the query efficiency, the platform adopts the dynamic query path optimization technology in the present invention, and dynamically adjusts the execution order of the query tree and the connection method between nodes by real-time monitoring the load conditions, data transmission delay, processing time, and access frequency of query nodes.
[0197] During the implementation process, first in the construction stage of the query tree, the platform analyzes the structured information of each node, including data type, field dependency, and access mode, to generate a preliminary query tree topology structure. Based on this information, the system automatically evaluates the execution performance of each node, including the load, processing time, and data transmission delay of the node. By real-time collecting the execution data of each node, the platform uses an intelligent scheduling algorithm to dynamically adjust the query path. Specifically, the system calculates the execution efficiency of each query path, preferentially processes the paths with higher execution efficiency, and adjusts the node order in the query path according to the real-time execution status to optimize the query response time.
[0198] To verify the effectiveness of this technology, the platform compared the query response time and system load before and after introducing the technology of the present invention. Before implementation, the query requests of the platform often had a response time exceeding 10 seconds during peak hours, resulting in the loss of some users. After introducing the dynamic optimization algorithm of the present invention, the system can dynamically adjust the query path according to the execution efficiency of the nodes, and successfully shorten the query response time to within 5 seconds. In addition, the platform also greatly improved the concurrent processing ability of the system, reduced the load of the query path, and alleviated the pressure on the server by adjusting the node order and load distribution in the query path. The following is a data table comparing the query response time and system load before and after implementation:
[0199] Table 1 Comparison Table of Query Response Time and System Load
[0200] As can be seen from Table 1, after introducing the dynamic query path optimization technology of the present invention, the query response time has been significantly reduced, from the original 10.5 seconds to 4.8 seconds, and the query success rate has also increased from 92% to 98%. The system load has decreased significantly, and the CPU usage rate has dropped from 85% to 60%, indicating that after query optimization, the load distribution of the system is more balanced and the query processing is more efficient.
[0201] To further verify the application effect of the technology of the present invention, the platform compared the query request volume and query performance data before and after implementation. Especially during the peak period of query request volume, the system can dynamically adjust the query path optimization algorithm, real-time identify and adjust the bottleneck nodes and high-load nodes in the query path, thus avoiding the problems of slow response and processing failure caused by the overloading of the query system. The following is the query performance data of the platform during the query peak period:
[0202] Table 2 Comparison of Performance Data During Query Peak Period
[0203] As can be seen from Table 2, during the peak query request volume period, after introducing the technology of the present invention, the query system of the platform has successfully improved the processing ability of query requests per second, from 1500 times to 2200 times, the query response time has been significantly reduced, from the original 12.4 seconds to 5.2 seconds, and the stability of the system has also been greatly improved. The number of response failures or timeouts has dropped from 45 times to only 5 times. The CPU utilization rate of the system has been effectively controlled, dropping from 95% to 72%, indicating that the system can still maintain an efficient and stable query service under high load.
[0204] Through the above embodiments, the platform has successfully solved the query latency and load problems in the query system caused by the increase in data volume and the complexity of query paths. The dynamic query path optimization technology enables the system to efficiently process large-scale query requests and effectively allocate computing resources by real-time adjusting the query path and node order, improving the query response speed and system stability. Finally, after introducing the technology of the present invention, the query system of the platform has significantly improved its processing capacity and query efficiency, and the system runs more efficiently and reliably, greatly improving the user experience.
[0205] In summary, the dynamic query path optimization technology provided by the present invention can significantly improve the performance of large-scale data query systems. While ensuring query efficiency, it effectively reduces the system load, improves the query response speed and stability, and is applicable to large-scale distributed systems such as e-commerce and finance with high requirements for query performance.
[0206] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A data structuring method based on a structured query tree, characterized in that, It includes the following steps: S1. Based on the structured features of the input data and the metadata description information, automatically generate a preliminary query tree framework, and adaptively adjust the topological structure, node configuration, and query path of the query tree; S2. Introduce a query tree dynamic optimization algorithm based on data dimension and query load analysis, and use the access frequency, storage location, and processing requirement information of the data source to adjust the connection method and query path between query tree nodes in real time; S3. Combine the adaptive conversion mechanism of data format and data structure, and dynamically design and adjust the data format conversion rules by analyzing the format features and hierarchical relationships of nodes; S4. During the execution of the query tree, monitor and analyze the execution performance of the query path in real time, and intelligently adjust the query path and execution strategy through the dynamic evaluation of data flow and query load; S5. Adopt a cross-level data interaction mechanism based on the hierarchical structure of the query tree, automatically identify and manage the data transmission requirements between different systems, and adopt an intelligent data mapping and format conversion strategy; S6. Use the structured information extraction in the query tree and the intelligent reasoning mechanism based on the Bayesian inference algorithm to automatically identify the internal associations between data, infer the conditional dependence relationship between multi-dimensional data by constructing a probability model, and dynamically adjust the data processing order and execution strategy of relevant nodes in the query tree; S7. Based on the dynamic query path optimization technology of the query tree, automatically identify and adjust the execution path of the query tree through the intelligent scheduling and optimization of the query execution path.
2. The data structuring method based on a structured query tree according to claim 1, characterized in that The specific content of S1 includes: S11. Extract the structured features of the input data, analyze the hierarchical relationship, field type, and interdependence information between fields of the data, construct preliminary nodes, and assign a unique identifier to each node; S12. Extract the metadata description information of the input data, obtain the storage mode, distribution characteristics, and access mode of the data source, including the physical storage location, data update frequency, data access frequency, and their corresponding query load distribution of the data source, and associate them with the nodes; S13. Based on the structured features of the input data and the metadata description information, construct a preliminary query tree framework. The topological structure of the preliminary query tree is determined according to the relevance and storage location between nodes, and the preliminary node configuration and query path of the query tree are generated through the distribution characteristics of the data source; S14. Adjust the preliminary query tree framework, and dynamically optimize the topological structure of the query tree according to the storage mode and access mode of the data source: ; Among them, represents the query path priority between node and node is the query frequency of node and is the access latency of node . The function describes the impact of the storage mode and access mode on the query path optimization; S15. Through the node configuration of the query tree, adjust the connection method between nodes, determine the execution order of nodes, and perform adaptive adjustment according to the access mode and load prediction to achieve flexible configuration of the query path; S16. Minimize the total cost of the query path through the following objective function to optimize the order of the query path: ; Among them, represents the query frequency of the node , is the data transmission load value from node to node , and are the access latencies of nodes and respectively; is the total number of nodes in the query tree. S17. Dynamically adjust the path formation rule of the query tree according to the mutual relationship between the metadata description information and the input data to adapt to different data query scenarios; S18. Based on the dynamic adjustment, generate an adaptive query tree framework that meets the characteristics of the data source and query requirements, and optimize the topological structure, node configuration, and query path of the query tree.
3. The data structuring method based on a structured query tree according to claim 2, wherein The specific content of S2 includes: S21. Based on the access frequency, storage location, and processing requirement information of the data source, construct a multi-dimensional feature model of the query tree nodes, including the access frequency, storage location, data size, and data source attributes of processing requirements for each node; S22. According to the multi-dimensional feature model, calculate the load value of the data source node: ; Among them, represents the load value of the th node, is the access frequency of node , is the storage location attribute of node , is the data processing requirement of node , is the calculation function of the node load value; S23. Based on the load value of the node, apply the dynamic optimization algorithm of the query tree to calculate and adjust the connection weights between the nodes in the query tree: ; Among them, represents the connection weight between node and node and are weight adjustment factors, represents the load value of the th node, represents the load value of the th node, is the data transmission requirement between node and node and are the data processing requirements for node and node S24. Through the adjustment of the connection weights between the nodes, optimize the query path in real time, dynamically adjust the connection mode of the nodes in the query tree, and optimize the data transmission delay and the execution efficiency of the query tree; S25. According to the load value and the connection weights, recalculate the query path of the query tree, adjust the data transmission path between the nodes and the execution order of the nodes, and optimize the topological structure of the query path; S26. Minimize the total delay and data transmission load of the query path through the following objective function: ; Among them, is the connection weight between and ; is the data transmission delay between nodes, and are the load values of and ; is the total number of nodes in the query tree; S27. By adjusting the query path and node configuration in real time, optimize the execution order and query path of the query tree, so that the execution delay and calculation load of the query path reach the optimum.
4. The data structuring method based on a structured query tree according to claim 1, characterized in that The specific steps of S3 are as follows: S31. Extract the format features of each node in the query tree, including data type, field characteristics, hierarchical relationship, and dependencies between fields. Construct a node formatting model through the hierarchical relationship and relevance between the nodes, and assign a unique formatting identifier to each node; S32. According to the formatting requirements of the nodes, establish a data format conversion rule set, where each rule is dynamically generated according to the structural characteristics, data volume, and task requirements of the nodes to meet the processing requirements of different query tasks; S33. Define the hierarchical structure of node format conversion. Combine the storage mode, distribution characteristics, and access mode of the data source to analyze and determine the format conversion path between the nodes, and automatically select the optimal conversion path to minimize the resource consumption of data format conversion during the execution of the query tree; S34. Calculate the format conversion complexity of each node through the following formula: ; Among them, represents the format conversion complexity of the th node, is the data volume of node , is the field hierarchy feature of node , is the storage mode of node , is a function representing the format conversion complexity of the node; S35. Based on the format conversion complexity of the nodes, dynamically optimize the data format conversion rules, adjust the format conversion order and path between the nodes, and optimize the data conversion overhead and execution efficiency of each node during the execution of the query tree; S36. Through the dynamic adjustment of the format conversion rules between the nodes, optimize the parsing and formatting processing methods of the nodes in the query tree in real time, and automatically select a suitable formatting method according to the task requirements, so that the data format of the nodes can be processed under different query tasks; S37. By analyzing the data flow and load distribution during the execution of the query tree, based on the hierarchical relationship and formatting requirements of the nodes, adjust the format conversion strategy in real time, and adjust the data format conversion rules according to different query requirements during the execution process.
5. The data structuring method based on a structured query tree according to claim 1, characterized in that The specific steps of S4 are as follows: S41. During the execution of the query tree, monitor the execution performance of the query path in real time. By monitoring the data flow volume, processing time, access frequency, and delay of each query node, construct a query path performance evaluation model; S42. By analyzing the load conditions and data flow conditions of each node during the execution of the query path, identify the bottleneck nodes in the query path, determine the impact of the bottleneck nodes on the overall query performance, and calculate the performance metrics of the bottleneck nodes: ; Among them, represents the bottleneck performance index of the th node, is the processing time of node , is the load value of node , is the data flow frequency of node , is the bottleneck node performance evaluation function; S43. Based on the performance metrics of the bottleneck nodes, dynamically adjust the connection methods between the nodes in the query path, reconfigure the execution order of the query path, and adjust the query path by calculating the optimization weights of each node: ; Among them, represents the optimization weight of the node , and is the optimization adjustment factor, is the maximum processing time of all nodes in the query tree, is the total number of nodes in the query tree, represents the bottleneck performance index of the th node; S44. According to the optimization weights of the nodes, optimize the execution order of the query tree through an intelligent scheduling algorithm, dynamically adjust the data transmission path and processing order in the query path, and optimize the latency and data transmission load in the query path; S45. Based on the node performance data and load information collected in real time during the execution of the query tree, adjust the execution strategy of the query tree adaptively, adjust the data processing strategies and query paths of each query node to optimize the execution efficiency of the query tree: ; Among them, is the processing load of the node , is the optimization weight of the node , and are respectively the processing time and load value of the node , is the total number of nodes in the query tree; S46. By adjusting the path and execution strategy of the query tree in real time, make the latency, load, and execution order of the query path reach the optimal state.
6. The data structuring method based on a structured query tree according to claim 1, wherein The specific steps of S5 are as follows: S51. Based on the hierarchical structure of the query tree, by analyzing the hierarchical relationships of the nodes in the query tree, identify the data transmission requirements between different systems, and construct a cross-system data transmission requirements model. The cross-system data transmission requirements model considers the storage modes, access methods, and data types of different systems where the nodes are located, and automatically identifies the data transmission paths and data transmission volumes between different levels; S52. According to the cross-system data transmission requirements model, dynamically generate a data mapping strategy, map the formats and structures of the nodes at different levels in the query tree, determine the mapping rules between each node, and ensure the consistency of the structure and format during cross-system data transmission; S53. According to the data transmission path and the mapping relationship of the nodes, dynamically establish cross-platform data format conversion rules. The format conversion rules combine the data types, data formats, and query task requirements of the query tree nodes, adjust the format conversion methods between each node in real time, and automatically adjust the encoding method and storage form of the data according to the data processing requirements of the target system; S54. Calculate the transmission load of the cross-system data transmission requirements through the following formula: ; Among them, represents the transmission load of the cross-system data transmission requirement between node and node ; is the data transmission delay between node and node ; and are the load values of node and node respectively; and are the storage modes of node and node respectively; is the calculation function of the cross-system transmission load; S55. Based on the calculation results of the cross-system data transmission load, adjust the connection methods between the query tree nodes and the data transmission path in real time, and optimize the data flow and load distribution during cross-platform data interaction in the query tree; S56. Through intelligent data mapping and format conversion strategies, during the execution of the query tree, adjust the format conversion rules of cross-system data according to the dynamic changes of the query tasks, enable the data on different platforms to be collaboratively processed in the query tree, and achieve cross-system data interaction and transmission; S57. By analyzing the cross-system data interaction performance data during the execution of the query tree, adjust the data transmission path, format conversion strategy, and data mapping rules in real time to optimize the efficiency of cross-platform data transmission.
7. The data structuring method based on a structured query tree according to claim 1, characterized in that The specific steps of S6 are as follows: S61. Extract the features of each node based on the structured information in the query tree, including data type, hierarchical relationship, field dependency, and access pattern, construct a relationship graph between the nodes, and identify the inherent associations existing between the nodes; S62. Use the Bayesian inference algorithm to construct a conditional probability model based on the structured information between the nodes: ; Among them, represents the conditional probability of node when a given parent node is considered, is the prior probability of node , is the conditional probability of the parent node when a given node is considered, is the prior probability of the parent node ; S63. Based on the conditional probability model, infer the dependency relationships between the nodes in the query tree, calculate the conditional dependency probabilities between multi-dimensional nodes, and identify the strength of the correlation under different query tasks, and model the dependency relationships between the nodes in an automated manner; S64. According to the inference results, dynamically adjust the processing order and execution strategy of the relevant nodes in the query tree, and adjust the execution order of the nodes to optimize the query path, which is achieved by calculating the inference weights of the nodes: ; Among them, represents the inference weight of the node , and is the weight adjustment factor, is the load of the node , is the conditional probability between the node and its parent node; S65. According to the inference weight, use the greedy algorithm to optimize the scheduling of query tree nodes by calculating the inference weight of each node and the load , sort the nodes by priority, and give priority to processing those with an inference weight higher than other nodes and a load lower than other nodes: ; Among them, represents the scheduling priority of the node , is the inference weight of the node, and is the load of the node; S66. According to the node order optimized by the scheduling algorithm, adjust the execution order of the nodes in the query path, and optimize the execution strategy of the query tree in real time.
8. The data structuring method based on a structured query tree according to claim 1, wherein The specific steps of S7 are as follows: S71. Based on the dynamic query path optimization technology of the query tree, by evaluating each node in the query tree and its corresponding data processing requirements, construct an execution performance model for each query node, and generate a preliminary query path topology structure; S72. By collecting the execution performance data of each query node in real time, including the load of the node, data transmission delay, processing time, and access frequency, use an intelligent scheduling algorithm to dynamically adjust the execution path of the query tree, automatically identify the bottleneck nodes and high-load nodes in the query path, and optimize the data transmission and processing order in the path; S73. During the execution path optimization process, based on the execution performance, data traffic, and network bandwidth of each node, calculate the execution efficiency of each query path, and optimize the topology structure of the query path: ; Among them, represents the execution efficiency of the query path , is the data traffic of the query path , and are the load values of node and node , is the data transmission delay between node and node . S74. According to the calculation results of the execution efficiency, sort the query paths by priority, give priority to processing the path with the largest execution efficiency value compared to other paths, and optimize the priority of the path with the smallest execution efficiency value compared to other paths; S75. During the execution of the query tree, by continuously monitoring the running state of the query path, dynamically adjust the node order in the query path, and dynamically adjust the connection method between the nodes in the query path according to the load of the nodes, so as to optimize the allocation of data flow and computing resources; S76. Combining the characteristics of large-scale distributed systems, schedule and optimize the execution of the query path through distributed load balancing technology, and dynamically adjust the node configuration and data transmission path of the query path according to the load of the nodes, data traffic, and the network topology structure of the system.
9. A data structuring processing system based on a structured query tree, which is applied to the data structuring processing method based on a structured query tree according to any one of claims 1-8, characterized in that, It includes the following modules: Query tree construction module: Automatically generate a query tree framework based on the input data source and its structured features, and dynamically adjust the topology structure of the query tree according to the type of nodes, query load, and system requirements; Node analysis module: By introducing a data dependency analysis algorithm, analyze the field dependencies, hierarchical relationships, and data formats of the nodes in the query tree, construct a relationship graph between the nodes, and model their structured features; Data Processing Module: Automatically construct a data processing flow based on the relationship graph between nodes, and optimize the data processing order according to the characteristics and dependencies of each node in the query tree; Query Path Optimization Module: Based on the load conditions of nodes and the multi-dimensional requirements of query tasks, collect the execution performance data of query nodes in real time, dynamically adjust the query path through an intelligent scheduling algorithm, automatically identify bottleneck nodes, and optimize the data transmission order and processing order in the path; Execution Efficiency Calculation Module: Optimize the topological structure of the query path according to the execution efficiency calculation formula of the query path, through the load of nodes, data transmission delay and traffic data; Query Path Priority Sorting Module: Sort the query paths by priority through the greedy algorithm, according to the execution efficiency calculation results of the query paths, and give priority to processing the path with the largest execution efficiency value compared with other paths; Data Format Conversion Module: Adjust the conversion rules of data format and structure according to the characteristics of query tree nodes and cross-platform data processing requirements; System Load Balancing Module: Combining the characteristics of large-scale distributed systems, monitor the load conditions, data traffic and network topological structure of nodes in the query path through load balancing technology, and dynamically adjust the node configuration and data transmission path in the query path.
Citation Information
Patent Citations
Calculation method based on dynamic tree structured expression
CN118036720A
Routing method and device based on intention-aware path, computer equipment, readable storage medium and program product
CN119788587A
Service proxy method and system based on Dores front-end node
CN119938335A
Optimized queries for file path indexing in a content repository
US20140122499A1
First futamura projection in the context of SQL expression evaluation
US20210064619A1