Performance optimization method and system for cross-data-source paging query
By calculating the data distribution weight and building the optimal data access path, dynamically adjusting the number of parallel query threads and connection pool connections, determining the shard size based on the data correlation coefficient, building a cache access mode map and a minimum cost spanning tree, and using reinforcement learning optimization performance strategy, solving the problems of low efficiency of paging query across data sources, unbalanced load and complex result merging, and achieving efficient and stable query optimization.
Patent Information
- Application Number
- CN202510350316.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The problems of inefficient pagination query across data sources, unbalanced load and complex results merging are especially difficult to effectively optimize the existing technology in large data volume and multi-data source scenarios.
By calculating the data distribution weight, building the optimal data access path, dynamically adjusting the number of parallel query threads and connection pool connections, determining the shard size based on the data correlation coefficient, building a cache access pattern map and a minimum cost spanning tree, and using reinforcement learning to optimize the performance strategy set to achieve efficient query across data sources.
Reduces cross-data source query latency, improves query speed, avoids resource waste, optimizes system performance and stability, and improves system resource utilization.
Smart Images

Figure CN120277102A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing technologies, and particularly to a method and system for optimizing the performance of cross-data source paging queries. Background Art
[0002] Cross-data source paging query is a common requirement in modern data management systems. Existing technologies usually adopt some basic strategies, such as performing independent paging queries on each data source and then merging the results. However, this method has some defects and deficiencies:
[0003] Low query efficiency: When the data volume is large and distributed across multiple data sources, performing queries on each data source separately will result in a large amount of network overhead and data transmission, thus reducing the query efficiency.
[0004] Load balancing problem: Different data sources may have different load capacities. Simple paging query methods are difficult to effectively balance the load, which may cause some data sources to be overloaded while other data sources are not fully utilized.
[0005] Complex result merging: The paging results obtained from multiple data sources need to be merged and sorted, which increases the complexity and time cost of processing, especially when complex data transformation or aggregation operations are required. Summary of the Invention
[0006] Embodiments of the present invention provide a method and system for optimizing the performance of cross-data source paging queries, which can solve the problems in the prior art.
[0007] In the first aspect of the embodiments of the present invention,
[0008] A method for optimizing the performance of cross-data source paging queries is provided, including:
[0009] After receiving a paging query request, extract the query conditions and paging parameters, calculate the data distribution weights based on the query conditions and generate a globally unique query identifier, construct a data source routing table, calculate the optimal data access path according to the data distribution weights and the data source routing table, split the paging query request into sub-query tasks, monitor the load status of the data sources, adjust the number of parallel query threads based on the load status, and allocate the sub-query tasks according to the optimal data access path and the number of parallel query threads;
[0010] Establish local indexes based on the sub-query task allocation results, calculate the index access frequencies and adaptively adjust the index structure; monitor the usage rate of the connection pool and dynamically adjust the number of connections, batch obtain data through the connection pool, generate a query execution plan and associate it with the globally unique query identifier and store it in the query plan cache, and obtain the query results of each data source;
[0011] Perform feature analysis based on the query results, calculate the data correlation coefficient, determine the data shard size based on the data correlation coefficient, construct a feature matrix including data distribution characteristics, numerical density, and cardinality information, and select the optimal merging strategy to merge the data; construct a cache access pattern map based on the feature matrix, calculate the data access association degree, aggregate and store data with an access association degree exceeding a preset association degree threshold based on the access association degree, and construct a minimum-cost spanning tree for data synchronization; establish a mapping relationship between the query load characteristics and the system performance including the data synchronization cost of the minimum-cost spanning tree, optimize the performance policy set through reinforcement learning, and adjust the cache update interval, data access path, and the feature matrix according to the performance policy set to obtain the final query result.
[0012] In an alternative embodiment,
[0013] After receiving a paging query request, extract the query conditions and paging parameters, calculate the data distribution weight based on the query conditions and generate a globally unique query identifier, construct a data source routing table, calculate the optimal data access path according to the data distribution weight and the data source routing table, split the paging query request into sub-query tasks, monitor the load status of the data source, adjust the number of parallel query threads based on the load status, and allocate the sub-query tasks according to the optimal data access path and the number of parallel query threads. The steps include:
[0014] Convert the query conditions into conditional feature vectors, calculate the ratio of the number of records satisfying the query conditions to the total number of records to obtain the conditional selection degree, perform exponential smoothing calculation on the current query frequency and the historical query frequency to obtain the column usage frequency, calculate the exponential decay value based on the time difference between the data update time and the current time to obtain the data freshness, perform weighted normalization on the conditional selection degree, column usage frequency, and data freshness to obtain the data distribution weight, and perform consistent hashing operation on the conditional feature vector to generate a globally unique query identifier;
[0015] Construct a two-layer data source routing table including a data source basic information layer and a weight distribution matrix layer. The data source basic information layer includes data source identifiers, connection parameters, and data volume statistical information, and the weight distribution matrix layer records the data distribution weights of each query condition in different data sources; construct a weighted directed graph with data sources as nodes, construct a path cost function based on node weights, edge weights, and cross-source query penalty terms, and calculate the optimal data access path including the minimum path cost through dynamic programming;
[0016] Calculate the processing time and transmission time of a single data source based on the optimal data access path to obtain the query cost, and split the paging query request into multiple sub-query tasks according to the query cost; collect the processor utilization rate, memory utilization rate, input / output waiting time, query response time, and query queue length to obtain the data source load status, calculate the load score based on the load status, decrease the number of parallel query threads when the load score is greater than the preset load threshold, and increase the number of parallel query threads when the load score is less than the preset load threshold. The increment and decrement step sizes of the number of parallel query threads are determined according to the difference between the load score and the preset load threshold;
[0017] Generate a task execution plan including query condition decomposition, data source selection, parallelism configuration, and result merging strategy based on the optimal data access path and the number of parallel query threads, and allocate the sub-query tasks based on the task execution plan.
[0018] In an alternative embodiment,
[0019] The steps of constructing a weighted directed graph with data sources as nodes and constructing a path cost function based on node weights, edge weights, and cross-source query penalty terms, and calculating the optimal data access path including the minimum path cost through dynamic programming include:
[0020] Construct a weighted directed graph with the data source set, where the weighted directed graph includes a node set and an edge set. Calculate the processing capacity coefficient based on the processor capacity, memory capacity, and disk input / output performance of the data source, and use the product of the processing capacity coefficient and the data distribution weight as the node weight. Calculate the edge weight based on the exponentially weighted moving average of the network latency and the data transmission cost, and the data transmission cost is determined according to the amount of transmitted data and the network bandwidth;
[0021] Construct a path cost function including node weight terms, edge weight terms, cross-source query penalty terms, data consistency cost terms, and concurrent query cost terms. The cross-source query penalty term grows piecewise with the number of data sources. The data consistency cost term is calculated based on the consistency level and the number of data sources. The concurrent query cost term is calculated based on the resource conflict rate and the lock waiting time;
[0022] The path optimization is performed using an improved dynamic programming algorithm with query awareness. The improvements include: extracting features from the query request and performing pattern recognition, calculating the query pattern weight based on the complexity of the query statement, the condition selection degree, the query frequency, and the historical response time, constructing an access pattern probability matrix based on the query historical data, and integrating the query feature information into the state transition process of dynamic programming; designing an early pruning strategy based on query cost estimation, maintaining a historical optimal path cache for filtering suboptimal paths, and dynamically adjusting the pruning threshold; executing the dynamic programming process, initializing the state matrix, assigning zero to the diagonal elements, assigning the corresponding edge weights to the directly connected nodes, and assigning infinity to the unconnected nodes; performing path search using an improved state transition equation, where the improved state transition equation adjusts the node weight by introducing the query pattern weight; when a better path is obtained through an intermediate node, updating the path information and synchronously updating the access pattern probability matrix and the optimal path cache; periodically collecting the path execution status information, calculating the deviation value between the actual path cost and the estimated cost, adjusting the query pattern weight coefficient according to the deviation value, and triggering the path recalculation process, and finally obtaining the optimal data access path with the minimum path cost.
[0023] In an alternative embodiment,
[0024] The steps of performing feature analysis based on the query result, calculating the data correlation coefficient, determining the data shard size based on the data correlation coefficient, and constructing a feature matrix including data distribution characteristics, numerical density, and cardinality information, and selecting the optimal merging strategy to merge the data include:
[0025] Performing feature extraction on the query result, including calculating the statistical features of numerical fields and character fields to generate a data feature vector; calculating the data correlation coefficient based on the data feature vector, including calculating the modified Pearson coefficient for numerical fields, calculating the similarity combination value for character fields, and calculating the data type balance weight coefficient for mixed fields to obtain the mixed correlation coefficient;
[0026] Determining the data shard size according to the mixed correlation coefficient, adopting the benchmark shard size when the mixed correlation coefficient is less than or equal to the first threshold, adjusting the shard size based on the logarithmic function when the mixed correlation coefficient is between the first threshold and the second threshold, and triggering the shard merging operation when the mixed correlation coefficient is greater than the second threshold; performing fragment optimization based on the shard variance, calculating the sum of squared deviations of each shard size from the average shard size, and performing shard rebalancing when the shard variance exceeds the shard balance determination threshold, splitting the shards larger than 1.5 times the average shard size or merging the shards smaller than 0.5 times the average shard size; real-time statistics of the shard quantity distribution and access frequency, and dynamically adjusting the shard size calculation parameters according to the monitoring data;
[0027] Construct a feature matrix that includes data distribution characteristics, numerical density, and cardinality information. The data distribution characteristics include data type identification, distribution type, data skewness, dispersion degree, and outlier ratio. The numerical density includes the amount of information stored per unit, compression ratio, null value ratio, and duplication degree. The cardinality information includes the number of unique values, cardinality estimation value, value range, and frequency distribution. Select a merging strategy based on the feature matrix, and determine the merging method by calculating the weighted combination of memory cost, computational cost, and input / output cost. The memory cost is calculated based on the shard size and buffer size. The computational cost is calculated based on the number of comparison operations and sorting operations. The input / output cost is calculated based on the number of read / write operations. When the skewness of the feature matrix is lower than the data distribution uniformity determination threshold, use multi-way balanced merging. When the feature matrix shows data skew, use segmented merging. When the feature matrix indicates memory constraints, use external merging. Calculate the cost deviation based on the ratio of the actual execution cost to the estimated cost, dynamically adjust the weight coefficients of the cost calculation terms according to the deviation, and update the feature matrix by introducing a learning rate to continuously optimize the merging strategy.
[0028] In an alternative embodiment,
[0029] The steps of constructing a cache access pattern map based on the feature matrix, calculating the data access correlation degree, and aggregating and storing data with an access correlation degree exceeding a preset correlation degree threshold based on the access correlation degree to construct a minimum-cost spanning tree for data synchronization include:
[0030] Construct a cache access pattern map based on the data skewness and dispersion degree in the feature matrix. The cache access pattern map is obtained by calculating the temporal locality strength, access frequency correlation, and spatial locality strength between data items.
[0031] Based on the cache access pattern map, introduce the compression ratio and duplication degree information in the numerical density feature dimension to construct a multi-level data access correlation network. Use breadth-first search to calculate the k-order neighborhood of each data item. The k-order neighborhood is a set of data items whose temporal correlation degree with this data item exceeds the corresponding threshold in a decreasing threshold sequence. Record the access level during the breadth-first search process and add the data items that meet the correlation degree requirements to the corresponding order neighborhood set. Calculate the compression ratio weight of the data item through the sigmoid function. The compression ratio weight is negatively correlated with the difference between the compression ratio of the data item and the average compression ratio. Calculate the duplication degree weight of the data item through the sigmoid function. The duplication degree weight is positively correlated with the difference between the duplication degree of the data item and the average duplication degree. Sum the products of the temporal correlation degree of each data item in the k-order neighborhood and the corresponding compression ratio weight and duplication degree weight, and perform normalization processing to obtain the k-order correlation degree. Aggregate and store the data based on the k-order correlation degree, and store the data items with a correlation degree exceeding the preset correlation degree threshold in adjacent storage spaces.
[0032] Construct a minimum-cost spanning tree for data synchronization. Calculate the data transmission cost between nodes based on the data transmission volume, calculate the compression / decompression cost based on the compression ratio, calculate the dirty data detection cost based on the data difference rate, and use the weighted combination of the data transmission cost, compression / decompression cost, and dirty data detection cost as the synchronization cost between nodes; Use the Prim algorithm to construct a minimum-cost spanning tree, select the node with the highest correlation degree as the starting node, use a priority queue to store candidate edges, and use the weighted adjustment of the synchronization cost of the candidate edge and the correlation degree of the target node as the priority of the edge, and update the priorities of all candidate edges related to the new node when adding a new node; Dynamically optimize the minimum-cost spanning tree based on the actual synchronization performance, adjust the cost parameters by monitoring the synchronization performance, and update the synchronization tree structure according to the locality score.
[0033] In an alternative embodiment,
[0034] The steps of establishing the mapping relationship between the query load characteristics and the system performance including the data synchronization cost of the minimum-cost spanning tree, optimizing the performance policy set through reinforcement learning, and adjusting the cache update interval, data access path, and the feature matrix according to the performance policy set to obtain the final query result include:
[0035] Construct a query load feature vector including dimensions of data access pattern, resource utilization, and cache hit rate. Calculate the temporal locality strength and spatial locality strength included in the data access pattern dimension based on the cache access pattern map. The resource utilization dimension includes CPU utilization, memory utilization, and network bandwidth utilization. The cache hit rate dimension includes the first-level cache hit rate and the second-level cache hit rate;
[0036] Define the weighted combination of the query response time and the data synchronization overhead as the system performance metric, where the data synchronization overhead is calculated based on the total edge weight of the data synchronization minimum-cost spanning tree. Establish a mapping function from the query load feature vector to the system performance metric through a multi-layer perceptron. The input layer of the multi-layer perceptron corresponds to the dimension of the query load feature vector, the hidden layer uses the ReLU activation function, and the output layer generates the system performance metric;
[0037] Construct a branched Deep Q-Network network structure that includes a backbone network and three expert branch networks. The backbone network includes an attention enhancement layer, which contains feature-level attention that assigns dynamic weights to the features of each dimension of the state vector and action-level attention that calculates the correlation between the current state and candidate actions. The three expert branch networks are a cache update interval prediction branch, a graph attention-based data access path selection branch, and a feature matrix update amount prediction branch. Use adaptive weights based on prediction accuracy for weighted integration of branch outputs. Construct a composite reward function that includes a performance improvement reward, a cache efficiency reward based on cache hit rate, a balance reward based on the entropy value of resource utilization, and a synchronization overhead penalty term based on the change in the edge weights of the minimum cost spanning tree. Introduce L2,1-norm feature selection regularization and temporal smoothing regularization to optimize network parameters. Use a hierarchical sampling strategy based on the magnitude of the reward value to maintain the experience replay pool, and introduce importance sampling correction to optimize weight updates. Use an exponential decay exploration strategy based on the number of time steps to improve the ε-greedy strategy, and combine average reward, value estimation error, and policy stability to evaluate the training process, and output the optimized network parameters as a performance policy set.
[0038] In an alternative embodiment,
[0039] The method further includes:
[0040] Discretize the value range of the cache update interval, the set of candidate data access paths generated based on the minimum cost spanning tree, and the value range of the feature matrix update amount to construct a parameter space. Maintain a pair of Beta distribution parameters for each discrete value of the parameters in the parameter space. The pair of Beta distribution parameters includes a first parameter that reflects the historical success degree of the parameter value and a second parameter that reflects the historical failure degree of the parameter value. Construct a comprehensive reward metric system based on the change rate of the system performance metrics before and after adjustment, the change in cache hit rate, and the change in synchronization cost.
[0041] When the comprehensive reward metric exceeds the positive threshold, increase the corresponding first parameter and record the current configuration. When the comprehensive reward metric is below the negative threshold, increase the corresponding second parameter and roll back to the historical optimal configuration. When the comprehensive reward metric is between the positive threshold and the negative threshold, maintain the current configuration and update the parameter importance.
[0042] Perform Beta distribution sampling on the parameters to be adjusted in the parameter space, determine the parameter adjustment priority based on the sampling values, and apply parameter adjustments one by one according to the parameter adjustment priority; initialize the Beta distribution parameter pair using historical optimization experience, and establish a parameter adjustment confidence evaluation mechanism. The parameter adjustment confidence evaluation mechanism includes: calculating the historical success rate of parameter adjustment, the stability coefficient of parameter values, and the environmental similarity, and generating the parameter adjustment confidence based on the weighted combination of the historical success rate, stability coefficient, and environmental similarity; the historical success rate is calculated based on the ratio of the number of times the system performance is improved after parameter adjustment to the total number of adjustments, the stability coefficient is calculated based on the variance of parameter value fluctuations, and the environmental similarity is calculated based on the cosine similarity between the current query load feature vector and the feature vector at the historical optimal configuration; dynamically adjust the exploration probability according to the parameter adjustment confidence to obtain the optimized final query result.
[0043] In the second aspect of the embodiments of the present invention,
[0044] Provide a performance optimization system for cross-data source paging queries, including:
[0045] The first unit is used to extract query conditions and paging parameters after receiving a paging query request, calculate the data distribution weight based on the query conditions and generate a globally unique query identifier, construct a data source routing table, calculate the optimal data access path according to the data distribution weight and the data source routing table, split the paging query request into sub-query tasks, monitor the data source load status, adjust the number of parallel query threads based on the load status, and allocate the sub-query tasks according to the optimal data access path and the number of parallel query threads;
[0046] The second unit is used to establish a local index based on the sub-query task allocation result, calculate the index access frequency and adaptively adjust the index structure; monitor the connection pool usage rate and dynamically adjust the number of connections, batch obtain data through the connection pool, generate a query execution plan and associate it with the globally unique query identifier and store it in the query plan cache, and obtain the query results of each data source;
[0047] A third unit, configured to perform feature analysis based on the query result, calculate a data correlation coefficient, determine a data shard size based on the data correlation coefficient, construct a feature matrix including data distribution characteristics, numerical density, and cardinality information, select an optimal merging strategy to merge data; construct a cache access pattern map based on the feature matrix, calculate a data access association degree, aggregate and store data with an access association degree exceeding a preset association degree threshold based on the access association degree, and construct a minimum-cost spanning tree for data synchronization; establish a mapping relationship between the query load characteristics and the system performance including the data synchronization cost of the minimum-cost spanning tree, optimize the performance policy set through reinforcement learning, and adjust the cache update interval, data access path, and the feature matrix according to the performance policy set to obtain a final query result.
[0048] In a third aspect of the embodiments of the present invention,
[0049] there is provided an electronic device, including:
[0050] a processor;
[0051] a memory for storing instructions executable by the processor;
[0052] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0053] In a fourth aspect of the embodiments of the present invention,
[0054] there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0055] Through strategies such as calculating data distribution weights, constructing an optimal data access path, parallel query, and adaptive index adjustment, the present invention effectively reduces the latency of cross-data source queries and improves the query speed.
[0056] Mechanisms such as dynamically adjusting the number of connections in the connection pool, monitoring the load status of the data source, and adjusting the number of parallel query threads in the present invention can effectively avoid resource waste and improve the utilization rate of system resources.
[0057] The intelligent management means such as data feature analysis, optimizing the performance policy set through reinforcement learning, and adaptively adjusting the cache update interval in the present invention realize the continuous optimization of system performance, and can dynamically adjust the policy according to the change of the data access pattern, improving the overall performance and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic flowchart of a method for optimizing the performance of cross-data source paging query according to an embodiment of the present invention;
[0059] Figure 2It is a schematic structural diagram of a performance optimization system for cross-data source paging query according to an embodiment of the present invention. Detailed implementation manners
[0060] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0061] The technical solutions of the present invention will be described in detail below with specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0062] Figure 1 It is a schematic flowchart of a performance optimization method for cross-data source paging query according to an embodiment of the present invention. As Figure 1 shown, the method includes:
[0063] S1. After receiving a paging query request, extract query conditions and paging parameters, calculate data distribution weights based on the query conditions and generate a globally unique query identifier, construct a data source routing table, calculate an optimal data access path according to the data distribution weights and the data source routing table, split the paging query request into sub-query tasks, monitor the data source load status, adjust the number of parallel query threads based on the load status, and allocate the sub-query tasks according to the optimal data access path and the number of parallel query threads;
[0064] S2. Establish a local index based on the sub-query task allocation result, calculate the index access frequency and adaptively adjust the index structure; monitor the connection pool usage rate and dynamically adjust the number of connections, batch obtain data through the connection pool, generate a query execution plan and associate it with the globally unique query identifier and store it in the query plan cache, and obtain the query results of each data source;
[0065] S3. Perform feature analysis based on the query results, calculate the data correlation coefficient, determine the data shard size based on the data correlation coefficient, construct a feature matrix including data distribution characteristics, numerical density, and cardinality information, select the optimal merging strategy to merge the data; construct a cache access pattern graph based on the feature matrix, calculate the data access correlation degree, aggregate and store data with an access correlation degree exceeding a preset correlation degree threshold based on the access correlation degree, and construct a minimum-cost spanning tree for data synchronization; establish a mapping relationship between the query load characteristics and the system performance including the data synchronization cost of the minimum-cost spanning tree, optimize the performance policy set through reinforcement learning, and adjust the cache update interval, data access path, and the feature matrix according to the performance policy set to obtain the final query result.
[0066] In an alternative embodiment,
[0067] After receiving a paged query request, extract the query conditions and paging parameters, calculate the data distribution weight based on the query conditions and generate a globally unique query identifier, construct a data source routing table, calculate the optimal data access path according to the data distribution weight and the data source routing table, split the paged query request into sub-query tasks, monitor the load status of the data source, adjust the number of parallel query threads based on the load status, and allocate the sub-query tasks according to the optimal data access path and the number of parallel query threads. The steps include:
[0068] Convert the query conditions into conditional feature vectors, calculate the ratio of the number of records satisfying the query conditions to the total number of records to obtain the conditional selection degree, perform exponential smoothing calculation on the current query frequency and the historical query frequency to obtain the column usage frequency, calculate the exponential decay value based on the time difference between the data update time and the current time to obtain the data freshness, perform weighted normalization on the conditional selection degree, column usage frequency, and data freshness to obtain the data distribution weight, and perform consistent hashing operation on the conditional feature vectors to generate a globally unique query identifier;
[0069] Construct a two-layer data source routing table including a data source basic information layer and a weight distribution matrix layer. The data source basic information layer includes data source identifiers, connection parameters, and data volume statistics information, and the weight distribution matrix layer records the data distribution weights of each query condition in different data sources; construct a weighted directed graph with data sources as nodes, construct a path cost function based on node weights, edge weights, and cross-source query penalty terms, and calculate the optimal data access path including the minimum path cost through dynamic programming;
[0070] Calculate the processing time and transmission time of a single data source based on the optimal data access path to obtain the query cost, and split the paging query request into multiple sub-query tasks according to the query cost; collect the processor utilization rate, memory utilization rate, input / output waiting time, query response time, and query queue length to obtain the data source load status, calculate the load score based on the load status, decrease the number of parallel query threads when the load score is greater than the preset load threshold, and increase the number of parallel query threads when the load score is less than the preset load threshold. The increment and decrement step sizes of the number of parallel query threads are determined according to the difference between the load score and the preset load threshold.
[0071] Generate a task execution plan including query condition decomposition, data source selection, parallelism configuration, and result merging strategy based on the optimal data access path and the number of parallel query threads, and allocate the sub-query tasks based on the task execution plan.
[0072] Exemplarily, first, parse the incoming paging query request, extract the query conditions and paging parameters, such as extracting the corresponding filtering conditions and paging range information from a paging query containing age and city conditions.
[0073] Next, calculate the data distribution weight and generate a globally unique query identifier. Convert the query conditions into condition feature vectors, and form the feature vectors by recording whether the conditions are satisfied or not. Suppose the total number of records in a certain table is 1000, and the proportion of records that meet the conditions gives a condition selectivity of 0.1. The current query frequency and historical frequency are calculated through exponential smoothing to obtain a column usage frequency of 90. The data update interval is calculated through an exponential decay function to obtain a data freshness of 0.86. Weighted normalization is performed on these three metrics to obtain the data distribution weight, and a consistent hashing operation is performed on the condition feature vectors to generate the query identifier.
[0074] Then, construct a data source routing table. The data source basic information layer records the identifier, connection information, and data statistics information of the data source. The weight distribution matrix layer records the distribution weights of each query condition in different data sources, such as the distribution ratio of a certain condition in different data sources.
[0075] Next, calculate the optimal data access path. Construct a weighted directed graph with the data sources as nodes, where the node weights are the load scores, the edge weights are the network latencies, and a cross-source query penalty value is set at the same time. Calculate the minimum-cost path through the path cost function, and the path cost comprehensively considers the node weights, edge weights, and the number of cross-source queries.
[0076] After that, split the paging query request into sub-query tasks. Calculate the processing and transmission times of each data source based on the optimal access path to obtain the query cost, and accordingly split the original paging request into multiple sub-tasks with smaller ranges.
[0077] Meanwhile, monitor the load status of the data source. Calculate the load score by collecting metrics such as processor utilization rate, memory utilization rate, input / output wait time, query response time, and query queue length. When the load score exceeds the preset threshold, decrease the number of parallel query threads; conversely, increase the number of parallel query threads. The step sizes for increase and decrease are dynamically determined according to the difference between the load score and the preset threshold. For example, when the difference is larger, the adjustment step size is larger, ensuring that the system can quickly respond to load changes.
[0078] Finally, allocate sub-query tasks according to the optimal data access path and the number of parallel query threads. Generate a complete task execution plan, which details: how to decompose the original query conditions, select appropriate data sources, configure the degree of parallel execution for each sub-task, and how to effectively merge the query results of each sub-task. Based on this execution plan, allocate the sub-query tasks to the corresponding data sources for execution to ensure the efficient completion of the query tasks.
[0079] The present invention guides query requests to the optimal data source through data distribution weights, reduces cross-source queries, reduces query latency, dynamically adjusts the number of parallel query threads according to the load status of the data source, avoids data source overload, ensures the stable operation of the system, splits paged query requests into sub-tasks according to query costs, and executes them in parallel, making full use of data source resources and improving query throughput.
[0080] In an alternative embodiment,
[0081] The steps of constructing a weighted directed graph with data sources as nodes and constructing a path cost function based on node weights, edge weights, and cross-source query penalty terms, and calculating the optimal data access path with the minimum path cost through dynamic programming include:
[0082] Construct a weighted directed graph with the data source set, where the weighted directed graph includes a node set and an edge set. Calculate the processing capacity coefficient based on the processor capacity, memory capacity, and disk input / output performance of the data source, and use the product of the processing capacity coefficient and the data distribution weight as the node weight. Calculate the edge weight based on the exponentially weighted moving average of network latency and data transmission cost, and the data transmission cost is determined according to the amount of transmitted data and network bandwidth;
[0083] Construct a path cost function including a node weight term, an edge weight term, a cross-source query penalty term, a data consistency cost term, and a concurrent query cost term. The cross-source query penalty term increases in segments with the number of data sources. The data consistency cost term is calculated based on the consistency level and the number of data sources. The concurrent query cost term is calculated based on the resource conflict rate and the lock wait time;
[0084] The query-aware improved dynamic programming algorithm is used for path optimization. The improvements include: extracting features from the query request and performing pattern recognition, calculating the query pattern weight based on the complexity of the query statement, the condition selection degree, the query frequency, and the historical response time, constructing an access pattern probability matrix based on the query historical data, and integrating the query feature information into the state transition process of dynamic programming; designing an early pruning strategy based on query cost estimation, maintaining a historical optimal path cache for filtering suboptimal paths, and dynamically adjusting the pruning threshold; executing the dynamic programming process, initializing the state matrix, assigning zero to the diagonal elements, assigning the corresponding edge weight to the directly connected nodes, and assigning infinity to the unconnected nodes; using an improved state transition equation for path search, where the improved state transition equation introduces the query pattern weight to adjust the node weight; when a better path is obtained through an intermediate node, updating the path information and synchronously updating the access pattern probability matrix and the optimal path cache; periodically collecting path execution status information, calculating the deviation value between the actual path cost and the estimated cost, adjusting the query pattern weight coefficient according to the deviation value, and triggering a path recalculation process, and finally obtaining the optimal data access path with the minimum path cost.
[0085] Exemplarily, first, the system collects information on all available data sources, including the processor capabilities, memory capacities, and disk input / output performance of each data source. Then, the system calculates the processing capacity coefficient of each data source based on these performance metrics. At the same time, the system obtains the data distribution weight of each data source, which reflects the distribution of the data required by the query on different data sources. Multiplying the processing capacity coefficient by the data distribution weight gives the node weight of each data source. For example, assuming that the processing capacity coefficient of data source A is 0.8 and the data distribution weight is 0.5, then its node weight is 0.8 × 0.5 = 0.4.
[0086] Next, the system needs to calculate the edge weights between the data sources. The system continuously monitors the network latency between the data sources and uses the exponentially weighted moving average algorithm to smooth the fluctuations in the network latency. In addition, the system calculates the data transmission cost based on the amount of data transmitted and the network bandwidth. Finally, the edge weight is jointly determined by the exponentially weighted moving average value of the network latency and the data transmission cost. For example, if the exponentially weighted moving average value of the network latency between data sources A and B is 10 ms and the data transmission cost is 2, then the edge weight between A and B can be set to 10 + 2 = 12.
[0087] Based on this, the system constructs a path cost function. This function consists of multiple components: node weight term, edge weight term, cross-source query penalty term, data consistency cost term, and concurrent query cost term. The cross-source query penalty term is designed such that as the number of data sources increases, the penalty term also increases, and it adopts a piecewise growth manner to more finely control the penalty intensity. The data consistency cost term is calculated based on the required consistency level and the number of data sources involved. The concurrent query cost term is calculated based on the resource conflict rate and the lock waiting time. For example, for a query involving three data sources, if strong consistency is required, the data consistency cost may be 5; if the resource conflict rate is 0.2 and the lock waiting time is 10ms, the concurrent query cost may be 2.
[0088] The system uses a query-aware improved dynamic programming algorithm for path optimization. First, the system performs feature extraction and pattern recognition on the incoming query requests, such as analyzing the complexity of the query statement, the condition selectivity, the query frequency, and the historical response time. Then, based on these features, the query pattern weight is calculated. At the same time, the system maintains an access pattern probability matrix, which is constructed based on historical query data and records the probabilities of access between different data sources. During the state transition process of dynamic programming, the system integrates the query feature information and the access pattern probability matrix, and dynamically adjusts the node weights using the query pattern weight.
[0089] To improve efficiency, the system designs an early pruning strategy based on query cost estimation. The system maintains a historical optimal path cache to quickly filter out suboptimal paths and dynamically adjusts the pruning threshold to adapt to different query loads. During the execution of dynamic programming, the system initializes the state matrix, assigns zero to the diagonal elements, assigns the corresponding edge weights to directly connected nodes, and assigns infinity to unconnected nodes. Then, the system uses an improved state transition equation for path search. When a better path can be obtained through a certain intermediate node, the system updates the path information and synchronously updates the access pattern probability matrix and the optimal path cache.
[0090] Finally, the system periodically collects path execution status information and calculates the deviation value between the actual path cost and the estimated cost. Based on the deviation value, the system adjusts the query pattern weight coefficient and triggers the path recalculation process. Eventually, the system obtains the optimal data access path with the minimum path cost. For example, if the actual path cost is much higher than the estimated cost, the system will correspondingly increase the query pattern weight coefficient and pay more attention to the influence of query features in the next path calculation.
[0091] By optimizing the data access path, the present invention reduces the data transmission time and processing time, thereby reducing the overall data access latency, improving the query response speed, and being able to dynamically select the optimal data access path according to the load conditions and query characteristics of the data source, thus balancing the load of each data source and improving the overall resource utilization rate. By adopting the dynamic programming algorithm and the early pruning strategy, it can efficiently process large-scale data sources and complex query requests, enhancing the scalability of the system and enabling it to adapt to the growing data volume and user requirements.
[0092] In an alternative embodiment,
[0093] The steps of performing feature analysis on the query result, calculating the data correlation coefficient, determining the data shard size based on the data correlation coefficient, constructing a feature matrix including data distribution characteristics, numerical density, and cardinality information, and selecting the optimal merging strategy to merge the data include:
[0094] Performing feature extraction on the query result, including calculating the statistical features of numerical fields and character fields to generate a data feature vector; calculating the data correlation coefficient based on the data feature vector, including calculating the modified Pearson coefficient for numerical fields, calculating the similarity combination value for character fields, and calculating the data type balance weight coefficient for mixed fields to obtain the mixed correlation coefficient;
[0095] Determining the data shard size according to the mixed correlation coefficient, adopting the benchmark shard size when the mixed correlation coefficient is less than or equal to the first threshold, adjusting the shard size based on the logarithmic function when the mixed correlation coefficient is between the first threshold and the second threshold, and triggering the shard merging operation when the mixed correlation coefficient is greater than the second threshold; performing fragment optimization based on the shard variance, calculating the sum of the squared deviations of each shard size from the average shard size, and performing shard rebalancing when the shard variance exceeds the shard balance determination threshold, splitting the shards that exceed 1.5 times the average shard size or merging the shards that are less than 0.5 times the average shard size; statistically counting the shard quantity distribution and access frequency in real time, and dynamically adjusting the shard size calculation parameters according to the monitoring data;
[0096] Construct a feature matrix that includes data distribution characteristics, numerical density, and cardinality information. The data distribution characteristics include data type identification, distribution type, data skewness, dispersion degree, and outlier ratio. The numerical density includes the amount of information stored per unit, compression ratio, null value ratio, and duplication degree. The cardinality information includes the number of unique values, cardinality estimate value, value range, and frequency distribution. Select a merging strategy based on the feature matrix, and determine the merging method by calculating the weighted combination of memory cost, computational cost, and input / output cost. The memory cost is calculated based on the shard size and buffer size. The computational cost is calculated based on the number of comparison operations and sorting operations. The input / output cost is calculated based on the number of read / write operations. When the skewness of the feature matrix is lower than the data distribution uniformity determination threshold, use multi-way balanced merging. When the feature matrix shows data skew, use segmented merging. When the feature matrix indicates memory constraints, use external merging. Calculate the cost deviation based on the ratio of the actual execution cost to the estimated cost, dynamically adjust the weight coefficients of the cost calculation terms according to the deviation, and update the feature matrix by introducing a learning rate to continuously optimize the merging strategy.
[0097] Exemplarily, first, perform feature extraction on the query results. For numerical fields, statistics such as the maximum value, minimum value, average value, variance, quantiles, etc. are calculated. For example, for a dataset containing user ages, the maximum age, minimum age, average age, age distribution range, and the number of users in different age groups can be calculated. For character fields, statistics such as length, character frequency, and occurrence times are calculated. For example, for a dataset containing user comments, the average comment length, the frequency of different characters, and the occurrence times of specific keywords can be calculated. Combine these statistical features into a data feature vector.
[0098] Next, calculate the data correlation coefficient based on the data feature vector. For numerical fields, a method similar to the Pearson correlation coefficient is used to calculate the correlation. For example, analyze the relationship between user age and purchasing power. For character fields, calculate the string similarity. For example, analyze the sentiment tendency of user comments. For mixed data containing numerical and character fields, weights are assigned according to the proportion of data types, and a mixed correlation coefficient is calculated. For example, comprehensively consider the relationship between user age, comment content, and purchasing power.
[0099] Then, determine the data shard size according to the mixed correlation coefficient. Two thresholds are preset in advance. If the mixed correlation coefficient is less than or equal to the first threshold, use the preset benchmark shard size. For example, divide the dataset into 100 shards. If the mixed correlation coefficient is between the two thresholds, adjust the shard size according to the size of the mixed correlation coefficient. For example, the larger the mixed correlation coefficient, the smaller the shard. If the mixed correlation coefficient is greater than the second threshold, trigger the shard merging operation. For example, merge the shards with higher correlation.
[0100] To ensure the balance of shard sizes, the shards are optimized. The sum of the squares of the differences between the size of each shard and the average shard size is calculated as the shard variance. If the shard variance exceeds a preset balance threshold, shard rebalancing is performed. For example, a shard larger than 1.5 times the average shard size is split, or multiple shards smaller than 0.5 times the average shard size are merged. At the same time, the shard quantity distribution and access frequency are statistically monitored in real time, and the calculation parameters for shard sizes are dynamically adjusted according to the monitoring data. For example, shards with a high access frequency can be appropriately reduced in size.
[0101] Next, a feature matrix containing data distribution characteristics, numerical density, and cardinality information is constructed. The data distribution characteristics include data type, data distribution type (such as uniform distribution, normal distribution), data skewness, dispersion degree, and outlier ratio. The numerical density includes the amount of information stored per unit, compression ratio, null value ratio, and duplication degree. The cardinality information includes the number of unique values, cardinality estimate, value range, and frequency distribution. For example, for a dataset containing user incomes, the distribution type of incomes, the proportions of high-income and low-income users, the compression ratio of the data, the number of null values, and the number of users at different income levels can be statistically analyzed.
[0102] Based on the feature matrix, a merging strategy is selected. The memory cost, computational cost, and input / output cost are calculated respectively. The memory cost is calculated according to the shard size and buffer size. The computational cost is calculated according to the number of comparison operations and sorting operations. The input / output cost is calculated according to the number of read / write operations. These three costs are weighted and combined to obtain the final cost evaluation value. An appropriate merging strategy is selected according to the characteristics of the feature matrix. For example, if the data skewness is lower than a preset uniformity threshold, multi-way balanced merging is adopted. If the data is skewed, segmented merging is adopted. If memory is limited, external merging is adopted.
[0103] Finally, the cost deviation is calculated based on the ratio of the actual execution cost to the estimated cost, and the weight coefficient of the cost calculation item is dynamically adjusted according to the deviation. For example, if the actual execution cost is much higher than the estimated cost, the weight of the computational cost is increased. At the same time, a learning rate is introduced, and the feature matrix is updated according to the cost deviation, thereby continuously optimizing the merging strategy.
[0104] By adaptively selecting a merging strategy and adjusting the shard size according to data characteristics, the present invention can effectively reduce the computational cost and input / output cost of data processing, thereby improving data processing efficiency. By dynamically adjusting the shard size and merging strategy, the memory consumption can be effectively controlled, and an appropriate merging method can be selected according to the data distribution, thereby optimizing resource utilization. By statistically monitoring the shard quantity distribution and access frequency in real time and dynamically adjusting the parameters, data skewness and resource bottlenecks can be effectively avoided, thereby enhancing system stability.
[0105] In an alternative embodiment,
[0106] The steps of constructing a cache access pattern map based on the feature matrix, calculating the data access correlation degree, and clustering and storing data with an access correlation degree exceeding a preset correlation degree threshold based on the access correlation degree, and constructing a minimum-cost spanning tree for data synchronization include:
[0107] Construct a cache access pattern map based on the data skewness and dispersion degree in the feature matrix. The cache access pattern map is obtained by calculating the temporal locality strength, access frequency correlation, and spatial locality strength between data items;
[0108] Based on the cache access pattern map, introduce the compression ratio and repetition degree information in the numerical density feature dimension to construct a multi-level data access association network. Use the breadth-first search method to calculate the k-order neighborhood of each data item. The k-order neighborhood is a set of data items whose temporal correlation degree with this data item exceeds the corresponding threshold in the decreasing threshold sequence. Record the access level during the breadth-first search process and add the data items that meet the correlation degree requirements to the corresponding order neighborhood set; calculate the compression ratio weight of the data item through the sigmoid function. The compression ratio weight is negatively correlated with the difference between the compression ratio of the data item and the average compression ratio; calculate the repetition degree weight of the data item through the sigmoid function. The repetition degree weight is positively correlated with the difference between the repetition degree of the data item and the average repetition degree; sum the product of the temporal correlation degree of each data item in the k-order neighborhood and the corresponding compression ratio weight and repetition degree weight, and perform normalization processing to obtain the k-order correlation degree; cluster and store the data based on the k-order correlation degree, and store the data items with a correlation degree exceeding the preset correlation degree threshold in adjacent storage spaces;
[0109] Construct a minimum-cost spanning tree for data synchronization. Calculate the data transmission cost between nodes based on the data transmission volume, calculate the compression and decompression cost based on the compression ratio, calculate the dirty data detection cost based on the data difference rate, and use the weighted combination of the data transmission cost, compression and decompression cost, and dirty data detection cost as the synchronization cost between nodes; use the Prim algorithm to construct a minimum-cost spanning tree, select the node with the maximum correlation degree as the starting node, use a priority queue to store candidate edges, adjust the synchronization cost of the candidate edge and the correlation degree of the target node as the priority of the edge, and update the priority of all candidate edges related to the new node when adding a new node; perform dynamic optimization on the minimum-cost spanning tree based on the actual synchronization performance, adjust the cost parameters by monitoring the synchronization performance, and update the synchronization tree structure according to the locality score.
[0110] Exemplarily, to optimize the efficiency and cost of data synchronization, a data synchronization optimization method based on a cache access pattern graph is proposed. First, construct a cache access pattern graph, then calculate the data access correlation based on this graph, aggregate and store the data with high correlation, and finally construct a minimum-cost spanning tree for data synchronization.
[0111] Construct a cache access pattern graph. The input is a feature matrix that records the access features of each data item, such as access time, access frequency, access location, etc. Based on the data skew (i.e., the degree of uneven distribution of data access) and dispersion (i.e., the degree of dispersion of data access) in the feature matrix, calculate the temporal locality strength between data items (i.e., within a period of time, if a data item is accessed, then the data items nearby are also likely to be accessed), the access frequency correlation (i.e., data items with similar access frequencies may have higher correlation), and the spatial locality strength (i.e., physically adjacent data items may be accessed together). For example, if data items A and B are frequently accessed consecutively in time, then their temporal locality strength is high; if the access frequencies of data items C and D are both high, then their access frequency correlation is high; if data items E and F are adjacent in storage space, then their spatial locality strength is high. Use these locality strengths and correlations as the weights of the edges to construct a cache access pattern graph.
[0112] Calculate the data access correlation and aggregate and store the relevant data. Based on the cache access pattern graph, introduce the compression ratio (i.e., the ratio of the size of the data after compression to the size before compression) and the repeatability (i.e., the proportion of repeated elements in the data) information in the numerical density feature dimension to construct a multi-level data access correlation network. For each data item, use the breadth-first search method to calculate its k-order neighborhood. The k-order neighborhood refers to the set of data items whose temporal correlation with this data item exceeds the corresponding threshold in the decreasing threshold sequence. For example, set the decreasing threshold sequence to 0.9, 0.8, 0.7, then the 1st-order neighborhood contains the data items whose temporal correlation with this data item exceeds 0.9, the 2nd-order neighborhood contains the data items whose temporal correlation with this data item exceeds 0.8, and so on. During the breadth-first search process, record the access level and add the data items that meet the correlation requirements to the corresponding order neighborhood set.
[0113] Next, calculate the compression ratio weight and the repeatability weight of each data item. The compression ratio weight of a data item is negatively correlated with the difference between its compression ratio and the average compression ratio, that is, the higher the compression ratio of a data item, the lower its weight. The repeatability weight of a data item is positively correlated with the difference between its repeatability and the average repeatability, that is, the higher the repeatability of a data item, the higher its weight. The calculation method is to map the difference to between 0 and 1 through the sigmoid function.
[0114] Then, sum up the products of the temporal correlation degrees of each data item in the k-order neighborhood and the corresponding compression ratio weights and duplication weights, and perform normalization processing to obtain the k-order correlation degree. Finally, based on the k-order correlation degree, cluster and store the data, and store the data items whose correlation degrees exceed the preset correlation degree threshold in adjacent storage spaces. For example, if the correlation degrees of data items A and B exceed the preset threshold, store them in adjacent storage spaces.
[0115] Construct a minimum-cost spanning tree for data synchronization. First, calculate the data transmission cost between nodes based on the data transmission volume, calculate the compression and decompression cost based on the compression ratio, and calculate the dirty data detection cost based on the data difference rate. Take the weighted combination of the data transmission cost, compression and decompression cost, and dirty data detection cost as the synchronization cost between nodes. For example, if a large amount of data needs to be transmitted between nodes A and B, the data transmission cost between them is high; if the data compression ratio of node C is very high, the compression and decompression cost is high; if the data difference rate between nodes D and E is very high, the dirty data detection cost is high.
[0116] Then, use the Prim algorithm to construct a minimum-cost spanning tree. Select the node with the maximum correlation degree as the starting node, use a priority queue to store candidate edges, take the weighted adjustment of the synchronization cost of the candidate edge and the correlation degree of the target node as the priority of the edge, and update the priorities of all candidate edges related to the new node when adding a new node. For example, if the synchronization cost between nodes A and B is low and the correlation degree of node B is high, the priority of edge AB is high.
[0117] Finally, perform dynamic optimization on the minimum-cost spanning tree based on the actual synchronization performance. Adjust the cost parameters by monitoring the synchronization performance, and update the synchronization tree structure according to the locality score. For example, if it is found that the synchronization cost of some edges is too high during the actual synchronization process, the corresponding cost parameters can be adjusted; if it is found that the locality scores of some nodes are high, they can be used as new starting nodes to construct a synchronization tree.
[0118] The present invention reduces the data transmission volume and dirty data detection cost by clustering and storing highly correlated data, thereby improving the data synchronization efficiency; optimizes the data synchronization path by constructing a minimum-cost spanning tree, reduces the data transmission cost, compression and decompression cost, and dirty data detection cost, thereby reducing the overall synchronization cost; and can adjust the cost parameters and synchronization tree structure according to the actual synchronization performance by dynamically optimizing the minimum-cost spanning tree, thereby enhancing the stability of the system.
[0119] In an alternative embodiment,
[0120] Establishing a mapping relationship between query load characteristics and system performance including the cost of data synchronization in the minimum cost spanning tree, optimizing the performance policy set through reinforcement learning, and adjusting the cache update interval, data access path, and the feature matrix according to the performance policy set to obtain the final query result includes the steps of:
[0121] Construct a query load feature vector including dimensions of data access pattern, resource utilization rate, and cache hit rate. Calculate the temporal locality strength and spatial locality strength included in the data access pattern dimension based on the cache access pattern map. The resource utilization rate dimension includes CPU utilization rate, memory utilization rate, and network bandwidth utilization rate. The cache hit rate dimension includes the first-level cache hit rate and the second-level cache hit rate;
[0122] Define the weighted combination of query response time and data synchronization overhead as the system performance metric, where the data synchronization overhead is calculated based on the total edge weight of the minimum cost spanning tree for data synchronization. Establish a mapping function from the query load feature vector to the system performance metric through a multi-layer perceptron. The input layer of the multi-layer perceptron corresponds to the dimension of the query load feature vector, the hidden layer uses the ReLU activation function, and the output layer generates the system performance metric;
[0123] Construct a branched Deep Q-Network network structure including a backbone network and three expert branch networks. The backbone network includes an attention enhancement layer, which includes feature-level attention for assigning dynamic weights to the features of each dimension of the state vector and action-level attention for calculating the correlation between the current state and candidate actions. The three expert branch networks are the cache update interval prediction branch, the data access path selection branch based on graph attention, and the feature matrix update amount prediction branch; perform weighted integration of the branch outputs using adaptive weights based on prediction accuracy; construct a composite reward function including a performance improvement reward, a cache efficiency reward based on cache hit rate, a balance reward based on the entropy value of resource utilization rate, and a synchronization overhead penalty term based on the change in the edge weight of the minimum cost spanning tree; introduce L2,1 norm feature selection regularization and temporal smoothing regularization to optimize network parameters; maintain an experience replay pool using a hierarchical sampling strategy based on the size of the reward value, and introduce importance sampling correction to optimize weight updates; improve the ε-greedy strategy using an exponential decay exploration strategy based on the number of time steps, and evaluate the training process in combination with average reward, value estimation error, and policy stability, and output the optimized network parameters as the performance policy set.
[0124] Exemplarily, first, construct a query workload feature vector. This vector consists of three dimensions: data access pattern, resource utilization, and cache hit rate. To obtain the data access pattern dimension, it is necessary to analyze the pattern of query access to data. For example, the temporal locality intensity and spatial locality intensity can be calculated based on the cache access pattern graph. The temporal locality intensity refers to the frequency of repeated access to the same data within a period of time, and the spatial locality intensity refers to whether the accessed data is clustered in physical storage. Suppose a query accesses data blocks A, B, and C. If A is frequently accessed in the next period of time, the temporal locality intensity is high; if B and C are adjacent in storage, the spatial locality intensity is high. The resource utilization dimension includes CPU utilization, memory utilization, and network bandwidth utilization, and these metrics can be obtained through system monitoring tools. The cache hit rate dimension includes the first-level cache hit rate and the second-level cache hit rate, which can be calculated from the cache access times and hit times. For example, if the first-level cache is accessed 100 times and hit 80 times, the first-level cache hit rate is 80%.
[0125] Next, establish the mapping relationship between query workload features and system performance. The system performance metric is defined as a weighted combination of query response time and data synchronization overhead. The data synchronization overhead is calculated based on the total edge weight of the minimum cost spanning tree for data synchronization. For example, if the data to be synchronized is distributed across three nodes, and the synchronization costs between the nodes are 3, 4, and 5 respectively, and the edge weight of the minimum spanning tree is 3 + 4 = 7, then the data synchronization overhead is 7. A multi-layer perceptron is used to establish the mapping function from the feature vector to the performance metric. The input layer of the multi-layer perceptron corresponds to each dimension of the query workload feature vector, the hidden layer uses the ReLU activation function, and the output layer generates the system performance metric. For example, a feature vector [0.8, 0.7, 0.9] containing temporal locality intensity, CPU utilization, and first-level cache hit rate, after passing through the multi-layer perceptron, outputs a system performance metric value of 0.5.
[0126] Then, use reinforcement learning to optimize the performance policy set. Construct a branched Deep Q-Network network structure, which consists of a backbone network and three expert branch networks. The backbone network includes an attention enhancement layer, which contains feature-level attention and action-level attention. Feature-level attention assigns dynamic weights to the features of each dimension of the state vector. For example, if the CPU utilization is too high, a higher weight is assigned to the CPU utilization dimension. Action-level attention calculates the correlation between the current state and the candidate actions. The three expert branch networks are respectively used to predict the cache update interval, select the data access path, and predict the feature matrix update amount. Adaptive weights based on prediction accuracy are used to weight and integrate the branch outputs. For example, if the prediction accuracy of the cache update interval prediction branch is high, a higher weight is given to it.
[0127] To train the reinforcement learning network, a composite reward function is designed. This function includes a performance improvement reward, a cache efficiency reward based on the cache hit rate, a balance reward based on the entropy value of resource utilization, and a synchronization overhead penalty term based on the change in the edge weights of the minimum cost spanning tree. For example, if the cache hit rate increases, a cache efficiency reward is given; if the resource utilization is unbalanced, a penalty is given. To optimize the network parameters, L2,1-norm feature selection regularization and temporal smoothing regularization are introduced. A hierarchical sampling strategy based on the magnitude of the reward value is adopted to maintain the experience replay pool, and importance sampling correction is introduced to optimize the weight update. The traditional ε-greedy strategy is improved by exponentially decreasing the exploration probability with the increase of the time step, that is, the exploration probability ε starts from the initial value of 0.9 and decreases exponentially with the increase of the time step t, and the decay rate is reduced to 1 / e of the previous value every 1000 steps, where e represents the natural constant, so that the agent explores new actions with a high probability in the initial stage of training and utilizes the known optimal actions with a high probability in the later stage of training, and the training process is evaluated by combining the average reward, value estimation error, and policy stability, and finally the optimized network parameters are output as the performance policy set.
[0128] According to the optimized performance policy set, the cache update interval, data access path, and feature matrix are adjusted, and finally the query result is obtained. For example, if the policy set indicates shortening the cache update interval, the system will update the cache data more frequently.
[0129] Through the reinforcement learning to optimize the performance policy set, the present invention can dynamically adjust the system configuration according to the query load characteristics, thereby minimizing the query response time and data synchronization overhead, significantly improving the overall performance of the system, being able to adaptively adjust the system parameters according to different query load characteristics, enabling the system to operate efficiently under different load conditions, enhancing the adaptability of the system, and reducing the need for manual intervention by automatically learning and adjusting the system configuration, thereby reducing the system operation and maintenance cost.
[0130] In an alternative embodiment,
[0131] The method further includes:
[0132] Discretize the value range of the cache update interval, the set of candidate data access paths generated based on the minimum cost spanning tree, and the value range of the feature matrix update amount to construct a parameter space; maintain a pair of Beta distribution parameters for each discrete value of the parameters in the parameter space, and the pair of Beta distribution parameters includes a first parameter reflecting the historical success degree of the parameter value and a second parameter reflecting the historical failure degree of the parameter value; construct a comprehensive reward metric system based on the change rate of the system performance index before and after adjustment, the change in the cache hit rate, and the change in the synchronization cost.
[0133] When the comprehensive reward metric exceeds the positive threshold, increase the corresponding first parameter and record the current configuration; when the comprehensive reward metric is lower than the negative threshold, increase the corresponding second parameter and roll back to the historical optimal configuration; when the comprehensive reward metric is between the positive threshold and the negative threshold, maintain the current configuration and update the parameter importance;
[0134] Perform Beta distribution sampling on the parameters to be adjusted in the parameter space, determine the parameter adjustment priority based on the sampling values, and apply parameter adjustments one by one according to the parameter adjustment priority; initialize the Beta distribution parameter pair using historical optimization experience, and establish a parameter adjustment confidence evaluation mechanism. The parameter adjustment confidence evaluation mechanism includes: calculating the historical success rate of parameter adjustment, the stability coefficient of parameter values, and the environmental similarity, and generating the parameter adjustment confidence based on the weighted combination of the historical success rate, stability coefficient, and environmental similarity; the historical success rate is calculated based on the ratio of the number of times the system performance improves after parameter adjustment to the total number of adjustments, the stability coefficient is calculated based on the variance of the fluctuations of parameter values, and the environmental similarity is calculated based on the cosine similarity between the current query load feature vector and the feature vector at the historical optimal configuration; dynamically adjust the exploration probability according to the parameter adjustment confidence to obtain the optimized final query result.
[0135] Exemplarily, first, determine the key cache parameters to be adjusted, including the cache update interval, the set of candidate data access paths, and the feature matrix update amount. For example, the cache update interval can be set in the range of 1 minute to 1 hour, the set of candidate data access paths can be generated from the table structure of the database based on the minimum cost spanning tree algorithm, and the feature matrix update amount can be set in the range of 10 to 1000.
[0136] Then, discretize the value ranges of these parameters to construct a parameter space. For example, discretize the cache update interval into five values: 1 minute, 5 minutes, 10 minutes, 30 minutes, and 1 hour, and discretize the feature matrix update amount into four values: 10, 100, 500, and 1000. For the set of candidate data access paths, assuming there are 5 candidate paths, discretize it into 5 independent options.
[0137] Next, maintain a Beta distribution parameter pair for each discrete value in the parameter space. The Beta distribution parameter pair contains two parameters: the first parameter reflects the historical success degree of the parameter value, and the second parameter reflects the historical failure degree of the parameter value. In the initial state, both parameters can be set to 1, indicating a neutral attitude towards the prior knowledge of all values. For example, the initial value of the Beta distribution parameter pair for the cache update interval of 5 minutes is (1, 1).
[0138] Subsequently, a comprehensive reward metric system is constructed to evaluate the effect of parameter adjustment. This system includes three metrics: the change rate of the system performance metric before and after adjustment, the change in cache hit rate, and the change in synchronization cost. For example, the system performance metric can be the average query response time, the cache hit rate can be calculated as the ratio of the number of cache accesses to the total number of accesses, and the synchronization cost can measure the additional overhead brought by cache updates. Suppose after a certain adjustment, the average query response time decreases by 10%, the cache hit rate increases by 5%, and the synchronization cost increases by 2%, then these three metrics can be combined into a comprehensive reward value according to the pre-set weights.
[0139] During the system operation, database query operations are performed according to the current configuration, and the change of the comprehensive reward metric is observed. When the comprehensive reward metric exceeds the pre-set positive threshold, for example, it increases by 10%, then increase the first parameter of the corresponding parameter value and record the current configuration as the new optimal configuration. When the comprehensive reward metric is lower than the pre-set negative threshold, for example, it decreases by 5%, then increase the second parameter of the corresponding parameter value and roll back the system configuration to the historical optimal configuration. When the comprehensive reward metric is between the positive threshold and the negative threshold, for example, the change is between -5% and 10%, then keep the current configuration unchanged and update the parameter importance, that is, adjust its weight according to the contribution degree of the parameter to the comprehensive reward metric.
[0140] To explore new parameter configurations, Beta distribution sampling is performed on the parameters to be adjusted in the parameter space. According to the Beta distribution of each parameter value, a sampling value is randomly generated. For example, if the Beta distribution parameters for the cache update interval of 5 minutes are (3, 2), then a value is randomly sampled according to this distribution as the new cache update interval.
[0141] Determine the parameter adjustment priority according to the sampling value. The parameters with high confidence in parameter adjustment are adjusted first. The parameter adjustment confidence evaluation mechanism includes calculating the historical success rate of parameter adjustment, the stability coefficient of parameter values, and the environmental similarity. The historical success rate is calculated as the ratio of the number of times the system performance improves after parameter adjustment to the total number of adjustments. The stability coefficient is calculated according to the variance of the parameter value fluctuations. The environmental similarity is calculated according to the cosine similarity between the current query load feature vector and the feature vector at the historical optimal configuration. For example, if the current query load feature vector is (1, 2, 3) and the feature vector at the historical optimal configuration is (2, 4, 6), then the cosine similarity between them is 1, indicating that the environment is very similar. The historical success rate, stability coefficient, and environmental similarity are weighted and combined to generate the parameter adjustment confidence. Dynamically adjust the exploration probability according to the parameter adjustment confidence. For example, the higher the confidence, the lower the exploration probability, so as to balance the utilization of historical experience and the exploration of new configurations. Finally, perform database queries according to the optimized parameter configuration to obtain the optimized final query result.
[0142] By dynamically adjusting cache parameters, the present invention achieves a higher cache hit rate, thereby reducing the number of database accesses and improving query efficiency. By intelligently adjusting the cache update strategy, unnecessary cache update operations are reduced, and the synchronization cost is lowered, thus reducing the system overhead. It can dynamically adjust cache parameters according to the changes in query load, adapt to different application scenarios, and maintain the stability of system performance.
[0143] Figure 2 It is a schematic structural diagram of a performance optimization system for cross-data-source paging queries in an embodiment of the present invention, as Figure 2 shown, the system includes:
[0144] A first unit, which, after receiving a paging query request, extracts query conditions and paging parameters, calculates data distribution weights based on the query conditions and generates a globally unique query identifier, constructs a data source routing table, calculates an optimal data access path according to the data distribution weights and the data source routing table, splits the paging query request into sub-query tasks, monitors the load status of the data source, adjusts the number of parallel query threads based on the load status, and allocates the sub-query tasks according to the optimal data access path and the number of parallel query threads;
[0145] A second unit, which establishes a local index based on the sub-query task allocation result, calculates the index access frequency and adaptively adjusts the index structure; monitors the connection pool usage rate and dynamically adjusts the number of connections, obtains data in batches through the connection pool, generates a query execution plan and associates it with the globally unique query identifier and stores it in the query plan cache, and obtains the query results of each data source;
[0146] A third unit, which performs feature analysis based on the query results, calculates the data correlation coefficient, determines the data shard size based on the data correlation coefficient, constructs a feature matrix including data distribution characteristics, numerical density, and cardinality information, selects an optimal merging strategy to merge data; constructs a cache access pattern map based on the feature matrix, calculates the data access association degree, aggregates and stores data exceeding a preset association degree threshold based on the access association degree, and constructs a minimum-cost spanning tree for data synchronization; establishes a mapping relationship between the query load characteristics and the system performance including the data synchronization cost of the minimum-cost spanning tree, optimizes the performance policy set through reinforcement learning, and adjusts the cache update interval, data access path, and the feature matrix according to the performance policy set to obtain the final query result.
[0147] In the third aspect of the embodiment of the present invention,
[0148] A kind of electronic device is provided, including:
[0149] A processor;
[0150] A memory for storing instructions executable by the processor;
[0151] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0152] In a fourth aspect of the embodiments of the present invention,
[0153] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0154] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0155] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing the performance of cross-data source paging queries, characterized in that Including: After receiving a paging query request, extract the query conditions and paging parameters, calculate the data distribution weight based on the query conditions and generate a globally unique query identifier, construct a data source routing table, calculate the optimal data access path according to the data distribution weight and the data source routing table, split the paging query request into sub-query tasks, monitor the data source load status, adjust the number of parallel query threads based on the load status, and allocate the sub-query tasks according to the optimal data access path and the number of parallel query threads; Establish a local index based on the sub-query task allocation result, calculate the index access frequency and adaptively adjust the index structure; monitor the connection pool usage rate and dynamically adjust the number of connections, batch obtain data through the connection pool, generate a query execution plan and associate it with the globally unique query identifier and store it in the query plan cache, and obtain the query results of each data source; Conduct feature analysis based on the query results, calculate the data correlation coefficient, determine the data shard size based on the data correlation coefficient, construct a feature matrix containing data distribution characteristics, numerical density, and cardinality information, and select the optimal merging strategy to merge the data; construct a cache access pattern graph based on the feature matrix, calculate the data access association degree, aggregate and store data exceeding the preset association degree threshold based on the access association degree, and construct a minimum-cost spanning tree for data synchronization; Establish a mapping relationship between the query load characteristics and the system performance including the data synchronization cost of the minimum-cost spanning tree, optimize the performance policy set through reinforcement learning, and adjust the cache update interval, data access path, and the feature matrix according to the performance policy set to obtain the final query result.
2. The method according to claim 1, wherein The steps of, after receiving a paging query request, extracting the query conditions and paging parameters, calculating the data distribution weight based on the query conditions and generating a globally unique query identifier, constructing a data source routing table, calculating the optimal data access path according to the data distribution weight and the data source routing table, splitting the paging query request into sub-query tasks, monitoring the data source load status, adjusting the number of parallel query threads based on the load status, and allocating the sub-query tasks according to the optimal data access path and the number of parallel query threads include: Convert the query conditions into conditional feature vectors, calculate the ratio of the number of records satisfying the query conditions to the total number of records to obtain the conditional selection degree, perform exponential smoothing calculation on the current query frequency and the historical query frequency to obtain the column usage frequency, calculate the exponential decay value based on the time difference between the data update time and the current time to obtain the data freshness, perform weighted normalization on the conditional selection degree, column usage frequency, and data freshness to obtain the data distribution weight, and perform consistent hashing operation on the conditional feature vector to generate a globally unique query identifier; Construct a two - layer data source routing table including a basic information layer of data sources and a weight distribution matrix layer. The basic information layer of data sources includes data source identifiers, connection parameters, and data volume statistical information. The weight distribution matrix layer records the data distribution weights of each query condition in different data sources; construct a weighted directed graph with data sources as nodes, construct a path cost function based on node weights, edge weights, and cross - source query penalty terms, and calculate the optimal data access path with the minimum path cost through dynamic programming; Calculate the processing time and transmission time of a single data source based on the optimal data access path to obtain the query cost, and split the paged query request into multiple sub - query tasks according to the query cost; collect the processor utilization rate, memory utilization rate, input / output waiting time, query response time, and query queue length to obtain the data source load status, calculate the load score based on the load status, decrease the number of parallel query threads when the load score is greater than the preset load threshold, and increase the number of parallel query threads when the load score is less than the preset load threshold. The increment and decrement steps of the number of parallel query threads are determined according to the difference between the load score and the preset load threshold; Generate a task execution plan including query condition decomposition, data source selection, parallelism configuration, and result merging strategy based on the optimal data access path and the number of parallel query threads, and allocate the sub - query tasks based on the task execution plan.
3. The method according to claim 2, wherein The steps of constructing a weighted directed graph with data sources as nodes, constructing a path cost function based on node weights, edge weights, and cross - source query penalty terms, and calculating the optimal data access path with the minimum path cost through dynamic programming include: Construct a weighted directed graph from the data source set. The weighted directed graph includes a node set and an edge set. Calculate the processing capacity coefficient based on the processor capacity, memory capacity, and disk input / output performance of the data source, take the product of the processing capacity coefficient and the data distribution weight as the node weight, calculate the edge weight based on the exponentially weighted moving average of network latency and data transmission cost, and the data transmission cost is determined according to the amount of transmitted data and network bandwidth; Construct a path cost function including node weight terms, edge weight terms, cross - source query penalty terms, data consistency cost terms, and concurrent query cost terms. The cross - source query penalty term increases in segments with the number of data sources. The data consistency cost term is calculated based on the consistency level and the number of data sources. The concurrent query cost term is calculated based on the resource conflict rate and lock waiting time; The query-aware improved dynamic programming algorithm is used for path optimization. The improvements include: extracting features from the query request and performing pattern recognition, calculating the query pattern weight based on the complexity of the query statement, the condition selection degree, the query frequency, and the historical response time, constructing the access pattern probability matrix based on the query historical data, and integrating the query feature information into the state transition process of dynamic programming; designing an early pruning strategy based on query cost estimation, maintaining a historical optimal path cache for filtering suboptimal paths, and dynamically adjusting the pruning threshold; executing the dynamic programming process, initializing the state matrix, assigning zero to the diagonal elements, assigning the corresponding edge weights to the directly connected nodes, and assigning infinity to the unconnected nodes; using an improved state transition equation for path search, where the improved state transition equation introduces the query pattern weight to adjust the node weight; when a better path is obtained through an intermediate node, updating the path information and synchronously updating the access pattern probability matrix and the optimal path cache; periodically collecting path execution status information, calculating the deviation value between the actual path cost and the estimated cost, adjusting the query pattern weight coefficient according to the deviation value, and triggering the path recalculation process, and finally obtaining the optimal data access path with the minimum path cost.
4. The method according to claim 1, wherein The steps of performing feature analysis on the query result, calculating the data correlation coefficient, determining the data shard size based on the data correlation coefficient, and constructing a feature matrix including data distribution characteristics, numerical density, and cardinality information, and selecting the optimal merging strategy to merge the data include: Extracting features from the query result, including calculating the statistical features of numerical fields and character fields to generate a data feature vector; calculating the data correlation coefficient based on the data feature vector, including calculating the corrected Pearson coefficient for numerical fields, calculating the similarity combination value for character fields, and calculating the data type balance weight coefficient for mixed fields to obtain the mixed correlation coefficient; Determining the data shard size according to the mixed correlation coefficient. When the mixed correlation coefficient is less than or equal to the first threshold, the benchmark shard size is used. When the mixed correlation coefficient is between the first threshold and the second threshold, the shard size is adjusted based on the logarithmic function. When the mixed correlation coefficient is greater than the second threshold, a shard merging operation is triggered; performing fragment optimization based on the shard variance, calculating the sum of the squared deviations of each shard size from the average shard size. When the shard variance exceeds the shard balance determination threshold, perform shard rebalancing, splitting the shards larger than 1.5 times the average shard size or merging the shards smaller than 0.5 times the average shard size; real-time statistics of the shard quantity distribution and access frequency, and dynamically adjusting the shard size calculation parameters according to the monitoring data; Construct a feature matrix that includes data distribution characteristics, numerical density, and cardinality information. The data distribution characteristics include data type identification, distribution type, data skewness, dispersion degree, and outlier ratio. The numerical density includes the amount of information stored per unit, compression ratio, null value ratio, and duplication degree. The cardinality information includes the number of unique values, cardinality estimation value, value range, and frequency distribution. Select a merging strategy based on the feature matrix, and determine the merging method by calculating the weighted combination of memory cost, computational cost, and input / output cost. The memory cost is calculated based on the shard size and buffer size. The computational cost is calculated based on the number of comparison operations and sorting operations. The input / output cost is calculated based on the number of read / write operations. When the skewness of the feature matrix is lower than the data distribution uniformity determination threshold, use multi-way balanced merging. When the feature matrix shows data skewness, use segmented merging. When the feature matrix indicates memory constraints, use external merging. Calculate the cost deviation based on the ratio of the actual execution cost to the estimated cost, dynamically adjust the weight coefficients of the cost calculation terms according to the deviation, and update the feature matrix by introducing a learning rate to continuously optimize the merging strategy.
5. The method according to claim 1, wherein Based on the feature matrix, construct a cache access pattern graph, calculate the data access correlation degree, and the steps of aggregating and storing data with an access correlation degree exceeding a preset correlation degree threshold based on the access correlation degree to construct a minimum-cost spanning tree for data synchronization include: Construct a cache access pattern graph based on the data skewness and dispersion degree in the feature matrix. The cache access pattern graph is obtained by calculating the temporal locality strength, access frequency correlation, and spatial locality strength between data items. Based on the cache access pattern graph, introduce the compression ratio and duplication degree information in the numerical density feature dimension to construct a multi-level data access correlation network. Use the breadth-first search method to calculate the k-order neighborhood for each data item. The k-order neighborhood is a set of data items whose temporal correlation degree with the data item exceeds the corresponding threshold in the decreasing threshold sequence. Record the access level during the breadth-first search process and add the data items that meet the correlation degree requirements to the corresponding order neighborhood set. Calculate the compression ratio weight of the data item through the sigmoid function. The compression ratio weight is negatively correlated with the difference between the compression ratio of the data item and the average compression ratio. Calculate the duplication degree weight of the data item through the sigmoid function. The duplication degree weight is positively correlated with the difference between the duplication degree of the data item and the average duplication degree. Sum the product of the temporal correlation degree of each data item in the k-order neighborhood and the corresponding compression ratio weight and duplication degree weight, and perform normalization processing to obtain the k-order correlation degree. Aggregate and store the data based on the k-order correlation degree, and store the data items with a correlation degree exceeding the preset correlation degree threshold in adjacent storage spaces. Construct a minimum-cost spanning tree for data synchronization, calculate the data transmission cost between nodes based on the data transmission volume, calculate the compression and decompression cost based on the compression ratio, calculate the dirty data detection cost based on the data difference rate, and use the weighted combination of the data transmission cost, compression and decompression cost, and dirty data detection cost as the synchronization cost between nodes; use the Prim algorithm to construct the minimum-cost spanning tree, select the node with the highest degree of association as the starting node, use a priority queue to store candidate edges, and use the weighted adjustment of the synchronization cost of the candidate edge and the degree of association of the target node as the priority of the edge, and update the priorities of all candidate edges related to the new node when adding a new node; dynamically optimize the minimum-cost spanning tree based on the actual synchronization performance, adjust the cost parameters by monitoring the synchronization performance, and update the synchronization tree structure according to the locality score.
6. The method according to claim 1, characterized in that, The steps of establishing the mapping relationship between the query load characteristics and the system performance including the data synchronization cost of the minimum-cost spanning tree, optimizing the performance policy set through reinforcement learning, and adjusting the cache update interval, data access path, and the feature matrix according to the performance policy set to obtain the final query result include: Construct a query load feature vector including dimensions of data access pattern, resource utilization, and cache hit rate. Calculate the temporal locality strength and spatial locality strength included in the data access pattern dimension based on the cache access pattern map. The resource utilization dimension includes CPU utilization, memory utilization, and network bandwidth utilization. The cache hit rate dimension includes the first-level cache hit rate and the second-level cache hit rate. Define the weighted combination of the query response time and the data synchronization overhead as the system performance metric, where the data synchronization overhead is calculated based on the total edge weight of the minimum-cost spanning tree for data synchronization. Establish a mapping function from the query load feature vector to the system performance metric through a multi-layer perceptron. The input layer of the multi-layer perceptron corresponds to the dimension of the query load feature vector, the hidden layer uses the ReLU activation function, and the output layer generates the system performance metric. Construct a branched Deep Q-Network network structure that includes a backbone network and three expert branch networks. The backbone network includes an attention enhancement layer, which contains feature-level attention that assigns dynamic weights to the features of each dimension of the state vector and action-level attention that calculates the association degree between the current state and candidate actions. The three expert branch networks are a cache update interval prediction branch, a data access path selection branch based on graph attention, and a feature matrix update amount prediction branch. Use adaptive weights based on prediction accuracy for weighted integration of branch outputs. Construct a composite reward function that includes a performance improvement reward, a cache efficiency reward based on cache hit rate, a balance reward based on the entropy value of resource utilization, and a synchronization overhead penalty term based on the change in the edge weights of the minimum cost spanning tree. Introduce L2,1 norm feature selection regularization and temporal smoothing regularization to optimize network parameters. Use a hierarchical sampling strategy based on the magnitude of the reward value to maintain the experience replay pool, and introduce importance sampling correction to optimize weight updates. Use an exponential decay exploration strategy based on the number of time steps to improve the ε-greedy strategy, and combine average reward, value estimation error, and policy stability to evaluate the training process, and output the optimized network parameters as the performance policy set.
7. The method according to claim 6, wherein The method further includes: Discretize the value range of the cache update interval, the set of candidate data access paths generated based on the minimum cost spanning tree, and the value range of the feature matrix update amount to construct a parameter space. Maintain a pair of Beta distribution parameters for each discrete value of the parameters in the parameter space. The pair of Beta distribution parameters includes a first parameter that reflects the historical success degree of the parameter value and a second parameter that reflects the historical failure degree of the parameter value. Construct a comprehensive reward metric system based on the change rate of the system performance metrics before and after adjustment, the change in cache hit rate, and the change in synchronization cost. When the comprehensive reward metric exceeds the positive threshold, increase the corresponding first parameter and record the current configuration. When the comprehensive reward metric is lower than the negative threshold, increase the corresponding second parameter and roll back to the historical optimal configuration. When the comprehensive reward metric is between the positive threshold and the negative threshold, maintain the current configuration and update the parameter importance. Perform Beta distribution sampling on the parameters to be adjusted in the parameter space, determine the parameter adjustment priority based on the sampling values, and apply parameter adjustments one by one according to the parameter adjustment priority; initialize the Beta distribution parameter pair using historical optimization experience, and establish a parameter adjustment confidence evaluation mechanism. The parameter adjustment confidence evaluation mechanism includes: calculating the historical success rate of parameter adjustment, the stability coefficient of parameter values, and the environmental similarity, and generating the parameter adjustment confidence based on the weighted combination of the historical success rate, the stability coefficient, and the environmental similarity; the historical success rate is calculated based on the ratio of the number of times the system performance is improved after parameter adjustment to the total number of adjustments, the stability coefficient is calculated based on the fluctuation variance of parameter values, and the environmental similarity is calculated based on the cosine similarity between the current query load feature vector and the feature vector at the historical optimal configuration; dynamically adjust the exploration probability according to the parameter adjustment confidence to obtain the optimized final query result.
8. A performance optimization system for cross-data-source paged queries, which is used to implement the method described in any one of the foregoing claims 1-7, characterized in that, Including: The first unit is configured to, after receiving a paged query request, extract query conditions and paging parameters, calculate data distribution weights based on the query conditions and generate a globally unique query identifier, construct a data source routing table, calculate the optimal data access path according to the data distribution weights and the data source routing table, split the paged query request into sub-query tasks, monitor the data source load status, adjust the number of parallel query threads based on the load status, and allocate the sub-query tasks according to the optimal data access path and the number of parallel query threads; The second unit is configured to establish a local index based on the sub-query task allocation result, calculate the index access frequency and perform adaptive adjustment on the index structure; monitor the connection pool usage rate and dynamically adjust the number of connections, batch obtain data through the connection pool, generate a query execution plan and associate it with the globally unique query identifier and store it in the query plan cache, and obtain the query results of each data source; The third unit is configured to perform feature analysis based on the query results, calculate the data correlation coefficient, determine the data shard size based on the data correlation coefficient, construct a feature matrix including data distribution characteristics, numerical density, and cardinality information, and select the optimal merging strategy to merge the data; construct a cache access pattern graph based on the feature matrix, calculate the data access correlation degree, aggregate and store data exceeding the preset correlation degree threshold based on the access correlation degree, and construct a minimum-cost spanning tree for data synchronization; Establish a mapping relationship between the query load characteristics and the system performance including the data synchronization cost of the minimum-cost spanning tree, optimize the performance policy set through reinforcement learning, and adjust the cache update interval, the data access path, and the feature matrix according to the performance policy set to obtain the final query result.
9. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Fragmented storage and query optimization method and system for high-concurrency database
CN120492489A
Sharded storage and query optimization method and system for high-concurrency database
CN120492489B
Database dynamic query optimization and resource scheduling method, equipment and medium
CN120849458A
Database query strategy selection method, system, equipment, product and medium
CN120872997A
Sparse direct solution method based on data out-of-core storage
CN120950809A