Cache-based data query analysis result optimization method and system
Through deep semantic understanding and dynamic index optimization, combined with resource scheduling, the problem of insufficient understanding in data queries and insane cache strategies is solved, efficient data queries and cache management are realized, and system performance and stability are improved.
Patent Information
- Application Number
- CN202510318748.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, data query understanding is not deep enough, data distribution perception is insufficient, and cache strategies are not intelligent enough, resulting in inaccurate query efficiency.
By receiving user query requests, semantic understanding and intent recognition, building query feature matrix, predicting data distribution feature maps, dynamically adjusting index structure, combining exception detection and resource scheduling, and optimizing query paths and cache strategies.
It improves query efficiency, reduces database access pressure, improves system performance and resource utilization, and ensures stable system operation.
Smart Images

Figure CN120296041A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of optimizing data query parsing results, and in particular, to a method and system for optimizing data query parsing results based on caching. Background Art
[0002] Data query is a crucial part of modern information systems. Efficient data query is crucial for improving user experience and system performance. With the continuous growth of data scale and the increase in query complexity, traditional database query methods are facing increasing challenges.
[0003] In the prior art, there are still problems such as insufficient understanding of queries, lack of awareness of data distribution, and non-intelligent caching strategies.
[0004] Therefore, there is an urgent need for a solution to solve the problems existing in the prior art. Summary of the Invention
[0005] The embodiments of the present invention provide a method and system for optimizing data query parsing results based on caching, which can at least solve some of the problems existing in the prior art.
[0006] In the first aspect of the embodiments of the present invention, a method for optimizing data query parsing results based on caching is provided, including:
[0007] Receiving a user data query request and extracting an original query statement, performing word segmentation processing on the original query statement to generate a query token set and adding it to a semantic understanding model to extract association relationships, obtaining a first query semantic network, identifying path information of the first query semantic network through an intention recognition model to generate a second query semantic network, constructing a query feature matrix based on the second query semantic network and adding it to a feature combination model, generating fusion feature data by calculating the association weight coefficient between matrices, extracting high-dimensional feature information from the fusion feature data through a deep residual network and generating a query feature identifier, calculating the similarity with a pre-set template feature identifier and generating a similarity score table, and performing hierarchical matching and determining a set of query rules corresponding to the successfully matched template feature identifier.
[0008] Add the query feature matrix to the data distribution prediction network, predict the spatio-temporal distribution law according to the historical access records to obtain the data distribution feature map, construct a data association model and generate a probability distribution table by calculating the association probability values between fields, construct a multi-level combined index based on the probability distribution table and construct an index structure tree in combination with the information benefit ratio, calculate the access timing features and construct a time decay calculation model to obtain a dynamic index adjustment scheme, add the execution paths in the query rule set to the path optimization network based on the dynamic index adjustment scheme, determine the optimal execution path in combination with the path search algorithm and generate a task execution plan, schedule the database executor according to the task execution plan and obtain the query result stream for real-time analysis, calculate the data value coefficient and write it into the multi-level cache system;
[0009] Collect the execution performance data of the user data query request and add it to the anomaly recognition model, identify the anomaly data points in combination with the clustering operation and generate an anomaly label set, calculate the deviation degree of the performance index based on the anomaly detection model based on the recurrent neural network and determine the temporal dependence relationship in combination with the causal inference network to generate a bottleneck location map, generate a performance score based on the bottleneck location map and the deviation degree of the performance index and generate a data value prediction sequence according to the data value coefficient, determine the load balancing coefficient of the storage area in combination with the resource scheduling network and the competitive learning algorithm to generate a preloading strategy, perform performance verification through the Monte Carlo simulation model and select the preloading strategy with the highest confidence as the optimal strategy.
[0010] In an alternative embodiment,
[0011] Receive the user data query request and extract the original query statement, perform word segmentation processing on the original query statement to generate a query token set and add it to the semantic understanding model to extract the association relationship to obtain the first query semantic network, and identify the path information of the first query semantic network through the intent recognition model to generate the second query semantic network, including:
[0012] Receive the user data query request, extract the original query statement in the user data query request, construct a token recognition tree, set the node attributes of the token recognition tree as token content, part of speech, and position offset, start from the beginning of the original query statement to perform forward maximum matching to find the longest matching token in the dictionary, start from the end of the original query statement to perform backward maximum matching to find the longest matching token in the dictionary, compare the word segmentation results of the forward maximum matching and the backward maximum matching, select the word segmentation result with the least number of tokens when the number of tokens is different, and select the word segmentation result with the least number of single characters when the number of tokens is the same to generate a query token set;
[0013] Construct a deep bidirectional long short-term memory network, input the word vector sequence of the query token set into the input layer of the deep bidirectional long short-term memory network, extract context features through multiple bidirectional long short-term memory layers of the deep bidirectional long short-term memory network, input the extracted context features into a conditional random field for part-of-speech tagging, identify institutional words, time words, measurement words, numerical words, object words, statistical index words, judgment words, and mark the part-of-speech attributes and semantic roles of the query token set;
[0014] Construct an eight-head self-attention model. The first attention head captures temporal associations, the second attention head captures conditional associations, the third attention head captures object associations, the fourth attention head captures index associations, the fifth attention head captures statistical associations, the sixth attention head captures subject associations, the seventh attention head captures attribution associations, and the eighth attention head captures combination associations. Input the labeled query token set into the eight-head self-attention model and calculate the attention scores between tokens;
[0015] Construct a first query semantic network based on the attention scores. Set the tokens as the nodes of the first query semantic network, set the semantic associations as the edges of the first query semantic network, and set the attention scores as the weights of the edges; set a graph structure pruning threshold, delete the edges with weights lower than the graph structure pruning threshold, and extract the main query path and branch query paths;
[0016] Construct a graph attention model. Input the first query semantic network into the graph attention layer of the graph attention model, calculate the attention scores between a node and its adjacent nodes based on the spatial attention mechanism, update the node semantic representation, input the updated node semantic representation into the path extraction layer of the graph attention model, extract the key query path based on the depth-first search algorithm, mark the attention scores of the nodes on the path, and generate a second query semantic network.
[0017] In an alternative embodiment,
[0018] Construct a query feature matrix based on the second query semantic network and add it to the feature combination model. Generate fused feature data by calculating the correlation weight coefficients between matrices, extract high-dimensional feature information from the fused feature data through a deep residual network and generate a query feature identifier, calculate the similarity with a pre-set template feature identifier and generate a similarity score table, and perform hierarchical matching and determine that the query rule set corresponding to the successfully matched template feature identifier includes:
[0019] Construct a query feature matrix based on the second query semantic network, extract the temporal association features, numerical comparison features, and data type features between nodes in the query semantic network, and construct a query feature matrix including table association dimension, field correlation dimension, filtering condition dimension, sorting rule dimension, and time range dimension;
[0020] Add the query feature matrix to the feature combination model, set a matrix correlation calculation unit in the feature combination model, calculate the correlation weight coefficient between matrices through the matrix correlation calculation unit, and according to the correlation weight coefficient, retain and transfer the feature information with a correlation weight coefficient higher than the weight threshold, and attenuate and filter the feature information with a correlation weight coefficient lower than the weight threshold to generate fused feature data;
[0021] Construct a multi-scale residual network, which includes a feature decomposition unit, a feature reconstruction unit and a feature optimization unit. The feature decomposition unit decomposes the fused feature data into multi-scale feature components through wavelet transform. The feature reconstruction unit sets an adaptive residual module to perform independent feature reconstruction on feature components of different scales. The feature optimization unit enhances the feature expression through a cross-scale feature interaction mechanism to generate a query feature identifier;
[0022] Construct a multi-view feature matching network, which includes a local feature matching sub-network, a global feature matching sub-network and a feature calibration sub-network. The local feature matching sub-network adopts a hierarchical attention mechanism to adaptively select key local features for matching. The global feature matching sub-network constructs a semantic consistency constraint in the feature space based on a contrastive learning strategy. The feature calibration sub-network optimizes the discriminability of the feature representation through the mutual information maximization criterion. Input the query feature identifier and the template feature identifier into the multi-view feature matching network to fuse the local structure similarity and the global semantic similarity to generate a similarity score table;
[0023] Perform hierarchical matching, dynamically adjust the matching thresholds of different query areas according to the system resource status. When the similarity exceeds the matching threshold of the corresponding query area, determine the template feature identifier with successful matching, and extract the query rule set corresponding to the template feature identifier.
[0024] In an alternative embodiment,
[0025] Add the query feature matrix to the data distribution prediction network, predict the spatio-temporal distribution law according to the historical access records to obtain the data distribution feature map, construct a data association model and generate a probability distribution table by calculating the association probability values between fields, construct a multi-layer composite index based on the probability distribution table and construct an index structure tree in combination with the information gain ratio, calculate the access timing features and construct a time decay calculation model to obtain a dynamic index adjustment scheme, add the execution paths in the query rule set to the path optimization network based on the dynamic index adjustment scheme, determine the optimal execution path in combination with the path search algorithm and generate a task execution plan, schedule the database executor according to the task execution plan and obtain the query result stream for real-time analysis, calculate the data value coefficient and write it into the multi-level cache system, including:
[0026] Add the query feature matrix to the data distribution prediction network, and perform feature extraction using the main branch of time series feature extraction and the auxiliary branch of spatial feature extraction. Extract access time, access object, access user, and access behavior features from the historical access records through the main branch of time series feature extraction. Capture short-term access features through the sliding convolutional layer, segment the access sequence and count the access metrics, extract periodic access patterns and trend access features through the recurrent processing unit. Construct an access object association graph through the auxiliary branch of spatial feature extraction, and extract node propagation features using the graph structure network. Integrate the time series features and the spatial features to generate a data distribution feature map;
[0027] Construct a data association model, divide it into an attribute association layer, a record association layer, and a table association layer for data association analysis. Calculate the correlation coefficient of numerical fields through feature correlation analysis in the attribute association layer, calculate the association value of categorical fields through information entropy analysis, and generate a field association matrix. Extract multi-field combination features using a pattern mining method in the record association layer, identify high-frequency combination patterns and key combination rules. Construct a data lineage graph based on the field reference relationship in the table association layer, analyze the data flow path and data dependency intensity, and generate a global data dependency graph. Integrate the field association matrix, the key combination rules, and the global data dependency graph to calculate the association probability values between fields and generate a probability distribution table;
[0028] Construct a multi-layer composite index based on the probability distribution table, divide the index into a global index layer, a partition index layer, and a local index layer. Select the main index field through information gain ratio calculation in the global index layer, construct a composite index based on the pre-determined high-frequency access field combination. Determine the partition boundary based on the data distribution characteristics in the partition index layer, partition the data and establish a sub-index. Calculate the access heat of the data partition in the local index layer and construct an index structure tree;
[0029] Calculate access timing features based on the composite index, the sub-index, and the index structure tree, construct a time decay calculation model, divide multiple time windows, set a time decay function for the accessed data within each time window, calculate the weights of the accessed data, count the index usage frequency, index hit rate, and index maintenance overhead, generate index evaluation metrics, and when the index evaluation metrics are lower than the threshold, trigger an index optimization operation to generate a dynamic index adjustment plan;
[0030] Add the execution paths in the query rule set to the path optimization network based on the dynamic index adjustment plan, divide into an operator layer, a resource layer, and a scheduling layer for path optimization. In the operator layer, decompose the query operation into basic operators and construct an operator dependency graph. In the resource layer, monitor the system resource status and generate a resource status vector. In the scheduling layer, combine a path search algorithm to identify parallelizable operator groups, evaluate the cost of the execution path, determine the optimal execution path, and generate a task execution plan;
[0031] Schedule the database executor according to the task execution plan, obtain the query result stream for real-time analysis, calculate the data value coefficient, write the data value coefficient into a multi-level cache system, divide into a hot cache layer, a warm data cache layer, and a cold data cache layer for data storage, and perform hierarchical storage of the data based on the data value coefficient.
[0032] In an alternative embodiment,
[0033] Calculating access timing features based on the composite index, the sub-index, and the index structure tree, constructing a time decay calculation model, dividing multiple time windows, setting a time decay function for the accessed data within each time window, calculating the weights of the accessed data, counting the index usage frequency, index hit rate, and index maintenance overhead, generating index evaluation metrics, and when the index evaluation metrics are lower than the threshold, triggering an index optimization operation to generate a dynamic index adjustment plan includes:
[0034] Collect the access logs of the composite index, the sub-index, and the index structure tree within a preset time period and extract the access time points and operation instructions in the access logs. Obtain an access frequency sequence by statistically counting the access time points at small time intervals, obtain an operation count sequence by statistically counting the operation instructions according to the instruction type. Perform a sliding window calculation on the access frequency sequence to extract the fluctuation value and the period value, perform pattern recognition on the operation count sequence to extract the behavior feature value, and calculate the access timing features based on the composite index, the sub-index, and the index structure tree;
[0035] Divide the access timing characteristics into multiple time windows at time intervals of 1 hour, 24 hours, and 168 hours, and construct a time decay calculation model to obtain a time trend component representing the change in access volume, a periodic fluctuation component representing repeated access, and a random fluctuation component representing burst access;
[0036] Execute the time decay function setting for the access data within each time window and obtain the time interval and access frequency distribution characteristics of the access data within each time window. Determine the function type of the time decay function based on the time interval and access frequency distribution characteristics and set the parameters of the time decay function according to the time span of each time window. Multiply the original weight of the access data within each time window by the calculation result of the corresponding time decay function to obtain the time-weighted access volume, and combine the time-weighted access volume with the access operation type to calculate the access data weight;
[0037] Statistically calculate the index usage frequency, index hit rate, and index maintenance overhead. Convert the index usage frequency, the index hit rate, and the index maintenance overhead to between 0 and 1 respectively according to preset rules. Multiply the converted metrics by the preset metric weights and sum them to obtain an index evaluation metric. Compare the index evaluation metric with a preset performance evaluation threshold. When the index evaluation metric is lower than the performance evaluation threshold, trigger an index optimization operation;
[0038] Generate a dynamic index adjustment plan and statistically calculate the number of co-accesses between single-column index pairs. When the number of co-accesses exceeds a preset co-access threshold, generate an index merge instruction to merge the single-column indexes with a co-access relationship to construct a composite index. Statistically calculate the access volume of the composite index. When the access volume is lower than a preset access volume threshold for 30 consecutive days, generate an index split instruction to split the composite index into multiple single-column indexes. Statistically calculate the number of days the index has not been accessed and the resource occupancy. When the number of days the index has not been accessed exceeds 90 days and the resource occupancy exceeds a preset resource threshold, generate an index deletion instruction to delete the index and release the storage space.
[0039] In an alternative embodiment,
[0040] Collect the execution performance data of user data query requests and add it to the anomaly recognition model. Combine clustering operations to identify anomaly data points and generate an anomaly mark set. Calculate the deviation degree of performance metrics based on an anomaly detection model based on a recurrent neural network and combine a causal inference network to determine the temporal dependence relationship, generate a bottleneck location map. Generate a performance score based on the bottleneck location map and the deviation degree of performance metrics and generate a data value prediction sequence according to the data value coefficient. Combine a resource scheduling network and a competitive learning algorithm to determine the load balancing coefficient of the storage area, generate a preloading strategy, and perform performance verification through a Monte Carlo simulation model and select the preloading strategy with the highest confidence as the optimal strategy, including:
[0041] Collect the execution performance data of the user data query request, record the application layer execution response time, the system layer resource usage status, and the network layer transmission metrics, collect and organize them into a time series data stream at preset time intervals, calculate the weighted coefficient matrix based on the historical data sequence using the exponentially weighted moving average algorithm, multiply the weighted coefficient matrix by the time series data stream to obtain a smoothed data sequence, use the linear interpolation algorithm to fill in the missing data through the weighted combination of adjacent valid data points, subtract the mean from the data sequence and divide by the standard deviation to obtain the normalized data and add it to the anomaly recognition model;
[0042] Calculate the local density value by calculating the number of neighboring points within a specified radius for each data point in the normalized data, construct a distance matrix to calculate the relative distance between data points, multiply the normalized local density value and the relative distance to obtain the anomaly degree matrix, screen out the anomaly data points from the anomaly degree matrix based on a set threshold, calculate the time series autocorrelation coefficient for the anomaly data points and verify their duration to generate an anomaly mark set;
[0043] For the anomaly mark set, calculate the deviation degree of the performance metric based on the anomaly detection model of the recurrent neural network, extract features from the performance metric sequence using the input gate, forget gate, and output gate of the long short-term memory network to obtain a feature sequence, construct a gated recurrent unit to calculate the hidden state sequence of the feature sequence, input the hidden state sequence into the conditional random field model to obtain the state transition probability matrix, construct a causal inference network based on the state transition probability matrix to calculate the conditional probability distribution between metrics, and generate a time series dependency graph;
[0044] Construct a bottleneck location map based on the time series dependency graph, convert the system components in the bottleneck location map into vector representations through random walk in the graph embedding algorithm, calculate the component similarity matrix based on the vector representations, apply the community discovery algorithm to identify the component subgraphs with close associations, construct a performance score by combining the deviation degree of the performance metric, and generate a data value prediction sequence based on the performance score and the data value coefficient;
[0045] Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network to predict the resource requirements, construct a state-action value function based on the resource requirements, iteratively update the action strategy through the reinforcement learning algorithm, calculate the load variance of each storage area in combination with the resource scheduling network, and obtain the load balancing coefficient by searching for an equilibrium scheme that minimizes the load variance according to the genetic algorithm's crossover and mutation operations and competitive learning;
[0046] Based on the load balancing coefficient, construct a policy tree according to Monte Carlo tree search and select expansion nodes through the Upper Confidence Bound (UCB) algorithm. Use a random forest to extract features from the historical load sequence and predict the fluctuation range. Set the search boundary of the particle swarm based on the fluctuation range, and iteratively optimize the policy parameters through the velocity-position update formula. Calculate the policy confidence using support vector regression, and select the preloading policy with the highest confidence as the optimal policy.
[0047] In an alternative embodiment,
[0048] Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network to predict resource requirements. Based on the resource requirements, construct a state-action value function, and iteratively update the action policy through a reinforcement learning algorithm. Combine a resource scheduling network to calculate the load variance of each storage area, and obtain the load balancing coefficient according to the genetic algorithm's crossover and mutation operations and competitive learning to search for an equilibrium solution that minimizes the load variance, including:
[0049] Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network. The hidden layer of the deep neural network introduces a non-linear transformation through the rectified linear unit function and sequentially reduces the number of neurons to achieve feature dimension compression. Use the Adaptive Moment Estimation (Adam) optimizer to update the network parameters, eliminate feature scale differences through a batch normalization layer, and use a dropout mechanism to prevent overfitting, and output the resource requirements.
[0050] Based on the resource requirements, construct a state-action value function. The state space of the state-action value function includes the current allocation ratio of various resources, the length of the task waiting queue, and the load level of the storage area. The action space of the state-action value function includes proportionally adjusting resource quotas, reallocating tasks, and migrating data blocks. Construct an experience replay buffer to store historical decision samples. The historical decision samples include the system state before executing an action, the specific action content, the system benefits obtained, and the state transition results after execution.
[0051] Iteratively update the action policy based on the historical decision samples through a reinforcement learning algorithm. Randomly sample the historical decision samples for parameter update in each iteration, and introduce a target network to synchronize the main network parameters to the target network regularly.
[0052] Combine a resource scheduling network to calculate the load variance of each storage area. The underlying physical resource pool of the resource scheduling network is responsible for managing hardware devices and recording the real-time status of the devices. The middle-layer virtual resource manager of the resource scheduling network virtualizes the resource pool and provides a unified resource view. The upper-layer load balancing scheduler of the resource scheduling network calculates the dispersion degree of the current load of each storage area from the global mean based on the real-time monitoring data of the storage area to obtain the load variance.
[0053] According to the crossover and mutation operations of the genetic algorithm and competitive learning, the mapping relationship between data blocks and storage areas is defined through chromosome coding. The gene position values of the chromosome coding represent the target storage area identifiers. Through the selection operation, the allocation scheme with the highest load balancing performance is retained. Through the crossover and mutation operations, partial data block allocation information is exchanged between parental chromosomes and the storage locations of multiple data blocks are randomly changed. Competitive learning is used to determine the locally optimal individual to obtain the load balancing coefficient.
[0054] In the second aspect of the embodiments of the present invention, a data query parsing result optimization system based on a cache is provided, including:
[0055] A first unit, configured to receive a user data query request and extract an original query statement, perform word segmentation processing on the original query statement to generate a query token set and add it to a semantic understanding model to extract an association relationship, obtain a first query semantic network, identify path information of the first query semantic network through an intent recognition model to generate a second query semantic network, construct a query feature matrix based on the second query semantic network and add it to a feature combination model, generate fusion feature data by calculating the association weight coefficient between matrices, extract high-dimensional feature information from the fusion feature data through a deep residual network and generate a query feature identifier, calculate the similarity with a pre-set template feature identifier and generate a similarity score table, perform hierarchical matching and determine a set of query rules corresponding to the successfully matched template feature identifier;
[0056] A second unit, configured to add the query feature matrix to a data distribution prediction network, predict the spatio-temporal distribution law according to historical access records to obtain a data distribution feature map, construct a data association model and generate a probability distribution table by calculating the association probability value between fields, construct a multi-layer combined index based on the probability distribution table and combine it with the information gain ratio to construct an index structure tree, calculate the access timing feature and construct a time decay calculation model to obtain a dynamic index adjustment scheme, add the execution path in the set of query rules to a path optimization network based on the dynamic index adjustment scheme, determine the optimal execution path in combination with a path search algorithm and generate a task execution scheme, schedule a database executor according to the task execution scheme and obtain a query result stream for real-time analysis, calculate a data value coefficient and write it into a multi-level cache system;
[0057] The third unit is used to collect the execution performance data of user data query requests and add it to the anomaly recognition model, identify anomaly data points through clustering operations and generate an anomaly label set, calculate the deviation degree of performance metrics based on a recurrent neural network-based anomaly detection model for anomaly causes and determine the temporal dependence relationship in combination with a causal inference network to generate a bottleneck location map, generate a performance score based on the bottleneck location map and the deviation degree of performance metrics, generate a data value prediction sequence according to the data value coefficient, determine the load balancing coefficient of the storage area in combination with a resource scheduling network and a competitive learning algorithm to generate a preloading strategy, perform performance verification through a Monte Carlo simulation model and select the preloading strategy with the highest confidence as the optimal strategy.
[0058] In the third aspect of the embodiments of the present invention,
[0059] a kind of electronic device is provided, including:
[0060] a processor;
[0061] a memory for storing instructions executable by the processor;
[0062] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0063] In the fourth aspect of the embodiments of the present invention,
[0064] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0065] In the present invention, through technical means such as semantic understanding, feature combination, and deep learning, the user's query intention is accurately understood, and based on a preset template, efficient matching is performed to quickly determine the query rule. In combination with data distribution prediction and dynamic index adjustment, the database execution path is optimized, thereby shortening the query response time and improving the query efficiency. According to the data value coefficient and the user access pattern, high-value data is cached in a multi-level cache system, and in combination with anomaly detection and preloading strategies, the cache content is dynamically adjusted to improve the cache hit rate, reduce the database access pressure, further improve the system performance, and based on performance data analysis and bottleneck location, in combination with a resource scheduling network and a competitive learning algorithm, the load balancing of the storage area is achieved, avoiding resource competition and congestion, maximizing the use of system resources, and ensuring the stable operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a schematic flowchart of the method for optimizing the data query parsing result based on caching in the embodiments of the present invention;
[0067] Figure 2This is a schematic structural diagram of a system for optimizing data query parsing results based on a cache according to an embodiment of the present invention. Detailed implementation manners
[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0069] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0070] Figure 1 This is a schematic flowchart of a method for optimizing data query parsing results based on a cache according to an embodiment of the present invention. As Figure 1 shown, the method includes:
[0071] S1. Receive a user data query request and extract an original query statement, perform word segmentation processing on the original query statement to generate a query token set and add it to a semantic understanding model to extract association relationships, obtain a first query semantic network, identify path information of the first query semantic network through an intention recognition model to generate a second query semantic network, construct a query feature matrix based on the second query semantic network and add it to a feature combination model, generate fusion feature data by calculating the association weight coefficient between matrices, extract high-dimensional feature information from the fusion feature data through a deep residual network and generate a query feature identifier, calculate the similarity with a pre-set template feature identifier and generate a similarity score table, and perform hierarchical matching and determine a set of query rules corresponding to the successfully matched template feature identifier;
[0072] The original query statement refers to the unprocessed query text input by the user, which usually contains some sentences or phrases in natural language and is the input of the retrieval system. The word segmentation process is the process of dividing the text in the query statement into individual words or phrases. The query token set refers to the set of words obtained by performing word segmentation on the query statement. A token is the basic unit in a query and is usually used for subsequent indexing, matching, and calculation. The association weight coefficient refers to the coefficient used to measure the degree of association between a query token and a target document in information retrieval, and is usually determined according to word frequency, inverse document frequency, word vectors, or other similarity calculation methods. The deep residual network is a deep neural network structure that solves the problem of gradient disappearance in deep networks by introducing residual connections (skip connections). The high-dimensional feature information refers to the high-dimensional feature vectors extracted from the query through a machine learning model or feature extraction method. The query feature identifier refers to the marker or encoding used to identify and represent the features of the query statement in an information retrieval system. Hierarchical matching refers to the process of performing matching at different levels on a query according to a hierarchical structure in information retrieval or natural language processing tasks.
[0073] In an alternative embodiment,
[0074] Receiving a user data query request and extracting the original query statement, performing word segmentation processing on the original query statement to generate a query token set and adding it to a semantic understanding model to extract association relationships, obtaining a first query semantic network, and identifying path information of the first query semantic network through an intent recognition model to generate a second query semantic network includes:
[0075] Receiving a user data query request, extracting the original query statement in the user data query request, constructing a token recognition tree, setting the node attributes of the token recognition tree to token content, part of speech, and position offset, performing forward maximum matching starting from the beginning of the original query statement to find the longest matching token in the dictionary, performing backward maximum matching starting from the end of the original query statement to find the longest matching token in the dictionary, comparing the word segmentation results of the forward maximum matching and the backward maximum matching, selecting the word segmentation result with the fewest number of tokens when the number of tokens is different, and selecting the word segmentation result with the fewest single-character words when the number of tokens is the same to generate a query token set;
[0076] Constructing a deep bidirectional long short-term memory network, inputting the character vector sequence of the query token set into the input layer of the deep bidirectional long short-term memory network, extracting context features through multiple bidirectional long short-term memory layers of the deep bidirectional long short-term memory network, inputting the extracted context features into a conditional random field for part-of-speech tagging to identify organizational words, time words, measurement words, numerical words, object words, statistical index words, and judgment words, and marking the part-of-speech attributes and semantic roles of the query token set.
[0077] Construct an eight - head self - attention model. The first attention head captures temporal associations, the second attention head captures conditional associations, the third attention head captures object associations, the fourth attention head captures metric associations, the fifth attention head captures statistical associations, the sixth attention head captures subject associations, the seventh attention head captures attribution associations, and the eighth attention head captures combination associations. Input the set of labeled query tokens into the eight - head self - attention model to calculate the attention scores between tokens.
[0078] Construct a first query semantic network based on the attention scores. Set the tokens as the nodes of the first query semantic network, set the semantic associations as the edges of the first query semantic network, and set the attention scores as the weights of the edges. Set the graph structure pruning threshold and delete the edges with weights lower than the graph structure pruning threshold to extract the main query path and the branch query paths.
[0079] Construct a graph attention model. Input the first query semantic network into the graph attention layer of the graph attention model. Calculate the attention scores between a node and its adjacent nodes based on the spatial attention mechanism to update the node semantic representations. Input the updated node semantic representations into the path extraction layer of the graph attention model to extract the key query path based on the depth - first search algorithm, and mark the attention scores of the nodes on the path to generate a second query semantic network.
[0080] The token recognition tree is a tree - like structure used to represent the tokens and their hierarchical relationships in a query statement. The forward maximum matching is a commonly used word segmentation algorithm. Based on the principle of dictionary matching, it matches the longest token word - by - word from the start position of the query statement until no match can be found. The longest - matched token refers to the word with the longest length found during the forward maximum matching process. The deep bidirectional long short - term memory network is a deep neural network model based on long short - term memory units, which can capture forward and backward time - dependent relationships in sequence data. The key query path refers to the optimal path that can effectively match the target information or obtain the query result during the query processing.
[0081] Receive the data query request submitted by the user. For example, the query statement entered by the user in the search box: "Mobile phone sales in the first quarter of 2023". Extract the original query statement: "Mobile phone sales in the first quarter of 2023" from this user data query request.
[0082] Construct a token recognition tree. Each node of this token recognition tree contains three attributes: token content, part of speech, and position offset. For example, for the token "2023", its token content is "2023", the part of speech is "time word", and the position offset is 0. Then, perform forward maximum matching and backward maximum matching on the original query statement. Forward maximum matching starts from the beginning of the sentence and looks for the longest matching token in the dictionary. For example, for "Mobile phone sales in the first quarter of 2023", starting from "2", the longest match in the dictionary is "2023". Backward maximum matching starts from the end of the sentence and looks for the longest matching token in the dictionary. For example, starting from "quantity", the longest match in the dictionary is "sales volume". Suppose the forward maximum matching results are: "2023", "first quarter", "mobile phone", "sales volume", and the backward maximum matching results are: "2023", "year", "first quarter", "mobile phone", "sales volume". Comparing the two word segmentation results, the number of tokens is the same, but the forward maximum matching result has fewer single-character words. Therefore, select the forward maximum matching result to generate the query token set {"2023", "first quarter", "mobile phone", "sales volume"}.
[0083] Construct a deep bidirectional long short-term memory network. Input the character vector sequence of the query token set, such as "2", "0", "2", "3", "year", "first", "quarter", "mobile phone", "sales volume", into the input layer of the network. The network extracts context features through multiple layers of bidirectional long short-term memory layers, such as the combined features of "2023" and "year". Input the extracted context features into a conditional random field for part-of-speech tagging. Identify that "2023" is a time word, "first quarter" is also a time word, "mobile phone" is an object word, and "sales volume" is a statistical indicator word. Mark the part-of-speech attributes and semantic roles of the query token set. Finally, obtain {"2023", time word}, {"first quarter", time word}, {"mobile phone", object word}, {"sales volume", statistical indicator word}.
[0084] Construct an eight-head self-attention model. Each attention head captures different semantic associations. For example, the first attention head captures temporal associations and focuses on the relationship between "2023" and "first quarter"; the second attention head captures conditional associations; the third attention head captures object associations and focuses on the relationship between "mobile phone" and "sales volume", and so on. Input the labeled query token set into the eight-head self-attention model to calculate the attention scores between tokens. For example, the attention scores between "2023" and "first quarter" are relatively high, and the attention scores between "mobile phone" and "sales volume" are also relatively high.
[0085] Based on the calculated attention scores, construct the first query semantic network. The tokens serve as the nodes of the network, the semantic associations as the edges of the network, and the attention scores as the weights of the edges. For example, there is an edge between "2023" and "the first quarter" with a relatively high weight. There is also an edge between "mobile phone" and "sales volume" with a relatively high weight. Set a graph structure pruning threshold to delete the edges with weights lower than the threshold. For example, if the attention score between "2023" and "mobile phone" is very low, then delete the edge between them. Extract the main query path and the branch query paths. For example, "2023" - "the first quarter" - "mobile phone" - "sales volume" is the main query path.
[0086] Construct a graph attention model. Input the first query semantic network into the graph attention layer of the graph attention model. Calculate the attention scores between a node and its adjacent nodes based on the spatial attention mechanism, and update the node semantic representations. Input the updated node semantic representations into the path extraction layer, extract the key query paths based on the depth-first search algorithm, mark the attention scores of the nodes on the paths, and generate the second query semantic network. For example, the finally extracted key query path is "2023 first quarter mobile phone sales volume", and the attention scores of each node are marked.
[0087] In this embodiment, through the multi-layer semantic understanding model and the graph attention model, the user's query intention can be understood more accurately, thereby improving the accuracy of the query results. Through graph structure pruning and path extraction, the key query paths can be quickly located, unnecessary calculations can be reduced, and thus the query efficiency can be improved. It can handle various complex query statements, including ambiguous queries and long-tail queries, improving the robustness of the query.
[0088] In an alternative implementation
[0089] Based on the second query semantic network, construct a query feature matrix and add it to the feature combination model. Generate fused feature data by calculating the correlation weight coefficients between the matrices. Extract high-dimensional feature information from the fused feature data through a deep residual network and generate a query feature identifier. Calculate the similarity with a pre-set template feature identifier and generate a similarity score table. Perform hierarchical matching and determine the set of query rules corresponding to the successfully matched template feature identifiers, including:
[0090] Based on the second query semantic network, construct a query feature matrix, extract the temporal association features, numerical comparison features, and data type features between the nodes in the query semantic network, and construct a query feature matrix including dimensions of table association, field correlation, filtering condition, sorting rule, and time range;
[0091] Add the query feature matrix to the feature combination model, set a matrix correlation calculation unit in the feature combination model, calculate the correlation weight coefficients between matrices through the matrix correlation calculation unit, and according to the correlation weight coefficients, retain and transmit the feature information with the correlation weight coefficient higher than the weight threshold, and attenuate and filter the feature information with the correlation weight coefficient lower than the weight threshold to generate fused feature data;
[0092] Construct a multi-scale residual network, which includes a feature decomposition unit, a feature reconstruction unit and a feature optimization unit. The feature decomposition unit decomposes the fused feature data into multi-scale feature components through wavelet transform. The feature reconstruction unit sets an adaptive residual module to independently reconstruct the feature components of different scales. The feature optimization unit enhances the feature expression through a cross-scale feature interaction mechanism to generate a query feature identifier;
[0093] Construct a multi-view feature matching network, which includes a local feature matching sub-network, a global feature matching sub-network and a feature calibration sub-network. The local feature matching sub-network adopts a hierarchical attention mechanism to adaptively select key local features for matching. The global feature matching sub-network constructs a semantic consistency constraint in the feature space based on a contrastive learning strategy. The feature calibration sub-network optimizes the discriminability of the feature representation through the mutual information maximization criterion. Input the query feature identifier and the template feature identifier into the multi-view feature matching network to fuse the local structural similarity and the global semantic similarity to generate a similarity score table;
[0094] Perform hierarchical matching, dynamically adjust the matching thresholds of different query areas according to the system resource status. When the similarity exceeds the matching threshold of the corresponding query area, determine the template feature identifier with successful matching, and extract the query rule set corresponding to the template feature identifier.
[0095] The attenuation filtering is a method of information filtering based on time or other factors. By setting lower weights for older or irrelevant query information, or removing it from the query processing, it ensures that the query information processed by the system is the latest and most relevant. The independent feature reconstruction refers to decomposing or reconstructing features into a set of independent features that can better represent the latent structure of the data during the data processing. The mutual information maximization criterion is an optimization criterion aiming to maximize the mutual information between two sets of data, and is usually used in tasks such as feature selection and image alignment.
[0096] Construct a query feature matrix. Extract features from the second query semantic network, which includes the structured representation of the query statement, such as entities, relationships, and attributes. Specifically, extract the temporal association features between nodes, such as the order of appearance of two entities in time; extract numerical comparison features, such as the range of a certain attribute value; extract data type features, such as the type of a certain entity. Organize these features into a matrix, which contains multiple dimensions, such as the table association dimension, used to represent which tables the query involves and the association relationships between the tables; the field-related dimension, used to represent which fields the query involves and the relationships between the fields; the filter condition dimension, used to represent the filter conditions of the query, such as the conditions in the WHERE clause; the sorting rule dimension, used to represent the sorting rules of the query, such as the rules in the ORDER BY clause; the time range dimension, used to represent the time range of the query.
[0097] For example, for a query "Find users who have purchased product A in the past week", the table association dimension may include the "user table" and the "order table", the field-related dimension may include "user ID", "product ID", and "purchase time", the filter condition dimension may include "product ID = A", and the time range dimension may include "the past week".
[0098] Add the query feature matrix to the feature combination model. The model contains a matrix association calculation unit. This unit calculates the association weight coefficients between the query feature matrix and other feature matrices in the model. If the association between two matrices is strong, the weight coefficient is high; otherwise, it is low. Set a weight threshold. Feature information above the threshold is retained and passed to the next stage, while feature information below the threshold is attenuated or filtered out, and finally, fused feature data is generated. For example, if the "user table" and the "order table" often appear together in queries, the association weight coefficient will be very high.
[0099] Use a multi-scale residual network to extract high-dimensional feature information. The network includes a feature decomposition unit, a feature reconstruction unit, and a feature optimization unit. The feature decomposition unit uses wavelet transform to decompose the fused feature data into feature components at multiple scales. The feature reconstruction unit uses an adaptive residual module to perform independent feature reconstruction on the feature components at different scales. The feature optimization unit uses a cross-scale feature interaction mechanism to enhance the feature representation, and finally, a query feature identifier is generated. For example, wavelet transform can decompose the "purchase time" feature into feature components at different time granularities, such as days, weeks, months, etc.
[0100] Calculate the similarity using a multi - perspective feature matching network. This network includes a local feature matching sub - network, a global feature matching sub - network, and a feature calibration sub - network. The local feature matching sub - network uses a hierarchical attention mechanism to adaptively select key local features for matching. The global feature matching sub - network constructs semantic consistency constraints in the feature space based on a contrastive learning strategy. The feature calibration sub - network uses the mutual information maximization criterion to optimize the discriminability of the feature representation. Input the generated query feature identifier and the pre - set template feature identifier into the multi - perspective feature matching network, fuse the local structural similarity and the global semantic similarity, and generate a similarity score table.
[0101] Perform hierarchical matching. Dynamically adjust the matching thresholds of different query areas according to the system resource status. If the similarity between a query feature identifier and a certain template feature identifier exceeds the matching threshold of the corresponding query area, it is considered a successful match, and the query rule set corresponding to the template feature identifier is extracted. For example, if the current system load is high, the matching threshold can be increased to reduce the number of queries to be processed.
[0102] In this embodiment, through the multi - scale residual network and the multi - perspective feature matching network, more comprehensive and refined feature information can be extracted, thereby improving the matching accuracy between queries and rules. Through hierarchical matching and dynamically adjusting the matching threshold, the matching strategy can be optimized according to the system resource status, thereby improving the matching efficiency. Through the feature combination model and the cross - scale feature interaction mechanism, noise and abnormal data can be effectively processed, thereby enhancing the robustness of the system.
[0103] S2. Add the query feature matrix to the data distribution prediction network, predict the spatio - temporal distribution law according to the historical access records to obtain a data distribution feature map, construct a data association model and generate a probability distribution table by calculating the association probability values between fields, construct a multi - layer combined index based on the probability distribution table and construct an index structure tree in combination with the information benefit ratio, calculate the access timing features and construct a time decay calculation model to obtain a dynamic index adjustment scheme, add the execution paths in the query rule set to the path optimization network based on the dynamic index adjustment scheme, determine the optimal execution path in combination with the path search algorithm and generate a task execution plan, schedule the database executor according to the task execution plan and obtain the query result stream for real - time analysis, calculate the data value coefficient and write it into the multi - level cache system;
[0104] The spatio-temporal distribution law refers to the distribution characteristics of data or events in the time and space dimensions, which is usually used to analyze data patterns, predict trends, or optimize resource scheduling. The multi-layer combined index is a database optimization strategy that builds different levels of indexes on multiple fields to improve query efficiency. The information benefit ratio is an indicator to measure the contribution degree of data or features in the decision-making process, which is usually used for feature selection and data analysis. The index structure tree is a tree-shaped data structure used to accelerate data query. The path optimization network is a computational model for optimal path search, which is usually applied to scenarios such as traffic planning, logistics scheduling, and computer networks. The path search algorithm is a class of methods for finding the optimal path in a network or graph structure. The query result stream refers to the result set returned in a streaming manner when a database or search engine processes a query.
[0105] In an alternative embodiment,
[0106] Add the query feature matrix to the data distribution prediction network, predict the spatio-temporal distribution law based on historical access records to obtain a data distribution feature map, construct a data association model and generate a probability distribution table by calculating the association probability values between fields, construct a multi-layer combined index based on the probability distribution table and construct an index structure tree in combination with the information benefit ratio, calculate the access timing features and construct a time decay calculation model to obtain a dynamic index adjustment scheme, add the execution paths in the query rule set to the path optimization network based on the dynamic index adjustment scheme, determine the optimal execution path in combination with the path search algorithm and generate a task execution plan, schedule the database executor according to the task execution plan and obtain the query result stream for real-time analysis, calculate the data value coefficient and write it into the multi-level cache system, including:
[0107] Add the query feature matrix to the data distribution prediction network, and use the main branch of temporal feature extraction and the auxiliary branch of spatial feature extraction for feature extraction. Extract the access time, access object, access user, and access behavior features from the historical access records through the main branch of temporal feature extraction. Capture short-term access features through the sliding convolutional layer, segment the access sequence and count the access metrics. Extract the periodic access pattern and trend access features through the recurrent processing unit. Construct an access object association graph through the auxiliary branch of spatial feature extraction, and extract the node propagation features using the graph structure network. Fuse the temporal features and the spatial features to generate a data distribution feature map;
[0108] Build a data association model, divide it into an attribute association layer, a record association layer, and a table association layer for data association analysis. In the attribute association layer, calculate the correlation coefficient of numerical fields through feature correlation analysis, calculate the association value of categorical fields through information entropy analysis, and generate a field association matrix. In the record association layer, use a pattern mining method to extract multi-field combination features, identify high-frequency combination patterns and key combination rules. In the table association layer, build a data lineage graph based on field reference relationships, analyze the data flow path and data dependency strength, and generate a global data dependency graph. Integrate the field association matrix, the key combination rules, and the global data dependency graph to calculate the association probability value between fields and generate a probability distribution table;
[0109] Based on the probability distribution table, build a multi-level composite index, divide the index into a global index layer, a partition index layer, and a local index layer. In the global index layer, select the main index field through information gain ratio calculation, and build a composite index based on a pre-determined high-frequency access field combination. In the partition index layer, determine the partition boundary based on the data distribution characteristics, partition the data, and establish a sub-index. In the local index layer, calculate the access heat of data partitions, and build an index structure tree;
[0110] Based on the composite index, the sub-index, and the index structure tree, calculate the access time series characteristics and build a time decay calculation model. Divide it into multiple time windows, set a time decay function for the accessed data within each time window, calculate the access data weight, count the index usage frequency, index hit rate, and index maintenance overhead, and generate index evaluation metrics. When the index evaluation metrics are lower than the threshold, trigger an index optimization operation to generate a dynamic index adjustment plan;
[0111] Based on the dynamic index adjustment plan, add the execution paths in the query rule set to the path optimization network, divide it into an operator layer, a resource layer, and a scheduling layer for path optimization. In the operator layer, decompose the query operation into basic operators and build an operator dependency graph. In the resource layer, monitor the system resource status and generate a resource status vector. In the scheduling layer, combine a path search algorithm to identify parallelizable operator groups, evaluate the execution path cost, determine the optimal execution path, and generate a task execution plan;
[0112] Schedule the database executor according to the task execution plan, obtain the query result stream for real-time analysis, calculate the data value coefficient, write the data value coefficient into a multi-level cache system, divide it into a hot cache layer, a warm data cache layer, and a cold data cache layer for data storage, and perform hierarchical storage of data based on the data value coefficient.
[0113] The access object association graph is a graph structure used to represent data access patterns, where nodes represent data objects and edges represent access relationships. The graph structure network refers to a computing network based on the graph data structure, which is used to represent and process complex relational data. The field association matrix is a matrix structure used to represent the association relationships between fields in a database table. The data lineage graph is a graph structure used to trace data flow and dependency relationships, and is commonly used in data governance, data quality analysis, and ETL (Extract, Transform, Load) process management. The operator dependency graph is a graph structure that represents the dependency relationships between operators during the execution of a computing task or query, which helps to optimize query execution plans, parallel computing scheduling, and task splitting, thereby improving computing efficiency. The resource status vector is a vectorized representation of the current status of computing resources (such as CPU, memory, storage, network bandwidth, etc.). The group of operators that can be executed in parallel refers to a group of operators that can be executed in parallel during the optimization of a computing task or database query. The path cost refers to the total overhead required from the starting point to the ending point during path search or optimization.
[0114] Construct a data distribution prediction network to predict the spatio-temporal distribution law of data. This network includes a main branch for extracting temporal features and an auxiliary branch for extracting spatial features. The main branch for extracting temporal features extracts features such as access time, access object, access user, and access behavior from historical access records. For example, record information such as user A accessed product X at 10:00, user B accessed product Y at 10:05, and user A accessed product Z at 10:10. Use a sliding convolutional layer to capture short-term access features. For example, count the number of accesses to each product in the last 10 minutes. Segment the access sequence, for example, divide it by day or by week, and count access metrics, for example, calculate the average number of accesses to each product per day or per week. Use a recurrent processing unit to extract periodic access patterns and trend access features. For example, identify the weekly peak access period or long-term access trend of a product. The auxiliary branch for extracting spatial features constructs an access object association graph. For example, if a user often accesses product X and product Y simultaneously, an edge is established between product X and product Y. Use the graph structure network to extract node propagation features. For example, calculate the centrality of each product in the association graph. Finally, fuse the temporal features and spatial features to generate a data distribution feature map, which reflects the access probabilities of different data at different times and in different spaces.
[0115] Build a data association model, which includes an attribute association layer, a record association layer, and a table association layer. In the attribute association layer, calculate the correlation coefficient of numerical fields through feature correlation analysis. For example, calculate the correlation coefficient between product price and sales volume. Calculate the association value of categorical fields through information entropy analysis. For example, calculate the degree of association between product category and user gender. Generate a field association matrix to represent the association strength between different fields. In the record association layer, adopt a pattern mining method to extract multi-field combination features. For example, discover that users often purchase products X, Y, and Z at the same time. Identify high-frequency combination patterns and key combination rules. For example, determine that the combination purchase frequency of products X, Y, and Z is very high. In the table association layer, build a data lineage graph based on field reference relationships. For example, if a certain field in table A references a certain field in table B, then establish an edge between table A and table B. Analyze the data flow path and data dependency strength, generate a global data dependency graph, and integrate the field association matrix, key combination rules, and global data dependency graph to calculate the association probability value between fields, and generate a probability distribution table, which reflects the data association probability between different fields.
[0116] Build a multi-layer composite index based on the probability distribution table. Divide the index into a global index layer, a partition index layer, and a local index layer. In the global index layer, select the main index field through the information gain ratio calculation. For example, select the field with the highest access frequency and the largest discrimination degree as the main index field. Build a composite index based on a pre-determined high-frequency access field combination. For example, combine the product ID and user ID to build a composite index. In the partition index layer, determine the partition boundary based on the data distribution characteristics. For example, partition according to the geographical location of the product. Partition the data and establish a sub-index. For example, establish a sub-index for each geographical area. In the local index layer, calculate the access heat of the data partition. For example, count the access frequency of the data in each partition. Build an index structure tree to organize and manage indexes at different levels.
[0117] Calculate the access timing characteristics based on the composite index, sub-index, and index structure tree, and build a time decay calculation model. Divide multiple time windows. For example, divide a day into 24 hours. Set a time decay function for the accessed data in each time window. For example, the closer the access time, the higher the weight. Calculate the access data weight. For example, the weight of the data accessed in the last hour is 1, and the weight of the data accessed in the previous hour is 0.5, and so on. Count the index usage frequency, index hit rate, and index maintenance overhead, and generate index evaluation metrics. For example, count the number of usage times, hit times, and maintenance time of each index. When the index evaluation metric is lower than the threshold. For example, the index hit rate is lower than 80%, trigger an index optimization operation to generate a dynamic index adjustment plan. For example, rebuild the index or adjust the index structure.
[0118] Add the execution paths in the query rule set to the path optimization network based on the dynamic index adjustment scheme. Divide the operator layer, resource layer, and scheduling layer for path optimization. In the operator layer, decompose the query operation into basic operators. For example, decompose an SQL query into operations such as selection, projection, and join. Construct an operator dependency graph. For example, if the output of operator A is the input of operator B, establish an edge between operator A and operator B. In the resource layer, monitor the system resource status, such as CPU utilization, memory usage, and disk I / O. Generate a resource status vector. For example, represent the utilization of each resource as a vector. In the scheduling layer, combine the path search algorithm to identify operator groups that can be executed in parallel, such as identifying operators that can be executed simultaneously. Evaluate the cost of the execution path. For example, calculate the execution time and resource consumption of each execution path. Determine the optimal execution path and generate a task execution plan. Schedule the database executor according to the task execution plan, obtain the query result stream for real-time analysis, calculate the data value coefficient. For example, calculate the data value coefficient according to the importance of the query result. Write the data value coefficient into the multi-level cache system, divide the hot cache layer, warm data cache layer, and cold data cache layer for data storage, and perform hierarchical storage of data based on the data value coefficient. For example, store high-value data in the hot cache layer and low-value data in the cold data cache layer.
[0119] In this embodiment, through dynamic index adjustment and path optimization, the query response time can be effectively reduced and the query efficiency can be improved. Through data distribution prediction and cache mechanism, the access pressure on the database can be reduced and the system load can be lowered. Through parallel execution and resource scheduling, the system resources can be fully utilized and the resource utilization rate can be improved.
[0120] In an alternative embodiment,
[0121] Calculate the access timing characteristics based on the composite index, the sub-index, and the index structure tree, and construct a time decay calculation model. Divide multiple time windows, set a time decay function for the accessed data in each time window, calculate the access data weight, count the index usage frequency, index hit rate, and index maintenance overhead, generate index evaluation metrics. When the index evaluation metrics are lower than the threshold, trigger the index optimization operation to generate a dynamic index adjustment scheme, including:
[0122] Collect the access logs of the composite index, the sub-index, and the index structure tree within a preset time period, and extract the access time points and operation instructions in the access logs. Statistically obtain the access frequency sequence by the small time interval for the access time points, and statistically obtain the operation count sequence by the instruction type for the operation instructions. Perform a sliding window calculation on the access frequency sequence to extract the fluctuation value and the period value, perform pattern recognition on the operation count sequence to extract the behavior feature value, and calculate the access timing feature based on the composite index, the sub-index, and the index structure tree;
[0123] Divide the access timing feature into multiple time windows at time intervals of 1 hour, 24 hours, and 168 hours, and construct a time decay calculation model to obtain a time trend component representing the change in access volume, a periodic fluctuation component representing repeated access, and a random fluctuation component representing burst access;
[0124] Perform a time decay function setting on the access data within each time window and obtain the time interval and access frequency distribution characteristics of the access data within each time window. Determine the function type of the time decay function based on the time interval and access frequency distribution characteristics and set the parameters of the time decay function according to the time span of each time window. Multiply the original weight of the access data within each time window by the calculation result of the corresponding time decay function to obtain the time-weighted access volume, and calculate the access data weight by combining the time-weighted access volume and the access operation type;
[0125] Statistically obtain the index usage frequency, index hit rate, and index maintenance overhead. Convert the index usage frequency, the index hit rate, and the index maintenance overhead to between 0 and 1 respectively according to preset rules. Multiply the converted metrics by the preset metric weights and sum them to obtain the index evaluation metric. Compare the index evaluation metric with the preset performance evaluation threshold, and trigger an index optimization operation when the index evaluation metric is lower than the performance evaluation threshold;
[0126] Generate a dynamic index adjustment plan and statistically obtain the collaborative access times between single-column index pairs. When the collaborative access times exceed the preset collaborative threshold, generate an index merge instruction to merge the single-column indexes with a collaborative access relationship to construct a composite index. Statistically obtain the access volume of the composite index. When the access volume is lower than the preset access volume threshold for 30 consecutive days, generate an index split instruction to split the composite index into multiple single-column indexes. Statistically obtain the unaccessed days and resource occupancy of the index. When the unaccessed days exceed 90 days and the resource occupancy exceeds the preset resource threshold, generate an index deletion instruction to delete the index and release the storage space.
[0127] The time decay function is a function used to model the gradual reduction of the influence of data or events over time. The index evaluation metric is a metric used to measure the performance of a database index, such as index hit rate, query speedup ratio, storage overhead, etc. The collaborative access relationship refers to the access patterns and relationships among multiple users or processes when accessing a database or shared resources. The single-column index refers to an index created for a single column in a database.
[0128] Collect the access logs of the composite index, sub-index, and index structure tree within a preset time period (e.g., the past month). Extract the access time points and operation instructions from the access logs. For example, an access log records that a user executed a query operation on the composite index (A, B) at 10:30:00 on October 27, 2024. Count the access frequency sequence at small time intervals according to the access time points. For example, count the number of accesses to the index (A, B) within each hour. Count the operation count sequence according to the instruction type for the operation instructions. For example, count the number of query, insert, update, and delete operations. Perform a sliding window calculation on the access frequency sequence (e.g., the window size is 24 hours), and extract the fluctuation value (e.g., the difference in access frequency between adjacent time windows) and the periodic value (e.g., the periodicity of the access frequency sequence). Perform pattern recognition on the operation count sequence and extract the behavior feature values (e.g., the proportion of different operation types). Calculate the access timing features based on the composite index, sub-index, and index structure tree. For example, the usage frequency of a certain sub-index in the composite index query.
[0129] Divide the access timing features into multiple time windows at time intervals of 1 hour, 24 hours, and 168 hours. Construct a time decay calculation model that decomposes the access volume into a time trend component, a periodic fluctuation component, and a random fluctuation component. For example, the exponential smoothing method can be used to extract the time trend component, the Fourier transform to extract the periodic fluctuation component, and the ARIMA model to extract the random fluctuation component.
[0130] Execute the time decay function setting for the access data within each time window. Obtain the time intervals and access frequency distribution characteristics of the access data within each time window. For example, count the average time interval between two accesses and the histogram of access frequencies within each time window. Determine the function type of the time decay function based on these characteristics. For example, if the access frequency follows a Poisson distribution, an exponential decay function can be selected. Set the parameters of the time decay function according to the time span of each time window. For example, the longer the time window, the faster the decay rate. Multiply the original weight of the access data within each time window (e.g., the number of accesses) by the calculation result of the corresponding time decay function to obtain the time-weighted access volume. Combine the time-weighted access volume with the access operation type to calculate the access data weight. For example, the weight of a query operation is higher than that of an update operation. Suppose within a certain time window, the access volume of index (A) is 1000 times and the calculation result of the time decay function is 0.8, then the time-weighted access volume is 800.
[0131] Count the index usage frequency, index hit rate, and index maintenance overhead. Convert these metrics to between 0 and 1 respectively according to preset rules. For example, the higher the usage frequency, the closer the converted value is to 1. Multiply the converted metrics by the preset metric weights and sum them to obtain the index evaluation metric. For example, assume the weights of index usage frequency, hit rate, and maintenance overhead are 0.5, 0.3, and 0.2 respectively, and the converted values are 0.9, 0.8, and 0.7 respectively, then the index evaluation metric is 0.5 * 0.9 + 0.3 * 0.8 + 0.2 * 0.7 = 0.85. Compare the index evaluation metric with the preset performance evaluation threshold (e.g., 0.8). When the index evaluation metric is lower than the performance evaluation threshold, trigger the index optimization operation.
[0132] Generate a dynamic index adjustment plan. Count the co-access times between single-column index pairs. For example, count the number of times index (A) and index (B) are used in the same query. When the co-access times exceed the preset co-threshold (e.g., 1000 times), generate an index merge instruction to merge the single-column indexes with co-access relationships to construct a composite index (A, B). Count the access volume of the composite index. When the access volume is lower than the preset access volume threshold (e.g., 100 times) for 30 consecutive days, generate an index split instruction to split the composite index (A, B) into multiple single-column indexes (A) and (B). Count the number of days the index has not been accessed and the resource occupancy. When the number of days without access exceeds 90 days and the resource occupancy exceeds the preset resource threshold (e.g., 1GB), generate an index deletion instruction to delete the index and release the storage space.
[0133] In this embodiment, by dynamically adjusting the index structure, the index hit rate is improved, thereby accelerating the query speed. By deleting unnecessary indexes, storage space is released, and the storage cost is reduced. The index is dynamically adjusted according to changes in the access pattern, so that the index always remains in the best state to meet the needs of business development.
[0134] S3. Collect the execution performance data of the user data query request and add it to the anomaly recognition model. Combine clustering operations to identify anomaly data points and generate an anomaly label set. Calculate the deviation degree of the performance index based on the anomaly detection model of the recurrent neural network and determine the time-series dependence relationship in combination with the causal inference network to generate a bottleneck location map. Generate a performance score based on the bottleneck location map and the deviation degree of the performance index, and generate a data value prediction sequence according to the data value coefficient. Combine the resource scheduling network and the competitive learning algorithm to determine the load balancing coefficient of the storage area and generate a preloading strategy. Perform performance verification through the Monte Carlo simulation model and select the preloading strategy with the highest confidence as the optimal strategy.
[0135] The clustering operation refers to dividing the objects in the data set into multiple categories (or clusters) through a certain algorithm, so that the objects in the same cluster are as similar as possible in certain features, while the objects in different clusters are as different as possible. The deviation degree of the performance index is a measure used to measure the difference between the actual system performance and the expected target. The causal inference network is a graph model based on causal relationships, used to describe and reason about the causal relationships between different variables. The bottleneck location map is a visual map obtained by analyzing each link in the system or process to find performance bottlenecks or limiting factors. The data value prediction sequence refers to predicting the future data value or trend in time-series data based on the existing data. The resource scheduling network refers to a network model used to allocate and manage limited computing or storage resources. The competitive learning algorithm is an algorithm based on group or multi-agent learning, in which multiple learning units optimize their respective strategies through competition or cooperation. The Monte Carlo simulation model is a numerical simulation method based on random sampling, used to solve the simulation and optimization problems of complex systems.
[0136] In an alternative embodiment,
[0137] Collect the execution performance data of the user data query request and add it to the anomaly recognition model. Combine clustering operations to identify anomaly data points and generate an anomaly label set. Based on the anomaly cause, calculate the deviation degree of the performance index through the anomaly detection model based on the recurrent neural network and determine the temporal dependence relationship by combining the causal inference network to generate a bottleneck location map. Generate a performance score based on the bottleneck location map and the deviation degree of the performance index and generate a data value prediction sequence according to the data value coefficient. Combine the resource scheduling network and the competitive learning algorithm to determine the load balancing coefficient of the storage area, generate a preloading strategy, perform performance verification through the Monte Carlo simulation model, and select the preloading strategy with the highest confidence as the optimal strategy, including:
[0138] Collect the execution performance data of the user data query request, record the application layer execution response time, the system layer resource usage status, and the network layer transmission metrics. Regularly collect and organize them into a time series data stream at preset time intervals. Use the exponentially weighted moving average algorithm to calculate the weighted coefficient matrix based on the historical data sequence. Multiply the weighted coefficient matrix by the time series data stream to obtain the smoothed data sequence. Use the linear interpolation algorithm to fill in the missing data through the weighted combination of adjacent valid data points. Subtract the mean from the data sequence and divide by the standard deviation to obtain the normalized data and add it to the anomaly recognition model;
[0139] Calculate the number of neighboring points within a specified radius for each data point in the normalized data to obtain the local density value. Construct a distance matrix to calculate the relative distance between data points. Multiply the normalized local density value and the relative distance to obtain the anomaly degree matrix. Screen out the anomaly data points from the anomaly degree matrix based on the set threshold. Calculate the temporal autocorrelation coefficient of the anomaly data points and verify their duration to generate an anomaly label set;
[0140] For the anomaly label set, calculate the deviation degree of the performance index through the anomaly detection model based on the recurrent neural network. Use the input gate, forget gate, and output gate of the long short-term memory network to extract features from the performance index sequence to obtain a feature sequence. Construct a gated recurrent unit to calculate the hidden state sequence of the feature sequence. Input the hidden state sequence into the conditional random field model to obtain the state transition probability matrix. Based on the state transition probability matrix, construct a causal inference network to calculate the conditional probability distribution between indicators and generate a temporal dependence relationship graph;
[0141] Construct a bottleneck location map based on the temporal dependence relationship graph. Convert the system components in the bottleneck location map into vector representations through random walk in the graph embedding algorithm. Calculate the component similarity matrix based on the vector representations. Apply the community discovery algorithm to identify the component subgraphs with close associations. Combine the deviation degree of the performance index to construct a performance score. Generate a data value prediction sequence based on the performance score and the data value coefficient;
[0142] Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network to predict resource requirements, construct a state-action value function based on the resource requirements, iteratively update the action policy through a reinforcement learning algorithm, calculate the load variance of each storage area in combination with a resource scheduling network, and obtain the load balancing coefficient according to the crossover and mutation operations of the genetic algorithm and competitive learning to search for an equilibrium solution that minimizes the load variance;
[0143] Based on the load balancing coefficient, construct a policy tree according to Monte Carlo tree search and select expansion nodes through the Upper Confidence Bound (UCB) algorithm. Use a random forest to extract features from the historical load sequence and predict the fluctuation range. Set the search boundary of the particle swarm based on the fluctuation range, iteratively optimize the policy parameters through the velocity-position update formula, calculate the policy confidence using support vector regression, and select the preloading policy with the highest confidence as the optimal policy.
[0144] The linear interpolation algorithm is a mathematical method for estimating the values of new data points through known data points. The community discovery algorithm is an algorithm used to identify subsets with close connections (i.e., "communities") between nodes in graph or network data. The Support Vector Regression (SVR) is a regression analysis method based on the Support Vector Machine (SVM). By constructing a regression plane that maximizes the margin, it can handle non-linear regression problems and is commonly used in tasks such as time series prediction and stock price prediction.
[0145] Collect the execution performance data of user data query requests, including the application layer execution response time, system layer resource usage status (such as CPU utilization, memory usage, disk I / O), and network layer transmission metrics (such as network latency, packet loss rate). Data collection is performed regularly at a preset time interval (such as every minute), and the collected data is organized into a time series data stream. For example, collect the application response time at the minute level and continuously collect data for a week to form a time series containing 10,080 data points. To smooth data fluctuations, use the exponential weighted moving average algorithm to process the time series data stream. For example, set the decay factor to 0.9 and perform weighted averaging on the response time data for the past week to obtain a smoothed data sequence. For missing data that may exist in the data sequence, use the linear interpolation algorithm to fill it. For example, if the response time data for a certain minute is missing, use the weighted average of the response time data for the two minutes before and after it for filling. Finally, standardize the data sequence, that is, subtract the mean from the data sequence and divide by the standard deviation, and add the standardized data to the anomaly recognition model.
[0146] Perform anomaly detection on standardized data. Calculate the number of neighboring points within a specified radius for each data point to obtain the local density value. For example, set the radius to 1, count the number of other data points whose distance from each data point is less than or equal to 1, construct a distance matrix, and calculate the relative distance between data points. For example, use the Euclidean distance to calculate the distance between data points. Multiply the normalized local density value and relative distance respectively to obtain the anomaly degree matrix. Set a threshold, such as 0.8, and filter out anomaly data points from the anomaly degree matrix. For the filtered anomaly data points, calculate the time series autocorrelation coefficient and verify its duration. For example, if the anomaly data points continuously appear for more than 5 minutes, generate an anomaly marker set.
[0147] Analyze the anomaly marker set to determine the deviation degree of performance metrics and the time series dependence relationship. Use the anomaly detection model of the recurrent neural network to calculate the deviation degree of performance metrics. For example, use the long short-term memory network to extract features from the performance metric sequence, and input the extracted feature sequence into the gated recurrent unit to calculate the hidden state sequence. Input the hidden state sequence into the conditional random field model to obtain the state transition probability matrix. Based on the state transition probability matrix, construct a causal inference network, calculate the conditional probability distribution between metrics, and generate a time series dependence graph. For example, if there is a strong correlation between the increase in CPU utilization and the increase in response time, establish the corresponding connection in the time series dependence graph.
[0148] Construct a bottleneck location map based on the time series dependence graph. Through the graph embedding algorithm, such as the random walk algorithm, convert the system components in the bottleneck location map into vector representations. Calculate the component similarity matrix based on the vector representations. Apply the community discovery algorithm to identify the component subgraphs with close associations. For example, if the similarity between the database server, application server, and cache server is high, identify them as a component subgraph. Combine the deviation degree of performance metrics to construct a performance score. For example, if the deviation degree of the response time of a certain component subgraph is high, give it a lower performance score. Generate a data value prediction sequence based on the performance score and the data value coefficient. For example, if the access frequency of a certain data block is high and the performance score is high, predict that its data value is high.
[0149] Generate a preloading strategy. Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network to predict resource requirements. Construct a state-action value function based on the resource requirements. Iteratively update the action strategy through the reinforcement learning algorithm. For example, use the Q-learning algorithm to learn the optimal preloading strategy. Combine the resource scheduling network to calculate the load variance of each storage area. Search for an equilibrium scheme that minimizes the load variance according to the genetic algorithm and the competitive learning algorithm to obtain the load balancing coefficient. For example, by adjusting the allocation ratio of data blocks between different storage areas, minimize the load variance of each storage area.
[0150] Perform performance verification on the generated preloading strategy. Based on the load balancing coefficient, construct a strategy tree according to Monte Carlo tree search, and select the expansion node through the Upper Confidence Bound (UCB) algorithm. Use random forest to extract features from the historical load sequence and predict the fluctuation range. Set the search boundary of the particle swarm based on the fluctuation range. Iteratively optimize the strategy parameters through the velocity-position update formula. Calculate the strategy confidence using support vector regression. Select the preloading strategy with the highest confidence as the optimal strategy. For example, by simulating the system load conditions under different preloading strategies, select the strategy that can minimize the response time and has the highest confidence as the optimal strategy.
[0151] In this embodiment, by predicting the data value and optimizing the load balancing of the storage area, high-value data can be preloaded into the storage area with lower load, thereby reducing data access latency and improving query efficiency. It can dynamically adjust the resource scheduling strategy according to real-time performance data and predicted resource requirements, achieve reasonable allocation and utilization of resources, avoid resource waste and bottlenecks, and can adaptively adjust the preloading strategy according to the change of system load to ensure the stability and high performance of the system under different load conditions.
[0152] In an alternative embodiment,
[0153] Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network to predict resource requirements. Based on the resource requirements, construct a state-action value function, and iteratively update the action strategy through a reinforcement learning algorithm. Combine the resource scheduling network to calculate the load variance of each storage area and obtain the load balancing coefficient according to the genetic algorithm's crossover and mutation operations and competitive learning to search for an equilibrium scheme that minimizes the load variance, including:
[0154] Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network. The hidden layer of the deep neural network introduces a non-linear transformation through the rectified linear unit function and sequentially reduces the number of neurons to achieve feature dimension compression. Use the Adaptive Moment Estimation (Adam) optimizer to update the network parameters, eliminate the feature scale difference through the batch normalization layer, and use the dropout mechanism to prevent overfitting, and output the resource requirements;
[0155] Based on the resource requirements, construct a state-action value function. The state space of the state-action value function includes the current allocation ratio of various resources, the length of the task waiting queue, and the load level of the storage area. The action space of the state-action value function includes proportionally adjusting the resource quota, task reallocation, and data block migration. Construct an experience replay buffer to store historical decision samples. The historical decision samples include the system state before executing the action, the specific action content, the obtained system benefits, and the state transition result after execution;
[0156] Iteratively update the action policy based on the historical decision samples through a reinforcement learning algorithm. In each round of iteration, randomly sample the historical decision samples for parameter update, and introduce a target network to synchronize the main network parameters to the target network regularly;
[0157] Calculate the load variance of each storage area in combination with a resource scheduling network. The underlying physical resource pool of the resource scheduling network is responsible for managing hardware devices and recording the real-time status of the devices. The middle-layer virtual resource manager of the resource scheduling network realizes the virtualization of the resource pool and provides a unified resource view. The upper-layer load balancing scheduler of the resource scheduling network calculates the degree of dispersion between the current load of each storage area and the global mean based on the real-time monitoring data of the storage area to obtain the load variance;
[0158] Define the mapping relationship from data blocks to storage areas according to the crossover and mutation operations of the genetic algorithm and competitive learning. The gene position values of the chromosome encoding represent the target storage area identifier. Retain the allocation scheme with the highest load balancing performance through the selection operation. Exchange partial data block allocation information between the parent chromosomes and randomly change the storage locations of multiple data blocks through the crossover and mutation operations. Use competitive learning to determine the locally optimal individual to obtain the load balancing coefficient.
[0159] The rectified linear unit function is a commonly used activation function in neural networks, which directly maps negative inputs to zero and keeps positive inputs unchanged. The adaptive moment estimation optimizer is an optimization method for neural network training, which accelerates convergence by calculating an adaptive learning rate for each parameter. The dropout mechanism is a regularization method to prevent neural network overfitting. During training, randomly select some neurons to be "deactivated", that is, not participate in the calculation, forcing the network to learn more robust features. The chromosome encoding is an important concept in the genetic algorithm, which refers to representing the solution of the problem as a "chromosome" (i.e., a string or sequence), and then generating new solutions through operations such as crossover and mutation.
[0160] Predict the data value in the storage system. Considering the importance differences of different data, for example, some data needs to be accessed frequently, while others are rarely accessed, so it is necessary to evaluate the data value. Adopt the time series analysis method, and according to information such as the historical access frequency and modification time of the data, predict the access probability of the data in the next period of time, and use the prediction result as an evaluation index of the data value. For example, for a storage system containing 1000 data blocks, according to the access records in the past week, predict the access probability of each data block in the next week, ranging from 0 to 1.
[0161] Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network to predict resource requirements. The hidden layers of this deep neural network use the rectified linear unit function to introduce non-linear transformations and successively reduce the number of neurons to achieve feature dimension compression. The adaptive moment estimation optimizer is used to update the network parameters, and batch normalization layers are used to eliminate feature scale differences. At the same time, a dropout mechanism is used to prevent overfitting. The network inputs include the data value prediction sequence, CPU usage rate, memory usage rate, network bandwidth usage rate, and the load distribution of each storage area. The output is the CPU, memory, and network bandwidth resources required in the future for a period of time. For example, the input data value prediction sequence is the access probability of 1000 data blocks, the resource usage status is the current CPU usage rate of 70%, memory usage rate of 80%, network bandwidth usage rate of 50%, and the load distributions of 5 storage areas are 60%, 70%, 80%, 90%, and 50% respectively. The resource requirements output by the network are that the CPU resources required in the next hour increase by 20%, memory resources increase by 10%, and network bandwidth resources increase by 15%.
[0162] Construct a state-action value function based on the predicted resource requirements. The state space includes the current allocation ratios of various resources, the length of the task waiting queue, and the load levels of the storage areas. The action space includes proportionally adjusting resource quotas, reallocating tasks, and migrating data blocks. Construct an experience replay buffer to store historical decision samples, including the system state before performing an action, the specific action content, the system benefits obtained, and the state transition results after execution. For example, the state space can be defined as the CPU allocation ratio, memory allocation ratio, network bandwidth allocation ratio, the length of the task waiting queue, and the load levels of 5 storage areas. The action space can be defined as increasing / decreasing the CPU allocation ratio, increasing / decreasing the memory allocation ratio, increasing / decreasing the network bandwidth allocation ratio, reallocating tasks to different storage areas, and migrating data blocks from one storage area to another storage area.
[0163] Through a reinforcement learning algorithm, iteratively update the action policy based on historical decision samples. In each iteration, randomly sample historical decision samples for parameter update, and introduce a target network to synchronize the main network parameters to the target network regularly. For example, using the deep Q-learning algorithm, in each iteration, randomly extract a batch of samples from the experience replay buffer, calculate the target Q value, and update the network parameters according to the difference between the target Q value and the current Q value.
[0164] Calculate the load variance of each storage area in combination with the resource scheduling network. The underlying physical resource pool of the resource scheduling network is responsible for managing hardware devices and recording the real-time status of the devices. The middle-layer virtual resource manager realizes the virtualization of the resource pool and provides a unified resource view. The upper-layer load balancing scheduler calculates the degree of dispersion between the current load of each storage area and the global mean based on the real-time monitoring data of the storage area to obtain the load variance. For example, if the loads of 5 storage areas are 60%, 70%, 80%, 90%, and 50% respectively, and the global mean is 70%, then the load variance is the average value of (60 - 70)^2 + (70 - 70)^2 + (80 - 70)^2 + (90 - 70)^2 + (50 - 70)^2. According to the crossover and mutation operations of the genetic algorithm and competitive learning, define the mapping relationship between data blocks and storage areas through chromosome encoding. The value of the gene position of the chromosome encoding represents the target storage area identifier. Retain the allocation scheme with the highest load balancing performance through the selection operation. Exchange partial data block allocation information between parent chromosomes and randomly change the storage locations of multiple data blocks through the crossover and mutation operations. Use competitive learning to determine the locally optimal individual to obtain the load balancing coefficient. For example, a chromosome can represent the storage locations of 1000 data blocks, and each gene position represents the target storage area identifier of a data block, with a value range of 1 to 5.
[0165] In this embodiment, accurately predicting resource requirements through a deep neural network and dynamically adjusting the resource allocation strategy in combination with reinforcement learning can effectively improve resource utilization and avoid resource waste. Load balancing can avoid system bottlenecks caused by excessive load in individual storage areas, improve the stability and reliability of the system. Optimizing the storage locations of data blocks through the genetic algorithm can effectively reduce the load difference between storage areas, achieve load balancing, and improve the overall performance of the system.
[0166] Figure 2 It is a schematic structural diagram of the system for optimizing the data query parsing result based on the cache according to the embodiment of the present invention. As Figure 2 shown, the system includes:
[0167] The first unit is used to receive a user data query request, extract the original query statement, perform word segmentation processing on the original query statement to generate a query token set, add it to the semantic understanding model to extract the association relationship, obtain the first query semantic network, identify the path information of the first query semantic network through the intent recognition model to generate the second query semantic network, construct a query feature matrix based on the second query semantic network and add it to the feature combination model, generate fusion feature data by calculating the association weight coefficient between matrices, extract high-dimensional feature information from the fusion feature data through a deep residual network and generate a query feature identifier, calculate the similarity with a pre-set template feature identifier and generate a similarity score table, perform hierarchical matching, and determine the query rule set corresponding to the successfully matched template feature identifier;
[0168] A second unit, configured to add the query feature matrix to a data distribution prediction network, predict the spatio-temporal distribution law according to historical access records to obtain a data distribution feature map, construct a data association model and generate a probability distribution table by calculating the association probability values between fields, construct a multi-level combined index based on the probability distribution table and construct an index structure tree in combination with the information gain ratio, calculate the access timing features and construct a time decay calculation model to obtain a dynamic index adjustment scheme, add the execution paths in the query rule set to a path optimization network based on the dynamic index adjustment scheme, determine the optimal execution path in combination with a path search algorithm and generate a task execution scheme, schedule a database executor according to the task execution scheme and obtain a query result stream for real-time analysis, calculate a data value coefficient and write it into a multi-level cache system;
[0169] A third unit, configured to collect the execution performance data of a user data query request and add it to an anomaly recognition model, identify anomaly data points in combination with clustering operations and generate an anomaly label set, calculate the deviation degree of performance indicators based on a recurrent neural network-based anomaly detection model for anomaly causes and determine the temporal dependence relationship in combination with a causal inference network to generate a bottleneck location map, generate a performance score based on the bottleneck location map and the deviation degree of performance indicators and generate a data value prediction sequence according to the data value coefficient, determine the load balancing coefficient of a storage area in combination with a resource scheduling network and a competitive learning algorithm, generate a preloading strategy, perform performance verification through a Monte Carlo simulation model and select the preloading strategy with the highest confidence as the optimal strategy.
[0170] In a third aspect of the embodiments of the present invention,
[0171] provided is an electronic device, including:
[0172] a processor;
[0173] a memory for storing instructions executable by the processor;
[0174] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0175] In a fourth aspect of the embodiments of the present invention,
[0176] provided is a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0177] The present invention can be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.
[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for optimizing the parsing result of data query based on cache, characterized in that Including: Receiving a user data query request and extracting an original query statement, performing word segmentation processing on the original query statement to generate a query token set and adding it to a semantic understanding model to extract association relationships, obtaining a first query semantic network, identifying path information of the first query semantic network through an intent recognition model to generate a second query semantic network, constructing a query feature matrix based on the second query semantic network and adding it to a feature combination model, generating fused feature data by calculating the association weight coefficient between matrices, extracting high-dimensional feature information from the fused feature data through a deep residual network and generating a query feature identifier, calculating the similarity with a pre-set template feature identifier and generating a similarity score table, performing hierarchical matching and determining a set of query rules corresponding to the successfully matched template feature identifier; Adding the query feature matrix to a data distribution prediction network, predicting the spatio-temporal distribution law according to historical access records to obtain a data distribution feature map, constructing a data association model and generating a probability distribution table by calculating the association probability value between fields, constructing a multi-layer combined index based on the probability distribution table and combining the information benefit ratio to construct an index structure tree, calculating the access time sequence feature and constructing a time decay calculation model to obtain a dynamic index adjustment scheme, adding the execution path in the set of query rules to a path optimization network based on the dynamic index adjustment scheme, determining the optimal execution path by combining a path search algorithm and generating a task execution scheme, scheduling a database executor according to the task execution scheme and obtaining a query result stream for real-time analysis, calculating a data value coefficient and writing it into a multi-level cache system; Collecting the execution performance data of the user data query request and adding it to an anomaly recognition model, identifying anomaly data points by combining clustering operations and generating an anomaly marker set, calculating the deviation degree of performance indicators through an anomaly detection model based on a recurrent neural network and determining the time sequence dependence relationship by combining a causal inference network to generate a bottleneck location map, generating a performance score based on the bottleneck location map and the deviation degree of performance indicators and generating a data value prediction sequence according to the data value coefficient, determining the load balancing coefficient of the storage area by combining a resource scheduling network and a competitive learning algorithm to generate a preloading strategy, performing performance verification through a Monte Carlo simulation model and selecting the preloading strategy with the highest confidence as the optimal strategy.
2. The method according to claim 1, wherein Receiving a user data query request and extracting an original query statement, performing word segmentation processing on the original query statement to generate a query token set and adding it to a semantic understanding model to extract association relationships, obtaining a first query semantic network, identifying path information of the first query semantic network through an intent recognition model to generate a second query semantic network including: Receive a user data query request, extract the original query statement in the user data query request, construct a word element recognition tree, set the node attributes of the word element recognition tree as word element content, part of speech, and position offset, start forward maximum matching from the beginning of the original query statement to find the longest matching word element in the dictionary, start backward maximum matching from the end of the original query statement to find the longest matching word element in the dictionary, compare the word segmentation results of the forward maximum matching and the backward maximum matching, select the word segmentation result with the fewest number of word elements when the number of word elements is different, and select the word segmentation result with the fewest single-character words when the number of word elements is the same, and generate a query word element set; Construct a deep bidirectional long short-term memory network, input the word vector sequence of the query word element set into the input layer of the deep bidirectional long short-term memory network, extract context features through multiple bidirectional long short-term memory layers of the deep bidirectional long short-term memory network, input the extracted context features into a conditional random field for part-of-speech tagging, identify organization words, time words, measurement words, numerical words, object words, statistical index words, judgment words, and mark the part-of-speech attributes and semantic roles of the query word element set; Construct an eight-head self-attention model. The first attention head captures temporal association, the second attention head captures conditional association, the third attention head captures object association, the fourth attention head captures index association, the fifth attention head captures statistical association, the sixth attention head captures subject association, the seventh attention head captures attribution association, and the eighth attention head captures combination association. Input the labeled query word element set into the eight-head self-attention model to calculate the attention scores between word elements; Construct a first query semantic network based on the attention scores. Set the word elements as the nodes of the first query semantic network, set the semantic associations as the edges of the first query semantic network, and set the attention scores as the weights of the edges; set a graph structure pruning threshold, delete the edges with weights lower than the graph structure pruning threshold, and extract the main query path and branch query paths; Construct a graph attention model. Input the first query semantic network into the graph attention layer of the graph attention model, calculate the attention scores between a node and its adjacent nodes based on the spatial attention mechanism, update the node semantic representation, input the updated node semantic representation into the path extraction layer of the graph attention model, extract the key query path based on the depth-first search algorithm, and mark the attention scores of the nodes on the path to generate a second query semantic network.
3. The method according to claim 2, wherein Construct a query feature matrix based on the second query semantic network and add it to the feature combination model. Generate fused feature data by calculating the correlation weight coefficients between matrices, extract high-dimensional feature information from the fused feature data through a deep residual network and generate a query feature identifier, calculate the similarity with a pre-set template feature identifier and generate a similarity score table, and perform hierarchical matching and determine that the query rule set corresponding to the successfully matched template feature identifier includes: Construct a query feature matrix based on the second query semantic network, extract the temporal association features, numerical comparison features, and data type features among the nodes in the query semantic network, and construct a query feature matrix including dimensions of table association, field correlation, filtering condition, sorting rule, and time range; Add the query feature matrix to the feature combination model, set a matrix association calculation unit in the feature combination model, calculate the association weight coefficient between matrices through the matrix association calculation unit, and according to the association weight coefficient, retain and transmit the feature information with an association weight coefficient higher than the weight threshold, and attenuate and filter the feature information with an association weight coefficient lower than the weight threshold to generate fused feature data; Construct a multi-scale residual network, which includes a feature decomposition unit, a feature reconstruction unit, and a feature optimization unit. The feature decomposition unit decomposes the fused feature data into multi-scale feature components through wavelet transform. The feature reconstruction unit sets an adaptive residual module to perform independent feature reconstruction on different-scale feature components. The feature optimization unit enhances the feature expression through a cross-scale feature interaction mechanism to generate a query feature identifier; Construct a multi-view feature matching network, which includes a local feature matching sub-network, a global feature matching sub-network, and a feature calibration sub-network. The local feature matching sub-network adopts a hierarchical attention mechanism to adaptively select key local features for matching. The global feature matching sub-network constructs a semantic consistency constraint in the feature space based on a contrastive learning strategy. The feature calibration sub-network optimizes the discriminability of the feature representation through the mutual information maximization criterion. Input the query feature identifier and the template feature identifier into the multi-view feature matching network, fuse the local structural similarity and the global semantic similarity to generate a similarity score table; Perform hierarchical matching, dynamically adjust the matching threshold of different query areas according to the system resource status. When the similarity exceeds the matching threshold of the corresponding query area, determine the template feature identifier with successful matching, and extract the query rule set corresponding to the template feature identifier.
4. The method according to claim 1, wherein Add the query feature matrix to the data distribution prediction network, predict the spatio-temporal distribution law based on historical access records to obtain a data distribution feature map, construct a data association model and generate a probability distribution table by calculating the association probability value between fields, construct a multi-layer combined index based on the probability distribution table and construct an index structure tree in combination with the information benefit ratio, calculate the access timing feature and construct a time decay calculation model to obtain a dynamic index adjustment scheme, add the execution path in the query rule set to the path optimization network based on the dynamic index adjustment scheme, determine the optimal execution path in combination with the path search algorithm and generate a task execution scheme, schedule the database executor according to the task execution scheme and obtain a query result stream for real-time analysis, calculate the data value coefficient and write it into the multi-level cache system, including: Add the query feature matrix to the data distribution prediction network, and use the time series feature extraction main branch and the spatial feature extraction auxiliary branch for feature extraction. Extract the access time, access object, access user, and access behavior features from the historical access records through the time series feature extraction main branch, capture the short-term access features through the sliding convolutional layer, segment the access sequence and count the access metrics, extract the periodic access patterns and trend access features through the recurrent processing unit, construct an access object association graph through the spatial feature extraction auxiliary branch, extract the node propagation features using the graph structure network, and fuse the time series features and the spatial features to generate a data distribution feature map; Construct a data association model, divide it into an attribute association layer, a record association layer, and a table association layer for data association analysis. In the attribute association layer, calculate the correlation coefficient of numerical fields through feature correlation analysis, calculate the association value of categorical fields through information entropy analysis, and generate a field association matrix. In the record association layer, use the pattern mining method to extract multi-field combination features, identify high-frequency combination patterns and key combination rules. In the table association layer, construct a data lineage graph based on the field reference relationship, analyze the data flow path and data dependency strength, and generate a global data dependency graph. Integrate the field association matrix, the key combination rules, and the global data dependency graph to calculate the association probability value between fields and generate a probability distribution table; Construct a multi-level composite index based on the probability distribution table. Divide the index into a global index layer, a partition index layer, and a local index layer. In the global index layer, select the main index field through the information gain ratio calculation, and construct a composite index based on the pre-determined high-frequency access field combination. In the partition index layer, determine the partition boundary based on the data distribution characteristics, partition the data, and establish a sub-index. In the local index layer, calculate the access heat of the data partition and construct an index structure tree; Calculate the access time series features based on the composite index, the sub-index, and the index structure tree, and construct a time decay calculation model. Divide it into multiple time windows, set a time decay function for the access data in each time window, calculate the access data weight, count the index usage frequency, index hit rate, and index maintenance overhead, and generate index evaluation metrics. When the index evaluation metrics are lower than the threshold, trigger an index optimization operation to generate a dynamic index adjustment plan; Add the execution paths in the query rule set to the path optimization network based on the dynamic index adjustment plan, and divide it into an operator layer, a resource layer, and a scheduling layer for path optimization. In the operator layer, decompose the query operation into basic operators and construct an operator dependency graph. In the resource layer, monitor the system resource status and generate a resource status vector. In the scheduling layer, combine the path search algorithm to identify the operator groups that can be executed in parallel, evaluate the execution path cost, determine the optimal execution path, and generate a task execution plan; Schedule the database executor according to the task execution plan, obtain the query result stream for real-time analysis, calculate the data value coefficient, write the data value coefficient into the multi-level cache system, divide it into a hot cache layer, a warm data cache layer and a cold data cache layer for data storage, and perform hierarchical storage of data based on the data value coefficient.
5. The method according to claim 4, characterized in that Calculate the access timing characteristics and construct a time decay calculation model based on the composite index, the sub-index and the index structure tree, divide into multiple time windows, set a time decay function for the accessed data within each time window, calculate the access data weight, count the index usage frequency, index hit rate and index maintenance overhead, generate index evaluation metrics, and when the index evaluation metrics are lower than the threshold, trigger an index optimization operation to generate a dynamic index adjustment plan, including: Collect the access logs of the composite index, the sub-index and the index structure tree within a preset time period and extract the access time points and operation instructions in the access logs, statistically obtain the access frequency sequence of the access time points at small time intervals, statistically obtain the operation count sequence of the operation instructions according to the instruction type, perform a sliding window calculation on the access frequency sequence to extract the fluctuation value and the period value, perform pattern recognition on the operation count sequence to extract the behavior feature value, and calculate the access timing characteristics based on the composite index, the sub-index and the index structure tree; Divide the access timing characteristics into multiple time windows at time intervals of 1 hour, 24 hours, and 168 hours, construct a time decay calculation model to obtain a time trend component representing the change in access volume, a periodic fluctuation component representing repeated access, and a random fluctuation component representing burst access; Execute the time decay function setting for the accessed data within each time window and obtain the time interval and access frequency distribution characteristics of the accessed data within each time window, determine the function type of the time decay function based on the time interval and access frequency distribution characteristics and set the parameters of the time decay function according to the time span of each time window, multiply the original weight of the accessed data within each time window by the calculation result of the corresponding time decay function to obtain the time-weighted access volume, and calculate the access data weight by combining the time-weighted access volume with the access operation type; Count the index usage frequency, index hit rate and index maintenance overhead, convert the index usage frequency, the index hit rate, and the index maintenance overhead to between 0 and 1 respectively according to a preset rule, multiply the converted metrics by the preset metric weights and sum them to obtain the index evaluation metric, compare the index evaluation metric with the preset performance evaluation threshold, and trigger an index optimization operation when the index evaluation metric is lower than the performance evaluation threshold; Generate a dynamic index adjustment plan and count the number of collaborative accesses between single-column indexes. When the number of collaborative accesses exceeds a preset collaborative threshold, generate an index merging instruction to merge the single-column indexes with a collaborative access relationship to construct a composite index. Count the access volume of the composite index. When the access volume is lower than a preset access volume threshold for 30 consecutive days, generate an index splitting instruction to split the composite index into multiple single-column indexes. Count the number of days the index has not been accessed and the resource occupancy. When the number of days the index has not been accessed exceeds 90 days and the resource occupancy exceeds a preset resource threshold, generate an index deletion instruction to delete the index and release the storage space.
6. The method according to claim 1, characterized in that, Collect the execution performance data of user data query requests and add it to the anomaly recognition model. Combine clustering operations to identify anomaly data points and generate an anomaly mark set. Calculate the deviation degree of performance metrics based on an anomaly detection model based on a recurrent neural network and determine the temporal dependence relationship by combining a causal inference network to generate a bottleneck location map. Generate a performance score based on the bottleneck location map and the deviation degree of performance metrics and generate a data value prediction sequence based on the data value coefficient. Combine a resource scheduling network and a competitive learning algorithm to determine the load balancing coefficient of the storage area, generate a preloading strategy, perform performance verification through a Monte Carlo simulation model, and select the preloading strategy with the highest confidence as the optimal strategy, including: Collect the execution performance data of user data query requests, record the application layer execution response time, the system layer resource usage status, and the network layer transmission metrics, collect them regularly at preset time intervals and organize them into a temporal data stream. Use the exponentially weighted moving average algorithm to calculate the weighted coefficient matrix based on the historical data sequence, multiply the weighted coefficient matrix by the temporal data stream to obtain a smoothed data sequence, use the linear interpolation algorithm to fill in the missing data through the weighted combination of adjacent valid data points, subtract the mean from the data sequence and divide by the standard deviation to obtain the normalized data and add it to the anomaly recognition model; Calculate the number of neighboring points within a specified radius for each data point in the normalized data to obtain the local density value, construct a distance matrix to calculate the relative distance between data points, multiply the local density value and the relative distance after normalizing them respectively to obtain the anomaly degree matrix, screen out anomaly data points from the anomaly degree matrix based on a set threshold, calculate the temporal autocorrelation coefficient of the anomaly data points and verify its duration to generate an anomaly mark set; For the anomaly mark set, calculate the deviation degree of performance metrics based on an anomaly detection model based on a recurrent neural network. Use the input gate, forget gate, and output gate of the long short-term memory network to extract features from the performance metric sequence to obtain a feature sequence, construct a gated recurrent unit to calculate the hidden state sequence of the feature sequence, input the hidden state sequence into a conditional random field model to obtain a state transition probability matrix, and construct a causal inference network based on the state transition probability matrix to calculate the conditional probability distribution between metrics to generate a temporal dependence relationship graph; Construct a bottleneck location map based on the time series dependency graph. Convert the system components in the bottleneck location map into vector representations through random walk in the graph embedding algorithm. Calculate the component similarity matrix based on the vector representations. Apply the community discovery algorithm to identify component subgraphs with close associations. Construct a performance score by combining the degree of deviation of the performance metrics. Generate a data value prediction sequence based on the performance score and the data value coefficient; Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network to predict resource requirements. Construct a state-action value function based on the resource requirements. Iteratively update the action policy through the reinforcement learning algorithm. Combine the resource scheduling network to calculate the load variance of each storage area and obtain the load balancing coefficient by searching for an equilibrium solution that minimizes the load variance according to the genetic algorithm's crossover and mutation operations and competitive learning; Based on the load balancing coefficient, construct a policy tree according to Monte Carlo tree search and select expansion nodes through the upper confidence bound algorithm. Use random forest to extract features from the historical load sequence and predict the fluctuation range. Set the search boundary of the particle swarm based on the fluctuation range. Iteratively optimize the policy parameters through the velocity-position update formula. Calculate the policy confidence using support vector regression and select the preloading policy with the highest confidence as the optimal policy.
7. The method according to claim 6, wherein Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network to predict resource requirements. Construct a state-action value function based on the resource requirements. Iteratively update the action policy through the reinforcement learning algorithm. Combine the resource scheduling network to calculate the load variance of each storage area and obtain the load balancing coefficient by searching for an equilibrium solution that minimizes the load variance according to the genetic algorithm's crossover and mutation operations and competitive learning includes: Input the data value prediction sequence, resource usage status, and load distribution into a pre-constructed deep neural network. The hidden layer of the deep neural network introduces nonlinear transformation through the rectified linear unit function and sequentially reduces the number of neurons to achieve feature dimension compression. Update the network parameters using the adaptive moment estimation optimizer. Eliminate the feature scale difference through the batch normalization layer and use the dropout mechanism to prevent overfitting, and output the resource requirements; Construct a state-action value function based on the resource requirements. The state space of the state-action value function includes the current allocation ratio of various resources, the length of the task waiting queue, and the load level of the storage area. The action space of the state-action value function includes proportionally adjusting the resource quota, reallocating tasks, and migrating data blocks. Construct an experience replay buffer to store historical decision samples. The historical decision samples include the system state before executing the action, the specific action content, the system benefits obtained, and the state transition result after execution; Iteratively update the action policy based on the historical decision samples through the reinforcement learning algorithm. Randomly sample the historical decision samples for parameter update in each iteration, and introduce a target network to periodically synchronize the main network parameters to the target network; Calculate the load variance of each storage area in combination with the resource scheduling network. The underlying physical resource pool of the resource scheduling network is responsible for managing hardware devices and recording the real-time status of the devices. The middle-layer virtual resource manager of the resource scheduling network realizes resource pool virtualization and provides a unified resource view. The upper-layer load balancing scheduler of the resource scheduling network calculates the degree of dispersion between the current load of each storage area and the global mean based on the real-time monitoring data of the storage area to obtain the load variance; According to the crossover and mutation operations of the genetic algorithm and competitive learning, define the mapping relationship from data blocks to storage areas through chromosome encoding. The gene position value of the chromosome encoding represents the target storage area identifier. Retain the allocation scheme with the highest load balancing performance through the selection operation. Exchange partial data block allocation information between parent chromosomes and randomly change the storage locations of multiple data blocks through the crossover and mutation operations. Use competitive learning to determine the local optimal individual to obtain the load balancing coefficient.
8. A data query parsing result optimization system based on caching, for implementing the method described in any one of the foregoing claims 1-7, characterized in that, Including: The first unit is used to receive a user data query request, extract the original query statement, perform word segmentation processing on the original query statement to generate a query token set, add it to the semantic understanding model to extract the association relationship, obtain the first query semantic network, identify the path information of the first query semantic network through the intention recognition model to generate the second query semantic network, construct a query feature matrix based on the second query semantic network and add it to the feature combination model, generate fusion feature data by calculating the association weight coefficient between matrices, extract high-dimensional feature information from the fusion feature data through the deep residual network and generate a query feature identifier, calculate the similarity with the pre-set template feature identifier and generate a similarity score table, perform hierarchical matching and determine the query rule set corresponding to the successfully matched template feature identifier; The second unit is used to add the query feature matrix to the data distribution prediction network, predict the spatio-temporal distribution law based on the historical access records to obtain the data distribution feature map, construct a data association model and generate a probability distribution table by calculating the association probability value between fields, construct a multi-layer combined index based on the probability distribution table and combine the information benefit ratio to construct an index structure tree, calculate the access time sequence feature and construct a time decay calculation model to obtain a dynamic index adjustment scheme, add the execution path in the query rule set to the path optimization network based on the dynamic index adjustment scheme, determine the optimal execution path in combination with the path search algorithm and generate a task execution scheme, schedule the database executor according to the task execution scheme and obtain the query result stream for real-time analysis, calculate the data value coefficient and write it into the multi-level cache system; The third unit is used to collect the execution performance data of the user data query request and add it to the anomaly recognition model, identify the anomaly data points through clustering operation and generate an anomaly label set, calculate the deviation degree of the performance index based on the anomaly detection model of the recurrent neural network and determine the temporal dependence relationship in combination with the causal inference network to generate a bottleneck location map, generate a performance score based on the bottleneck location map and the deviation degree of the performance index and generate a data value prediction sequence according to the data value coefficient, determine the load balancing coefficient of the storage area in combination with the resource scheduling network and the competitive learning algorithm to generate a preloading strategy, perform performance verification through the Monte Carlo simulation model and select the preloading strategy with the highest confidence as the optimal strategy.
9. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Query request processing method and device, electronic equipment, medium and program product
CN120631941A
Data multi-centralized management method and system based on DOA
CN120892477A
Bitmap index container optimization method and system based on multi-dimensional dynamic decision
CN120910028A
Energy transformation and digital economy collaborative analysis method and system based on artificial intelligence
CN121210427A
Storage strategy optimization method and device based on data popularity and data consanguinity
CN121278004A