Query method and system for generating SQL statements based on natural language
By introducing distributed lock management, query path evaluation and multi-level fault tolerance mechanisms into the query system that generates SQL statements in natural language, concurrent query and resource occupation management problems are solved in high concurrency scenarios, and efficient query and system stability are achieved.
Patent Information
- Application Number
- CN202411960744.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The existing query methods for generating SQL statements in natural languages are difficult to effectively manage concurrent queries and resource occupation in high concurrency scenarios, and lack fault tolerance mechanisms and adaptive adjustment capabilities, resulting in low query efficiency and poor system stability.
The multi-level lock identification is obtained through the distributed lock management module, the timeout time of the lock identification is dynamically adjusted, the data template is generated based on the query mode knowledge base of concurrency control, query path evaluation and select the optimal query path, predict resource occupation and set monitoring thresholds, monitor query status in real time and trigger multi-level fault tolerance mechanisms.
Effectively manage concurrent queries, reduce resource competition and wait time, and improve query efficiency; realize multi-level fault tolerance mechanism to enhance the stability and fault tolerance of the system; improve the robustness of the system through adaptive adjustment and optimization of the query process.
Smart Images

Figure CN119377241B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to programming technology, and in particular to a query method and system for generating SQL statements based on natural language. Background Art
[0002] The rapid development of natural language processing technology and database technology has made it possible to query databases using natural language. Users can use natural language to express query requirements without writing complex SQL statements, which reduces the threshold for using databases. However, the existing query methods for generating SQL statements using natural language have the following defects and deficiencies:
[0003] First, most existing methods lack refined control over the query process, making it difficult to effectively manage concurrent queries and resource usage. In high-concurrency scenarios, query congestion and resource competition are prone to occur, leading to low query efficiency and even system crashes.
[0004] Secondly, existing methods generally lack a complete fault-tolerant mechanism. When an exception occurs during the query process, it is often difficult to quickly locate and solve the problem, resulting in query failure or returning incorrect results. This is unacceptable for application scenarios that require high data consistency and reliability.
[0005] Finally, existing methods are usually difficult to adaptively adjust based on the user's query history and data characteristics. This makes it difficult to guarantee query efficiency and accuracy when facing complex query scenarios or changes in data distribution. The lack of continuous learning and optimization capabilities limits the practicality and universality of existing methods. Summary of the invention
[0006] The embodiments of the present invention provide a query method and system for generating SQL statements based on natural language, which can solve the problems in the prior art.
[0007] According to a first aspect of the embodiments of the present invention,
[0008] Provides a query method for generating SQL statements based on natural language, including:
[0009] Receive a user query request and perform identity authentication to obtain data source configuration information including data source type, database name and table name; based on the identity authentication result and data source configuration information, obtain a multi-level lock identifier through a distributed lock management module in a preset order, wherein the multi-level lock identifier includes a user session-level lock identifier, a database connection pool lock identifier and an SQL execution lock identifier, and the distributed lock management module dynamically adjusts the timeout time of each level of lock identifier based on the historical query load; extract a query pattern from a query pattern knowledge base of concurrent control based on the data source configuration information and the user's historical query records, and generate a data template, wherein the data template includes a data source identifier, a model request identifier, database information and user query content;
[0010] Performing query path evaluation on the data template, wherein the query path evaluation includes query complexity evaluation and data scale evaluation, and selecting an optimal query path according to the evaluation result; calling a large model server through an HTTP request according to the optimal query path to parse the data template and generate an SQL statement;
[0011] Based on the structure of the SQL statement, resource occupancy is predicted and monitoring thresholds are set; query status is monitored in real time during execution; when an exception is detected, a multi-level fault-tolerant mechanism including breakpoint resumption, service degradation, and standby switching is triggered; exception information is recorded and the cause of the error is analyzed in combination with the SQL execution plan; query strategies are adjusted according to the analysis results to complete query recovery; the multi-level locking identifier is released, the query results are returned to the user, and the query information is updated to the query pattern knowledge base in an asynchronous manner.
[0012] In an optional embodiment,
[0013] The steps of obtaining a multi-level lock identifier through a distributed lock management module in a preset order, wherein the multi-level lock identifier includes a user session level lock identifier, a database connection pool lock identifier, and an SQL execution lock identifier, and the distributed lock management module dynamically adjusts the timeout time of each level of lock identifier based on the historical query load includes:
[0014] Acquire multi-level lock identifiers in a preset order, and acquire user session-level lock identifiers through a user session management unit using distributed key-value storage. The user session-level lock identifier generates a lock key value by combining a session number and a timestamp, and ensures mutually exclusive access through atomic operations;
[0015] Based on the user session-level lock identifier, a temporary sequential node is established through a database connection pool management unit to obtain a database connection pool lock identifier, wherein the database connection pool lock identifier uses a fair lock strategy to allocate connection resources and maintains resource status through a node monitoring mechanism;
[0016] Based on the database connection pool lock identifier, a lease mechanism is established through the SQL execution management unit to obtain the SQL execution lock identifier, the SQL execution lock identifier ensures concurrency safety through comparison and exchange operations, and sets a lease validity period;
[0017] Using a sliding time window to count query load information of lock identifiers at all levels, the query load information includes query frequency, execution time distribution and resource usage peak, and calculating the benchmark execution time of lock identifiers at all levels based on the query load information;
[0018] A feedback control mechanism is used to dynamically adjust the timeout time of each level of locking identification. The feedback control mechanism adjusts the parameters by calculating the execution time error and triggers gradient parameter adjustment in the case of a sudden load change.
[0019] In an optional embodiment,
[0020] The steps of extracting a query pattern from a query pattern knowledge base of concurrent control based on the data source configuration information and the user's historical query records, and generating a data template, wherein the data template includes a data source identifier, a model request identifier, database information, and user query content include:
[0021] Parsing the data source configuration information to form data source description information; constructing a multidimensional feature vector space based on the data source description information, wherein the multidimensional feature vector space includes database structure features, load features, timing features, and business features, wherein the database structure features include table node information and inter-table association information, the load features include resource density information, the timing features include execution time mode information and periodicity intensity information, and the business features include business domain information;
[0022] Feature extraction is performed on the user's historical query records, and the extracted query features are mapped to the multidimensional feature vector space. The improved DBSCAN clustering algorithm is used for clustering analysis. The improved DBSCAN clustering algorithm combines the edit distance and the feature vector cosine similarity to define the distance function, dynamically adjusts the density threshold based on the query distribution, and supports online mode evolution through incremental updates; based on the clustering analysis results, a query mode knowledge base for concurrency control is constructed. The query mode knowledge base adopts a knowledge graph structure, and the mode applicability condition information, performance characteristic information, and resource requirement information are respectively marked on the nodes and edges;
[0023] Extracting query patterns from the query pattern knowledge base according to the data source description information and the knowledge graph structure; using a context-aware pattern matching algorithm in the extraction process; and performing a comprehensive evaluation by calculating a time correlation score, a load compatibility score, and a business alignment score; wherein the time correlation score is calculated based on the execution time window and periodic characteristics of the query pattern to determine the degree of matching with the current query; the load compatibility score is calculated based on the degree of matching between the resource requirements of the query pattern and the current system load state; and the business alignment score is calculated based on the degree of conformity between the query pattern and the current business scenario constraints; and selecting the optimal query pattern based on the comprehensive evaluation result;
[0024] A data template is generated according to the extracted query pattern, and an adaptive template generation strategy is adopted to construct a basic template and select a concurrent control strategy; a feedback mechanism is established for the execution process of the data template, and the evaluation result is used to update the query pattern knowledge base.
[0025] In an optional embodiment,
[0026] The query path evaluation is performed on the data template, wherein the query path evaluation includes query complexity evaluation and data scale evaluation. The step of selecting the optimal query path according to the evaluation result includes:
[0027] Constructing a complexity calculation model based on a query syntax tree, wherein the complexity calculation model performs complexity scoring on nodes in the query syntax tree, constructing a complexity feature vector based on a node nesting depth feature value, a table connection quantity feature value, and a subquery complexity feature value, and obtaining the complexity score by performing an inner product operation of a feature weight vector and a complexity feature vector;
[0028] Establishing a multidimensional data scale assessment model, wherein the multidimensional data scale assessment model calculates table data volume information, filter condition selection rate information, and association relationship duplication information based on histogram data distribution, and maintains table-level data statistics information and column-level data distribution information;
[0029] Based on the evaluation results of the complexity score and the multidimensional data scale evaluation model, a cost calculation function for the query path is constructed, wherein the cost calculation function integrates the path complexity score information, the data scale information and the historical execution time information, and an improved dynamic programming algorithm is used to perform path search on the calculation results of the cost calculation function, wherein the improved dynamic programming algorithm calculates the local optimal solution through the state transfer equation and obtains the optimal query path through the backtracking operation; during the path search process, a path pruning operation exceeding a threshold, a local optimal solution caching operation and a multi-path parallel evaluation operation are performed;
[0030] A query execution plan is generated according to the optimal query path, wherein the query execution plan includes node execution sequence information, data access mode information and resource allocation information.
[0031] In an optional embodiment,
[0032] Constructing a cost calculation function for a query path, the cost calculation function integrates path complexity score information, data scale information, and historical execution time information, using an improved dynamic programming algorithm to perform path search on the calculation result of the cost calculation function, the improved dynamic programming algorithm calculates a local optimal solution through a state transfer equation, and obtaining an optimal query path through a backtracking operation includes the following steps:
[0033] Constructing a hierarchical and progressive multi-factor cost calculation model, the multi-factor cost calculation model includes a path complexity scoring module, a data scale evaluation module and a historical execution statistics module, wherein the path complexity scoring module includes processor calculation complexity evaluation and connection operation cost evaluation, the data scale evaluation module includes intermediate result set evaluation and data skew degree quantification, the historical execution statistics module includes execution time analysis and resource consumption evaluation, and generating initial cost evaluation information based on the multi-factor cost calculation model;
[0034] According to the initial cost evaluation information, an influence factor regulator based on deep Q learning is designed, and the influence factor regulator includes a state space, an action space and a reward function, wherein the state space includes current load information and historical load information, the action space includes adjustment range information of each influence factor, and the reward function is composed of query execution time, resource utilization rate and query success rate, and the reward value is obtained by weighted combination calculation; the deep Q learning adopts a dual network structure, including a policy network and a target network, stores and samples training data through an experience replay mechanism, and uses the processor utilization rate, memory occupancy and IO operation number based on the initial cost evaluation information as network input, and determines that the training converges when the Q value change rate is less than a preset threshold and the number of consecutive training rounds exceeds a preset number of rounds;
[0035] Executing a hierarchical dynamic programming algorithm based on the adjustment result of the influencing factor adjuster, including hierarchically constructing an initial execution plan based on the query syntax tree to generate a set of candidate plans, constructing a local cost model and performing path search based on an ant colony optimization algorithm based on pheromone concentration and heuristic factors, establishing a path dependency graph and identifying a target path;
[0036] The candidate plan set is optimized according to the path dependency graph, including using heuristic rules for fast pruning, decomposing the pruned plan to obtain a sub-path set, optimizing each sub-path and updating global state information using local update rules and global update rules of an ant colony algorithm, dynamically adjusting the resource allocation strategy according to the system load state, selecting the path with the highest score as the optimal query path by comparing the comprehensive performance scores of the optimized paths, and generating an optimized execution plan based on the optimal query path.
[0037] In an optional embodiment,
[0038] The steps of predicting resource usage based on the structure of the SQL statement and setting a monitoring threshold, monitoring the query status in real time during execution, and triggering a multi-level fault tolerance mechanism including breakpoint resumption, service degradation, and standby switching when an abnormality is detected include:
[0039] Constructing a SQL structure parsing model based on a syntax tree, wherein the SQL structure parsing model decomposes an SQL statement into an operation operator sequence, extracts table connection depth, predicate condition complexity, and aggregate function nesting level to form an SQL structure feature vector, quantitatively analyzes the structural features in the SQL structure feature vector, and establishes a resource occupancy prediction model in combination with the operation operator execution cost, and outputs a resource prediction value based on the resource occupancy prediction model;
[0040] Constructing a dynamic threshold adjustment algorithm to obtain a threshold adjustment result, wherein the dynamic threshold adjustment algorithm calculates a monitoring indicator alarm threshold by comprehensively considering resource prediction values and system load status, wherein the system load status is quantified by a load impact function including CPU utilization, memory occupancy, and IO waiting time, and the monitoring indicator alarm threshold is dynamically adjusted with query complexity and system load status;
[0041] Establishing a multi-level fault tolerance mechanism according to the threshold adjustment result, wherein the multi-level fault tolerance mechanism includes a breakpoint resume level, a service degradation level, and a standby switching level, and triggering corresponding levels of fault tolerance processing for different resource prediction value intervals;
[0042] The breakpoint resume level constructs recovery point information by recording the current execution state, intermediate result set and progress information, and performs query recovery based on the recovery point information; the service degradation level adjusts the execution parameters by calculating the degradation degree, so that the resource demand is reduced to below the target threshold, and re-executes the query based on the degraded execution parameters; the standby switching level selects an alternate execution plan from a preset set of alternate plans based on the current resource prediction value, and switches the query to the alternate execution plan;
[0043] A real-time status monitoring module is constructed to continuously collect performance indicators during the query execution process. When it is detected that the performance indicator exceeds the monitoring indicator alarm threshold, the abnormal information is fed back to the multi-level fault tolerance mechanism and the corresponding level of fault tolerance processing is triggered.
[0044] In an optional embodiment,
[0045] A multi-level fault tolerance mechanism is established according to the threshold adjustment result, wherein the multi-level fault tolerance mechanism includes a breakpoint resume level, a service degradation level, and a standby switching level. The steps of triggering corresponding levels of fault tolerance processing for different resource prediction value intervals include:
[0046] An adaptive interval division module is constructed based on the historical load data of the system. The adaptive interval division module divides the resource prediction value into a low load interval, a medium load interval and a high load interval by using the quartile method, and introduces a dynamic adjustment factor including the current load and performance index of the system. The interval boundary is adjusted in real time according to the dynamic adjustment factor, and the dynamic interval division result is output;
[0047] A fault tolerance level trigger module is constructed based on the dynamic interval division result, and the fault tolerance level trigger module includes a breakpoint resume trigger unit, a service degradation trigger unit and a standby switching trigger unit, wherein the breakpoint resume trigger unit is triggered when the resource prediction value is in a low load interval and the query state score exceeds a preset threshold, the service degradation trigger unit is triggered when the resource prediction value is in a medium load interval and the performance index is lower than the target value, and the standby switching trigger unit is triggered when the resource prediction value is in a high load interval or an abnormality is detected;
[0048] Construct a state machine-based collaborative scheduling module, the collaborative scheduling module includes a normal state, a breakpoint resume state, a service degradation state and a standby switching state, establishes a state transition matrix based on the trigger signal of the fault tolerance level trigger module, and defines the transition conditions between states; the collaborative scheduling module allocates resource quotas to different fault tolerance states according to the query priority weight and the system load state, and calculates the state transition cost based on resource overhead, switching delay and performance loss; after receiving the fault tolerance level trigger signal, the collaborative scheduling module determines the target state according to the state transition matrix, selects the optimal transition path based on the state transition cost, and executes the state switching according to the resource quota allocation result;
[0049] A fault-tolerance effect evaluation module is constructed, which collects performance data during the state switching process and feeds the performance data back to the adaptive interval division module for optimizing the interval boundary and the dynamic adjustment factor.
[0050] According to a second aspect of the embodiments of the present invention,
[0051] Provides a query system that generates SQL statements based on natural language, including:
[0052] The first unit is used to receive a user query request and perform identity authentication, and obtain data source configuration information including a data source type, a database name, and a table name; based on the identity authentication result and the data source configuration information, a multi-level lock identifier is obtained through a distributed lock management module in a preset order, wherein the multi-level lock identifier includes a user session-level lock identifier, a database connection pool lock identifier, and an SQL execution lock identifier, and the distributed lock management module dynamically adjusts the timeout time of each level of lock identifier based on the historical query load; based on the data source configuration information and the user's historical query records, a query pattern is extracted from a query pattern knowledge base of concurrent control to generate a data template, wherein the data template includes a data source identifier, a model request identifier, database information, and user query content;
[0053] The second unit is used to perform query path evaluation on the data template, wherein the query path evaluation includes query complexity evaluation and data scale evaluation, and select an optimal query path according to the evaluation result; according to the optimal query path, a large model server is called through an HTTP request to parse the data template to generate an SQL statement;
[0054] The third unit is used to predict resource occupancy and set monitoring thresholds based on the structure of the SQL statement, monitor the query status in real time during execution, and trigger a multi-level fault-tolerant mechanism including breakpoint resumption, service degradation and standby switching when an abnormality is detected; record the abnormal information and analyze the cause of the error in combination with the SQL execution plan, adjust the query strategy according to the analysis results, and complete query recovery; release the multi-level locking identifier, return the query result to the user, and update the query information to the query pattern knowledge base in an asynchronous manner.
[0055] According to a third aspect of the embodiments of the present invention,
[0056] An electronic device is provided, comprising:
[0057] processor;
[0058] a memory for storing processor-executable instructions;
[0059] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0060] According to a fourth aspect of the embodiments of the present invention,
[0061] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0062] The present invention effectively manages concurrent queries, reduces resource competition and waiting time, and thus improves query efficiency through a multi-level locking mechanism, query path evaluation, and selection of an optimal query path.
[0063] The present invention can predict resource occupancy through query path evaluation and monitoring threshold setting, and perform resource allocation and adjustment according to actual conditions to avoid resource waste.
[0064] The distributed lock management and multi-level fault-tolerant mechanism of the present invention can effectively respond to query requests in high-concurrency scenarios, ensure stable system operation, record and analyze abnormal information, and dynamically adjust query strategies, which can continuously optimize the query process and improve the system's fault tolerance and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 A flowchart of a query method for generating SQL statements based on natural language according to an embodiment of the present invention;
[0066] Figure 2 The diagram is a structural diagram of a query system for generating SQL statements based on natural language according to an embodiment of the present invention. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0068] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0069] Figure 1 FIG. 1 is a flow chart of a query method for generating SQL statements based on natural language according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0070] S1. Receive a user query request and perform identity authentication to obtain data source configuration information including data source type, database name and table name; based on the identity authentication result and data source configuration information, obtain a multi-level lock identifier through a distributed lock management module in a preset order, wherein the multi-level lock identifier includes a user session-level lock identifier, a database connection pool lock identifier and an SQL execution lock identifier, and the distributed lock management module dynamically adjusts the timeout time of each level of lock identifier based on the historical query load; extract a query pattern from a query pattern knowledge base of concurrent control based on the data source configuration information and the user's historical query records, and generate a data template, wherein the data template includes a data source identifier, a model request identifier, database information and user query content;
[0071] S2. Perform query path evaluation on the data template, the query path evaluation includes query complexity evaluation and data scale evaluation, and select the optimal query path according to the evaluation result; call the large model server through HTTP request according to the optimal query path to parse the data template and generate SQL statements;
[0072] S3. Predict resource usage based on the structure of the SQL statement and set monitoring thresholds, monitor query status in real time during execution, and trigger a multi-level fault-tolerance mechanism including breakpoint resumption, service degradation, and standby switching when an anomaly is detected; record the exception information and analyze the cause of the error in combination with the SQL execution plan, adjust the query strategy based on the analysis results, and complete query recovery; release the multi-level locking identifier, return the query result to the user, and update the query information to the query pattern knowledge base in an asynchronous manner.
[0073] In an optional embodiment,
[0074] The steps of obtaining a multi-level lock identifier through a distributed lock management module in a preset order, wherein the multi-level lock identifier includes a user session level lock identifier, a database connection pool lock identifier, and an SQL execution lock identifier, and the distributed lock management module dynamically adjusts the timeout time of each level of lock identifier based on the historical query load includes:
[0075] Acquire multi-level lock identifiers in a preset order, and acquire user session-level lock identifiers through a user session management unit using distributed key-value storage. The user session-level lock identifier generates a lock key value by combining a session number and a timestamp, and ensures mutually exclusive access through atomic operations;
[0076] Based on the user session-level lock identifier, a temporary sequential node is established through a database connection pool management unit to obtain a database connection pool lock identifier, the database connection pool lock identifier uses a fair lock strategy to allocate connection resources, and maintains resource status through a node monitoring mechanism;
[0077] Based on the database connection pool lock identifier, a lease mechanism is established through the SQL execution management unit to obtain the SQL execution lock identifier, the SQL execution lock identifier ensures concurrency safety through comparison and exchange operations, and sets a lease validity period;
[0078] Using a sliding time window to count query load information of lock identifiers at all levels, the query load information includes query frequency, execution time distribution and resource usage peak, and calculating the benchmark execution time of lock identifiers at all levels based on the query load information;
[0079] A feedback control mechanism is used to dynamically adjust the timeout time of each level of locking identification. The feedback control mechanism adjusts the parameters by calculating the execution time error and triggers gradient parameter adjustment in the case of a sudden load change.
[0080] Exemplarily, first, obtain the user session-level lock identifier. The user session management unit uses a distributed key-value storage, such as Redis or Etcd, to implement locking. The lock key value is generated by combining the session number and the timestamp, such as "session_12345_1678886400". Through atomic operations, such as Redis's SETNX command, it is guaranteed that only one request in the same session at the same time can obtain the lock. Assuming that user A's session number is 12345 and the current timestamp is 1678886400, the key value that user A attempts to obtain the lock is "session_12345_1678886400". If the key value does not exist, the SETNX operation is successful and user A obtains the lock.
[0081] After obtaining the user session-level lock identifier, the database connection pool lock identifier is then obtained. The database connection pool management unit is implemented by establishing a temporary sequence node in the distributed key-value storage. For example, using the temporary sequence node function of ZooKeeper, all requests to obtain a connection pool lock will create a temporary sequence node. The order of the nodes determines the order in which the locks are acquired, and a fair lock strategy is implemented. When a node releases the lock, ZooKeeper notifies the next sequence node to maintain the resource status. Assume that after user A obtains the user session-level lock identifier, a temporary sequence node " / lock / connection_pool / 0000000001" is created on ZooKeeper.
[0082] After obtaining the database connection pool lock identifier, continue to obtain the SQL execution lock identifier. The SQL execution management unit is implemented using a lease mechanism. Concurrency safety is ensured by comparing and exchanging operations, such as Redis's CAS operation or the database's optimistic lock mechanism. At the same time, the lease validity period is set to prevent the lock holder from accidentally crashing and causing a deadlock. Assume that after user A obtains the database connection pool lock identifier, he uses Redis's CAS operation to set the key value "sql_lock_123" to the current timestamp plus the lease validity period, for example "1678886400_30", which means that the lock is valid for 30 seconds.
[0083] In order to dynamically adjust the timeout period, the system uses a sliding time window to count the query load information of each level of lock identification. For example, using a 1-minute sliding time window, the query frequency, execution time distribution, and resource usage peak are counted every 10 seconds. Assume that in the past minute, the average acquisition time of the user session-level lock identification is 10 milliseconds, the average acquisition time of the database connection pool lock identification is 5 milliseconds, and the average execution time of the SQL execution lock identification is 100 milliseconds.
[0084] Finally, the system uses a feedback control mechanism to dynamically adjust the timeouts of the lock flags at all levels. The execution time error is calculated by comparing the actual execution time with the benchmark execution time. The timeout is adjusted based on the error. For example, if the actual execution time is 10% greater than the benchmark execution time, the timeout is increased by 5%. In the case of a sudden load change, such as a sudden large number of requests, a gradient parameter adjustment is triggered to adjust the timeout more quickly. Assuming that the SQL execution time of user A is 110 milliseconds, which is 10% greater than the benchmark execution time of 100 milliseconds, the timeout of the SQL execution lock flag is increased from 30 seconds to 31.5 seconds.
[0085] The present invention avoids occupying resources due to excessive locking time by dynamically adjusting the timeout period, thereby improving system throughput and resource utilization; the multi-level locking mechanism and lease mechanism can effectively prevent deadlock and resource competition, and improve the stability and reliability of the system; the dynamic adjustment of the timeout period can be adjusted according to the actual load situation, avoiding users from waiting for a long time and improving the user experience.
[0086] In an optional embodiment,
[0087] The steps of extracting a query pattern from a query pattern knowledge base of concurrent control based on the data source configuration information and the user's historical query records, and generating a data template, wherein the data template includes a data source identifier, a model request identifier, database information, and user query content include:
[0088] Parsing the data source configuration information to form data source description information; constructing a multidimensional feature vector space based on the data source description information, wherein the multidimensional feature vector space includes database structure features, load features, timing features, and business features, wherein the database structure features include table node information and inter-table association information, the load features include resource density information, the timing features include execution time mode information and periodicity intensity information, and the business features include business domain information;
[0089] Feature extraction is performed on the user's historical query records, and the extracted query features are mapped to the multidimensional feature vector space. The improved DBSCAN clustering algorithm is used for clustering analysis. The improved DBSCAN clustering algorithm combines the edit distance and the feature vector cosine similarity to define the distance function, dynamically adjusts the density threshold based on the query distribution, and supports online mode evolution through incremental updates; based on the clustering analysis results, a query mode knowledge base for concurrency control is constructed. The query mode knowledge base adopts a knowledge graph structure, and the mode applicability condition information, performance characteristic information, and resource requirement information are respectively marked on the nodes and edges;
[0090] Extracting query patterns from the query pattern knowledge base according to the data source description information and the knowledge graph structure; using a context-aware pattern matching algorithm in the extraction process; and performing a comprehensive evaluation by calculating a time correlation score, a load compatibility score, and a business alignment score; wherein the time correlation score is calculated based on the execution time window and periodic characteristics of the query pattern to determine the degree of matching with the current query; the load compatibility score is calculated based on the degree of matching between the resource requirements of the query pattern and the current system load state; and the business alignment score is calculated based on the degree of conformity between the query pattern and the current business scenario constraints; and selecting the optimal query pattern based on the comprehensive evaluation result;
[0091] A data template is generated according to the extracted query pattern, and an adaptive template generation strategy is adopted to construct a basic template and select a concurrent control strategy; a feedback mechanism is established for the execution process of the data template, and the evaluation result is used to update the query pattern knowledge base.
[0092] Exemplarily, first, obtain data source configuration information, such as database connection information, table structure definition, business metadata, etc. Taking a relational database as an example, the configuration information can be stored in an XML file or a database table. Parse the configuration information, extract the database name, table name, field name, data type, primary key, foreign key, index and other information, and form structured data source description information.
[0093] Next, a multidimensional feature vector space is constructed based on the data source description information. This space contains multiple dimensions and is used to characterize the characteristics of the query pattern. In terms of database structure characteristics, table node information is extracted, such as table size, number of fields, etc.; and inter-table association information, such as foreign key relationships, connection types, etc. In terms of load characteristics, resource intensity information is collected, such as CPU usage, memory occupancy, number of I / O operations, etc. In terms of time series characteristics, historical query records are analyzed to extract execution time pattern information, such as query execution time, peak time period, etc.; and periodic intensity information, such as daily, weekly, and monthly query frequency changes. In terms of business characteristics, business domain information is extracted in combination with business metadata, such as the business module to which the query belongs and the key business indicators involved. For example, for an order query on an e-commerce platform, its business domain information may include "order management" and "transaction payment".
[0094] User historical query record feature extraction and cluster analysis. Obtain user historical query records, which can be extracted from database logs or query audit systems. Perform feature extraction on each query record, such as the keywords of the SQL statement, the operation type (SELECT, INSERT, UPDATE, DELETE), the table name involved, the field name, etc. Map the extracted query features to the multidimensional feature vector space constructed previously.
[0095] The improved DBSCAN clustering algorithm is used for cluster analysis. The algorithm combines the edit distance and the feature vector cosine similarity to define the distance function. The edit distance is used to measure the text similarity of SQL statements, and the feature vector cosine similarity is used to measure the similarity of query features in multidimensional space. For example, if the edit distance of two query SQL statements is small and their cosine similarity in the multidimensional feature vector space is high, then the two queries are considered to be relatively similar. The density threshold is dynamically adjusted based on the query distribution to avoid poor clustering results due to uneven query distribution. Online schema evolution is supported through incremental updates. For example, when a new query record arrives, its features can be mapped to the multidimensional space, and the cluster to which it belongs can be determined based on the distance function, or a new cluster can be created.
[0096] Query pattern knowledge base construction and query pattern extraction. Construct a query pattern knowledge base for concurrency control based on cluster analysis results. The knowledge base adopts a knowledge graph structure, where nodes represent query patterns and edges represent relationships between patterns, such as parent-child relationships, similarity relationships, etc. Mark the nodes and edges with pattern applicability information, such as data volume range, number of concurrent users, etc.; performance characteristic information, such as average execution time, maximum response time, etc.; and resource requirement information, such as CPU, memory, I / O, etc.
[0097] According to the data source description information, query patterns are extracted from the query pattern knowledge base according to the knowledge graph structure. In the extraction process, a context-aware pattern matching algorithm is used. For example, the time context, load context, and business context of the current query are considered. A comprehensive evaluation is performed by calculating the time correlation score, load compatibility score, and business alignment score. The time correlation score is calculated based on the execution time window and periodic characteristics of the query pattern to determine the degree of match with the current query. The load compatibility score is calculated based on the degree of match between the query pattern's resource requirements and the current system load status. The business alignment score is calculated based on the degree of compliance between the query pattern and the current business scenario constraints. For example, if a query pattern is suitable for a high-load scenario, but the current system load is low, the load compatibility score of the pattern is low. The optimal query pattern is selected based on the comprehensive evaluation results.
[0098] Data template generation and feedback mechanism establishment. Generate data templates based on the extracted query patterns. Adopt adaptive template generation strategies to build basic templates and select concurrency control strategies. For example, select appropriate concurrency control strategies based on the characteristics of the query patterns, such as optimistic locking, pessimistic locking, read-write separation, etc.
[0099] Establish a feedback mechanism for the execution process of the data template. Collect performance data during the template execution process, such as execution time, resource consumption, etc. Use the evaluation results to update the query pattern knowledge base. For example, if the actual execution effect of a query pattern is not good, you can reduce its weight in the knowledge base or update its applicable condition information.
[0100] For example, in an order query scenario of an e-commerce platform, query patterns such as "query orders in the last month", "query orders for a specified user", and "query orders for a specified product" can be extracted based on historical query records. When a new query request arrives, such as "query orders for a certain user in the last week", the system can extract the two patterns of "query orders for a specified user" and "query orders for the last month" from the knowledge base based on context information, and conduct a comprehensive evaluation based on indicators such as time correlation, load compatibility, and business alignment, and finally select the most appropriate pattern and generate the corresponding data template.
[0101] The present invention can reduce the time of query optimization and thus improve query efficiency by pre-building a query pattern knowledge base and selecting the optimal query pattern according to context information; through the selection of concurrency control strategies and the annotation of resource demand information, the occupation of system resources by concurrent queries can be effectively controlled to avoid resource competition and waste; by monitoring and providing feedback on the execution process of the query pattern, potential performance problems can be discovered in a timely manner, and adjustments and optimizations can be made, thereby enhancing the stability of the system.
[0102] In an optional embodiment,
[0103] The query path evaluation is performed on the data template, wherein the query path evaluation includes query complexity evaluation and data scale evaluation. The step of selecting the optimal query path according to the evaluation result includes:
[0104] Constructing a complexity calculation model based on a query syntax tree, wherein the complexity calculation model performs complexity scoring on nodes in the query syntax tree, constructing a complexity feature vector based on a node nesting depth feature value, a table connection quantity feature value, and a subquery complexity feature value, and obtaining the complexity score by performing an inner product operation of a feature weight vector and a complexity feature vector;
[0105] Establishing a multidimensional data scale assessment model, wherein the multidimensional data scale assessment model calculates table data volume information, filter condition selection rate information, and association relationship duplication information based on histogram data distribution, and maintains table-level data statistics information and column-level data distribution information;
[0106] Based on the evaluation results of the complexity score and the multidimensional data scale evaluation model, a cost calculation function for the query path is constructed, wherein the cost calculation function integrates the path complexity score information, the data scale information and the historical execution time information, and an improved dynamic programming algorithm is used to perform path search on the calculation results of the cost calculation function, wherein the improved dynamic programming algorithm calculates the local optimal solution through the state transfer equation and obtains the optimal query path through the backtracking operation; during the path search process, a path pruning operation exceeding a threshold, a local optimal solution caching operation and a multi-path parallel evaluation operation are performed;
[0107] A query execution plan is generated according to the optimal query path, wherein the query execution plan includes node execution sequence information, data access mode information and resource allocation information.
[0108] Exemplarily, first, a complexity calculation model based on a query syntax tree is constructed. The model parses the query statement into a syntax tree and performs a complexity score on each node in the tree. For example, a simple query statement "SELECT *FROM table1 WHERE col1 = 1 AND col2 = 2" will be parsed into a syntax tree, which contains nodes such as "SELECT", "FROM", "WHERE", "AND", and "=". The complexity score of each node is determined by multiple features, such as the nesting depth of the node, the number of table connections involved, and the complexity of the subquery. For example, the nesting depth of the "WHERE" node is 1, and the nesting depth of the "AND" node is 2. Assume that the complexity of the "SELECT" node is 1, the complexity of the "FROM" node is 1, the complexity of the "WHERE" node is 2, the complexity of the "AND" node is 3, and the complexity of the "=" node is 1. Finally, the complexity scores of all nodes are summed up to obtain the complexity score of the entire query statement, for example, the complexity score in this example is 1+1+2+3+1+1=9.
[0109] Next, a multidimensional data scale assessment model is established. This model maintains table-level and column-level statistics, such as the amount of data in the table and the data distribution in the column. At the same time, the model can also estimate the amount of data that the query actually needs to access based on the filter conditions and associations. For example, for table table1, assuming that the total data volume is 1000 rows, the filter condition selection rate of col1=1 is 0.1, and the filter condition selection rate of col2=2 is 0.2. Since there is no association between col1 and col2, the amount of data that needs to be accessed is estimated to be 1000×0.1×0.2 = 20 rows. This model uses a histogram to represent the distribution of data, so that the amount of data can be estimated more accurately.
[0110] Then, based on the complexity score and data scale evaluation results, a cost calculation function for the query path is constructed. This function takes into account factors such as the complexity of the path, the data scale, and the historical execution time to calculate the cost of each candidate path. For example, suppose there are two candidate paths. The complexity score of path 1 is 9, the amount of data to be accessed is 20 rows, and the historical execution time is 10ms; the complexity score of path 2 is 12, the amount of data to be accessed is 10 rows, and the historical execution time is 15ms. The cost calculation function can be defined as: cost = complexity score × data scale + historical execution time. Therefore, the cost of path 1 is 9×20+10=190, and the cost of path 2 is 12×10+15 =135.
[0111] Finally, an improved dynamic programming algorithm is used to search all candidate paths to find the path with the lowest cost. During the search process, path pruning, local optimal solution caching, and multi-path parallel evaluation are performed to improve search efficiency. For example, if the intermediate cost of a path has exceeded the total cost of the currently found optimal path, the path can be directly pruned. After the search is completed, the final query execution plan is generated based on the optimal path, which contains information such as node execution order, data access method, and resource allocation. For example, if path 2 is finally selected, the generated execution plan will specify to access col2 first, then col1, and allocate corresponding computing resources.
[0112] The present invention can significantly reduce the query execution time and improve query efficiency by accurately evaluating query complexity and data size and selecting the optimal query path; selecting the optimal query path can reduce the CPU, memory, IO and other resource consumption of the database and improve the overall performance of the database; by optimizing the query path, it can avoid executing inefficient queries and prevent the database from being overloaded, thereby enhancing the stability of the system.
[0113] In an optional embodiment,
[0114] Constructing a cost calculation function for a query path, the cost calculation function integrates path complexity score information, data scale information, and historical execution time information, using an improved dynamic programming algorithm to perform path search on the calculation result of the cost calculation function, the improved dynamic programming algorithm calculates a local optimal solution through a state transfer equation, and obtaining an optimal query path through a backtracking operation includes the following steps:
[0115] Constructing a hierarchical and progressive multi-factor cost calculation model, the multi-factor cost calculation model includes a path complexity scoring module, a data scale evaluation module and a historical execution statistics module, wherein the path complexity scoring module includes processor calculation complexity evaluation and connection operation cost evaluation, the data scale evaluation module includes intermediate result set evaluation and data skew degree quantification, the historical execution statistics module includes execution time analysis and resource consumption evaluation, and generating initial cost evaluation information based on the multi-factor cost calculation model;
[0116] According to the initial cost evaluation information, an influence factor regulator based on deep Q learning is designed, and the influence factor regulator includes a state space, an action space and a reward function, wherein the state space includes current load information and historical load information, the action space includes adjustment range information of each influence factor, and the reward function is composed of query execution time, resource utilization rate and query success rate, and the reward value is obtained by weighted combination calculation; the deep Q learning adopts a dual network structure, including a policy network and a target network, stores and samples training data through an experience replay mechanism, and uses the processor utilization rate, memory occupancy and IO operation number based on the initial cost evaluation information as network input, and determines that the training converges when the Q value change rate is less than a preset threshold and the number of consecutive training rounds exceeds a preset number of rounds;
[0117] Executing a hierarchical dynamic programming algorithm based on the adjustment result of the influencing factor adjuster, including hierarchically constructing an initial execution plan based on the query syntax tree to generate a set of candidate plans, constructing a local cost model and performing path search based on an ant colony optimization algorithm based on pheromone concentration and heuristic factors, establishing a path dependency graph and identifying a target path;
[0118] The candidate plan set is optimized according to the path dependency graph, including using heuristic rules for fast pruning, decomposing the pruned plan to obtain a sub-path set, optimizing each sub-path and updating global state information using local update rules and global update rules of an ant colony algorithm, dynamically adjusting the resource allocation strategy according to the system load state, selecting the path with the highest score as the optimal query path by comparing the comprehensive performance scores of the optimized paths, and generating an optimized execution plan based on the optimal query path.
[0119] Exemplarily, first, a hierarchical and progressive multi-factor cost calculation model is constructed. The model consists of three modules: a path complexity scoring module, a data scale evaluation module, and a historical execution statistics module. The path complexity scoring module is further subdivided into processor computational complexity evaluation and join operation cost evaluation. For example, for a query containing a sorting operation, the processor computational complexity evaluation will consider the complexity of the sorting algorithm and the amount of data; the join operation cost evaluation will consider the type of join algorithm (such as hash join, nested loop join) and the size of the data table involved in the join. The data scale evaluation module includes intermediate result set evaluation and data skew degree quantification. For example, for a multi-table join query, the intermediate result set evaluation will estimate the size of the intermediate result set generated by each join operation; the data skew degree quantification will analyze the distribution of data on each node and quantify the degree of data skew. The historical execution statistics module includes execution time analysis and resource consumption evaluation. For example, this module will record the historical execution time, CPU usage, memory consumption and other information of the query. Based on the evaluation results of these modules, the initial cost evaluation information is generated. For example, the initial cost of a query can be expressed as the weighted sum of the evaluation values of each module.
[0120] Next, a deep Q-learning-based influencing factor regulator is designed based on the initial cost evaluation information. The regulator includes a state space, an action space, and a reward function. The state space contains current load information (such as CPU utilization, memory usage) and historical load information (such as the average CPU utilization over a period of time). The action space contains the adjustment range information of each influencing factor. For example, the weights of the three influencing factors of path complexity, data scale, and historical execution time in the overall cost evaluation can be adjusted. The reward function consists of query execution time, resource utilization, and query success rate, and the reward value is calculated by weighted combination. For example, a query with short execution time, low resource utilization, and successful execution will receive a higher reward value. Deep Q learning adopts a dual network structure, namely a policy network and a target network, and uses an experience replay mechanism to store and sample training data. The processor utilization, memory usage, and IO operation number of the initial cost evaluation information are used as network inputs. When the Q value change rate is less than a preset threshold (such as 0.001) and the number of consecutive training rounds exceeds a preset number of rounds (such as 100 rounds), the training converges. Assume that after training, the regulator adjusts the weight of path complexity to 0.4, the weight of data size to 0.3, and the weight of historical execution time to 0.3.
[0121] Then, a hierarchical dynamic programming algorithm is executed based on the adjustment results of the influencing factor regulator. First, the initial execution plan is hierarchically constructed based on the query syntax tree, and a set of candidate plans is generated. For example, for a query containing connection and filtering operations, execution plans with different connection orders and filtering conditions can be generated. Then, a local cost model is constructed, and a path search is performed based on an ant colony optimization algorithm based on pheromone concentration and heuristic factors. For example, pheromone concentration can reflect the quality of a path, and heuristic factors can reflect the length or complexity of a path. Through the ant colony algorithm, a local optimal path can be found. Finally, a path dependency graph is established and the target path is identified. For example, a path dependency graph can represent the dependency relationship between different execution plans, and the target path is the execution plan with the lowest cost.
[0122] Finally, the candidate plan set is optimized according to the path dependency graph. First, heuristic rules are used for fast pruning. For example, obviously unreasonable execution plans can be pruned, such as placing the filter operation after the join operation. Then, the pruned plan is decomposed to obtain a set of subpaths. For example, a complex execution plan can be decomposed into multiple simple subpaths. Next, the local update rules and global update rules of the ant colony algorithm are used to optimize each subpath and update the global state information. For example, the pheromone concentration can be updated according to the execution status of the subpath. Then, the resource allocation strategy is dynamically adjusted according to the system load status. For example, if the system load is high, the resources allocated to a query can be reduced. Finally, by comparing the comprehensive performance scores of each optimized path, the path with the highest score is selected as the optimal query path, and the optimized execution plan is generated based on the path. For example, if the execution time and resource consumption of an optimized path are both the lowest, then the path is selected as the optimal query path.
[0123] The present invention uses a multi-factor cost model and deep Q learning to more accurately evaluate query costs and select lower-cost execution plans, thereby reducing query execution time and resource consumption; uses a hierarchical dynamic programming algorithm and an ant colony optimization algorithm to find the optimal query path, thereby improving query execution efficiency; and dynamically adjusts resource allocation strategies to avoid system overload, thereby enhancing system stability.
[0124] In an optional embodiment,
[0125] The steps of predicting resource usage based on the structure of the SQL statement and setting a monitoring threshold, monitoring the query status in real time during execution, and triggering a multi-level fault tolerance mechanism including breakpoint resumption, service degradation, and standby switching when an abnormality is detected include:
[0126] Constructing a SQL structure parsing model based on a syntax tree, wherein the SQL structure parsing model decomposes an SQL statement into an operation operator sequence, extracts table connection depth, predicate condition complexity, and aggregate function nesting level to form an SQL structure feature vector, quantitatively analyzes the structural features in the SQL structure feature vector, and establishes a resource occupancy prediction model in combination with the operation operator execution cost, and outputs a resource prediction value based on the resource occupancy prediction model;
[0127] Constructing a dynamic threshold adjustment algorithm to obtain a threshold adjustment result, wherein the dynamic threshold adjustment algorithm calculates a monitoring indicator alarm threshold by comprehensively considering resource prediction values and system load status, wherein the system load status is quantified by a load impact function including CPU utilization, memory occupancy, and IO waiting time, and the monitoring indicator alarm threshold is dynamically adjusted with query complexity and system load status;
[0128] Establishing a multi-level fault tolerance mechanism according to the threshold adjustment result, wherein the multi-level fault tolerance mechanism includes a breakpoint resume level, a service degradation level, and a standby switching level, and triggering corresponding levels of fault tolerance processing for different resource prediction value intervals;
[0129] The breakpoint resume level constructs recovery point information by recording the current execution state, intermediate result set and progress information, and performs query recovery based on the recovery point information; the service degradation level adjusts the execution parameters by calculating the degradation degree, so that the resource demand is reduced to below the target threshold, and re-executes the query based on the degraded execution parameters; the standby switching level selects an alternate execution plan from a preset set of alternate plans based on the current resource prediction value, and switches the query to the alternate execution plan;
[0130] A real-time status monitoring module is constructed to continuously collect performance indicators during the query execution process. When it is detected that the performance indicator exceeds the monitoring indicator alarm threshold, the abnormal information is fed back to the multi-level fault tolerance mechanism and the corresponding level of fault tolerance processing is triggered.
[0131] Exemplarily, first, a SQL structure parsing model is constructed. The model analyzes SQL statements based on syntax trees and decomposes them into a series of operation operators, such as SELECT, FROM, WHERE, JOIN, GROUP BY, etc. Then, SQL structural features are extracted, including table join depth (the number of nesting levels of JOIN statements), predicate condition complexity (the complexity of conditional expressions in the WHERE clause, such as the number and nesting level of comparison operators and logical operators), and aggregate function nesting level (aggregate functions, such as SUM, AVG, MAX, MIN, etc., the number of nested calls). Suppose there is an SQL statement: "SELECT COUNT(*)FROM orders JOIN customers ON orders.customer_id = customers.id WHEREorders.order_date > '2023-01-01' AND customers.country = 'USA' GROUP BYorders.status", then its table join depth is 1, predicate condition complexity is 2 (two conditional expressions), and aggregate function nesting level is 1. These structural features are quantified, such as the table join depth is a value of 1, the predicate condition complexity is a value of 2, and the aggregate function nesting level is a value of 1, to form an SQL structural feature vector. Finally, combined with the execution cost of the operation operator (such as the CPU time and IO number corresponding to each operation operator), a resource occupancy prediction model is established. For example, based on historical query data, the resource consumption corresponding to different SQL structural feature vectors can be counted, and a regression model can be trained to predict resource occupancy. Assume that according to historical data, the predicted CPU time of this SQL statement is 100ms and the predicted IO number is 500 times.
[0132] Next, a dynamic threshold adjustment algorithm is constructed. The algorithm takes into account the resource prediction value and the system load status and calculates the alarm threshold of the monitoring indicator. The system load status is quantified by the load impact function, which includes indicators such as CPU utilization, memory occupancy, and IO waiting time. For example, a load impact function can be set to calculate the system load status based on the weighted average of CPU utilization, memory occupancy, and IO waiting time, assuming that the current system load status is 0.8. The dynamic threshold adjustment algorithm dynamically adjusts the alarm threshold of the monitoring indicator based on the resource prediction value and the system load status. For example, based on the predicted CPU time and the system load status, the CPU time alarm threshold can be set to 1.5 times the predicted CPU time, that is, 150ms.
[0133] Then, a multi-level fault-tolerance mechanism is established based on the threshold adjustment results. The mechanism includes the breakpoint resume level, the service degradation level, and the standby switching level, which triggers the corresponding level of fault-tolerance processing for different resource prediction value intervals. For example, when the predicted CPU time is between 100ms and 200ms, the breakpoint resume level is triggered; when the predicted CPU time is between 200ms and 500ms, the service degradation level is triggered; when the predicted CPU time exceeds 500ms, the standby switching level is triggered. The breakpoint resume level records the current execution status, intermediate result set, and progress information to facilitate query recovery in the event of a failure. The service degradation level reduces resource requirements to below the target threshold by adjusting execution parameters, such as reducing the amount of data returned or reducing query accuracy. The standby switching level selects a suitable standby execution plan from a preset set of standby plans and switches the query to that plan.
[0134] Finally, a real-time status monitoring module is built. This module continuously collects performance indicators during query execution, such as CPU time, IO times, etc. When it is detected that the performance indicator exceeds the monitoring indicator alarm threshold, for example, the actual CPU time exceeds 150ms, the abnormal information is fed back to the multi-level fault tolerance mechanism and the corresponding level of fault tolerance processing is triggered, such as starting the breakpoint resume mechanism.
[0135] The present invention can predict the execution time and resource consumption of queries in advance by accurately predicting resource occupancy, and optimize according to the prediction results, thereby improving query efficiency; by real-time monitoring of query status and triggering a multi-level fault-tolerant mechanism, it can effectively respond to various abnormal situations and prevent query failures or system crashes, thereby enhancing system stability; by dynamically adjusting monitoring thresholds and service degradation mechanisms, resource allocation can be dynamically adjusted according to system load status to avoid resource waste, thereby improving resource utilization.
[0136] In an optional embodiment,
[0137] A multi-level fault tolerance mechanism is established according to the threshold adjustment result, wherein the multi-level fault tolerance mechanism includes a breakpoint resume level, a service degradation level, and a standby switching level. The steps of triggering corresponding levels of fault tolerance processing for different resource prediction value intervals include:
[0138] An adaptive interval division module is constructed based on the historical load data of the system. The adaptive interval division module divides the resource prediction value into a low load interval, a medium load interval and a high load interval by using the quartile method, and introduces a dynamic adjustment factor including the current load and performance index of the system. The interval boundary is adjusted in real time according to the dynamic adjustment factor, and the dynamic interval division result is output;
[0139] A fault tolerance level trigger module is constructed based on the dynamic interval division result, and the fault tolerance level trigger module includes a breakpoint resume trigger unit, a service degradation trigger unit and a standby switching trigger unit, wherein the breakpoint resume trigger unit is triggered when the resource prediction value is in a low load interval and the query state score exceeds a preset threshold, the service degradation trigger unit is triggered when the resource prediction value is in a medium load interval and the performance index is lower than the target value, and the standby switching trigger unit is triggered when the resource prediction value is in a high load interval or an abnormality is detected;
[0140] Construct a state machine-based collaborative scheduling module, the collaborative scheduling module includes a normal state, a breakpoint resume state, a service degradation state and a standby switching state, establishes a state transition matrix based on the trigger signal of the fault tolerance level trigger module, and defines the transition conditions between states; the collaborative scheduling module allocates resource quotas to different fault tolerance states according to the query priority weight and the system load state, and calculates the state transition cost based on resource overhead, switching delay and performance loss; after receiving the fault tolerance level trigger signal, the collaborative scheduling module determines the target state according to the state transition matrix, selects the optimal transition path based on the state transition cost, and executes the state switching according to the resource quota allocation result;
[0141] A fault-tolerance effect evaluation module is constructed, which collects performance data during the state switching process and feeds the performance data back to the adaptive interval division module for optimizing the interval boundary and the dynamic adjustment factor.
[0142] Exemplarily, first, an adaptive interval division module is constructed. This module uses the system's historical load data and the quartile method to divide the resource prediction value into a low load interval, a medium load interval, and a high load interval. For example, the CPU utilization data of the past year is collected, sorted from small to large, and the data is divided into four equal parts to obtain three dividing points, namely the 25% quantile, the 50% quantile, and the 75% quantile, which correspond to the upper limits of the low, medium, and high load intervals, respectively. In order to adapt to the real-time load changes of the system, a dynamic adjustment factor is introduced. This factor comprehensively considers the current load and performance indicators of the system, such as the current CPU utilization, memory usage, response time, etc. Assuming that the current CPU utilization is higher than the historical average level and the response time has increased, the dynamic adjustment factor will adjust the interval boundary upward, so that the upper limit of the low load interval becomes higher, and the range of the medium load interval and the high load interval is also adjusted accordingly. For example, the original upper limit of the low load interval is 30%, which may become 40% after dynamic adjustment.
[0143] Next, build a fault tolerance level trigger module. This module contains a breakpoint resume trigger unit, a service degradation trigger unit, and a standby switching trigger unit. When the resource prediction value is in the low load range and the query status score exceeds the preset threshold (for example, the score exceeds 90, indicating that the query task complexity is high), breakpoint resume is triggered. When the resource prediction value is in the medium load range and the performance indicator (for example, the average response time) is lower than the target value (for example, the response time exceeds 200 milliseconds), service degradation is triggered. When the resource prediction value is in the high load range, or an abnormal situation is detected (for example, a database connection failure), a standby switching is triggered.
[0144] Subsequently, a collaborative scheduling module based on a state machine is constructed. This module includes a normal state, a breakpoint resume state, a service degradation state, and an alternate switching state. According to the trigger signal of the fault tolerance level trigger module, a state transition matrix is established to define the transition conditions between states. For example, the normal state can be converted to the breakpoint resume state, the service degradation state, or the alternate switching state, but the breakpoint resume state cannot be directly converted to the alternate switching state. The collaborative scheduling module allocates resource quotas for different fault tolerance states according to the query priority weight and the system load state. For example, under high load conditions, the resource quota of the core business is prioritized. At the same time, the module calculates the state transition cost based on resource overhead, switching delay, and performance loss. For example, the cost of switching from the normal state to the alternate switching state is high because the backup system needs to be started and data synchronization is performed. After receiving the fault tolerance level trigger signal, the collaborative scheduling module determines the target state according to the state transition matrix and selects the optimal transition path based on the state transition cost. For example, if the service degradation signal is triggered, but the current system load is high and the cost of switching to the alternate switching state is too high, it will choose to maintain the service degradation state. Finally, the module performs state switching according to the resource quota allocation results.
[0145] Finally, a fault tolerance effect evaluation module is constructed. This module collects performance data during the state switching process, such as response time, error rate, resource utilization, etc., and feeds this data back to the adaptive interval division module to optimize the interval boundary and dynamic adjustment factor. For example, if service degradation is frequently triggered, it means that the division of the medium load interval may be too narrow and the interval boundary needs to be adjusted.
[0146] The present invention can effectively cope with resource load fluctuations of varying degrees and avoid system crashes through a multi-level fault-tolerant mechanism and a dynamic adjustment strategy; dynamically adjust resource allocation based on state machine collaborative scheduling and resource quota allocation to avoid resource waste; and minimize service interruption time and reduce user-perceived delays and error rates through mechanisms such as breakpoint resumption and service degradation.
[0147] The present invention also provides an embodiment:
[0148] The first module: user interface interaction module, including the following functional units:
[0149] The data source configuration and display unit includes a configuration page that provides functions for adding, editing and deleting data sources, and a display page that dynamically generates and displays a data source list based on user configuration, wherein the data source information includes the data source type, connection information and configuration parameters; the database and data table selection unit includes a display of a database list under the selected data source and a data table list under the selected database, and supports users to perform hierarchical selection of databases and data tables by clicking; the SQL query and natural language conversion unit includes an SQL query input interface with SQL syntax highlighting and automatic completion functions, and a natural language to SQL function interface for non-technical personnel; the query result display unit includes presenting the query results in a table format, and providing operation functions such as result sorting and filtering.
[0150] The second module: User rights management module, including the following functional units:
[0151] The permission verification unit is used to filter and display the database and data table lists according to user permissions, and to perform permission verification during query operations; the permission configuration unit is used by administrators to configure data source access permissions, database access permissions, and data table access permissions for different users.
[0152] The third module: user lock and retry module, including the following functional units:
[0153] The user lock unit implements a lock mechanism based on the unique identifier of the user ID and the query request through Redis, sets a predetermined lock validity period, and prevents repeated queries; the retry mechanism unit provides up to three automatic retries when the query request fails, sets a predetermined waiting time between each retry, and issues a failure prompt to the user when the number of retries is exceeded.
[0154] The fourth module: data body construction module, including the following processing units:
[0155] First, a collecting unit is used to collect the basic configuration information selected by the user on the interface, including the data source type, data source unique identifier, target database name and data table name specified by the user; secondly, a constructing unit is used to construct a data template containing the basic configuration information, the data template includes a data source type identification item, a data source unique identification item, a model request unique identification item and a database information set, wherein each database information item in the database information set includes a database name and a data table information set, each data table information item in the data table information set includes a table name and a column information set, and each column information item in the column information set includes a column name and a column data type; then, a filling unit is used to fill the query requirement item in the data template with the query requirement in the natural language form input by the user; finally, a sending unit is used to send the filled data template to the large model server based on a preset network transmission protocol, wherein the preset network transmission protocol includes identity authentication information and request type information.
[0156] The fifth module: large model processing module, including the following processing units:
[0157] A receiving and parsing unit is used to receive and parse request data, and save the requested unique identifier, data source information, database information, data table information, user speech and other information to the system database; a query unit is used to query the vector database based on the data source information to obtain metadata information, and the vector database obtains the latest metadata information from the data management module through regular updates or real-time queries; a generating unit is used to obtain corresponding data templates according to different data source types, fill user speech and metadata information into the data template to generate SQL statements, and call the LLAVA large model to parse and optimize the SQL statements to generate output results including query result tables and explanation suggestions.
[0158] The sixth module: result return and display module, including the following processing units:
[0159] The return unit is used to return the processing results generated by the large model system, including the query result table, query statements and execution time information, to the front-end system; the display unit is used to present the processing results on the result display page of the front-end system and provide viewing and downloading functions; the recording unit is used to record the query log information for subsequent analysis and optimization.
[0160] Figure 2 FIG. 1 is a schematic diagram of a query system for generating SQL statements based on natural language according to an embodiment of the present invention. Figure 2 As shown, the system comprises:
[0161] The first unit is used to receive a user query request and perform identity authentication, and obtain data source configuration information including a data source type, a database name, and a table name; based on the identity authentication result and the data source configuration information, a multi-level lock identifier is obtained through a distributed lock management module in a preset order, wherein the multi-level lock identifier includes a user session-level lock identifier, a database connection pool lock identifier, and an SQL execution lock identifier, and the distributed lock management module dynamically adjusts the timeout time of each level of lock identifier based on the historical query load; based on the data source configuration information and the user's historical query records, a query pattern is extracted from a query pattern knowledge base of concurrent control to generate a data template, wherein the data template includes a data source identifier, a model request identifier, database information, and user query content;
[0162] The second unit is used to perform query path evaluation on the data template, wherein the query path evaluation includes query complexity evaluation and data scale evaluation, and select an optimal query path according to the evaluation result; according to the optimal query path, a large model server is called through an HTTP request to parse the data template to generate an SQL statement;
[0163] The third unit is used to predict resource occupancy and set monitoring thresholds based on the structure of the SQL statement, monitor the query status in real time during execution, and trigger a multi-level fault-tolerant mechanism including breakpoint resumption, service degradation and standby switching when an abnormality is detected; record the abnormal information and analyze the cause of the error in combination with the SQL execution plan, adjust the query strategy according to the analysis results, and complete query recovery; release the multi-level locking identifier, return the query result to the user, and update the query information to the query pattern knowledge base in an asynchronous manner.
[0164] According to a third aspect of the embodiments of the present invention,
[0165] An electronic device is provided, comprising:
[0166] processor;
[0167] a memory for storing processor-executable instructions;
[0168] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0169] According to a fourth aspect of the embodiments of the present invention,
[0170] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0171] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A query method for generating SQL statements based on natural language, characterized in that: include: Receive user query requests and perform identity authentication to obtain data source configuration information including data source type, database name, and table name; Based on the identity authentication result and the data source configuration information, a multi-level lock identifier is obtained through a distributed lock management module in a preset order, wherein the multi-level lock identifier includes a user session-level lock identifier, a database connection pool lock identifier, and an SQL execution lock identifier. The distributed lock management module dynamically adjusts the timeout time of each level of lock identifier based on the historical query load; Extracting a query pattern from a concurrent control query pattern knowledge base based on the data source configuration information and the user's historical query records, and generating a data template, wherein the data template includes a data source identifier, a model request identifier, database information, and user query content; Performing query path evaluation on the data template, wherein the query path evaluation includes query complexity evaluation and data scale evaluation, and selecting an optimal query path according to the evaluation result; According to the optimal query path, a large model server is called through an HTTP request to parse the data template and generate an SQL statement; Predict resource usage based on the structure of the SQL statement and set monitoring thresholds, monitor query status in real time during execution, and trigger a multi-level fault tolerance mechanism including breakpoint resumption, service degradation, and standby switching when an anomaly is detected; record the anomaly information and analyze the cause of the error in combination with the SQL execution plan, adjust the query strategy based on the analysis results, and complete query recovery; releasing the multi-level locking identifier, returning the query result to the user, and updating the query information to the query pattern knowledge base in an asynchronous manner; The steps of obtaining a multi-level lock identifier through a distributed lock management module in a preset order, wherein the multi-level lock identifier includes a user session level lock identifier, a database connection pool lock identifier, and an SQL execution lock identifier, and the distributed lock management module dynamically adjusts the timeout time of each level of lock identifier based on the historical query load includes: Acquire multi-level lock identifiers in a preset order, and acquire user session-level lock identifiers through a user session management unit using distributed key-value storage. The user session-level lock identifier generates a lock key value by combining a session number and a timestamp, and ensures mutually exclusive access through atomic operations; Based on the user session-level lock identifier, a temporary sequential node is established through a database connection pool management unit to obtain a database connection pool lock identifier, wherein the database connection pool lock identifier uses a fair lock strategy to allocate connection resources and maintains resource status through a node monitoring mechanism; Based on the database connection pool lock identifier, a lease mechanism is established through the SQL execution management unit to obtain the SQL execution lock identifier, the SQL execution lock identifier ensures concurrency safety through comparison and exchange operations, and sets a lease validity period; Using a sliding time window to count query load information of lock identifiers at all levels, the query load information includes query frequency, execution time distribution and resource usage peak, and calculating the benchmark execution time of lock identifiers at all levels based on the query load information; A feedback control mechanism is used to dynamically adjust the timeout time of each level of locking identification. The feedback control mechanism adjusts the parameters by calculating the execution time error and triggers gradient parameter adjustment in the case of a sudden load change.
2. The method according to claim 1, characterized in that The steps of extracting a query pattern from a query pattern knowledge base of concurrent control based on the data source configuration information and the user's historical query records, and generating a data template, wherein the data template includes a data source identifier, a model request identifier, database information, and user query content include: Parsing the data source configuration information to form data source description information; constructing a multidimensional feature vector space based on the data source description information, wherein the multidimensional feature vector space includes database structure features, load features, timing features, and business features, wherein the database structure features include table node information and inter-table association information, the load features include resource density information, the timing features include execution time mode information and periodicity intensity information, and the business features include business domain information; Feature extraction is performed on the user's historical query records, and the extracted query features are mapped to the multidimensional feature vector space. The improved DBSCAN clustering algorithm is used for clustering analysis. The improved DBSCAN clustering algorithm combines the edit distance and the feature vector cosine similarity to define the distance function, dynamically adjusts the density threshold based on the query distribution, and supports online mode evolution through incremental updates; based on the clustering analysis results, a query mode knowledge base for concurrency control is constructed. The query mode knowledge base adopts a knowledge graph structure, and the mode applicability condition information, performance characteristic information, and resource requirement information are respectively marked on the nodes and edges; Extracting query patterns from the query pattern knowledge base according to the data source description information and the knowledge graph structure; using a context-aware pattern matching algorithm in the extraction process; and performing a comprehensive evaluation by calculating a time correlation score, a load compatibility score, and a business alignment score; wherein the time correlation score is calculated based on the execution time window and periodic characteristics of the query pattern to determine the degree of matching with the current query; the load compatibility score is calculated based on the degree of matching between the resource requirements of the query pattern and the current system load state; and the business alignment score is calculated based on the degree of conformity between the query pattern and the current business scenario constraints; and selecting the optimal query pattern based on the comprehensive evaluation result; A data template is generated according to the extracted query pattern, and an adaptive template generation strategy is adopted to construct a basic template and select a concurrent control strategy; a feedback mechanism is established for the execution process of the data template, and the evaluation result is used to update the query pattern knowledge base.
3. The method according to claim 1, characterized in that The query path evaluation is performed on the data template, wherein the query path evaluation includes query complexity evaluation and data scale evaluation. The step of selecting the optimal query path according to the evaluation result includes: Constructing a complexity calculation model based on a query syntax tree, wherein the complexity calculation model performs complexity scoring on nodes in the query syntax tree, constructing a complexity feature vector based on a node nesting depth feature value, a table connection quantity feature value, and a subquery complexity feature value, and obtaining the complexity score by performing an inner product operation of a feature weight vector and a complexity feature vector; Establishing a multidimensional data scale assessment model, wherein the multidimensional data scale assessment model calculates table data volume information, filter condition selection rate information, and association relationship duplication information based on histogram data distribution, and maintains table-level data statistics information and column-level data distribution information; Based on the evaluation results of the complexity score and the multidimensional data scale evaluation model, a cost calculation function for the query path is constructed, wherein the cost calculation function integrates the path complexity score information, the data scale information and the historical execution time information, and an improved dynamic programming algorithm is used to perform path search on the calculation results of the cost calculation function, wherein the improved dynamic programming algorithm calculates the local optimal solution through the state transfer equation and obtains the optimal query path through the backtracking operation; during the path search process, a path pruning operation exceeding a threshold, a local optimal solution caching operation and a multi-path parallel evaluation operation are performed; A query execution plan is generated according to the optimal query path, wherein the query execution plan includes node execution sequence information, data access mode information and resource allocation information.
4. The method according to claim 3, characterized in that Constructing a cost calculation function for a query path, the cost calculation function integrates path complexity score information, data scale information, and historical execution time information, using an improved dynamic programming algorithm to perform path search on the calculation result of the cost calculation function, the improved dynamic programming algorithm calculates a local optimal solution through a state transfer equation, and obtaining an optimal query path through a backtracking operation includes the following steps: Constructing a hierarchical and progressive multi-factor cost calculation model, the multi-factor cost calculation model includes a path complexity scoring module, a data scale evaluation module and a historical execution statistics module, wherein the path complexity scoring module includes processor calculation complexity evaluation and connection operation cost evaluation, the data scale evaluation module includes intermediate result set evaluation and data skew degree quantification, the historical execution statistics module includes execution time analysis and resource consumption evaluation, and generating initial cost evaluation information based on the multi-factor cost calculation model; According to the initial cost evaluation information, an influence factor regulator based on deep Q learning is designed, and the influence factor regulator includes a state space, an action space and a reward function, wherein the state space includes current load information and historical load information, the action space includes adjustment range information of each influence factor, and the reward function is composed of query execution time, resource utilization rate and query success rate, and the reward value is obtained by weighted combination calculation; the deep Q learning adopts a dual network structure, including a policy network and a target network, stores and samples training data through an experience replay mechanism, and uses the processor utilization rate, memory occupancy and IO operation number based on the initial cost evaluation information as network input, and determines that the training converges when the Q value change rate is less than a preset threshold and the number of consecutive training rounds exceeds a preset number of rounds; Executing a hierarchical dynamic programming algorithm based on the adjustment result of the influencing factor adjuster, including hierarchically constructing an initial execution plan based on the query syntax tree to generate a set of candidate plans, constructing a local cost model and performing path search based on an ant colony optimization algorithm based on pheromone concentration and heuristic factors, establishing a path dependency graph and identifying a target path; The candidate plan set is optimized according to the path dependency graph, including using heuristic rules for fast pruning, decomposing the pruned plan to obtain a sub-path set, optimizing each sub-path and updating global state information using local update rules and global update rules of an ant colony algorithm, dynamically adjusting the resource allocation strategy according to the system load state, selecting the path with the highest score as the optimal query path by comparing the comprehensive performance scores of the optimized paths, and generating an optimized execution plan based on the optimal query path.
5. The method according to claim 1, characterized in that The steps of predicting resource usage based on the structure of the SQL statement and setting a monitoring threshold, monitoring the query status in real time during execution, and triggering a multi-level fault tolerance mechanism including breakpoint resumption, service degradation, and standby switching when an abnormality is detected include: Constructing a SQL structure parsing model based on a syntax tree, wherein the SQL structure parsing model decomposes an SQL statement into an operation operator sequence, extracts table connection depth, predicate condition complexity, and aggregate function nesting level to form an SQL structure feature vector, quantitatively analyzes the structural features in the SQL structure feature vector, and establishes a resource occupancy prediction model in combination with the operation operator execution cost, and outputs a resource prediction value based on the resource occupancy prediction model; Constructing a dynamic threshold adjustment algorithm to obtain a threshold adjustment result, wherein the dynamic threshold adjustment algorithm calculates a monitoring indicator alarm threshold by comprehensively considering resource prediction values and system load status, wherein the system load status is quantified by a load impact function including CPU utilization, memory occupancy, and IO waiting time, and the monitoring indicator alarm threshold is dynamically adjusted with query complexity and system load status; Establishing a multi-level fault tolerance mechanism according to the threshold adjustment result, wherein the multi-level fault tolerance mechanism includes a breakpoint resume level, a service degradation level, and a standby switching level, and triggering corresponding levels of fault tolerance processing for different resource prediction value intervals; The breakpoint resume level constructs recovery point information by recording the current execution state, intermediate result set and progress information, and performs query recovery based on the recovery point information; the service degradation level adjusts the execution parameters by calculating the degradation degree, so that the resource demand is reduced to below the target threshold, and re-executes the query based on the degraded execution parameters; the standby switching level selects an alternate execution plan from a preset set of alternate plans based on the current resource prediction value, and switches the query to the alternate execution plan; A real-time status monitoring module is constructed to continuously collect performance indicators during the query execution process. When it is detected that the performance indicator exceeds the monitoring indicator alarm threshold, the abnormal information is fed back to the multi-level fault tolerance mechanism and the corresponding level of fault tolerance processing is triggered.
6. The method according to claim 5, characterized in that A multi-level fault tolerance mechanism is established according to the threshold adjustment result, wherein the multi-level fault tolerance mechanism includes a breakpoint resume level, a service degradation level, and a standby switching level. The steps of triggering corresponding levels of fault tolerance processing for different resource prediction value intervals include: An adaptive interval division module is constructed based on the historical load data of the system. The adaptive interval division module divides the resource prediction value into a low load interval, a medium load interval and a high load interval by using the quartile method, and introduces a dynamic adjustment factor including the current load and performance index of the system. The interval boundary is adjusted in real time according to the dynamic adjustment factor, and the dynamic interval division result is output; A fault tolerance level trigger module is constructed based on the dynamic interval division result, and the fault tolerance level trigger module includes a breakpoint resume trigger unit, a service degradation trigger unit and a standby switching trigger unit, wherein the breakpoint resume trigger unit is triggered when the resource prediction value is in a low load interval and the query state score exceeds a preset threshold, the service degradation trigger unit is triggered when the resource prediction value is in a medium load interval and the performance index is lower than the target value, and the standby switching trigger unit is triggered when the resource prediction value is in a high load interval or an abnormality is detected; Construct a state machine-based collaborative scheduling module, the collaborative scheduling module includes a normal state, a breakpoint resume state, a service degradation state and a standby switching state, establishes a state transition matrix based on the trigger signal of the fault tolerance level trigger module, and defines the transition conditions between states; the collaborative scheduling module allocates resource quotas to different fault tolerance states according to the query priority weight and the system load state, and calculates the state transition cost based on resource overhead, switching delay and performance loss; after receiving the fault tolerance level trigger signal, the collaborative scheduling module determines the target state according to the state transition matrix, selects the optimal transition path based on the state transition cost, and executes the state switching according to the resource quota allocation result; A fault-tolerance effect evaluation module is constructed, which collects performance data during the state switching process and feeds the performance data back to the adaptive interval division module for optimizing the interval boundary and the dynamic adjustment factor.
7. A query system for generating SQL statements based on natural language, used to implement the method according to any one of claims 1 to 6, characterized in that: include: The first unit is used to receive a user query request and perform identity authentication, and obtain data source configuration information including a data source type, a database name, and a table name; Based on the identity authentication result and the data source configuration information, a multi-level lock identifier is obtained through a distributed lock management module in a preset order, wherein the multi-level lock identifier includes a user session-level lock identifier, a database connection pool lock identifier, and an SQL execution lock identifier. The distributed lock management module dynamically adjusts the timeout time of each level of lock identifier based on the historical query load; Extracting a query pattern from a concurrent control query pattern knowledge base based on the data source configuration information and the user's historical query records, and generating a data template, wherein the data template includes a data source identifier, a model request identifier, database information, and user query content; The second unit is used to perform query path evaluation on the data template, wherein the query path evaluation includes query complexity evaluation and data scale evaluation, and select an optimal query path according to the evaluation result; According to the optimal query path, a large model server is called through an HTTP request to parse the data template and generate an SQL statement; The third unit is used to predict resource usage and set monitoring thresholds based on the structure of the SQL statement, monitor the query status in real time during execution, and trigger a multi-level fault tolerance mechanism including breakpoint resumption, service degradation, and standby switching when an abnormality is detected; record abnormal information and analyze the cause of the error in combination with the SQL execution plan, adjust the query strategy according to the analysis results, and complete query recovery; The multi-level locking identifier is released, the query result is returned to the user, and the query information is updated to the query pattern knowledge base in an asynchronous manner.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Distributed multi-level storage system and method based on HDFS
CN109284258A
Distributed database parallel query method, system and device and storage medium
CN116662395A