Database cascade operation intelligent analysis execution method and system

By building a dependency weight matrix and multi-dimensional feature classification, combined with adaptive path optimization and neural network model, the problems of low execution efficiency, low resource utilization and improper exception handling in database cascading operations are solved, and efficient and reliable database cascading operations are achieved.

CN120256103APending Publication Date: 2025-07-04NORTH CHINA MUNICIPAL ENG DESIGN & RES INST
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510325355.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When handling database cascading operations, the prior art fails to fully consider the dynamic characteristics of the correlation between tables, resulting in lack of flexibility in execution plans, insufficient resource allocation, low execution efficiency, and lack of an effective exception handling mechanism, making it difficult to ensure data consistency.

Method used

The dependency weight matrix is ​​constructed through a recursive query algorithm, the circular dependency link is identified and decomposed into directed ring-free subgraphs, and task scheduling is used using multi-dimensional feature classification and adaptive path optimization, performance prediction and resource adjustment are combined with neural network models, and data consistency is ensured through link tracing and cascading rollback.

Benefits of technology

It improves the execution efficiency and stability of database cascading operations, optimizes resource utilization, ensures data consistency, quickly locates and handles exceptions, and improves system reliability and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256103A_ABST
    Figure CN120256103A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent analysis and execution method and system for database cascading operation, and relates to the technical field of database management.The intelligent analysis and execution method comprises the steps that a target table and an association table are extracted through a recursive query algorithm, field mapping analysis is carried out to generate a dependency weight matrix, cyclic dependency is recognized based on the matrix and decomposed into directed acyclic subgraphs, and the directed acyclic subgraphs are analyzed; node depth values and breadth values are calculated for topological sorting to generate an execution sequence; performing node classification on the execution sequence, organizing the execution sequence into a batch processing task group, calculating an optimal execution path through an adaptive path optimization algorithm, generating a distributed transaction control instruction, and allocating the task group to an execution thread pool for execution by using a consistent Hash algorithm; monitoring an execution process, dynamically adjusting thread pool resources according to a performance prediction value, executing cascade rollback when an exception occurs, and reconstructing a transaction execution path; according to the method, the execution efficiency and stability of the database cascade operation can be effectively improved, and the system resource consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field, and in particular to an intelligent parsing and execution method and system for database cascade operations. Background Art

[0002] Database cascade operations are a series of consecutive database operations executed based on the inter-table association relationship, and are widely used in scenarios such as distributed data processing and business process automation. With the increase in business complexity, the association relationships between database tables are becoming increasingly complex, and the setting of various referential integrity constraints makes a single database operation may trigger a series of chain reactions. Especially in the microservices architecture, cross-service database operations will form a complex dependency network, and how to efficiently parse and execute these cascade operations has become an important technical challenge. The technical fields involved include multiple aspects such as database management, distributed computing, and transaction processing.

[0003] However, the prior art usually adopts a static parsing method to process database cascade operations, and there are mainly the following problems: the dynamic characteristics of the inter-table association are not fully considered during the parsing process, resulting in the generated execution plan lacking flexibility and being unable to adapt to business changes; a fixed task allocation strategy is adopted, and the resource allocation cannot be dynamically adjusted according to the actual load situation, resulting in low execution efficiency; there is a lack of an effective exception handling mechanism, and it is difficult to accurately locate the source of the problem and take remedial measures in a timely manner when an execution failure occurs; the existing parallel processing solutions are often too conservative and cannot fully utilize hardware resources to improve processing performance. These problems are particularly prominent in large-scale data processing scenarios.

[0004] In summary, there is an urgent need for an intelligent parsing and execution method for database cascade operations, which realizes dynamic parsing by constructing a dependency relationship weight matrix, adopts multi-dimensional feature classification and adaptive path optimization to realize intelligent task scheduling, combines a neural network model for performance prediction and dynamic resource adjustment, and ensures data consistency through link tracing and cascade rollback. Solve the technical problems such as low execution efficiency, poor reliability, and low resource utilization in the prior art, and provide an efficient and reliable database cascade operation solution. The technical solution of the present invention can solve the problems in the prior art. Summary of the Invention

[0005] An embodiment of the present invention provides an intelligent parsing and execution method and system for database cascade operations, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiment of the present invention,

[0007] Provided is an intelligent parsing and execution method for database cascade operations, including:

[0008] Receive a database cascade operation request, parse the database cascade operation request into an operation instruction set, extract a target table and associated tables from the operation instruction set through a recursive query algorithm, perform field mapping analysis to generate a dependency relationship weight matrix, and determine a cyclic dependency link based on the dependency relationship weight matrix through a strongly connected component identification algorithm; decompose the cyclic dependency link to obtain a directed acyclic subgraph, calculate the node depth value and node breadth value according to the directed acyclic subgraph, and generate a cascade operation execution sequence through a topological sorting algorithm;

[0009] Read the cascade operation execution sequence, perform node classification processing on the cascade operation execution sequence through a multi-dimensional feature classification algorithm, organize the classified nodes into a batch processing task group, calculate the optimal execution path of the batch processing task group through an adaptive path optimization algorithm, generate a distributed transaction control instruction according to the optimal execution path, implant a transaction synchronization point in the distributed transaction control instruction, the transaction synchronization point is set based on the two-phase commit protocol, allocate the batch processing task group to an execution thread pool through a consistent hashing algorithm, and the execution thread pool executes the batch processing task group based on a pipeline parallel processing mechanism to generate execution status information;

[0010] Monitor the running process of the execution thread pool, continuously collect the execution status information through a sliding window algorithm to generate performance data, input the performance data into a pre-trained neural network model to obtain a performance prediction value, when the performance prediction value exceeds a preset performance threshold, calculate the thread load level based on the execution status information, adjust the resources of the execution thread pool according to the thread load level, if the execution status information contains an abnormal signal, locate the abnormal transaction link through a link tracing and positioning algorithm, perform cascade rollback according to the abnormal transaction link, and reconstruct the transaction execution path according to the cascade operation execution sequence to complete the database cascade operation.

[0011] In an alternative embodiment,

[0012] Receive a database cascade operation request, parse the database cascade operation request into an operation instruction set, extract a target table and associated tables from the operation instruction set through a recursive query algorithm, perform field mapping analysis to generate a dependency relationship weight matrix, and determining a cyclic dependency link based on the dependency relationship weight matrix includes:

[0013] Convert the database cascading operation request into an abstract syntax tree through a preset syntax parser, and extract an operation instruction set based on the abstract syntax tree; based on the operation instruction set, use a recursive query algorithm of breadth-first search, starting from the target table in the operation instruction set, recursively query the associated tables through foreign key constraint information and join condition information. During the recursive query process, add the target table to the queue, sequentially take out the head node of the queue for expansion, maintain access marker information for the head node, record the associated field information and association type information between the associated table and the target table, and add the unvisited adjacent nodes to the tail of the queue to generate an association graph structure including the target table and the associated tables;

[0014] Perform field mapping on each pair of associated tables in the association graph structure, extract the field mapping relationship of each pair of associated tables, calculate the field mapping weight value based on the field mapping relationship, summarize the field mapping weight values to obtain the table-level dependency weight value, and construct a dependency relationship weight matrix;

[0015] Perform the first depth-first search traversal on the dependency relationship weight matrix, push the currently visited node onto the traversal stack, record the push timestamp and completion timestamp of the currently visited node, and construct a timestamp pair for the currently visited node; calculate the reachable node set of the currently visited node based on the traversal stack, and associate the reachable node set with the timestamp pair to form a node feature vector;

[0016] Transpose the dependency relationship weight matrix to obtain an inverse adjacency matrix, normalize the weight values in the inverse adjacency matrix to generate a normalized inverse adjacency matrix; map the node feature vector to the corresponding row of the normalized inverse adjacency matrix to construct a node dependency feature matrix;

[0017] Based on the descending order of the completion timestamps in the timestamp pair, perform the second depth-first search traversal on the node dependency feature matrix, and calculate the dependency strength value between adjacent nodes on the traversal path according to the corresponding weight value in the node dependency feature matrix and the node feature vector,

[0018] Filter the node association relationships with the dependency strength value greater than the preset dependency threshold based on a preset dependency threshold; perform connectivity analysis on the node association relationships, identify a node subset with bidirectional dependency characteristics in the node association relationships, and determine it as a cyclic dependency link; calculate the dependency strength value of each node according to the in-degree weight and out-degree weight of the node in the cyclic dependency link and the weight value in the node dependency feature matrix, and perform normalization processing on the dependency strength value to obtain a normalized dependency strength value; classify the dependency degree of the cyclic dependency link based on the numerical distribution of the normalized dependency strength value to determine the dependency degree of the cyclic dependency link.

[0019] In an alternative embodiment,

[0020] Decompose the cyclic dependency link to obtain a directed acyclic subgraph, calculate the node depth value and the node breadth value according to the directed acyclic subgraph, and generate a cascaded operation execution sequence through a topological sorting algorithm, including:

[0021] Statistically analyze the dependency trigger times of each edge in the cyclic dependency link in the historical execution data, use the dependency trigger times as the dependency strength value of the edge, calculate the normalized dependency strength value of the edge, obtain the set of loops where each edge in the cyclic dependency link is located, calculate the loop strength of each loop according to the product of the weight values of the edges in the loop, divide the loop strength by the loop length to obtain the loop unit strength, accumulate the loop unit strengths of all the loops where the edge is located to obtain the criticality index of the edge, and perform a weighted combination of the criticality index of the edge and the normalized dependency strength value of the edge to obtain the comprehensive evaluation value of the edge;

[0022] Construct a minimum heap containing the comprehensive evaluation values, sequentially take out the current edge with the smallest comprehensive evaluation value from the minimum heap, determine whether the deletion of the current edge causes the graph to split, if it does not split, add the current edge to the feedback edge set and update the comprehensive evaluation values of the adjacent edges, and repeat the execution until a directed acyclic subgraph and the corresponding remaining edge set are obtained;

[0023] In the directed acyclic subgraph, set the depth value of the starting node with an in-degree of zero to zero, set the breadth value of the terminating node with an out-degree of zero to zero, and according to the connection relationship of the remaining edge set, set the depth value of the non-starting node to the maximum value of the depth values of all its previous nodes plus one through forward traversal, and set the breadth value of the non-terminating node to the maximum value of the breadth values of all its subsequent nodes plus one through backward traversal;

[0024] Perform a weighted combination calculation on the depth value, breadth value, normalized value of the historical execution time of each node, and the number of associated edges of the node in the feedback edge set to obtain the node priority value;

[0025] Based on the node priority value, construct a priority queue, update the priority value of the node according to the execution status of the node in each time window and adjust the sorting of the nodes in the queue, select multiple current nodes that meet the parallel conditions with the highest priority value from the priority queue, and add the current nodes to the cascaded operation execution sequence in the selected order.

[0026] In an alternative embodiment,

[0027] Read the cascaded operation execution sequence, perform node classification processing on the cascaded operation execution sequence through a multi-dimensional feature classification algorithm, and organize the classified nodes into batch processing task groups, including:

[0028] Read the cascade operation execution sequence and construct a multi-layer information network, where the weight of the edges in the first-layer network represents the data transmission volume between nodes, the weight of the edges in the second-layer network represents the resource competition degree between nodes, and the weight of the edges in the third-layer network represents the operation similarity between nodes;

[0029] Based on the multi-layer information network, calculate the degree centrality, betweenness centrality, and closeness centrality of each node, determine the centrality index, and combine the centrality indexes to form the network structure feature vector of the node; obtain the state sequence of each node in the cascade operation execution sequence, extract the mean, variance, skewness, and kurtosis of the state sequence, determine the statistical features, and combine the statistical features to form the time series feature vector of the node; based on the input data scale, computational complexity, and memory occupancy of the node in the cascade operation execution sequence, determine the performance index, and construct the performance feature vector of the node based on the performance index; combine the network structure feature vector, time series feature vector, and performance feature vector to construct the feature matrix of the node;

[0030] Calculate the Mahalanobis distance between nodes in the feature matrix, and construct the similarity network of nodes based on the Mahalanobis distance; in the similarity network, obtain the steady-state distribution of nodes by iteratively calculating the transition probability between node pairs; based on the steady-state distribution, identify the densely connected subgraphs in the similarity network, and divide the nodes belonging to the same densely connected subgraph into the same category;

[0031] Calculate the path length between nodes within each category, and split the category with a path length greater than the preset path length threshold to obtain the optimized node classification result; organize the nodes with the same category into a batch processing task group.

[0032] In an alternative embodiment,

[0033] Calculate the optimal execution path of the batch processing task group through an adaptive path optimization algorithm, and generate distributed transaction control instructions according to the optimal execution path, including:

[0034] Based on each task node in the batch processing task group, calculate the load volatility by taking the ratio of the cumulative absolute value of the load value difference between adjacent time points within the sampling period of the task node to the sampling period;

[0035] Based on the weighted combination of the load volatility difference and the time overlap degree, construct a task node distance matrix, and perform hierarchical clustering on the batch processing task group according to the task node distance matrix to obtain task grouping information;

[0036] Construct an adaptive weight calculation model based on the task grouping information, input the system load data of the task node into the sigmoid function to obtain a normalized load value, and calculate the timeliness adjustment coefficient and resource adjustment coefficient of the task node according to the normalized load value;

[0037] Determine the timeliness evaluation value, resource evaluation value, and volatility evaluation value according to the timeliness adjustment coefficient, resource adjustment coefficient, and load volatility, and calculate the evaluation score of the task node by weighted calculation;

[0038] Construct a reinforcement learning model using the evaluation score as the state feature, construct an action space and a state space based on the state feature, and determine the reward function according to the evaluation score, resource competition degree, and parallelism;

[0039] Use the deep Q network to iteratively update the action value function to generate the selection strategy of the task node;

[0040] Perform path search on the batch task group based on the selection strategy, randomly select execution nodes to construct candidate paths in the exploration phase, and select the optimal execution nodes based on the action value function to construct the optimal execution path in the exploitation phase;

[0041] Construct a distributed transaction control instruction according to the optimal execution path. The distributed transaction control instruction includes a pre-commit instruction and a commit instruction, where the pre-commit instruction is determined by collecting the status information of the task nodes on the optimal execution path;

[0042] Maintain the causal relationship of the task nodes by incrementally updating and comparing the clock values of adjacent task nodes, record the execution order of the task nodes, determine the vector clock instruction, and embed the vector clock instruction in the distributed transaction control instruction;

[0043] Input the execution result of the vector clock instruction into the distributed consensus module, and coordinate the transaction status of the task nodes through the distributed consensus module to generate a distributed transaction control instruction.

[0044] In an alternative embodiment,

[0045] The link tracing and positioning algorithm includes:

[0046] Collect transaction execution information to generate a transaction identification number, which is encoded by a timestamp field, a machine identification field, and an incrementing serial number field to mark all database operations on the same transaction link; record the operation execution information of the database operation, including the operation identification number, the parent operation identification number, the operation type information, the operation start time, the operation end time, and the operation status information;

[0047] Build a call relationship tree for the operation execution information, use the operation identification number as a tree node, and use the parent operation identification number as the parent-child node association relationship;

[0048] Based on each node in the call relationship tree, determine the anomaly score by calculating the weighted combination of the operation anomaly benchmark score, the operation latency anomaly score, and the operation state transition anomaly score, and calculate the anomaly attenuation coefficient according to the hierarchical depth of the tree where the node is located. Multiply the anomaly attenuation coefficient by the anomaly score to obtain the final anomaly score of the node;

[0049] Calculate the time overlap degree between adjacent database operations, use the time overlap degree as the operation association strength, and build an operation dependency network based on the operation association strength; analyze the operation execution order in the operation dependency network, and calculate the conditional probability value between adjacent operations to determine the operation propagation probability;

[0050] Filter the operation association edges whose final anomaly score is greater than a preset first anomaly threshold and whose operation propagation probability is greater than a preset second anomaly threshold, and build an anomaly propagation subgraph;

[0051] Calculate the minimum weight spanning tree in the anomaly propagation subgraph to obtain the core path of anomaly propagation;

[0052] Count the in-degree values of each node in the core path, add the nodes with an in-degree value of zero to the queue to be processed, continuously take out the nodes in the queue and update the in-degree values of the corresponding adjacent nodes, and generate the propagation link of the anomaly transaction according to the order of node dequeueing.

[0053] In an alternative embodiment,

[0054] The cascading rollback includes:

[0055] Build a multi-dimensional transaction dependency graph, store transaction operations in layers according to timestamps, establish a read-write version chain for data items, record the read-write timing information and data version numbers of data items, and generate a transaction version dependency relationship;

[0056] Based on the transaction version dependency relationship, build a conflict matrix, calculate the data access overlap degree between adjacent transactions, use the data access overlap degree as the conflict weight, and identify the high-frequency conflict areas where the conflict weight exceeds the preset conflict threshold;

[0057] According to the conflict weight, calculate the criticality of the transaction node by accumulating the weighted values of the data modification amount, resource occupation duration, the conflict weight, and the transaction nesting level corresponding to the transaction node;

[0058] Classify the transaction nodes based on the criticality level, and merge the transaction nodes whose criticality level meets the preset approximation threshold into rollback batches according to the preset approximation threshold. Calculate the resource consumption value and influence range of each rollback batch, and generate batch priorities.

[0059] Generate a rollback execution plan according to the batch priorities, convert the rollback batches into parallel execution tasks in the order of batch priorities, and allocate rollback threads to each parallel execution task based on the resource consumption value.

[0060] Start a rollback executor, submit the parallel execution tasks to a rollback thread pool, record the start timestamp of each rollback operation, and perform concurrent control by comparing the data version numbers.

[0061] Monitor the rollback executor. When a rollback failure is detected, locate the failure position according to the timestamp, construct a compensation transaction to retry the rollback operation until the rollback is successful.

[0062] Collect the operation logs output by the rollback executor, extract the version number sequence from the operation logs, reconstruct the transaction execution path based on the version number sequence, and verify data consistency.

[0063] In the second aspect of the embodiments of the present invention,

[0064] Provide an intelligent parsing and execution system for database cascade operations, including:

[0065] A first unit, configured to receive a database cascade operation request, parse the database cascade operation request into an operation instruction set, extract a target table and associated tables from the operation instruction set through a recursive query algorithm, perform field mapping analysis to generate a dependency relationship weight matrix, and determine a cyclic dependency link based on the dependency relationship weight matrix through a strongly connected component identification algorithm; decompose the cyclic dependency link to obtain a directed acyclic subgraph, calculate the node depth value and node breadth value according to the directed acyclic subgraph, and generate a cascade operation execution sequence through a topological sorting algorithm.

[0066] A second unit, configured to read the cascade operation execution sequence, perform node classification processing on the cascade operation execution sequence through a multi-dimensional feature classification algorithm, organize the classified nodes into a batch processing task group, calculate the optimal execution path of the batch processing task group through an adaptive path optimization algorithm, generate a distributed transaction control instruction according to the optimal execution path, implant a transaction synchronization point in the distributed transaction control instruction, the transaction synchronization point is set based on the two-phase commit protocol, and allocate the batch processing task group to an execution thread pool through a consistent hashing algorithm, and the execution thread pool executes the batch processing task group based on a pipeline parallel processing mechanism to generate execution status information.

[0067] The third unit is used to monitor the running process of the execution thread pool, continuously collect the execution status information through a sliding window algorithm to generate performance data, input the performance data into a pre-trained neural network model to obtain a performance prediction value. When the performance prediction value exceeds a preset performance threshold, calculate the thread load level based on the execution status information, adjust the resources of the execution thread pool according to the thread load level. If an abnormal signal is included in the execution status information, locate the abnormal transaction link through a link tracing and positioning algorithm, perform cascading rollback according to the abnormal transaction link, and reconstruct the transaction execution path according to the cascading operation execution sequence to complete the database cascading operation.

[0068] In the third aspect of the embodiments of the present invention,

[0069] a kind of electronic device is provided, including:

[0070] a processor;

[0071] a memory for storing instructions executable by the processor;

[0072] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0073] In the fourth aspect of the embodiments of the present invention,

[0074] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0075] In the embodiments of the present invention, through algorithms such as recursive query, field mapping analysis, and strongly connected component identification, the database cascading operation request is intelligently parsed, and an optimal execution sequence is generated based on the dependency relationship and topological sorting, avoiding circular dependencies and performance bottlenecks, and improving the execution efficiency; adopting a distributed transaction control mechanism and a two-phase commit protocol to ensure data consistency; based on the consistent hashing algorithm, tasks are assigned to the execution thread pool, and the pipeline parallel processing mechanism is utilized to give full play to the advantages of multi-core processors, greatly improving the execution speed of cascading operations; the execution process is monitored in real time, performance data is collected through a sliding window algorithm, and performance prediction is performed through a neural network model. The resources of the execution thread pool are dynamically adjusted according to the prediction results and the thread load level to ensure the stable operation of the system; at the same time, it has the ability to trace abnormal transaction links and perform cascading rollback, effectively handling abnormal situations and ensuring data security. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 It is a schematic flowchart of the intelligent parsing and execution method for database cascading operations in the embodiments of the present invention;

[0077] Figure 2This is a schematic structural diagram of the intelligent parsing and execution system for database cascade operations in an embodiment of the present invention. Detailed implementation manners

[0078] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0079] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0080] Figure 1 This is a schematic flowchart of the intelligent parsing and execution method for database cascade operations in an embodiment of the present invention. As Figure 1 shown, the method includes:

[0081] S101. Receive a database cascade operation request, parse the database cascade operation request into an operation instruction set, extract a target table and associated tables from the operation instruction set through a recursive query algorithm, perform field mapping analysis to generate a dependency relationship weight matrix, determine a cyclic dependency link based on the dependency relationship weight matrix through a strongly connected component identification algorithm; decompose the cyclic dependency link to obtain a directed acyclic subgraph, calculate the node depth value and node breadth value according to the directed acyclic subgraph, and generate a cascade operation execution sequence through a topological sorting algorithm;

[0082] In this embodiment, the target table and associated tables are extracted through a recursive query algorithm, and a dependency relationship weight matrix is generated in combination with field mapping analysis to ensure that the dependency relationships of all data tables are accurately identified, providing an accurate execution basis for subsequent cascade operations; a strongly connected component identification algorithm is used to detect cyclic dependencies, and a directed acyclic subgraph is obtained through decomposition, effectively avoiding cyclic dependency problems in database operations, ensuring smooth transaction execution, and avoiding operation failures caused by deadlocks or conflicts; the depth value and breadth value of each node in the directed acyclic subgraph are calculated, and an optimal cascade operation execution sequence is generated in combination with a topological sorting algorithm to ensure that database operations are executed along the optimal path, reducing unnecessary waiting and rollbacks, and improving the overall operation efficiency; through a reasonable cascade operation order and dependency relationship processing, it is ensured that the execution of transactions strictly follows the database consistency constraints, preventing data inconsistency problems caused by incorrect operation orders, and improving the stability and reliability of the database system.

[0083] S102. Read the cascade operation execution sequence, perform node classification processing on the cascade operation execution sequence through a multi-dimensional feature classification algorithm, organize the classified nodes into batch processing task groups, calculate the optimal execution path of the batch processing task groups through an adaptive path optimization algorithm, generate a distributed transaction control instruction according to the optimal execution path, implant a transaction synchronization point in the distributed transaction control instruction, the transaction synchronization point is set based on the two-phase commit protocol, and distribute the batch processing task groups to the execution thread pool through a consistent hashing algorithm. The execution thread pool executes the batch processing task groups based on a pipeline parallel processing mechanism to generate execution status information;

[0084] In a specific implementation, first, the system reads the cascade operation execution sequence generated in the previous step and extracts feature information for each operation node in the sequence. The extracted features include multiple dimensions such as operation type, data table scale, number of associated tables, and estimated execution duration. Based on this feature information, a multi-dimensional feature classification algorithm is used to perform clustering analysis on the operation nodes, and nodes with similar features are classified into the same category. According to the classification results, nodes of the same category are organized into batch processing task groups, and each task group contains a set of operation nodes that can be executed in parallel.

[0085] Next, for each batch processing task group, an adaptive path optimization algorithm is used to calculate the optimal execution path. This algorithm first constructs a task dependency network, analyzes the dependency relationships between the nodes in the task group, and then combines the execution costs and resource requirements of the nodes to search for the execution path with the minimum overall execution cost through an iterative optimization method. Factors such as data locality and load balancing are considered during the optimization process to ensure that the generated execution path can not only ensure correctness but also achieve high execution efficiency.

[0086] According to the calculated optimal execution path, the system generates corresponding distributed transaction control instructions. During the process of generating the control instructions, the system implants transaction synchronization points at key execution nodes. These synchronization points are set based on the two-phase commit protocol and include a preparation phase and a commit phase. In the preparation phase, the system checks whether the execution conditions of all participating nodes are met; in the commit phase, after confirming that all nodes are ready, the actual data modification operation is executed.

[0087] Subsequently, the system uses a consistent hashing algorithm to distribute the batch processing task groups to specific threads in the execution thread pool. This algorithm constructs a hash ring, maps tasks and threads to different positions on the hash ring, ensures that tasks are evenly distributed to each thread, and minimizes the impact of task reallocation when the number of threads changes. The current load status of each thread is considered during the distribution process to avoid load imbalance.

[0088] Finally, the execution thread pool executes the assigned tasks using a pipeline parallel processing mechanism. During the execution process, each thread maintains a task execution queue and processes the tasks in the queue in a pipeline manner. The pipeline processing is divided into multiple stages, including task preprocessing, data loading, execution operations, result verification, etc. Each stage collaborates asynchronously to achieve pipelining of the processing process and improve the overall execution efficiency.

[0089] In this embodiment, the multi-dimensional feature classification algorithm is used to classify nodes in the cascade operation execution sequence and organize them into a batch processing task group, reducing the execution redundancy of tasks and improving the overall execution efficiency; the adaptive path optimization algorithm is used to calculate the optimal execution path of the batch processing task group to ensure that tasks are executed in the optimal manner in a distributed environment, reducing the occupation of computing resources and improving the system throughput; transaction synchronization points based on the two-phase commit protocol are implanted in the distributed transaction control instructions to ensure the atomicity of cross-node transaction operations and reduce the risk of data inconsistency caused by transaction failures; the batch processing task group is reasonably allocated to the execution thread pool in combination with the consistent hashing algorithm, and the pipeline parallel processing mechanism is used to efficiently execute tasks, maximizing the system concurrency ability and improving the overall transaction processing performance.

[0090] S103. Monitor the running process of the execution thread pool, continuously collect the execution status information through the sliding window algorithm to generate performance data, input the performance data into a pre-trained neural network model to obtain a performance prediction value. When the performance prediction value exceeds the preset performance threshold, calculate the thread load level based on the execution status information, adjust the resources of the execution thread pool according to the thread load level. If the execution status information contains an abnormal signal, locate the abnormal transaction link through the link tracing and positioning algorithm, perform cascade rollback according to the abnormal transaction link, and reconstruct the transaction execution path according to the cascade operation execution sequence to complete the database cascade operation.

[0091] In a specific implementation manner, first, the system starts a dedicated monitoring module to monitor the running status of the execution thread pool in real time. The monitoring scope includes key indicators such as the task execution situation, resource occupation situation, and response time of each thread. Through the sliding window algorithm, the system continuously collects this execution status information within a fixed time window. The sliding window is updated over time to maintain the latest status data. At the same time, through data aggregation processing, various indicators reflecting the system performance are generated to form a complete performance data set.

[0092] Next, the system inputs the collected performance data into a pre-trained neural network model. By learning the characteristics of historical execution data, this model can predict the system's performance in the future for a period of time. The predicted metrics include multiple dimensions such as resource utilization rate, task processing latency, throughput, etc. The system compares the predicted performance values with the preset performance thresholds. When the predicted value exceeds the threshold, the resource adjustment mechanism is triggered.

[0093] Before making resource adjustments, the system first calculates the load level of each thread based on the current execution status information. The load calculation takes into account multiple factors such as the length of the task queue of the thread, CPU utilization rate, memory occupancy, etc. According to the calculated load level, the system determines whether to increase or decrease the number of threads, or adjust the task allocation strategy. Specific adjustment measures include dynamically creating new processing threads, suspending threads with too low load, reallocating tasks, etc.

[0094] Meanwhile, the system continuously checks whether the execution status information contains abnormal signals. Abnormal signals may come from multiple aspects such as task execution timeout, data access error, resource exhaustion, etc. Once an abnormal signal is detected, the system immediately starts the link tracing and positioning algorithm, traces back along the transaction execution path in reverse, and determines the specific location and scope of influence where the abnormality occurs.

[0095] After determining the abnormal transaction link through link tracing, the system starts the cascading rollback process. The rollback process first determines the scope of transactions that need to be rolled back, and then performs rollback operations in reverse order according to the dependency relationship. To ensure the reliability of the rollback, the system will establish checkpoints for each rollback operation to verify the execution results of the rollback operations. If new abnormalities occur during the rollback process, the system will record the relevant information and retry.

[0096] After completing the rollback of the abnormal transaction, the system reconstructs the transaction execution path according to the original cascading operation execution sequence. The reconstruction process will avoid known problem points and adjust the execution strategy if necessary, such as adjusting the parallelism, modifying the resource allocation plan, etc. The reconstructed execution path needs to ensure data consistency and integrity.

[0097] Finally, the system submits the reconstructed transaction execution path to the execution thread pool to re-execute the database cascading operation. During the re-execution process, the system will pay special attention to the links that had problems before, strengthen the monitoring intensity to ensure that the operation can be completed normally. At the same time, the system will record the experience data of this abnormal handling for optimizing the pre-trained model and improving the abnormal handling strategy.

[0098] In this embodiment, the execution status information is continuously collected through the sliding window algorithm, and performance data is generated to achieve real-time monitoring of the execution thread pool, timely detect performance fluctuations, and ensure the stable operation of the system. A pre-trained neural network model is used to predict the performance data. When the performance prediction value exceeds the threshold, resource adjustment is performed based on the thread load level to optimize the utilization rate of computing resources and prevent performance degradation caused by overload. The abnormal transaction link is quickly determined through the link tracing and positioning algorithm, the troubleshooting time is shortened, the impact of abnormal transactions on the overall database operation is avoided, and the reliability of transaction processing is improved. After an anomaly is detected, cascading rollback is performed according to the abnormal transaction link, and the transaction execution path is reconstructed according to the cascading operation execution sequence to ensure the consistency and integrity of database transactions and reduce the risk of data corruption or loss.

[0099] In an alternative embodiment, a database cascading operation request is received, the database cascading operation request is parsed into an operation instruction set, the target table and associated tables are extracted from the operation instruction set through a recursive query algorithm, and a dependency relationship weight matrix is generated by performing field mapping analysis. Based on the dependency relationship weight matrix, a strongly connected component identification algorithm is used to determine the cyclic dependency link, including:

[0100] The database cascading operation request is converted into an abstract syntax tree through a preset syntax parser, and the operation instruction set is extracted based on the abstract syntax tree. Based on the operation instruction set, a breadth-first search recursive query algorithm is used. Starting from the target table in the operation instruction set, recursive queries are made to the associated tables through foreign key constraint information and join condition information. During the recursive query process, the target table is added to the queue, the head node of the queue is taken out and expanded in turn, access marker information for the head node is maintained, the association field information and association type information between the associated table and the target table are recorded, and the unvisited adjacent nodes are added to the tail of the queue to generate an association graph structure including the target table and the associated tables.

[0101] Field mapping is performed on each pair of associated tables in the association graph structure, the field mapping relationship between each pair of associated tables is extracted, the field mapping weight value is calculated based on the field mapping relationship, the field mapping weight values are summarized to obtain the table-level dependency weight value, and a dependency relationship weight matrix is constructed.

[0102] The first depth-first search traversal is performed on the dependency relationship weight matrix. The currently visited node is pushed onto the traversal stack, the in-stack timestamp and completion timestamp of the currently visited node are recorded, and a timestamp pair for the currently visited node is constructed. Based on the traversal stack, the reachable node set of the currently visited node is calculated, and the reachable node set is associated with the timestamp pair to form a node feature vector.

[0103] Transpose the dependency weight matrix to obtain an inverse adjacency matrix, normalize the weight values in the inverse adjacency matrix, and generate a normalized inverse adjacency matrix; map the node feature vectors to the rows corresponding to the normalized inverse adjacency matrix to construct a node dependency feature matrix;

[0104] Based on the descending order of the completion timestamps in the timestamp pair, perform a second depth - first search traversal on the node dependency feature matrix, and calculate the dependency strength values between adjacent nodes on the traversal path according to the corresponding weight values in the node dependency feature matrix and the node feature vectors.

[0105] Based on a preset dependency threshold, filter out the node association relationships whose dependency strength values are greater than the dependency threshold; perform connectivity analysis on the node association relationships, identify the node subsets with bidirectional dependency features in the node association relationships, and determine them as cyclic dependency links; calculate the dependency strength values of each node according to the in - degree weight and out - degree weight of the node in the cyclic dependency link and the weight values in the node dependency feature matrix, and perform normalization processing on the dependency strength values to obtain normalized dependency strength values; classify the cyclic dependency links according to the numerical distribution of the normalized dependency strength values to determine the dependency degree of the cyclic dependency links.

[0106] In a specific embodiment, first, receive a database cascade operation request. For example, an SQL statement containing multiple update and delete operations. This request is input into a preset syntax parser, such as the ANTLR SQL parser. The parser converts the SQL statement into an abstract syntax tree (AST). For example, "UPDATE table1 SET col1 = 'value1' WHERE id = 1; DELETE FROM table2 WHERE id = 2;" will be parsed into a tree - like structure containing update and delete operation nodes. Extract the operation instruction set from the abstract syntax tree, which contains information such as the target table name, operation type (update or delete), and conditional expression of each operation. For example, extract operation instructions such as "UPDATE table1" and "DELETE table2".

[0107] Next, based on the extracted operation instruction set, a recursive query algorithm of breadth-first search is adopted to extract the target table and associated tables. Taking the target table in the operation instruction set as the starting point, for example, "table1". Through the metadata information of the database, such as foreign key constraints and join conditions, query the tables associated with "table1". Suppose "table1" is associated with "table2" through a foreign key, and "table2" is associated with "table3". Add "table1" to the queue. Take out the head node "table1" of the queue, query its associated table "table2", and add "table2" to the tail of the queue. At the same time, record the associated field information and association type (such as foreign key association) between "table1" and "table2". Continue to take out the head node "table2" of the queue, query its associated table "table3", and add "table3" to the tail of the queue, recording the association information. Finally, generate an association graph structure containing the target table and all associated tables.

[0108] Then, perform field mapping analysis on each pair of associated tables in the association graph structure. For example, "table1" and "table2" are associated through fields "id1" and "id2", and "table2" and "table3" are associated through fields "id2" and "id3". Extract the field mapping relationships of each pair of associated tables, such as ("table1.id1", "table2.id2") and ("table2.id2", "table3.id3"). Calculate the field mapping weight values based on the field mapping relationships. For example, if the data types of two fields are the same and the names are similar, a higher weight value is assigned; otherwise, a lower weight value is assigned. Summarize the field mapping weight values to obtain the table-level dependency weight values. For example, the dependency weight of "table1" on "table2" is 0.8, and the dependency weight of "table2" on "table3" is 0.6. Construct a dependency relationship weight matrix. The rows and columns of this matrix represent all the tables in the association graph, and the values of the matrix elements represent the dependency weights between the tables.

[0109] Perform a depth-first search traversal on the dependency relationship weight matrix. Taking any node as the starting node, for example, "table1". Push "table1" onto the traversal stack, recording its push timestamp and completion timestamp. Suppose the push timestamp of "table1" is 1 and the completion timestamp is 6. Calculate the reachable node set of "table1", such as {"table2", "table3"}. Associate the reachable node set with the timestamp pair to form a node feature vector, such as (1, 6, {"table2", "table3"}). Repeat this process for all nodes.

[0110] Transpose the dependency weight matrix to obtain the inverse adjacency matrix. Normalize the weight values in the inverse adjacency matrix. For example, divide all weight values by the sum of the weight values in each row to generate a normalized inverse adjacency matrix. Map the node feature vectors to the corresponding rows of the normalized inverse adjacency matrix to construct the node dependency feature matrix.

[0111] Perform a second depth-first search traversal on the node dependency feature matrix based on the descending order of the completion timestamps. Calculate the dependency strength values between adjacent nodes on the traversal path according to the corresponding weight values and node feature vectors in the node dependency feature matrix. For example, calculate the dependency strength value between "table1" and "table2" as 0.9. Filter the node association relationships with dependency strength values greater than a preset dependency threshold, such as 0.7. Perform connectivity analysis on the filtered node association relationships. Identify the node subsets with bidirectional dependency features. For example, if there is a bidirectional dependency relationship between "table2" and "table3", then determine {"table2", "table3"} as a cyclic dependency link. Calculate the dependency strength values of each node according to the in-degree weight and out-degree weight of the node in the cyclic dependency link, combined with the weight values in the node dependency feature matrix. Normalize the dependency strength values to obtain the normalized dependency strength values. Classify the cyclic dependency links according to the numerical distribution of the normalized dependency strength values. For example, divide the dependency degree into three levels: high, medium, and low.

[0112] In this embodiment, by identifying cyclic dependency links, potential data consistency problems can be discovered in advance, data errors or losses caused by cascading operations can be avoided, thereby improving the security of database operations; by identifying cyclic dependency links, the execution order of database operations can be optimized, unnecessary calculations and resource consumption can be reduced, thereby improving database performance; by providing tools for identifying and analyzing cyclic dependency links, the work of database administrators can be simplified and the database management efficiency can be improved.

[0113] In an alternative embodiment, decomposing the cyclic dependency link to obtain a directed acyclic subgraph, calculating the node depth value and node breadth value according to the directed acyclic subgraph, and generating a cascading operation execution sequence through a topological sorting algorithm includes:

[0114] Count the number of times each edge in the cyclic dependency link is triggered in the historical execution data, use the number of dependency triggers as the dependency strength value of the edge, calculate the normalized dependency strength value of the edge, obtain the set of loops where each edge in the cyclic dependency link is located, calculate the loop strength of each loop according to the product of the weight values of the edges in the loop, divide the loop strength by the loop length to obtain the loop unit strength, accumulate the loop unit strengths of all the loops where the edge is located to obtain the criticality index of the edge, and perform a weighted combination of the criticality index of the edge and the normalized dependency strength value of the edge to obtain the comprehensive evaluation value of the edge;

[0115] Construct a minimum heap containing the comprehensive evaluation values, sequentially take out the current edge with the smallest comprehensive evaluation value from the minimum heap, determine whether the deletion of the current edge causes the graph to split, if it does not split, add the current edge to the feedback edge set and update the comprehensive evaluation values of the adjacent edges, and repeat the execution until a directed acyclic subgraph and the corresponding remaining edge set are obtained;

[0116] In the directed acyclic subgraph, set the depth value of the starting node with an in-degree of zero to zero, set the breadth value of the terminating node with an out-degree of zero to zero, and according to the connection relationship of the remaining edge set, set the depth value of the non-starting nodes to the maximum value of the depth values of all its predecessor nodes plus one through forward traversal, and set the breadth value of the non-terminating nodes to the maximum value of the breadth values of all its successor nodes plus one through backward traversal;

[0117] Perform a weighted combination calculation on the depth value, breadth value, normalized value of the historical execution time of each node, and the number of associated edges of the node in the feedback edge set to obtain the node priority value;

[0118] Construct a priority queue based on the node priority value, update the node priority value according to the execution status of the node in each time window and adjust the sorting of the nodes in the queue, select multiple current nodes that meet the parallel conditions with the highest priority values from the priority queue, and add the current nodes to the cascade operation execution sequence in the selected order.

[0119] The number of dependency triggers specifically refers to the number of times a certain edge in the cyclic dependency link is triggered (executed) in the historical execution data. This indicator is used to measure the activity and importance of this dependency relationship in the actual operating environment. The higher the number of triggers, the more frequently this dependency relationship is used in the database operation process and the greater the impact on the transaction execution.

[0120] The loop strength specifically refers to an index that measures the overall dependence strength of a cyclic dependence path. Its calculation method is the product of the weight values of all edges in the loop. The higher the loop strength, the closer the dependence relationship on this cyclic dependence path, which may impose greater constraints or impacts on transaction execution. In dependence analysis, loop strength can be used to distinguish critical dependence paths from low-impact dependence paths, thereby optimizing the dependence resolution strategy.

[0121] In a specific implementation, first, count the historical dependence trigger times of each edge in the cyclic dependence link. For example, the historical trigger times of edge A→B are 100 times, and the trigger times of edge B→C are 50 times. These trigger times are used as the dependence strength values of the edges. Then, perform normalization processing on the dependence strength values of each edge. Assuming the total trigger times are 200 times, the normalized dependence strength value of edge A→B is 0.5, and the normalized dependence strength value of edge B→C is 0.25.

[0122] Next, identify the loop sets where each edge in the cyclic dependence link is located. For example, edge A→B and edge B→A form a loop, and edge B→C, edge C→D, and edge D→B form another loop. Then, calculate the loop strength of each loop. The loop strength is obtained by multiplying the normalized dependence strength values of the edges in the loop. Assuming the loop strength of loop A→B→A is 0.5×0.4 = 0.2. Divide the loop strength by the loop length (the number of edges) to obtain the loop unit strength. For example, the loop unit strength of loop A→B→A is 0.2 / 2 = 0.1. Accumulate the loop unit strengths of all loops where the edge is located to obtain the criticality index of the edge. Assuming the criticality index of edge A→B is 0.1 + 0.05 = 0.15. Finally, perform a weighted combination of the criticality index of the edge and the normalized dependence strength value to obtain the comprehensive evaluation value of the edge. For example, assuming the weights are 0.6 and 0.4 respectively, the comprehensive evaluation value of edge A→B is 0.15×0.6 + 0.5×0.4 = 0.29.

[0123] Construct a minimum heap containing all edges, where the comprehensive evaluation value of the edge is used as the sorting basis. Sequentially take out the edge with the smallest comprehensive evaluation value from the minimum heap. Determine whether the graph splits after deleting this edge. If the graph does not split after deletion, add this edge to the feedback edge set and update the comprehensive evaluation values of the adjacent edges. Repeat this process until a directed acyclic subgraph and the corresponding remaining edge set are obtained. For example, assuming the graph does not split after deleting edge B→A, then add B→A to the feedback edge set.

[0124] In the generated directed acyclic subgraph, set the depth value of the starting node with an in-degree of zero to zero. Set the breadth value of the terminating node with an out-degree of zero to zero. According to the connection relationship of the remaining edge set, calculate the depth values of non-starting nodes through forward traversal. The depth value of a node is equal to the maximum value of the depth values of all its predecessor nodes plus one. For example, if the depth value of the predecessor node A of node B is 0, then the depth value of node B is 1. Calculate the breadth values of non-terminating nodes through backward traversal. The breadth value of a node is equal to the maximum value of the breadth values of all its successor nodes plus one. For example, if the breadth value of the successor node C of node B is 0, then the breadth value of node B is 1.

[0125] Perform a weighted combination of the depth value, breadth value, normalized value of the historical execution time of each node, and the number of associated edges of the node in the feedback edge set to calculate the node priority value. For example, the depth value of node A is 0, the breadth value is 2, the normalized value of the historical execution time is 0.3, and the number of associated edges is 1. Assuming the weights are 0.2, 0.3, 0.4, and 0.1 respectively, then the priority value of node A is 0.2×0 + 0.3×2 + 0.4×0.3 + 0.1×1 = 0.82.

[0126] Construct a priority queue based on the node priority values. In each time window, update the priority values of the nodes according to their execution status and adjust the sorting of the nodes in the queue. Select multiple nodes that meet the parallel conditions and have the highest priority values from the priority queue. Add the selected nodes to the cascade operation execution sequence in the selected order. For example, assume that nodes A and C have the highest priority values and meet the parallel conditions, then add them to the cascade operation execution sequence.

[0127] In this embodiment, by decomposing the cyclic dependency link and generating an execution sequence according to the node priority, it is possible to effectively avoid cyclic waiting, reduce resource contention, and thus improve the overall execution efficiency of the system; by identifying and handling cyclic dependencies, potential deadlocks and errors can be avoided, and the stability and reliability of the system can be improved; through the priority queue and parallel execution, system resources can be fully utilized, resource waste can be avoided, and resource utilization can be improved.

[0128] In an alternative embodiment, read the cascade operation execution sequence, perform node classification processing on the cascade operation execution sequence through a multi-dimensional feature classification algorithm, and organize the classified nodes into batch processing task groups, including:

[0129] Read the cascade operation execution sequence and construct a multi-layer information network, where the weight of the edges in the first layer network represents the data transmission volume between nodes, the weight of the edges in the second layer network represents the resource contention degree between nodes, and the weight of the edges in the third layer network represents the operation similarity between nodes;

[0130] Based on the multi-layer information network, calculate the degree centrality, betweenness centrality, and closeness centrality of each node, determine the centrality index, and combine the centrality index to form the network structure feature vector of the node; obtain the state sequence of each node in the cascade operation execution sequence, extract the mean, variance, skewness, and kurtosis of the state sequence, determine the statistical features, and combine the statistical features to form the time series feature vector of the node; based on the input data scale, computational complexity, and memory occupancy of the node in the cascade operation execution sequence, determine the performance metrics, and construct the performance feature vector of the node based on the performance metrics; combine the network structure feature vector, time series feature vector, and performance feature vector to construct the feature matrix of the node;

[0131] Calculate the Mahalanobis distance between nodes in the feature matrix, and construct the similarity network of nodes based on the Mahalanobis distance; in the similarity network, obtain the steady-state distribution of nodes by iteratively calculating the transition probability between node pairs; based on the steady-state distribution, identify the densely connected subgraphs in the similarity network, and divide the nodes belonging to the same densely connected subgraph into the same category;

[0132] Calculate the path length between nodes within each category, and split the categories with a path length greater than the preset path length threshold to obtain the optimized node classification result; organize the nodes with the same category into a batch processing task group.

[0133] In a specific embodiment, first, read the cascade operation execution sequence. For example, a cascade operation execution sequence can be represented as a series of nodes, where each node represents an operation and there is a sequential dependency relationship between nodes. An example sequence can be A→B→C→D→E, indicating that operation A must be executed before B, B must be executed before C, and so on.

[0134] Then, construct a multi-layer information network. This network consists of three layers, representing the data transfer volume, resource competition degree, and operation similarity between nodes respectively. Taking nodes A, B, and C as an example, if A transfers 10MB of data to B, the weight of the edge from A to B in the first-layer network is 10. If A and B compete for the same CPU core with a competition degree of 0.8, the weight of the edge from A to B in the second-layer network is 0.8. If A and C are both database read operations with a similarity of 0.9, the weight of the edge from A to C in the third-layer network is 0.9.

[0135] Next, calculate the network structure feature vectors of each node. Based on the constructed multi-layer information network, calculate the degree centrality, betweenness centrality, and closeness centrality of each node. Degree centrality represents the number of direct connections of a node to other nodes. Betweenness centrality represents the number of times a node appears on the shortest paths between other nodes. Closeness centrality represents the reciprocal of the average distance from a node to other nodes. Combine these centrality metrics to form the network structure feature vector of the node. For example, if the degree centrality of node A is 2, the betweenness centrality is 1, and the closeness centrality is 0.5, then its network structure feature vector is [2, 1, 0.5].

[0136] At the same time, obtain the state sequence of each node in the cascade operation execution sequence, extract the mean, variance, skewness, and kurtosis of the state sequence, determine the statistical features, and combine these statistical features to form the time series feature vector of the node. For example, if the state sequence of node A is [1, 2, 3, 4, 5], its mean is 3, variance is 2, skewness is 0, and kurtosis is -1.2, then its time series feature vector is [3, 2, 0, -1.2].

[0137] In addition, based on the input data size, computational complexity, and memory occupancy of the node in the cascade operation execution sequence, determine the performance metrics, and construct the performance feature vector of the node based on these performance metrics. For example, if the input data size of node A is 10 MB, the computational complexity is 1000 floating-point operations, and the memory occupancy is 2 MB, then its performance feature vector is [10, 1000, 2].

[0138] Combine the network structure feature vector, time series feature vector, and performance feature vector to construct the feature matrix of the node. For example, the feature matrix of node A is [[2, 1, 0.5], [3, 2, 0, -1.2], [10, 1000, 2]].

[0139] Then, calculate the Mahalanobis distance between the nodes in the feature matrix, and construct the similarity network of the nodes based on the Mahalanobis distance. The Mahalanobis distance takes into account the correlation between different features. In the similarity network, obtain the steady-state distribution of the nodes by iteratively calculating the transition probabilities between node pairs. The steady-state distribution represents the final probability distribution of each node.

[0140] Based on the steady-state distribution, identify the densely connected subgraphs in the similarity network, and divide the nodes belonging to the same densely connected subgraph into the same category. A densely connected subgraph refers to a subgraph with close connections between nodes.

[0141] Calculate the path length between the nodes within each category, and split the categories with a path length greater than the preset path length threshold to obtain the optimized node classification result. The path length refers to the length of the shortest path between the nodes within the category. The preset path length threshold is set according to the actual situation.

[0142] Finally, organize the nodes with the same category into a batch task group. For example, if nodes A, B, and C belong to the same category, they are organized into a batch task group.

[0143] In this embodiment, by organizing similar nodes into a batch task group, computing resources can be fully utilized, task scheduling overhead can be reduced, and thus the execution efficiency of cascading operations can be improved; by considering the resource competition degree between nodes, resource conflicts can be avoided and resource utilization can be optimized; by considering the similarity and state sequence characteristics between nodes, nodes with similar behaviors can be organized together, improving the stability and predictability of the system.

[0144] In an alternative embodiment, the optimal execution path of the batch task group is calculated through an adaptive path optimization algorithm, and the distributed transaction control instruction is generated according to the optimal execution path, including:

[0145] Based on each task node in the batch task group, the load volatility is calculated by the ratio of the sum of the absolute values of the load value differences at adjacent time points within the sampling period of the task node to the sampling period;

[0146] Based on the weighted combination of the load volatility difference and the time overlap degree, a task node distance matrix is constructed, and hierarchical clustering is performed on the batch task group according to the task node distance matrix to obtain task grouping information;

[0147] Based on the task grouping information, an adaptive weight calculation model is constructed. The system load data of the task node is input into the sigmoid function to obtain a normalized load value, and the aging adjustment coefficient and resource adjustment coefficient of the task node are calculated according to the normalized load value;

[0148] According to the aging adjustment coefficient, resource adjustment coefficient, and load volatility, the aging evaluation value, resource evaluation value, and volatility evaluation value are determined, and the evaluation score of the task node is calculated by weighting;

[0149] The evaluation score is used as a state feature to construct a reinforcement learning model. An action space and a state space are constructed based on the state feature, and a reward function is determined according to the evaluation score, resource competition degree, and parallelism;

[0150] The action value function is iteratively updated using a deep Q-network to generate the selection strategy of the task node;

[0151] Based on the selection strategy, path search is performed on the batch task group. In the exploration stage, an execution node is randomly selected to construct a candidate path, and in the exploitation stage, the optimal execution node is selected based on the action value function to construct the optimal execution path;

[0152] Construct a distributed transaction control instruction according to the optimal execution path. The distributed transaction control instruction includes a pre-commit instruction and a commit instruction, where the pre-commit instruction is determined by collecting the status information of task nodes on the optimal execution path;

[0153] Incrementally update and maintain the causal relationship of task nodes by comparing the clock values of adjacent task nodes, record the execution order of task nodes, determine the vector clock instruction, and embed the vector clock instruction in the distributed transaction control instruction;

[0154] Input the execution result of the vector clock instruction into the distributed consensus module, and coordinate the transaction status of task nodes through the distributed consensus module to generate a distributed transaction control instruction.

[0155] The time overlap degree refers to the degree of overlap of the execution time periods of different task nodes within a certain time range. Usually in batch processing tasks or distributed systems, the execution time periods of task nodes may overlap, indicating that the execution of different tasks may affect each other. The time overlap degree can be used as a measure of the correlation between tasks, and is used to adjust task scheduling strategies, optimize resource allocation, etc.

[0156] The vector clock instruction refers to a mechanism based on vector clocks, which is used to track the execution order of tasks or events in a distributed system. Vector clocks are a technology used to solve the time order problem in distributed systems. It maintains an array (vector) of "timestamps" for each node and updates the timestamps when passing between different task nodes, so as to ensure the consistency of the event order of each node. The vector clock instruction is part of this mechanism and is usually used to represent the causal relationship or execution order between task nodes, and to ensure the synchronous execution of tasks in a distributed system.

[0157] The distributed consensus module is a mechanism in a distributed system that ensures that all nodes reach a consistent decision on a certain transaction or operation. It solves the problem of how each node reaches an agreement on certain issues without a central controller. Typical consensus algorithms include Paxos, Raft, etc., and are often used in fields such as distributed databases or blockchains. In this task description, the distributed consensus module is used to coordinate the transaction status of each task node, ensure that the task nodes can reach an agreement and correctly execute the distributed transaction control instruction, so as to ensure the high availability and consistency of the system.

[0158] In a specific embodiment, first, calculate the load volatility of the task nodes. Select a sampling period, for example, 1 minute, and record the load values of each task node at different time points within this period. For example, the load values of task node A within 1 minute are 5, 7, 6, and 8 respectively. Calculate the absolute value of the difference in load values between adjacent time points. For example, the absolute value of (7 - 5) is 2, the absolute value of (6 - 7) is 1, and the absolute value of (8 - 6) is 2. Accumulate these absolute values of the differences to get 5. Take the ratio of the accumulated value to the sampling period as the load volatility. For example, 5 / 1 = 5, so the load volatility of task node A is 5. Perform the same calculation for all task nodes in the batch task group to obtain the load volatility of each task node.

[0159] Next, construct the task node distance matrix. Calculate the distance between task nodes based on the difference in load volatility and the time overlap degree between task nodes. The time overlap degree refers to the proportion of time during which two task nodes execute overlappingly. For example, if the execution period of task node A is from 9 am to 10 am and the execution period of task node B is from 9:30 am to 10:30 am, then the time overlap degree between A and B is 0.5. Perform a weighted combination of the load volatility difference and the time overlap degree. For example, multiply the load volatility difference by the weight 0.7, multiply the time overlap degree by the weight 0.3, and then add the two to obtain the distance between task nodes. Construct a matrix of all the distances between task nodes, that is, the task node distance matrix.

[0160] Then, perform task grouping. Based on the task node distance matrix, use the hierarchical clustering algorithm to group the task nodes. The hierarchical clustering algorithm is an iterative algorithm that merges the two task nodes with the closest distance into a group, then recalculates the distance between groups and continues to merge until all task nodes are merged into one group. Through hierarchical clustering, the batch task group can be divided into multiple subgroups, and the distance between task nodes within each subgroup is relatively close, and the load volatility and time overlap degree are also relatively close.

[0161] Next, construct an adaptive weight calculation model. Input the system load data of the task nodes into the sigmoid function to normalize the load value to between 0 and 1. For example, if the load value of task node A is 10 and the output value of the sigmoid function is 0.9999, then the normalized load value is 0.9999. Calculate the timeliness adjustment coefficient and resource adjustment coefficient of the task node according to the normalized load value. The timeliness adjustment coefficient is used to adjust the execution time of the task, and the resource adjustment coefficient is used to adjust the resources required by the task.

[0162] Then, calculate the evaluation score of the task node. Determine the timeliness evaluation value, resource evaluation value, and volatility evaluation value according to the timeliness adjustment coefficient, resource adjustment coefficient, and load volatility. For example, the timeliness evaluation value is equal to the timeliness adjustment coefficient multiplied by a preset weight, the resource evaluation value is equal to the resource adjustment coefficient multiplied by a preset weight, and the volatility evaluation value is equal to the load volatility multiplied by a preset weight. Weightedly sum up the timeliness evaluation value, resource evaluation value, and volatility evaluation value to obtain the evaluation score of the task node.

[0163] Next, construct a reinforcement learning model. Use the evaluation score of the task node as the state feature to construct the action space and state space. The action space represents the execution nodes that can be selected, and the state space represents the state of the current task node. Determine the reward function according to the evaluation score, resource competition degree, and parallelism. The resource competition degree refers to the degree to which multiple task nodes compete for the same resource, and the parallelism refers to the number of task nodes that can be executed simultaneously.

[0164] Use the deep Q-network to iteratively update the action-value function and generate the selection strategy for the task node. The deep Q-network is a deep learning algorithm that can learn the optimal action-value function to guide the selection of task nodes.

[0165] Then, perform path search. In the exploration phase, randomly select execution nodes to construct candidate paths. In the exploitation phase, select the optimal execution node based on the action-value function to construct the optimal execution path.

[0166] Finally, generate distributed transaction control instructions. According to the optimal execution path, construct distributed transaction control instructions, including pre-commit instructions and commit instructions. Collect the state information of the task nodes on the optimal execution path to determine the pre-commit instructions. Update and maintain the causal relationship of the task nodes by comparing the clock values of adjacent task nodes in increasing order, record the execution order of the task nodes, and determine the vector clock instructions. Embed the vector clock instructions in the distributed transaction control instructions. Input the execution result of the vector clock instructions into the distributed consensus module, and coordinate the transaction states of the task nodes through the distributed consensus module to generate the final distributed transaction control instructions.

[0167] In this embodiment, through the adaptive path optimization algorithm, the optimal execution path can be found, the waiting time between task nodes can be reduced, the resource utilization rate can be improved, and thus the overall execution efficiency of batch tasks can be improved; by considering factors such as the load volatility, resource competition degree, and parallelism of task nodes, system overload can be avoided and the stable operation of the system can be ensured; through vector clocks and the distributed consensus module, the transaction states of task nodes can be effectively coordinated, the management of distributed transactions can be simplified, and the reliability of the system can be improved.

[0168] In an alternative embodiment, the link tracing and positioning algorithm includes:

[0169] Collect transaction execution information to generate a transaction identification number, which is encoded by a timestamp field, a machine identification field, and an incrementing sequence number field, and marks all database operations on the same transaction link; record the operation execution information of the database operation, including an operation identification number, a parent operation identification number, operation type information, an operation start time, an operation end time, and operation status information;

[0170] Build a call relationship tree of the operation execution information, use the operation identification number as a tree node, and use the parent operation identification number as the parent-child node association relationship;

[0171] Based on each node in the call relationship tree, determine an anomaly score by calculating a weighted combination of an operation anomaly baseline score, an operation latency anomaly score, and an operation status transition anomaly score, and calculate an anomaly attenuation coefficient according to the hierarchical depth of the tree where the node is located, and multiply the anomaly attenuation coefficient by the anomaly score to obtain the final anomaly score of the node;

[0172] Calculate the time overlap degree between adjacent database operations, use the time overlap degree as the operation association strength, and build an operation dependency network based on the operation association strength; analyze the operation execution order in the operation dependency network, and calculate the conditional probability value between adjacent operations to determine the operation propagation probability;

[0173] Filter the operation association edges whose final anomaly score is greater than a preset first anomaly threshold and whose operation propagation probability is greater than a preset second anomaly threshold, and build an anomaly propagation subgraph;

[0174] Calculate the minimum weight spanning tree in the anomaly propagation subgraph to obtain the core path of anomaly propagation;

[0175] Count the in-degree values of each node in the core path, add the nodes with an in-degree value of zero to the pending queue, continuously take out the nodes in the queue and update the in-degree values of the corresponding adjacent nodes, and generate the propagation link of the abnormal transaction according to the order of node dequeue.

[0176] The call relationship tree specifically refers to a tree structure used to represent the hierarchical call relationship between database operations. Among them, each database operation is a node of the tree, the parent node represents the upper-level operation that triggers the operation, and the child node represents the subsequent operation triggered by the operation. The call relationship tree can help analyze the execution path of the transaction, trace the dependency relationship between operations, and identify the source of anomaly propagation.

[0177] The operation propagation probability specifically refers to the probability that another related operation occurs after one operation occurs in the database operation dependency network. It reflects the association strength between operations and is usually obtained through historical data statistics.

[0178] The propagation link of the abnormal transaction is a sequence of abnormal operation arranged in the order of transaction execution, which is used to describe the propagation path of the exception in the transaction. It is based on the exception propagation subgraph. By screening the operations with higher propagation probability and higher exception scores, the transmission relationship of the exception in the transaction is constructed and arranged in the order of dependence to form a complete exception link. In the propagation link of the abnormal transaction, the starting node is usually the source of the exception, such as the key failure operation at the beginning of the transaction, and the subsequent nodes are the affected operations, such as queries, updates, or commit operations that cannot be executed due to the failure of the previous operation. By constructing the propagation link of the abnormal transaction, the root cause of the exception can be traced, the scope of the exception impact can be analyzed, and a decision-making basis for exception recovery can be provided.

[0179] In a specific embodiment, first, the transaction execution information is collected to generate a transaction identification number. The transaction identification number is encoded by a timestamp field, a machine identification field, and an incrementing sequence number field, and is used to mark all database operations on the same transaction link. For example, a transaction identification number can be 20240727100000-SERVER01-0001, where 20240727100000 represents the timestamp, SERVER01 represents the machine identification, and 0001 represents the incrementing sequence number.

[0180] Next, the execution information of the database operation is recorded. This information includes the operation identification number, the parent operation identification number, the operation type information, the operation start time, the operation end time, and the operation status information. For example, the execution information of a database operation can be: the operation identification number is OP0001, the parent operation identification number is OP0000, the operation type is SELECT, the operation start time is 10:00:00.001, the operation end time is 10:00:00.005, and the operation status is successful.

[0181] Then, a call relationship tree of the operation execution information is established. The operation identification number is used as the tree node, and the parent operation identification number is used as the parent-child node association relationship. For example, if the parent operation identification number of OP0002 is OP0001, then OP0002 is the child node of OP0001.

[0182] Calculate the anomaly scores of the nodes based on each node in the call relationship tree. First, calculate the operation anomaly baseline score, the operation latency anomaly score, and the operation state transition anomaly score respectively according to the operation execution information. The operation anomaly baseline score is determined according to the operation type. For example, the baseline score for a SELECT operation is 1, and the baseline score for an UPDATE operation is 2. The operation latency anomaly score is determined according to the comparison result between the actual operation duration and the preset threshold. For example, if the operation duration exceeds the preset threshold, the latency anomaly score is 1, otherwise it is 0. The operation state transition anomaly score is determined according to whether the operation state is an abnormal state. For example, if the operation state is a failure, the state transition anomaly score is 1, otherwise it is 0. Then, perform a weighted combination of the three scores to obtain the anomaly score. For example, assuming the weights are 0.2, 0.5, and 0.3 respectively, the anomaly score is 0.2 × baseline score + 0.5 × latency anomaly score + 0.3 × state transition anomaly score. Finally, calculate the anomaly attenuation coefficient according to the hierarchical depth of the tree where the node is located. The deeper the hierarchy, the smaller the attenuation coefficient. For example, the attenuation coefficient of the root node is 1, and the attenuation coefficient of its child node is 0.8, and so on. Multiply the anomaly attenuation coefficient by the anomaly score to obtain the final anomaly score of the node.

[0183] Calculate the time overlap degree between adjacent database operations, and use this as the operation association strength. For example, if the end time of operation A is later than the start time of operation B, and the start time of operation A is earlier than the end time of operation B, then there is a time overlap between operation A and operation B. The time overlap degree can be represented by the proportion of the overlap time to the total duration of the two operations. Construct an operation dependency network based on the operation association strength.

[0184] Analyze the operation execution order in the operation dependency network, and calculate the conditional probability value between adjacent operations to determine the operation propagation probability. For example, if the probability of operation B occurring is very high after operation A is completed, then the propagation probability from operation A to operation B is relatively high.

[0185] Filter the operation association edges whose final anomaly scores are greater than the preset first anomaly threshold and whose operation propagation probabilities are greater than the preset second anomaly threshold, and construct an anomaly propagation subgraph. For example, assume the first anomaly threshold is 0.5 and the second anomaly threshold is 0.8, then only the operation association edges that meet these two conditions will be selected.

[0186] Calculate the minimum weight spanning tree in the anomaly propagation subgraph to obtain the core path of the anomaly propagation. The weight of the edge can use the operation association strength or the operation propagation probability.

[0187] Count the in-degree values of each node in the core path, and add the nodes with in-degree values of zero to the processing queue. Continuously take out the nodes in the queue and update the in-degree values of the corresponding adjacent nodes, and generate the propagation link of the abnormal transaction according to the order of node dequeue.

[0188] In this embodiment, by analyzing the operation dependency relationship and the abnormal propagation path, the root cause of the abnormal transaction can be quickly located, avoiding the cumbersome and inefficient manual troubleshooting; by constructing an abnormal propagation subgraph and a core path, the propagation process of the abnormal transaction can be clearly shown, helping developers quickly understand the problem, thereby improving the troubleshooting efficiency; by automatically analyzing and locating the anomaly, manual intervention can be reduced, the system maintenance cost can be lowered, and the system stability can be improved.

[0189] In an alternative embodiment, the cascading rollback includes:

[0190] Construct a multi-dimensional transaction dependency graph, store transaction operations in layers according to timestamps, establish a read-write version chain of data items, record the read-write timing information and data version numbers of data items, and generate a transaction version dependency relationship;

[0191] Based on the transaction version dependency relationship, construct a conflict matrix, calculate the data access overlap degree between adjacent transactions, use the data access overlap degree as a conflict weight, and identify high-frequency conflict regions where the conflict weight exceeds a preset conflict threshold;

[0192] According to the conflict weight, calculate the criticality of transaction nodes by accumulating the weighted values of the data modification amount, resource occupation duration, the conflict weight, and the transaction nesting level corresponding to the transaction nodes;

[0193] Based on the criticality, classify the transaction nodes, merge the transaction nodes whose criticality meets the approximation threshold into rollback batches according to a preset approximation threshold, calculate the resource consumption value and influence range of each rollback batch, and generate a batch priority;

[0194] Generate a rollback execution plan according to the batch priority, convert the rollback batches into parallel execution tasks in the order of batch priority, and allocate rollback threads for each parallel execution task based on the resource consumption value;

[0195] Start a rollback executor, submit the parallel execution tasks to a rollback thread pool, record the start timestamp of each rollback operation, and perform concurrent control by comparing the data version numbers;

[0196] Monitor the rollback executor. When a rollback failure is detected, locate the failure position according to the timestamp, construct a compensation transaction to retry the rollback operation until the rollback is successful;

[0197] Collect the operation logs output by the rollback executor, extract the version number sequence from the operation logs, reconstruct the transaction execution path based on the version number sequence, and verify data consistency.

[0198] In a specific embodiment, first, a multi-dimensional transaction dependency graph is constructed. This dependency graph not only records the timestamp order of transaction operations but also records the data items read and written by each transaction operation, as well as the version numbers of each data item. For example, transaction T1 reads version 1 of data item A at timestamp 10 and writes version 2 of data item B at timestamp 20, and this information will be recorded in the dependency graph. At the same time, based on the read-write timing information of data items and data version numbers, transaction version dependencies are generated. For example, if T2 reads version 2 of data item B at timestamp 30, then T2 depends on T1.

[0199] Next, a conflict matrix is constructed based on the transaction version dependencies. The conflict matrix is used to represent the degree of data access overlap between transactions. For example, if both transaction T1 and T2 modify data item C, then the value of the corresponding cell in the conflict matrix will increase. This value represents the degree of data access overlap and can be used as a conflict weight. By setting a preset conflict threshold, high-frequency conflict regions with conflict weights exceeding the threshold can be identified. For example, if T1, T2, and T3 frequently modify data items C and D, then the region where C and D are located is a high-frequency conflict region.

[0200] Then, the criticality of transaction nodes is calculated. The criticality comprehensively considers the amount of data modified by the transaction, the resource occupation duration, the conflict weight, and the transaction nesting level. For example, a transaction that modifies a large amount of data, occupies a long time, is located in a high-frequency conflict region, and has a deep nesting level will have a relatively high criticality. Specifically, these factors can be weighted and summed to obtain the criticality value of the transaction. For example, assume that transaction T4 modifies 100 pieces of data, occupies 10 seconds, has a conflict weight of 5, and a nesting level of 3, then its criticality can be calculated as 100×w1 + 10×w2 + 5×w3 + 3×w4, where w1, w2, w3, and w4 are the weights of the corresponding factors respectively.

[0201] After that, the transaction nodes are classified based on the criticality, and the transaction nodes with similar criticality are merged into rollback batches. For example, set an approximation threshold of 10. If the criticalities of transaction T5 and T6 are 100 and 105 respectively, then they can be merged into a rollback batch. For each rollback batch, calculate its resource consumption value (such as CPU time, memory occupation, etc.) and influence range (such as the number of data items involved). Based on the resource consumption value and the influence range, batch priorities are generated. For example, batches with high resource consumption values and large influence ranges will have higher priorities.

[0202] Next, generate a rollback execution plan according to the batch priority. Convert the rollback batches into parallel execution tasks in the order of priority. For example, the batch with the highest priority will be converted into the first execution task. Based on the resource consumption value of each batch, allocate rollback threads to each parallel execution task. For example, tasks with high resource consumption values will be allocated more threads.

[0203] Start the rollback executor, and submit the parallel execution tasks to the rollback thread pool. Record the start timestamp of each rollback operation, and perform concurrent control by comparing data version numbers. For example, if a rollback operation attempts to modify a data item that has already been modified by other rollback operations, version number comparison will be performed to ensure data consistency.

[0204] Monitor the running status of the rollback executor. When a rollback failure is detected, locate the failure position according to the timestamp, construct a compensation transaction to retry the rollback operation until the rollback is successful. For example, if the rollback operation fails at timestamp 500, the specific operation that failed will be located according to the log, and a compensation transaction will be constructed to re-execute the operation.

[0205] Finally, collect the operation logs output by the rollback executor, extract the version number sequence in the logs, reconstruct the transaction execution path based on the version number sequence, and verify data consistency. For example, by checking whether the final data version number meets the expectation, it can be determined whether data consistency is guaranteed.

[0206] In this embodiment, by identifying high-frequency conflict areas and rolling back in batches according to the criticality, the resources and time required for rollback can be effectively reduced, thereby improving the rollback efficiency; by parallelly executing rollback tasks and the concurrent control mechanism, the risk of rollback failure can be reduced, and data consistency can be guaranteed; by automatically generating a rollback execution plan and monitoring the rollback execution process, rollback management can be simplified, and the cost of manual intervention can be reduced.

[0207] Figure 2 It is a schematic structural diagram of the database cascade operation intelligent parsing and execution system according to the embodiment of the present invention. As Figure 2 shown, the system includes:

[0208] The first unit is used to receive a database cascade operation request, parse the database cascade operation request into an operation instruction set, extract the target table and associated tables from the operation instruction set through a recursive query algorithm, perform field mapping analysis to generate a dependency relationship weight matrix, and determine a cyclic dependency link based on the dependency relationship weight matrix through a strongly connected component recognition algorithm; decompose the cyclic dependency link to obtain a directed acyclic subgraph, calculate the node depth value and node breadth value according to the directed acyclic subgraph, and generate a cascade operation execution sequence through a topological sorting algorithm;

[0209] A second unit, configured to read the cascade operation execution sequence, perform node classification processing on the cascade operation execution sequence through a multi-dimensional feature classification algorithm, organize the classified nodes into a batch processing task group, calculate an optimal execution path of the batch processing task group through an adaptive path optimization algorithm, generate a distributed transaction control instruction according to the optimal execution path, implant a transaction synchronization point in the distributed transaction control instruction, the transaction synchronization point being set based on a two-phase commit protocol, allocate the batch processing task group to an execution thread pool through a consistent hashing algorithm, and the execution thread pool executes the batch processing task group based on a pipeline parallel processing mechanism to generate execution status information;

[0210] A third unit, configured to monitor the running process of the execution thread pool, continuously collect the execution status information through a sliding window algorithm to generate performance data, input the performance data into a pre-trained neural network model to obtain a performance prediction value, when the performance prediction value exceeds a preset performance threshold, calculate a thread load level based on the execution status information, adjust resources of the execution thread pool according to the thread load level, if an abnormal signal is included in the execution status information, locate an abnormal transaction link through a link tracing and positioning algorithm, perform cascade rollback according to the abnormal transaction link, and reconstruct a transaction execution path according to the cascade operation execution sequence to complete a database cascade operation.

[0211] In a third aspect of the embodiments of the present invention,

[0212] There is provided an electronic device, including:

[0213] A processor;

[0214] A memory for storing instructions executable by the processor;

[0215] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0216] In a fourth aspect of the embodiments of the present invention,

[0217] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0218] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0219] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. Intelligent parsing and execution method for database cascading operations, characterized in that, Including: Receiving a database cascade operation request, parsing the database cascade operation request into an operation instruction set, extracting a target table and associated tables from the operation instruction set through a recursive query algorithm, performing field mapping analysis to generate a dependency relationship weight matrix, determining a cyclic dependency link based on the dependency relationship weight matrix through a strongly connected component identification algorithm; decomposing the cyclic dependency link to obtain a directed acyclic subgraph, calculating node depth values and node breadth values according to the directed acyclic subgraph, and generating a cascade operation execution sequence through a topological sorting algorithm; Reading the cascade operation execution sequence, performing node classification processing on the cascade operation execution sequence through a multi-dimensional feature classification algorithm, organizing the classified nodes into batch processing task groups, calculating an optimal execution path for the batch processing task groups through an adaptive path optimization algorithm, generating a distributed transaction control instruction according to the optimal execution path, implanting a transaction synchronization point in the distributed transaction control instruction, the transaction synchronization point being set based on a two-phase commit protocol, distributing the batch processing task groups to an execution thread pool through a consistent hashing algorithm, and the execution thread pool executing the batch processing task groups based on a pipeline parallel processing mechanism to generate execution status information; Monitoring the running process of the execution thread pool, continuously collecting the execution status information through a sliding window algorithm to generate performance data, inputting the performance data into a pre-trained neural network model to obtain a performance prediction value, when the performance prediction value exceeds a preset performance threshold, calculating a thread load level based on the execution status information, adjusting resources of the execution thread pool according to the thread load level, if an abnormal signal is included in the execution status information, locating an abnormal transaction link through a link tracing and positioning algorithm, performing cascade rollback according to the abnormal transaction link, and reconstructing a transaction execution path according to the cascade operation execution sequence to complete the database cascade operation.

2. The method according to claim 1, wherein Receiving a database cascade operation request, parsing the database cascade operation request into an operation instruction set, extracting a target table and associated tables from the operation instruction set through a recursive query algorithm, performing field mapping analysis to generate a dependency relationship weight matrix, and determining a cyclic dependency link based on the dependency relationship weight matrix includes: Converting the database cascade operation request into an abstract syntax tree through a preset syntax parser, and extracting an operation instruction set based on the abstract syntax tree; based on the operation instruction set, adopting a breadth-first search recursive query algorithm, starting from the target table in the operation instruction set, performing recursive query to associated tables through foreign key constraint information and join condition information, during the recursive query process, adding the target table to a queue, sequentially taking out the head node of the queue for expansion, maintaining access marker information for the head node, recording the association field information and association type information between the associated table and the target table, and adding unvisited adjacent nodes to the tail of the queue to generate an association graph structure including the target table and associated tables; Perform field mapping on each pair of associated tables in the associated graph structure, extract the field mapping relationships of each pair of associated tables, calculate the field mapping weight values based on the field mapping relationships, summarize the field mapping weight values to obtain the table-level dependency weight values, and construct a dependency relationship weight matrix; Perform the first depth-first search traversal on the dependency relationship weight matrix, push the currently visited node onto the traversal stack, record the in-stack timestamp and completion timestamp of the currently visited node, and construct a timestamp pair for the currently visited node; calculate the reachable node set of the currently visited node based on the traversal stack, and associate the reachable node set with the timestamp pair to form a node feature vector; Transpose the dependency relationship weight matrix to obtain an inverse adjacency matrix, normalize the weight values in the inverse adjacency matrix to generate a normalized inverse adjacency matrix; map the node feature vector to the corresponding row of the normalized inverse adjacency matrix to construct a node dependency feature matrix; Based on the descending order of the completion timestamps in the timestamp pair, perform the second depth-first search traversal on the node dependency feature matrix, and calculate the dependency strength value between adjacent nodes on the traversal path according to the corresponding weight values in the node dependency feature matrix and the node feature vector, Filter the node association relationships with the dependency strength value greater than the preset dependency threshold based on the preset dependency threshold; perform connectivity analysis on the node association relationships, identify the node subset with bidirectional dependency characteristics in the node association relationships, and determine it as a cyclic dependency link; calculate the dependency strength value of each node according to the in-degree weight and out-degree weight of the node in the cyclic dependency link and the weight values in the node dependency feature matrix, and perform normalization processing on the dependency strength value to obtain the normalized dependency strength value; classify the dependency degree of the cyclic dependency link based on the numerical distribution of the normalized dependency strength value to determine the dependency degree of the cyclic dependency link.

3. The method according to claim 1, characterized in that Decompose the cyclic dependency link to obtain a directed acyclic subgraph, calculate the node depth value and node breadth value according to the directed acyclic subgraph, and generate a cascaded operation execution sequence through a topological sorting algorithm, including: Count the dependency trigger times of each edge in the cyclic dependency link in the historical execution data, use the dependency trigger times as the dependency strength value of the edge, calculate the normalized dependency strength value of the edge, obtain the loop set where each edge in the cyclic dependency link is located, calculate the loop strength of each loop according to the product of the weight values of each edge in the loop, divide the loop strength by the loop length to obtain the loop unit strength, accumulate the loop unit strengths of all loops where the edge is located to obtain the criticality index of the edge, and perform weighted combination of the criticality index of the edge and the normalized dependency strength value of the edge to obtain the comprehensive evaluation value of the edge; Construct a minimum heap containing the comprehensive evaluation values, sequentially take out the current edge with the smallest comprehensive evaluation value from the minimum heap, judge whether the deletion of the current edge causes the graph to split, if it does not split, add the current edge to the feedback edge set and update the comprehensive evaluation values of the adjacent edges, and repeat the execution until a directed acyclic subgraph and the corresponding remaining edge set are obtained; In the directed acyclic sub - graph, set the depth value of the starting node with an in - degree of zero to zero, and set the breadth value of the terminating node with an out - degree of zero to zero. According to the connection relationships in the remaining edge set, through forward traversal, set the depth value of non - starting nodes to the maximum value of the depth values of all its predecessor nodes plus one, and through backward traversal, set the breadth value of non - terminating nodes to the maximum value of the breadth values of all its successor nodes plus one; Perform weighted combination calculation on the depth value, breadth value, normalized value of the historical execution time of each node, and the number of associated edges of the node in the feedback edge set to obtain the node priority value; Construct a priority queue based on the node priority value. In each time window, update the priority value of the node according to the execution status of the node and adjust the sorting of the nodes in the queue. Select multiple current nodes with the highest priority values that meet the parallel conditions from the priority queue, and add the current nodes to the cascade operation execution sequence in the selected order.

4. The method according to claim 1, wherein Read the cascade operation execution sequence, and perform node classification processing on the cascade operation execution sequence through a multi - dimensional feature classification algorithm. Organize the classified nodes into batch processing task groups, including: Read the cascade operation execution sequence and construct a multi - layer information network. Among them, the weight of the edges in the first - layer network represents the data transmission volume between nodes, the weight of the edges in the second - layer network represents the resource competition degree between nodes, and the weight of the edges in the third - layer network represents the operation similarity between nodes; Based on the multi - layer information network, calculate the degree centrality, betweenness centrality, and closeness centrality of each node, determine the centrality index, and combine the centrality index to form the network structure feature vector of the node; Obtain the state sequence of each node in the cascade operation execution sequence, extract the mean, variance, skewness, and kurtosis of the state sequence, determine the statistical features, and combine the statistical features to form the time - series feature vector of the node; Based on the input data scale, computational complexity, and memory occupancy of the node in the cascade operation execution sequence, determine the performance index, and construct the performance feature vector of the node based on the performance index; Combine the network structure feature vector, time - series feature vector, and performance feature vector to construct the feature matrix of the node; Calculate the Mahalanobis distance between the nodes in the feature matrix, and construct the similarity network of the nodes based on the Mahalanobis distance; In the similarity network, obtain the steady - state distribution of the nodes by iteratively calculating the transition probability between node pairs; Based on the steady - state distribution, identify the densely connected sub - graphs in the similarity network, and divide the nodes belonging to the same densely connected sub - graph into the same category; Calculate the path length between the nodes within each category, and split the category with a path length greater than the preset path length threshold to obtain the optimized node classification result; Organize the nodes with the same category into batch processing task groups.

5. The method according to claim 1, characterized in that, Calculate the optimal execution path of the batch processing task group through an adaptive path optimization algorithm, and generate distributed transaction control instructions according to the optimal execution path, including: Based on each task node in the batch task group, the load volatility is calculated by the ratio of the sum of the absolute values of the load value differences at adjacent time points within the sampling period of the task node to the sampling period. Based on the weighted combination of the load volatility difference and the time overlap degree, a task node distance matrix is constructed, and hierarchical clustering is performed on the batch task group according to the task node distance matrix to obtain task grouping information. An adaptive weight calculation model is constructed based on the task grouping information. The system load data of the task node is input into the sigmoid function to obtain a normalized load value, and the timeliness adjustment coefficient and resource adjustment coefficient of the task node are calculated according to the normalized load value. According to the timeliness adjustment coefficient, resource adjustment coefficient, and load volatility, the timeliness evaluation value, resource evaluation value, and volatility evaluation value are determined, and the evaluation score of the task node is calculated by weighting. The evaluation score is used as a state feature to construct a reinforcement learning model. An action space and a state space are constructed based on the state feature, and a reward function is determined according to the evaluation score, resource competition degree, and parallelism. The action-value function is iteratively updated using the deep Q-network to generate the selection strategy of the task node. Based on the selection strategy, path search is performed on the batch task group. In the exploration phase, an execution node is randomly selected to construct a candidate path, and in the exploitation phase, the optimal execution node is selected based on the action-value function to construct an optimal execution path. A distributed transaction control instruction is constructed according to the optimal execution path. The distributed transaction control instruction includes a pre-commit instruction and a commit instruction, where the pre-commit instruction is determined by collecting the status information of the task nodes on the optimal execution path. The causal relationship of the task nodes is maintained by incrementally updating and comparing the clock values of adjacent task nodes, the execution order of the task nodes is recorded, a vector clock instruction is determined, and the vector clock instruction is embedded in the distributed transaction control instruction. The execution result of the vector clock instruction is input into the distributed consensus module, and the distributed consensus module coordinates the transaction status of the task nodes to generate a distributed transaction control instruction.

6. The method according to claim 1, characterized in that, The link tracing and positioning algorithm includes: Collect transaction execution information to generate a transaction identification number, which is encoded by a timestamp field, a machine identification field, and an incrementing sequence number field, and marks all database operations on the same transaction link; record the operation execution information of the database operation, including the operation identification number, the parent operation identification number, the operation type information, the operation start time, the operation end time, and the operation status information. Establish a call relationship tree for the operation execution information, use the operation identification number as the tree node, and use the parent operation identification number as the parent-child node association relationship. Based on each node in the call relationship tree, an anomaly score is determined by calculating the weighted combination of the operation anomaly benchmark score, the operation delay anomaly score, and the operation status migration anomaly score, and an anomaly attenuation coefficient is calculated according to the hierarchical depth of the tree where the node is located. The final anomaly score of the node is obtained by multiplying the anomaly attenuation coefficient by the anomaly score. Calculate the time overlap degree between adjacent database operations, use the time overlap degree as the operation association strength, and construct an operation dependency network based on the operation association strength; analyze the operation execution order in the operation dependency network, calculate the conditional probability value between adjacent operations to determine the operation propagation probability; Filter the operation association edges whose final anomaly score is greater than a preset first anomaly threshold and the operation propagation probability is greater than a preset second anomaly threshold, and construct an anomaly propagation subgraph; Calculate the minimum weight spanning tree in the anomaly propagation subgraph to obtain the core path of anomaly propagation; Count the in-degree values of each node in the core path, add the nodes with in-degree value of zero to the processing queue, continuously take out the nodes in the queue and update the in-degree values of the corresponding adjacent nodes, and generate the propagation link of the abnormal transaction according to the node dequeue order.

7. The method according to claim 1, wherein The cascading rollback includes: Construct a multi-dimensional transaction dependency graph, store transaction operations in layers according to timestamps, establish a read-write version chain of data items, record the read-write timing information and data version numbers of data items, and generate transaction version dependency relationships; Construct a conflict matrix based on the transaction version dependency relationship, calculate the data access overlap degree between adjacent transactions, use the data access overlap degree as the conflict weight, and identify the high-frequency conflict areas where the conflict weight exceeds the preset conflict threshold; According to the conflict weight, calculate the criticality of the transaction node by accumulating the weighted values of the data modification amount, resource occupation duration, the conflict weight, and the transaction nesting level corresponding to the transaction node; Classify the transaction nodes based on the criticality, merge the transaction nodes whose criticality meets the approximation threshold into rollback batches according to a preset approximation threshold, calculate the resource consumption value and influence range of each rollback batch, and generate batch priorities; Generate a rollback execution plan according to the batch priorities, convert the rollback batches into parallel execution tasks in the order of batch priorities, and allocate rollback threads for each parallel execution task based on the resource consumption value; Start the rollback executor, submit the parallel execution tasks to the rollback thread pool, record the start timestamp of each rollback operation, and perform concurrent control by comparing the data version numbers; Monitor the rollback executor, when a rollback failure is detected, locate the failure position according to the timestamp, construct a compensation transaction to retry the rollback operation until the rollback is successful; Collect the operation logs output by the rollback executor, extract the version number sequence in the operation logs, reconstruct the transaction execution path based on the version number sequence, and verify data consistency.

8. Database cascade operation intelligent parsing and execution system for implementing the method described in any one of the foregoing claims 1-7, characterized in that Include: The first unit is used to receive a database cascading operation request, parse the database cascading operation request into an operation instruction set, extract the target table and associated tables from the operation instruction set through a recursive query algorithm, perform field mapping analysis to generate a dependency relationship weight matrix, and determine a cyclic dependency link based on the dependency relationship weight matrix through a strongly connected component identification algorithm; decompose the cyclic dependency link to obtain a directed acyclic subgraph, calculate the node depth value and node breadth value according to the directed acyclic subgraph, and generate a cascading operation execution sequence through a topological sorting algorithm; A second unit, configured to read the cascade operation execution sequence, perform node classification processing on the cascade operation execution sequence through a multi-dimensional feature classification algorithm, organize the classified nodes into a batch processing task group, calculate an optimal execution path of the batch processing task group through an adaptive path optimization algorithm, generate a distributed transaction control instruction according to the optimal execution path, implant a transaction synchronization point in the distributed transaction control instruction, the transaction synchronization point is set based on a two-phase commit protocol, allocate the batch processing task group to an execution thread pool through a consistent hashing algorithm, and the execution thread pool executes the batch processing task group based on a pipeline parallel processing mechanism to generate execution status information; A third unit, configured to monitor the running process of the execution thread pool, continuously collect the execution status information through a sliding window algorithm to generate performance data, input the performance data into a pre-trained neural network model to obtain a performance prediction value, when the performance prediction value exceeds a preset performance threshold, calculate a thread load level based on the execution status information, adjust resources of the execution thread pool according to the thread load level, if an abnormal signal is included in the execution status information, locate an abnormal transaction link through a link tracing and positioning algorithm, perform cascade rollback according to the abnormal transaction link, and reconstruct a transaction execution path according to the cascade operation execution sequence to complete a database cascade operation.

9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Platform operation performance evaluation method and system based on collaborative intelligent analysis

    CN120631723A

  • Comprehensive administrative law enforcement management service system

    CN120725627A

  • AI-combined multi-dimensional data set analysis processing method and system

    CN120744333A

  • A method and system for multidimensional dataset analysis and processing combined with AI

    CN120744333B

  • Optimization method for incremental migration mode of database synchronization tool

    CN120763256A