Cloud database operation and maintenance management method and system, electronic equipment and storage medium

By detecting early warning information and operating status data, combining the repair strategy library and intelligent model, the repair strategy is determined independently, and the problems of low repair efficiency and low accuracy in cloud database operation and maintenance are solved, and fast and accurate fault repair and stable operation are achieved.

CN120492424APending Publication Date: 2025-08-15SHENZHEN LIUXIN TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510346952.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The repair efficiency and low accuracy in the operation and maintenance of existing cloud databases lead to poor operational stability. The existing technology relies on manual inspection and self-healing rules to deal with complex failures, which may cause misrepair.

Method used

By detecting early warning information and operating status data, combining the first and second repair strategy libraries and policy recommendation models, the repair strategy is determined independently, and deep reinforcement learning and recurrent neural network models are used to predict and repair strategy optimization.

Benefits of technology

It realizes fast and accurate fault detection and repair, reduces fault detection time, improves repair efficiency and accuracy, and ensures the stable operation of cloud databases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492424A_ABST
    Figure CN120492424A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of cloud databases, and provides a cloud database operation and maintenance management method and system, electronic equipment and a storage medium. The cloud database operation and maintenance management method comprises the following steps: detecting instance nodes according to early warning information and operation state data; determining a target repair strategy from a first repair strategy library; determining a target fault type corresponding to the target operation state data; determining a target repair strategy from a second repair strategy library; determining a target repair strategy according to the target joint state data and a strategy recommendation model; and repairing the target instance node according to the determined target repairing strategy. According to the method, the abnormal instance node can be detected in time according to the early-warning potential fault information, the fault discovery time is effectively shortened, in addition, the applicable repair strategy is autonomously determined through the synergistic effect of multiple modes, the abnormal instance node in the cloud database is efficiently and accurately repaired, and stable operation of the cloud database is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of cloud database technology, and in particular relates to a cloud database operation and maintenance management method, system, electronic device and storage medium. Background Art

[0002] With the development of cloud computing technology, cloud databases have become widely used. Existing technologies for cloud database operation and maintenance primarily rely on manual labor after a failure occurs. When an anomaly occurs, operations and maintenance personnel are notified to go online to handle the issue. This manual, experience-based troubleshooting approach is inefficient and impacts repair efficiency. Currently, automated repairs based on pre-set self-healing rules are also available. However, these rules alone are inadequate for addressing the complex and ever-changing nature of cloud database failures and may even result in incorrect repairs, impacting repair efficiency and accuracy, and ultimately, the operational stability of the cloud database. Summary of the Invention

[0003] The embodiments of the present application provide a cloud database operation and maintenance management method, system, electronic device, and storage medium, which can solve the problem of poor cloud database operation stability caused by low cloud database repair efficiency and low accuracy.

[0004] In a first aspect, an embodiment of the present application provides a cloud database operation and maintenance management method, including: Detecting the corresponding instance node based on the warning information and the operating status data of each instance node in the cloud database; the warning information indicates a potential failure of the cloud database; When a target instance node with an abnormal operating state is detected, a target repair strategy whose target operating state data satisfies a state condition is determined from a first repair strategy library; wherein the target operating state data is the operating state data of the target instance node, the state condition is that the value of the indicator data in the target operating state data is within a corresponding abnormal indicator value range, and the first repair strategy library includes at least one abnormal indicator value range of the indicator data and a matching first repair strategy; When the target repair strategy does not exist in the first repair strategy library, determining a target fault type corresponding to the target operating status data; Determining a target repair strategy that matches the target fault type from a second repair strategy library; the second repair strategy library includes at least one fault type and a matching second repair strategy; When the target repair strategy does not exist in the second repair strategy library, determining the target repair strategy based on the target joint state data and a strategy recommendation model; the joint state data includes the target operating state data and the historical repair data of the target instance node; the strategy recommendation model is configured to take the joint state data as input and output the repair strategy that can obtain the maximum reward value under the input; Repair the target instance node according to the determined target repair strategy.

[0005] Compared with the prior art, the embodiments of the present application have the following beneficial effects: It can promptly detect abnormal instance nodes based on the potential fault information in the early warning, effectively shortening the time to discover the fault. In addition, it can independently determine the applicable repair strategy through the synergy of multiple methods, and efficiently and accurately repair the abnormal instance nodes in the cloud database to ensure the stable operation of the cloud database.

[0006] In a possible implementation manner of the first aspect, the step of determining a target fault type corresponding to the target operating status data includes: Determining the target fault type according to the target operating state data and a preset fault diagnosis rule; the fault diagnosis rule includes at least one operating state data and a matching fault type; and / or, The target fault type is determined based on the target operating status data and a fault diagnosis model; the fault diagnosis model is used to take the operating status data as input and output at least one fault type and a corresponding probability value, and the target fault type is the fault type corresponding to the maximum probability value.

[0007] In the above solution, the fault type can be quickly and accurately analyzed from the operating status data by combining rule matching and model diagnosis.

[0008] In a possible implementation of the first aspect, the training process of the policy recommendation model specifically includes: Constructing a state space, action space, and reward function for an instance node; the state space represents the operating state data and historical repair data of the instance node, the action space represents the repair strategy, and the reward function represents the action value of the repair strategy; Building a deep reinforcement learning model based on the state space, the action space, and the reward function, and constructing an experience recycling pool; For the current running state data, select and execute the repair strategy with the maximum predicted action value according to the deep reinforcement learning model to obtain the running state data of the next state and the corresponding reward value; The current operating status data, the current repair strategy, the operating status data of the next state and the corresponding reward value are stored as sample data in the experience recovery pool, and the sample data is extracted from the experience recovery pool. The deep reinforcement learning model is trained with the goal of minimizing the difference between the predicted action value and the preset target action value to obtain a trained strategy recommendation model.

[0009] In a possible implementation manner of the first aspect, after the step of repairing the target instance node according to the determined target repair strategy, the method further includes: Obtain the current running status data of the repaired target instance node; Evaluate the current operating status data using a preset repair evaluation model, and if at least one indicator data in the current operating status data does not conform to a preset normal indicator value range, return to the step of determining the target repair strategy; If all indicator data in the current operating status data are in compliance with the preset normal indicator value range, the repair is completed.

[0010] In the above solution, the use of intelligent models can more comprehensively and accurately judge the actual repair effect, ensure that the instance node is truly restored to normal, and adjust the repair strategy in a timely manner based on the actual repair effect, further improving the repair efficiency and accuracy.

[0011] In a possible implementation manner of the first aspect, before the step of repairing the target instance node according to the determined target repair strategy, the step further includes: Determine the risk assessment result of the target repair strategy according to a preset risk prediction model; Develop risk control strategies based on the risk assessment results.

[0012] In the above solution, intelligent models are used to pre-evaluate and control the risks that may be brought about by the repair strategy, avoiding new anomalies during the repair process and further ensuring the operational stability of the cloud database.

[0013] In a possible implementation of the first aspect, the cloud database operation and maintenance management method further includes: Obtaining real-time operating status data of each instance node in the cloud database during a preset time period; Input the real-time operating status data into the trained fault prediction model for prediction, and obtain the prediction result corresponding to each instance node; the prediction result includes the probability of potential fault occurrence, potential fault type and potential fault occurrence time; The warning information is generated according to the prediction result; the warning information includes identification information of the warning instance node, the potential fault type, the potential fault occurrence time and prevention strategy, and the warning instance node is an instance node whose probability of potential fault occurrence is greater than a preset probability threshold.

[0014] In the above solution, the model can be used to intelligently predict potential failures of instance nodes, and preventive measures can be taken before potential failures occur, greatly reducing the probability of actual failures. It can also facilitate subsequent key inspections of instance nodes that may fail, effectively shortening the time interval from fault discovery to repair processing, and ensuring the operational stability of the cloud database.

[0015] In a possible implementation of the first aspect, the training process of the fault prediction model specifically includes: Acquire a historical data set; the historical data set includes historical operating status data and historical fault data of the instance node; Performing data preprocessing on the historical data set to construct a sample set, and dividing the sample set into a training set and a validation set; The recurrent neural network model is iteratively trained based on the training set, and then the recurrent neural network model after each training is verified and evaluated based on the verification set until the recurrent neural network model converges or reaches a preset number of training rounds, thereby obtaining a trained fault prediction model.

[0016] In a second aspect, an embodiment of the present application provides a cloud database operation and maintenance management system, including: a detection module configured to detect corresponding instance nodes based on warning information and operating status data of each instance node in the cloud database; the warning information indicates a potential failure of the cloud database; A first repair strategy determination module is configured to, when a target instance node with an abnormal operating state is detected, determine from a first repair strategy library a target repair strategy whose target operating state data satisfies a state condition; wherein the target operating state data is operating state data of the target instance node, the state condition is that a value of indicator data in the target operating state data is within a corresponding abnormal indicator value range, and the first repair strategy library includes at least one abnormal indicator value range of indicator data and a matching first repair strategy; a fault type determination module, which determines a target fault type corresponding to the target operating status data when the target repair strategy does not exist in the first repair strategy library; A second repair strategy determination module is configured to determine a target repair strategy that matches the target fault type from a second repair strategy library; the second repair strategy library includes at least one fault type and a matching second repair strategy; a third repair strategy determination module, configured to, when the target repair strategy does not exist in the second repair strategy library, determine the target repair strategy based on target joint state data and a strategy recommendation model; the joint state data includes the target operating state data and historical repair data of the target instance node; the strategy recommendation model is configured to use the joint state data as input and output a repair strategy that can obtain a maximum reward value under the input; The repair module is used to repair the target instance node according to the determined target repair strategy.

[0017] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the cloud database operation and maintenance management method described in any one of the first aspects above is implemented.

[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the cloud database operation and maintenance management method described in any one of the above-mentioned first aspects is implemented.

[0019] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the cloud database operation and maintenance management method described in any one of the above-mentioned first aspects.

[0020] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] Figure 1 This is a flowchart of a cloud database operation and maintenance management method provided by an embodiment of the present application; Figure 2 This is a flow chart of generating warning information in a cloud database operation and maintenance management method provided by another embodiment of the present application; Figure 3 is a flowchart of a fault prediction model training process provided by another embodiment of the present application; Figure 4This is a flowchart of a training strategy recommendation model provided by another embodiment of the present application; Figure 5 This is a flowchart of evaluating the repair effect in the cloud database operation and maintenance management method provided by another embodiment of the present application; Figure 6 This is a flowchart of predicting repair risks in a cloud database operation and maintenance management method provided by another embodiment of the present application; Figure 7 This is a schematic diagram of the structure of the cloud database operation and maintenance management system provided by an embodiment of the present application; Figure 8 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0024] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0025] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0026] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0027] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0028] See also Figure 1, is a flowchart of a cloud database operation and maintenance management method provided in an embodiment of the present application. As an example and not a limitation, the method may include the following steps: S11 , detecting the corresponding instance nodes according to the warning information and the operating status data of each instance node in the cloud database.

[0029] Warning information indicates potential cloud database failures. Specifically, a cloud database instance node refers to the database operation unit deployed by a cloud service provider within a managed database service. Each instance node stores and processes corresponding data. There is no limit on the number of instance nodes; they can generally be configured based on cost considerations and actual data requirements. Node management is a critical component of cloud server operations and maintenance.

[0030] In this embodiment, the operating status data can be understood as multi-dimensional performance indicator data, such as CPU (Central Processing Unit) usage, memory occupancy, disk I / O (Input / Output) read and write rate, network bandwidth usage, database transaction processing volume, transaction response time, etc.

[0031] In a possible implementation, the detection priority of the instance nodes may be determined based on the early warning information, that is, instance nodes with potential faults may be monitored intensively and their operating status data may be analyzed preferentially.

[0032] S12. When a target instance node with an abnormal running state is detected, a target repair strategy whose target running state data meets the state condition is determined from the first repair strategy library, and then step S16 is executed.

[0033] Among them, the target operating status data is the operating status data of the target instance node, the status condition is that the value of the indicator data in the target operating status data is within the corresponding abnormal indicator value range, and the first repair strategy library includes at least one abnormal indicator value range of indicator data and a matching first repair strategy.

[0034] S13: When the target repair strategy does not exist in the first repair strategy library, determine the target fault type corresponding to the target operating status data.

[0035] S14. Determine a target repair strategy that matches the target fault type from the second repair strategy library, and then execute step S16.

[0036] The second repair strategy library includes at least one fault type and a matching second repair strategy.

[0037] S15. When the target repair strategy does not exist in the second repair strategy library, determine the target repair strategy according to the target joint state data and the strategy recommendation model, and then execute step S16.

[0038] The joint state data includes the target running state data and the historical repair data of the target instance node. The strategy recommendation model is used to take the joint state data as input and output the repair strategy that can obtain the maximum reward value under the input.

[0039] S16. Repair the target instance node according to the determined target repair strategy.

[0040] This embodiment provides a cloud database operation and maintenance management method that can promptly detect abnormal instance nodes based on early warning potential fault information, effectively shortening the time to fault discovery. In addition, through the synergistic effect of multiple methods, it can autonomously determine the applicable repair strategy, efficiently and accurately repair abnormal instance nodes in the cloud database, and ensure the stable operation of the cloud database.

[0041] In a possible implementation, the warning information is updated regularly at regular intervals. Figure 2 , the generation method of early warning information may include the following steps: S101: Acquire real-time operating status data of each instance node in a cloud database during a preset time period.

[0042] Specifically, the preset time period can be set according to the actual update frequency. For example, assuming that the warning information is updated once a day, that is, the preset time period is twenty-four hours, the operating status data within the twenty-four hours before the current moment can be used as real-time operating status data.

[0043] S102: Input the real-time operating status data into the trained fault prediction model for prediction, and obtain the prediction result corresponding to each instance node.

[0044] The prediction results include the probability of potential failure, the type of potential failure and the time of potential failure.

[0045] For example, if the fault prediction model predicts that a high CPU usage fault may occur within the next 24 hours, the output is "A high CPU usage fault may occur within the next 24 hours."

[0046] S103. Generate warning information based on the prediction results.

[0047] The warning information includes the identification information of the warning instance node, the potential fault type, the potential fault occurrence time and the prevention strategy. The warning instance node is the instance node whose potential fault occurrence probability is greater than the preset probability threshold.

[0048] It should be noted that when the probability of a potential fault occurring is greater than the preset probability threshold, the prediction of the potential fault type and the time of potential fault occurrence is entered, and the identification information of the corresponding instance node is obtained. If the probability of a potential fault occurring is lower than the preset probability threshold, the output is "no fault will occur in the future."

[0049] In one possible implementation, a preventative strategy can be developed based on the potential failure type and potential failure time. For example, if a certain instance node is predicted to fail due to excessive memory usage within the next few hours, the preventative strategy could be to optimize memory resource allocation, clean up memory, or expand memory capacity, potentially preventing future failures caused by excessive memory usage.

[0050] This embodiment provides a cloud database operation and maintenance management method that can use models to intelligently predict potential failures of instance nodes, and can take preventive measures before potential failures occur, greatly reducing the probability of actual failures. It also facilitates subsequent key detection of instance nodes that may fail, effectively shortening the time interval from fault discovery to repair processing, and ensuring the operational stability of the cloud database.

[0051] Optionally, an implementation of step S101 may include collecting the running status data of the instance node at a preset time interval. For example, the preset time interval is 5 minutes, that is, the running status data of the instance node is collected once with 5 minutes as a sampling point.

[0052] In a possible implementation, between step S101 and step S102 , data preprocessing is also included for the real-time operating status data, specifically including operations such as data cleaning and normalization.

[0053] Optionally, before step S102, the method further includes constructing and training a fault prediction model, such as Figure 3 As shown in Figure 2, the training process of the fault prediction model includes: S1001. Obtain a historical data set.

[0054] The historical data set includes the historical operating status data and historical fault data of the instance node. Specifically, the historical operating status data can represent the operating status data over a longer period of time, such as the operating status data within the past month, while the historical fault data includes the time of fault occurrence, fault type, and status data before the fault.

[0055] It should be noted that the method for obtaining historical running status data is the same as the method for obtaining real-time running status data, that is, the running status data of the instance node is collected at the same preset time interval.

[0056] In an optional implementation, operating status data and fault data can be continuously collected from each instance node of the cloud database, and these data can be stored in the data warehouse. These collected data can be selected from the data warehouse as historical data or real-time data, thereby providing data support for subsequent model training using historical data and prediction using real-time data.

[0057] Specifically, for example, historical data is data collected within thirty days before the current moment, with a sampling point every 5 minutes; real-time data is data collected within 24 hours before the current moment, also with a sampling point every 5 minutes.

[0058] S1002: Preprocess the historical data set, construct a sample set, and divide the sample set into a training set and a validation set.

[0059] In an optional implementation, preprocessing operations such as data cleaning and normalization are performed on the historical data set to remove noise and invalid data. Data cleaning also includes missing value processing and outlier processing. Specifically, the following examples illustrate the operation methods of missing value processing, outlier processing, and normalization processing.

[0060] Missing value processing: If a certain indicator data at a sampling point is missing, linear interpolation is used to fill it in. For example, if the CPU usage is missing at the i-th sampling point, the CPU usage at the i-1th sampling point is 30%, and the CPU usage at the i+1th sampling point is 32%, then the CPU usage at the i-th sampling point is filled in as 31%.

[0061] Outlier handling: An outlier is defined as a value that exceeds the normal range of a metric by three standard deviations. For example, if the normal mean of CPU usage is 50% and the standard deviation is 10%, then values outside the range of 20%-80% are considered outliers and replaced with the median of the metric.

[0062] Normalization: Use the Min-Max normalization method to scale the data of each indicator to the range [0,1]. Taking CPU usage as an example, assuming that the minimum CPU usage in historical data is 10% and the maximum is 90%, for the CPU usage x at a certain sampling point, the normalized CPU usage calculation formula is: .

[0063] It should be noted that the same preprocessing operations are performed on real-time data as on historical data to ensure consistency in data format and range.

[0064] S1003. Iteratively train the recurrent neural network model based on the training set, and then verify and evaluate the recurrent neural network model after each training based on the validation set until the recurrent neural network model converges or reaches a preset number of training rounds, thereby obtaining a trained fault prediction model.

[0065] Optionally, an implementation of step S1002 may include dividing the sample set into a training set and a validation set according to a preset ratio, such as the common 80% of the data as a training set and 20% of the data as a validation set. Of course, it can also be divided according to other ratios.

[0066] Specifically, this embodiment uses the LSTM model (Long Short-Term Memory) as an example of a recurrent neural network model. The network structure of the LSTM model includes an input layer: the input dimension is set to four, corresponding to the four indicators of CPU utilization, memory utilization, disk I / O read / write rate, and network bandwidth utilization. The LSTM layer (hidden layer) sets the number of hidden units to 128. The output layer sets the output dimension to one, that is, the probability value of a failure, which is used to predict whether a failure will occur at a certain time in the future. For example, assuming the probability threshold is set to 0.5, if the output probability value is greater than 0.5, it is predicted that the instance node will fail in the future; otherwise, it is predicted to be normal.

[0067] In this embodiment, the LSTM model is expanded by adding multiple neurons to the output layer, each corresponding to a fault type. Specifically, when the output of the output layer fails, each neuron outputs the probability of the corresponding fault type. The fault type with the highest output probability value can be used as the predicted potential fault type.

[0068] In this embodiment, time series analysis is combined with the predicted potential fault type and the temporal patterns of historical fault occurrences to estimate the time range within which potential faults may occur. Time series analysis methods include trend analysis and periodic analysis, which analyze the temporal changes in operating status data to reveal periodic patterns or trends for different fault types.

[0069] In each round of training, the RNN model randomly selects a batch of data from the training set for forward and backward propagation, updating the model parameters. After each round of training, the RNN model is evaluated using the validation set, recording the loss and accuracy of the validation set. Based on the performance of the validation set, the model parameters are further adjusted to avoid overfitting or underfitting.

[0070] For example, training parameters include but are not limited to learning rate, batch size, number of training rounds, activation function, and optimizer. Specifically: Learning rate: The learning rate is a hyperparameter that controls the update speed of the model weights and biases. A too high learning rate may cause model oscillation, while a too low learning rate may slow model convergence. The initial learning rate setting range is 0.001-0.01, and it can be dynamically adjusted during training. In this example, the learning rate is set to 0.001.

[0071] Batch size: The batch size is the number of samples processed in each iteration. Smaller batch sizes can better reduce fluctuations during gradient descent but increase noise. Larger batch sizes can improve computational efficiency but cannot fully utilize the parallel computing capabilities of the GPU. Typically, the batch size can be set to 32, 64, or 128. In this example, the batch size is set to 32, meaning 32 samples are used in each training run.

[0072] Training rounds: The number of training rounds refers to the number of iterations when training the model. Too few iterations will lead to underfitting of the model, while too many iterations will lead to overfitting of the model. It can usually be set to 100-200 times. In this embodiment, the number of training rounds is set to 100.

[0073] Activation function: Activation functions in LSTM models include but are not limited to sigmoid, tanh, and ReLU (common activation functions). The sigmoid function maps input values between 0 and 1, the tanh function maps input values between -1 and 1, and the ReLU function solves the vanishing gradient problem. This example uses the ReLU function as the activation function.

[0074] Optimizer: An optimizer represents a parameter update method, including but not limited to Adam, SGD, and RMSprop (common optimizers). This example uses the Adam optimizer.

[0075] Optionally, an implementation method of step S12 may include parsing the target operating status data, extracting various indicator data and corresponding values; comparing the value of each indicator data with the abnormal indicator value range of each indicator data in the first repair strategy library one by one; if there is an indicator data whose value meets the abnormal indicator value range of any indicator data, then determining the first repair strategy corresponding to the indicator data as the target repair strategy, and then executing step S16; otherwise, executing step S13.

[0076] Specifically, for example, it is assumed that the various indicator data and corresponding values in the target operating status data of the target instance node are as shown in Table 1.

[0077]

[0078] As an example but not a limitation, it is assumed that the first repair strategy library is as shown in Table 2.

[0079]

[0080] For the target instance node in the preceding example, since the memory usage is 93%, which meets the abnormal indicator value range of "memory usage greater than 90%", you can directly select "memory cleanup or expansion" as the target repair strategy for the target instance node.

[0081] In this embodiment, the direct matching method can quickly find corresponding repair methods for common faults, thereby improving the efficiency of handling common faults.

[0082] Optionally, an implementation of step S13 may include determining a target fault type according to the target operating state data and preset fault diagnosis rules.

[0083] When handling complex faults, in addition to the aforementioned indicator data, other relevant information must also be considered, including but not limited to database operation logs, the number of concurrent transactions, and the proportion of lock wait time, to more comprehensively assess the fault situation.

[0084] Fault diagnosis rules include at least one piece of operational status data and a matching fault type. Specifically, a series of clear rules are set to preliminarily determine the fault type. For example, if the average transaction response time exceeds 2500 milliseconds and the lock wait time accounts for more than 20%, it can be determined to be a database lock conflict fault.

[0085] Optionally, an implementation of step S13 may include determining a target fault type according to the target operating state data and a fault diagnosis model.

[0086] The fault diagnosis model is used to take the operating status data as input and at least one fault type and corresponding probability value as output, and the target fault type is the fault type corresponding to the maximum probability value.

[0087] Specifically, the fault diagnosis model is trained using a decision tree model, which is then used to further accurately determine the fault type. The decision tree features include operational status data from multiple dimensions, such as transaction response time, lock wait time percentage, number of database connections, and number of concurrent transactions. Using this operational status data as input, the decision tree model outputs a probability distribution for the fault type. Assuming the output probability of a database lock conflict fault is 0.85, and the probabilities of other fault types are all lower than 0.85, the fault type is ultimately determined to be a database lock conflict.

[0088] In one possible implementation, step S13 can also use the potential fault type provided by the early warning information as a reference to more accurately determine the actual fault type. Specifically, for example, if the average transaction response time exceeds 2500 milliseconds and the lock wait time accounts for more than 20%, it can be preliminarily determined to be a database lock conflict fault. At the same time, if the early warning information indicates that a database lock conflict fault may occur in the future, the fault type can be further determined to be a database lock conflict.

[0089] This embodiment provides a cloud database operation and maintenance management method that uses a combination of rule matching and model diagnosis to quickly and accurately analyze the fault type from the operating status data, avoid blindly attempting repair measures, and improve the accuracy of repairs.

[0090] Optionally, in step S14, if the second repair strategy library stores detailed and accurate second repair strategies for different fault types, then the corresponding repair strategy can be selected from the second repair strategy library based on the fault type determined in step S13, and then step S16 is executed. If the fault type is unclear and an accurate repair strategy cannot be selected from the second repair strategy library, step S15 is further executed.

[0091] As an example and not a limitation, it is assumed that the second repair strategy library is as shown in Table 3:

[0092] For the target instance node in the above example, since the fault type is determined to be a database lock conflict, "unlock operation, optimize lock mechanism" can be directly selected as the target repair strategy for the target instance node.

[0093] Optionally, an implementation method of step S15 may include constructing and training a strategy recommendation model; using the strategy recommendation model, combined with the target operating status data and historical repair data of the current target instance node, recommending the most appropriate repair strategy for the target instance node, and then executing step S16; the strategy recommendation model continuously adjusts the recommendation of the repair strategy according to the repair effect.

[0094] In one possible implementation, Figure 4 As shown in Figure 2, the training process of the strategy recommendation model specifically includes: S21. Construct the state space, action space and reward function of the instance node.

[0095] Among them, the state space represents the operating status data and historical repair data of the instance node, the action space represents the repair strategy, and the reward function represents the action value of the repair strategy.

[0096] State space: Contains the operating status data of the instance node, including but not limited to CPU usage, memory usage, disk I / O rate, network bandwidth usage, number of database connections, transaction response time, etc., as well as historical repair records, including but not limited to the repair strategies and repair results of several past repairs.

[0097] Action space: All repair strategies implemented so far.

[0098] Reward function: Accurate rewards are given based on the effectiveness of the repair. If the metrics of the instance node return to normal within a certain period of time after the repair, a positive reward is given. If the metrics remain abnormal after the repair, such as further extended transaction response times or a significant increase in CPU usage, a negative reward is given.

[0099] S22. Build a deep reinforcement learning model based on the state space, action space, and reward function, and construct an experience recycling pool.

[0100] In one possible implementation, the deep reinforcement learning model is trained using a deep Q-network (DQN). The DQN network consists of a first neural network for recommending repair strategies and a second neural network for calculating Q-values. Q-values can be understood as action values: the expected total reward that an agent can obtain from its current state to a future state by following a repair strategy in the action space, given operational state data in the state space. In this embodiment, the network parameters of the first and second neural networks are synchronized.

[0101] This embodiment provides a cloud database operation and maintenance management method, in which the experience replay pool enables the neural network model in deep reinforcement learning training to achieve rapid learning convergence based on a large amount of historical data; during the deep reinforcement learning training process, the experience replay pool will be continuously expanded and iterated, so that the deep reinforcement learning model can obtain comprehensive and valuable experience data.

[0102] S23. For the current running status data, select and execute the repair strategy with the largest predicted action value based on the deep reinforcement learning model to obtain the running status data of the next state and the corresponding reward value.

[0103] S24. Store the current operating status data, the current repair strategy, the operating status data of the next state and the corresponding reward value as sample data in the experience recovery pool, and extract sample data from the experience recovery pool to train the deep reinforcement learning model with the goal of minimizing the difference between the predicted action value and the preset target action value, and obtain a trained strategy recommendation model.

[0104] In one possible implementation, after executing a recommended repair strategy, the reward value is updated based on the actual repair effect. Using an experience replay mechanism, the new state, action, reward, and next state are stored in an experience replay pool. Periodically, a batch of data is randomly sampled from the experience replay pool to update the neural network parameters of the DQN model to optimize the recommended strategy.

[0105] This embodiment provides a cloud database operation and maintenance management method that improves the accuracy and success rate of repair strategy recommendations by continuously learning and adjusting the strategy recommendation model. For example, after repeated learning and adjustment, the model can more accurately recommend the most appropriate repair strategy for different fault conditions, increasing the repair success rate from an initial 70% to 85%.

[0106] In one possible implementation, Figure 5 As shown, after executing the step of repairing the target instance node according to the determined target repair strategy, the cloud database operation and maintenance management method further includes: S17. Obtain the current running status data of the repaired target instance node.

[0107] Specifically, after the repair is completed, relevant performance indicators and status data of the instance node are collected, such as memory usage, CPU usage, response time, transaction processing success rate, etc.

[0108] S18. Evaluate the current operating status data using a preset repair evaluation model. If at least one indicator data in the current operating status data does not conform to the preset normal indicator value range, return to the step of determining the target repair strategy; if all indicator data in the current operating status data conform to the preset normal indicator value range, the repair is completed.

[0109] In one possible implementation, a pre-trained repair assessment model is used to evaluate the effectiveness of repairs. This model, trained by studying a large amount of pre- and post-repair data and the corresponding repair results, accurately determines whether an instance node has truly returned to normal. Specifically, the model inputs the collected current operating status data. If all indicators are within the normal range and meet the model's criteria, the repair is considered a success; otherwise, the repair is considered a failure.

[0110] This embodiment provides a cloud database operation and maintenance management method that utilizes an intelligent model to more comprehensively and accurately determine the actual repair effect, ensuring that instance nodes are truly restored to normal. The repair strategy is adjusted in a timely manner based on the actual repair effect, further improving repair efficiency and accuracy.

[0111] Optionally, the cloud database operation and maintenance management method further includes recording the repair data of the target instance node during the execution of the repair strategy and storing it as historical repair data of the target instance node. Specifically, during the repair process, information such as the time when the instance node failure occurred, the type of failure that occurred, the repair strategy used, and the time when the repair was completed are recorded in real time.

[0112] In one possible implementation, Figure 6 As shown, before executing the step of repairing the target instance node according to the determined target repair strategy, the cloud database operation and maintenance management method further includes: S151. Determine the risk assessment result of the target repair strategy based on a preset risk prediction model.

[0113] Optionally, before performing a repair operation, an AI model is used to predict the risks that may be caused by the repair operation. The AI model can consider multiple factors, such as the impact of the repair operation on other nodes, changes in system resources, data consistency risks, etc., to quantitatively assess the risks.

[0114] S152. Develop risk control strategies based on risk assessment results.

[0115] Optionally, corresponding risk control measures are formulated based on the risk prediction results, and step S16 is performed after the risk control measures are implemented. For example, if it is predicted that the memory expansion operation may affect the performance of other instance nodes, the resource allocation of other instance nodes is adjusted before the operation.

[0116] In one possible implementation, the cloud database operation and maintenance management method further includes monitoring the cloud database instance nodes in real time during the execution of the risk control and repair strategies. If an anomaly occurs, the method returns to step S151. For example, when adjusting database parameters, the cloud database instance nodes are monitored step by step and in real time. If an anomaly occurs, the risk control measures are promptly stopped and the risk is reassessed.

[0117] This embodiment provides a cloud database operation and maintenance management method that uses an intelligent model to pre-evaluate and control risks that may be brought about by repair strategies, thereby avoiding new anomalies during the repair process and further ensuring the operational stability of the cloud database.

[0118] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0119] Corresponding to the cloud database operation and maintenance management method described in the above embodiment, Figure 7A structural block diagram of a cloud database operation and maintenance management system provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0120] Reference Figure 7 , the cloud database operation and maintenance management system includes: The detection module 11 is configured to detect the corresponding instance nodes according to the warning information and the running status data of each instance node in the cloud database; the warning information indicates a potential failure of the cloud database.

[0121] The first repair strategy determination module 12 is used to determine, from the first repair strategy library, a target repair strategy whose target operating status data satisfies a status condition when a target instance node with an abnormal operating status is detected; wherein the target operating status data is the operating status data of the target instance node, the status condition is that the value of the indicator data in the target operating status data is within the corresponding abnormal indicator value range, and the first repair strategy library includes at least one abnormal indicator value range of the indicator data and a matching first repair strategy.

[0122] The fault type determination module 13 is configured to determine a target fault type corresponding to the target operating status data when there is no repair strategy in the first repair strategy library whose target operating status data satisfies a corresponding status condition.

[0123] The second repair strategy determination module 14 is configured to determine a target repair strategy that matches the target fault type from a second repair strategy library; the second repair strategy library includes at least one fault type and a matching second repair strategy.

[0124] The third repair strategy determination module 15 is used to determine the target repair strategy based on the target joint state data and the strategy recommendation model when there is no target repair strategy matching the target fault type in the second repair strategy library; the joint state data includes the target operating state data and the historical repair data of the target instance node, and the strategy recommendation model is used to take the joint state data as input and take as output the repair strategy adopted when the maximum reward value can be obtained under the input.

[0125] The repair module 16 is configured to repair the target instance node according to the determined target repair strategy.

[0126] In some embodiments of the present application, the fault type determination module 13 may be specifically configured to determine the target fault type based on the target operating state data and preset fault diagnosis rules; the fault diagnosis rules include at least one operating state data and a matching fault type.

[0127] In some embodiments of the present application, the fault type determination module 13 can also be specifically used to determine the target fault type based on the target operating status data and the fault diagnosis model; the fault diagnosis model is used to take the operating status data as input and at least one fault type and corresponding probability value as output, and the target fault type is the fault type corresponding to the maximum probability value.

[0128] In some embodiments of the present application, the cloud database operation and maintenance management system further includes a first model training module for training a strategy recommendation model.

[0129] In some embodiments of the present application, the first model training module is specifically used to construct the state space, action space and reward function of the instance node; the state space represents the operating status data and historical repair data of the instance node, the action space represents the repair strategy, and the reward function represents the action value of the repair strategy; a deep reinforcement learning model is constructed based on the state space, action space and reward function, and an experience recovery pool is constructed; for the current operating status data, the repair strategy with the largest predicted action value is selected and executed according to the deep reinforcement learning model to obtain the operating status data of the next state and the corresponding reward value; the current operating status data, the current repair strategy, the operating status data of the next state and the corresponding reward value are stored as sample data in the experience recovery pool, and sample data is extracted from the experience recovery pool to train the deep reinforcement learning model with the goal of minimizing the difference between the predicted action value and the preset target action value to obtain a trained strategy recommendation model.

[0130] In some embodiments of the present application, the cloud database operation and maintenance management system also includes a repair effect evaluation module for obtaining the current operating status data of the target instance node after repair; the current operating status data is evaluated through a preset repair evaluation model. If there is at least one indicator data in the current operating status data that does not meet the preset normal indicator value range, the module returns to the step of determining the target repair strategy; if all indicator data in the current operating status data meet the preset normal indicator value range, the repair is completed.

[0131] In some embodiments of the present application, the cloud database operation and maintenance management system further includes a repair risk prediction module for determining a risk assessment result of a target repair strategy based on a preset risk prediction model; and formulating a risk control strategy based on the risk assessment result.

[0132] In some embodiments of the present application, the cloud database operation and maintenance management system also includes an early warning module for obtaining real-time operating status data of each instance node in the cloud database during a preset time period; inputting the real-time operating status data into a trained fault prediction model for prediction to obtain a prediction result corresponding to each instance node; the prediction result includes the probability of potential fault occurrence, the potential fault type and the potential fault occurrence time; generating early warning information based on the prediction result; the early warning information includes identification information of the early warning instance node, the potential fault type, the potential fault occurrence time and prevention strategy, and the early warning instance node is an instance node whose potential fault occurrence probability is greater than a preset probability threshold.

[0133] In some embodiments of the present application, the cloud database operation and maintenance management system further includes a second model training module for training a fault prediction model.

[0134] In some embodiments of the present application, the second model training module is specifically used to obtain a historical data set; the historical data set includes historical operating status data and historical fault data of instance nodes; data preprocessing is performed on the historical data set to construct a sample set, and the sample set is divided into a training set and a validation set; the recurrent neural network model is iteratively trained based on the training set, and then the recurrent neural network model after each training is verified and evaluated based on the validation set until the recurrent neural network model converges or reaches a preset number of training rounds, thereby obtaining a trained fault prediction model.

[0135] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0136] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0137] Figure 8This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. Figure 8 As shown, the electronic device 3 of the embodiment includes: at least one processor 30 ( Figure 8 Only one is shown in the figure) a processor, a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor 30. When the processor 30 executes the computer program 32, the steps in the above-mentioned cloud database operation and maintenance management method embodiment are implemented.

[0138] The electronic device 3 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic device 3 may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art will understand that Figure 8 This is merely an example of the electronic device 3 and does not constitute a limitation on the electronic device 3 . The electronic device 3 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device 3 may also include input and output devices, network access devices, etc.

[0139] The processor 30 may be a central processing unit (CPU), or may be another general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor.

[0140] In some embodiments, the memory 31 may be an internal storage unit of the electronic device 3, such as the hard drive or memory of the electronic device 3. In other embodiments, the memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the electronic device 3. Furthermore, the memory 31 may include both an internal storage unit of the electronic device 2 and an external storage device. The memory 31 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of a computer program. The memory 31 may also be used to temporarily store data that has been output or is about to be output.

[0141] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned cloud database operation and maintenance management method embodiment can be implemented.

[0142] An embodiment of the present application provides a computer program product. When the computer program product is run on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned cloud database operation and maintenance management method embodiment when executing the computer program product.

[0143] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a camera / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.

[0144] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0145] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0146] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0147] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0148] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A cloud database operation and maintenance management method, characterized in that: include: Detect the corresponding instance nodes based on the warning information and the operating status data of each instance node in the cloud database; The warning information indicates a potential failure of the cloud database; When a target instance node with an abnormal operating state is detected, a target repair strategy whose target operating state data satisfies a state condition is determined from a first repair strategy library; wherein the target operating state data is the operating state data of the target instance node, the state condition is that the value of the indicator data in the target operating state data is within a corresponding abnormal indicator value range, and the first repair strategy library includes at least one abnormal indicator value range of the indicator data and a matching first repair strategy; When the target repair strategy does not exist in the first repair strategy library, determining a target fault type corresponding to the target operating status data; Determining a target repair strategy that matches the target fault type from a second repair strategy library; the second repair strategy library includes at least one fault type and a matching second repair strategy; When the target repair strategy does not exist in the second repair strategy library, determining the target repair strategy based on the target joint state data and a strategy recommendation model; the joint state data includes the target operating state data and the historical repair data of the target instance node, and the strategy recommendation model is configured to use the joint state data as input and output a repair strategy that can obtain the maximum reward value under the input; Repair the target instance node according to the determined target repair strategy.

2. A cloud database operation and maintenance management method according to claim 1, characterized in that: The step of determining the target fault type corresponding to the target operating status data includes: Determining the target fault type according to the target operating state data and a preset fault diagnosis rule; the fault diagnosis rule includes at least one operating state data and a matching fault type; and / or, The target fault type is determined based on the target operating status data and a fault diagnosis model; the fault diagnosis model is used to take the operating status data as input and output at least one fault type and a corresponding probability value, and the target fault type is the fault type corresponding to the maximum probability value.

3. A cloud database operation and maintenance management method according to claim 1, characterized in that: The training process of the strategy recommendation model specifically includes: Constructing a state space, action space, and reward function for an instance node; the state space represents the operating state data and historical repair data of the instance node, the action space represents the repair strategy, and the reward function represents the action value of the repair strategy; Building a deep reinforcement learning model based on the state space, the action space, and the reward function, and constructing an experience recycling pool; For the current running state data, select and execute the repair strategy with the maximum predicted action value according to the deep reinforcement learning model to obtain the running state data of the next state and the corresponding reward value; The current operating status data, the current repair strategy, the operating status data of the next state and the corresponding reward value are stored as sample data in the experience recovery pool, and the sample data is extracted from the experience recovery pool. The deep reinforcement learning model is trained with the goal of minimizing the difference between the predicted action value and the preset target action value to obtain a trained strategy recommendation model.

4. A cloud database operation and maintenance management method according to claim 1, characterized in that: After the step of repairing the target instance node according to the determined target repair strategy, the following step further includes: Obtain the current running status data of the repaired target instance node; Evaluate the current operating status data using a preset repair evaluation model, and if at least one indicator data in the current operating status data does not conform to a preset normal indicator value range, return to the step of determining the target repair strategy; If all indicator data in the current operating status data are in compliance with the preset normal indicator value range, the repair is completed.

5. A cloud database operation and maintenance management method according to claim 1, characterized in that: Before the step of repairing the target instance node according to the determined target repair strategy, the following step further comprises: Determine the risk assessment result of the target repair strategy according to a preset risk prediction model; Develop risk control strategies based on the risk assessment results.

6. A cloud database operation and maintenance management method according to claim 1, characterized in that: The cloud database operation and maintenance management method further includes: Obtaining real-time operating status data of each instance node in the cloud database during a preset time period; Input the real-time operating status data into the trained fault prediction model for prediction, and obtain the prediction result corresponding to each instance node; the prediction result includes the probability of potential fault occurrence, potential fault type and potential fault occurrence time; The warning information is generated according to the prediction result; the warning information includes identification information of the warning instance node, the potential fault type, the potential fault occurrence time and prevention strategy, and the warning instance node is an instance node whose probability of potential fault occurrence is greater than a preset probability threshold.

7. A cloud database operation and maintenance management method according to claim 1, characterized in that: The training process of the fault prediction model specifically includes: Acquire a historical data set; the historical data set includes historical operating status data and historical fault data of the instance node; Performing data preprocessing on the historical data set to construct a sample set, and dividing the sample set into a training set and a validation set; The recurrent neural network model is iteratively trained based on the training set, and then the recurrent neural network model after each training is verified and evaluated based on the verification set until the recurrent neural network model converges or reaches a preset number of training rounds, thereby obtaining a trained fault prediction model.

8. A cloud database operation and maintenance management system, characterized in that: include: The detection module is used to detect the corresponding instance nodes based on the warning information and the operating status data of each instance node in the cloud database; The warning information indicates a potential failure of the cloud database; A first repair strategy determination module is configured to, when a target instance node with an abnormal operating state is detected, determine from a first repair strategy library a target repair strategy whose target operating state data satisfies a state condition; wherein the target operating state data is operating state data of the target instance node, the state condition is that a value of indicator data in the target operating state data is within a corresponding abnormal indicator value range, and the first repair strategy library includes at least one abnormal indicator value range of indicator data and a matching first repair strategy; a fault type determination module, which determines a target fault type corresponding to the target operating status data when the target repair strategy does not exist in the first repair strategy library; A second repair strategy determination module is configured to determine a target repair strategy that matches the target fault type from a second repair strategy library; the second repair strategy library includes at least one fault type and a matching second repair strategy; a third repair strategy determination module, configured to, when the target repair strategy does not exist in the second repair strategy library, determine a target repair strategy based on target joint state data and a strategy recommendation model; the joint state data includes the target operating state data and historical repair data of the target instance node; the strategy recommendation model is configured to use the joint state data as input and output a repair strategy that can obtain a maximum reward value under the input; The repair module is used to repair the target instance node according to the determined target repair strategy.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the cloud database operation and maintenance management method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the cloud database operation and maintenance management method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Self-checking and self-repairing method and device of electronic equipment, equipment and storage medium

    CN121277743A

  • Self-checking and self-repairing method and device of electronic equipment, equipment and storage medium

    CN121277743B