Software automation operation and maintenance method and system based on AI technology

By applying AI technology in software operation and maintenance, automated analysis and diagnosis of software operation abnormalities, the problem of inefficiency of traditional operation and maintenance methods is solved, and high-accurate fault diagnosis and preventive maintenance are achieved.

CN119987832AInactive Publication Date: 2025-05-13BEIJING TONGYU HUAZHOU TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510088703.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional software operation and maintenance methods rely on manual experience and are difficult to quickly and accurately analyze massive data, resulting in low fault diagnosis efficiency and cannot meet the high requirements of modern software systems for stability and reliability.

Method used

Using the software automation operation and maintenance method based on AI technology, we use historical data, identify software functional modules, extract risk parameters, use decision trees and abnormal detection algorithms to troubleshoot faults, and establish a prediction model to evaluate the future operation performance of the software.

Benefits of technology

It realizes rapid positioning and understanding of software operation abnormalities, improves the accuracy and efficiency of fault diagnosis, can detect potential risks in advance and take preventive maintenance measures to ensure the stable operation of the software system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987832A_ABST
    Figure CN119987832A_ABST
Patent Text Reader

Abstract

The invention discloses a software automatic operation and maintenance method and system based on an AI technology, and relates to the technical field of software automatic operation and maintenance. The method comprises the steps of establishing a software operation exception database; performing operation monitoring on the function modules, identifying the modules and parameters, generating a flow chart, verifying execution time and parameter use conditions, and processing according to results; risk parameters and function problems are extracted, partition is carried out through a decision tree, key parameters are determined through t-SNE and a clustering algorithm, and fault types are diagnosed according to an anomaly detection algorithm and an SVM model; monitoring risk parameters, calculating related indexes to obtain a comprehensive evaluation value and a risk index, and establishing and training a fault prediction model to evaluate the future performance of the software; and performing maintenance or optimization measures to prevent potential faults and realize automatic operation and maintenance of the software. According to the invention, the whole process management from data collection to automatic maintenance is realized, the accuracy and efficiency of fault diagnosis are improved, and the stability of the system and the capability of coping with a complex environment are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of software automation operation and maintenance, and specifically relates to a software automation operation and maintenance method and system based on AI technology. Background Art

[0002] With the rapid development of information technology, software systems are becoming increasingly complex, with increasing functions and scale. In this context, software operation and maintenance faces many challenges.

[0003] Traditional software operation and maintenance methods mainly rely on manual experience, and operation and maintenance personnel need to spend a lot of time and energy to monitor the software operation status and troubleshoot faults. When faced with massive amounts of software operation data, manual analysis is inefficient and it is difficult to quickly and accurately discover potential problems. For example, when monitoring software performance indicators (such as CPU usage and memory usage), it is difficult for humans to grasp the data change trend in real time, and it is also difficult to find key information related to faults from complex logs.

[0004] Moreover, traditional methods have great limitations in fault diagnosis. For some complex faults, manual judgment is prone to misjudgment and omission, and it is impossible to accurately locate the root cause of the fault. This not only prolongs the time to repair the fault, but may also have a serious impact on the business. At the same time, traditional operation and maintenance methods often handle faults only after they occur, lack the ability to predict and prevent potential risks, and cannot meet the high requirements of modern software systems for stability and reliability.

[0005] With the rise of artificial intelligence technology, new ideas have been provided for solving software operation and maintenance problems. Software automated operation and maintenance methods based on AI technology have emerged, aiming to utilize AI's powerful data analysis and processing capabilities to improve operation and maintenance efficiency, enhance fault diagnosis accuracy, and achieve preventive maintenance to cope with increasingly complex software operation and maintenance needs. Summary of the invention

[0006] In order to overcome the shortcomings and deficiencies of the above-mentioned prior art, the first purpose of the present invention is to provide a software automation operation and maintenance method based on AI technology; the second purpose of the present invention is to provide a software automation operation and maintenance system based on AI technology.

[0007] The first object of the present invention adopts the following technical solution:

[0008] The software automation operation and maintenance method based on AI technology has the following process:

[0009] Step 1: Collect historical data, including logs and performance indicators during normal operation, and known failure modes and corresponding log modes and performance indicator changes, and set association rules between various parameters;

[0010] Step 2: Identify software function modules and their operation parameters, generate a software operation flow chart, verify the execution status of the function modules and the use of parameters, and output an alarm signal or an alarm signal corresponding to the abnormal duration as the actual execution status based on the comparison result between the module execution interval and the preset time interval;

[0011] Step 3: Extract risk parameters and functional problems from the actual execution situation, use decision trees to implement risk partitioning, analyze software operation abnormality feature data to determine key parameters, and determine the fault type by comparing the anomaly detection algorithm with the software operation abnormality database;

[0012] Step 4: Monitor the risk parameters in the risk area. When the risk parameters exceed the threshold, generate a trend curve and calculate the fault probability change index, risk parameter sensitivity index coefficient and fault impact severity index, and then calculate the comprehensive evaluation value. Based on the comprehensive evaluation value and the fault frequency, the risk index is obtained, and a prediction model is established to evaluate the performance, including establishing a software operation fault prediction model through a classification algorithm, performing principal component analysis on the initial sample set, generating a new software fault feature data sample set, and training an integrated software operation fault prediction model to predict the future software operation performance;

[0013] Step 5: Based on the evaluation results of the prediction model, formulate and implement corresponding maintenance or optimization measures.

[0014] Preferably, in step one, the performance indicators collected include CPU usage and memory usage.

[0015] Preferably, in step 2, the specific method of verifying the corresponding execution time and parameter usage of each software function module under the execution state is:

[0016] Obtain the execution time and parameter usage of each software function module, determine the time difference between the current execution time and the previous execution time, calculate the module execution interval, and compare the module execution interval with the preset time interval;

[0017] If the module time interval exceeds the preset time interval, the parameter usage in the current module execution interval is obtained to determine whether the current parameter usage meets expectations. If so, an alarm message related to the current module execution interval is generated. If not, the current parameter usage is recorded and an alarm signal about the parameter usage is generated.

[0018] If the module time interval is less than the preset time interval, the start time and end time of the current software function module are verified, the corresponding start time and end time are set as the abnormal duration, and an alarm signal corresponding to the abnormal duration is generated, and the verified alarm signal is output as the actual execution status.

[0019] Preferably, in step three, the method of extracting risk parameters and functional problems from the actual execution situation is: mapping the obtained alarm signal to the functional problem, when one alarm signal is associated with multiple functional problems, associating the functional problem with the alarm signal according to the different scenarios corresponding to the functional problem, obtaining the parameters of the associated functional problem and the operating parameters corresponding to the alarm signal, comparing the parameters of the functional problem with the operating parameters corresponding to the alarm signal, and obtaining the risk parameters existing in the operating parameters.

[0020] Preferably, in step 3, the process of implementing risk partitioning using a decision tree is:

[0021] The fault cause, component feature and fault feature are combined to obtain the combined feature. The child nodes of the decision tree represent the combined feature, the root node represents all functional problems to be evaluated, and the leaf nodes represent the output risk partitioning results. The Gini impurity of the fault feature of each child node of the decision tree is calculated in turn, and the decision tree is divided according to the Gini impurity of the fault feature, and sorted according to the eigenvalue size of the component feature. The leaf nodes of the decision tree that have been finally sorted and divided are used as the current risk partitioning results to obtain the corresponding risk areas.

[0022] Preferably, in step 4, the calculation formula of the risk parameter sensitivity index coefficient is: in, represents the risk parameter sensitivity index coefficient; n represents the number of risk parameters; represents the expected value variance corresponding to the i-th risk parameter; It represents the total variance corresponding to the risk parameter. The calculation formula for the expected value variance corresponding to each risk parameter is: Among them, X j represents the i-th risk parameter; P(X j ) represents the probability of failure; Y[X=X j ] is expressed as P(X j ) conditional expected value; the fault impact severity index is expressed as the coefficient value corresponding to the maximum value of the current fault occurrence probability selected from the software operation anomaly database; according to the fault probability change index, the risk parameter sensitivity index coefficient and the fault impact severity index, a comprehensive evaluation value is calculated, and the comprehensive evaluation value is expressed as the weighted sum of the fault probability change index, the risk parameter sensitivity index coefficient and the fault impact severity index; the fault frequency corresponding to the comprehensive evaluation value is obtained, and based on the comprehensive evaluation value and the fault frequency, the formula is used Get the risk index and output the risk index as the risk situation of each risk area, where R is the risk index, represents the comprehensive evaluation value, f represents the fault frequency, ω1 represents the weight of the comprehensive evaluation value, and ω2 represents the weight of the fault frequency.

[0023] Preferably, step four includes: inputting an initial sample set of software fault feature data under different operating conditions, classifying all software fault feature data in the initial sample set to obtain several sample subsets with different key features, resampling each sample subset using stratified sampling, performing principal component analysis to obtain principal component coefficient vectors, repeating the execution to obtain several groups of principal component coefficient vectors, and merging the principal component coefficient vectors into a principal component coefficient matrix, rearranging the coefficient matrix and transforming the initial sample set based on it, generating a new sample set of software fault feature data, training a software operation fault prediction model based on the new sample set of software fault feature data to obtain a sub-model of the software operation fault prediction model, looping multiple times to obtain an integrated software operation fault prediction model, acquiring real-time software fault feature data during software operation, and inputting the real-time software fault feature data into the integrated software operation fault prediction model, analyzing and predicting the real-time software fault feature data through the software operation fault prediction model, and obtaining an evaluation of the future operation performance of the software.

[0024] Preferably, the maintenance or optimization measures include automatically restarting services and adjusting resource configurations.

[0025] The second object of the present invention adopts the following technical solution:

[0026] The software automation operation and maintenance system based on AI technology is used to implement the software automation operation and maintenance method based on AI technology. The system includes:

[0027] Software operation abnormality database module: responsible for collecting and storing historical data, including logs, performance indicators, known failure modes and their corresponding log modes and performance indicator changes during normal operation;

[0028] Functional module operation monitoring module: used to identify the functional modules of the software and the corresponding operating parameters of each module, and generate a software operation flow chart;

[0029] Risk and Fault Analysis Module: It is used to extract risk parameters and functional problems from actual execution situations, determine key parameters using decision trees and clustering algorithms, and diagnose fault types through anomaly detection algorithms;

[0030] Risk assessment module: used to monitor risk parameters in risk areas, calculate comprehensive assessment values ​​and risk indicators, establish prediction models and evaluate their performance;

[0031] Maintenance measures execution module: formulates and executes corresponding maintenance or optimization measures according to the evaluation results of the prediction model.

[0032] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0033] 1. The present invention can quickly locate and understand abnormal situations in software operation by establishing a software operation abnormality database, collecting and analyzing a large amount of historical data, and setting parameter association rules. In the functional module operation monitoring stage, the functional module and its operating parameters are automatically identified, an operation flow chart is generated, and the execution time and parameter usage are verified in real time. Once an abnormality occurs, an alarm can be issued and processed in time, which reduces the time and energy of manual troubleshooting and improves the efficiency of software operation and maintenance.

[0034] 2. The present invention uses advanced AI technologies, such as t-SNE method, clustering algorithm, anomaly detection algorithm and SVM model, to conduct in-depth analysis of software operation abnormality feature data. These technologies determine key parameters, diagnose and accurately determine the fault type. Compared with traditional methods, the root cause of software faults can be more accurately identified, misjudgment and missed judgment can be avoided, and the accuracy of fault diagnosis can be improved.

[0035] 3. The present invention evaluates and predicts the future performance of the software based on risk assessment and prediction models. By monitoring risk parameters and calculating relevant indexes to obtain risk indicators, potential risk areas and potential failure hazards can be discovered in advance. Preventive maintenance measures can be formulated and implemented based on the prediction results, such as automatically restarting services and adjusting resource configurations. Effective measures can be taken before a failure occurs to avoid the impact of software failures on the business, ensure the stable operation of the software system, and reduce operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0037] Figure 1 A flowchart of the software automation operation and maintenance method based on AI technology of the present invention is shown;

[0038] Figure 2 The module diagram of the software automation operation and maintenance system based on AI technology of the present invention is shown;

[0039] Figure 3 A flow chart of the functional module operation monitoring of the present invention is shown. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] In addition, the described features, structures or characteristics may be combined in one or more example embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the example embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, steps, etc. may be adopted. In other cases, well-known structures, methods, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0042] Embodiment 1:

[0043] See also Figure 1 As shown, the software automation operation and maintenance method based on AI technology in this embodiment has the following process:

[0044] Step 1: Establish a software operation abnormality database.

[0045] Establish a software operation anomaly database and collect historical data, including logs during normal operation, performance indicators (such as CPU usage, memory usage), known failure modes and their corresponding log patterns, and performance indicator changes. Set association rules between various parameters in the software operation anomaly database, such as the relationship between a specific log pattern and a specific type of failure.

[0046] Step 2: Functional module operation monitoring.

[0047] See also Figure 3 As shown in the figure, the function module operation monitoring is as follows:

[0048] S21. Identify the functional modules of the software and the corresponding operating parameters of each functional module, such as the response time and resource consumption of the module.

[0049] S22. Generate a software operation flow chart based on the identified software function modules, and determine the execution state (such as running, paused or error reporting state) and operation parameters corresponding to each function module in the operation flow chart.

[0050] S23, verify the execution time and parameter usage corresponding to each software function module in the execution state. Obtain the execution time and parameter usage of each software function module, determine the time difference between the current execution time and the last execution time, and calculate the module execution interval.

[0051] S24, comparing the module execution interval with the preset time interval:

[0052] If the module time interval exceeds the preset time interval, the parameter usage in the current module execution interval is obtained to determine whether the current parameter usage meets expectations. If it does, an alarm message related to the current module execution interval is generated; if it does not, the current parameter usage is recorded and an alarm signal about parameter usage is generated.

[0053] If the module time interval is less than the preset time interval, the start time and end time of the current software function module are verified, the corresponding start time and end time are set as the abnormal duration, and an alarm signal corresponding to the abnormal duration is generated, and the verified alarm signal is output as the actual execution status.

[0054] Step 3: Risk and failure analysis.

[0055] S31. Extract risk parameters and functional problems of the current software functional module from the actual execution situation (including the above-mentioned alarm signals, etc.). Map the acquired alarm signals to the functional problems. When one alarm signal is associated with multiple functional problems, the functional problems are associated with the alarm signals according to the different scenarios corresponding to the functional problems. Obtain the parameters of the associated functional problems and the operating parameters corresponding to the alarm signals, compare the parameters of the functional problems with the operating parameters corresponding to the alarm signals, and obtain the risk parameters existing in the operating parameters.

[0056] S32. Determine the risk parameters existing under each functional problem, identify the fault causes corresponding to the current risk parameters and functional problems, the fault characteristics corresponding to the risk parameters and functional problems, and the component characteristics of the software components corresponding to the location information. Input the identified fault causes, fault characteristics, and component characteristics into the decision tree, and use the decision tree to implement risk partitioning. Combine the fault causes, component characteristics, and fault characteristics to obtain combined characteristics. The child nodes of the decision tree represent the combined characteristics, the root node represents all functional problems to be evaluated, and the leaf nodes represent the output risk partitioning results. Calculate the Gini impurity of the fault characteristics of each child node of the decision tree in turn, divide the decision tree according to the Gini impurity of the fault characteristics, and sort them according to the eigenvalue size of the component characteristics. The leaf nodes of the decision tree that have been finally sorted and divided are used as the current risk partitioning results to obtain the corresponding risk areas.

[0057] S33. Analyze the software operation abnormality feature data to determine the key parameters: Use the t-SNE method to map the software operation abnormality feature data to the feature space, and apply the dimensionality reduction technology to construct the software abnormality feature point set. Use the clustering algorithm to analyze the feature point set, take several cluster centers as the initial solution, set several particle numbers and the optimal number of iterations, and randomly generate several initial solutions. According to the current cluster center position, calculate the fitness value of each particle through the fitness function, and update the current fitness value to the individual optimal value and the current position to the individual optimal position. Find the global optimal value and the global optimal position through the individual optimal values ​​of all particles. Determine whether the set optimal number of iterations is reached. If it is reached, the iteration is completed; if not, the particle speed and position are continued to be updated. According to the updated particle position, use the minimum distance principle to assign each software fault sample in the software fault feature point set to the corresponding several cluster centers. Recalculate the fitness value of each particle. If the fitness value of the particle is less than its individual optimal fitness value, update the individual optimal fitness value of the particle to the current fitness value, and update its individual optimal position. Compare the individual best fitness values ​​of all particles, find the minimum value as the global extreme value, update the global extreme value position, continue to iterate, and finally determine the key parameters that affect the abnormal operation of the software.

[0058] S34. Diagnose and determine the fault type: diagnose the key parameters through the anomaly detection algorithm, and compare them with the parameters in the software operation anomaly database to determine whether they conform to the known fault mode. Obtain the key parameter data set that affects the software operation failure, use the key parameter data set as the training sample of SVM, and obtain the optimal parameters of SVM parameters through the optimization algorithm. Use the obtained SVM optimal parameters to train the training set and establish an anomaly detection model for software fault diagnosis. Input the key parameters that affect the software operation failure into the anomaly detection model, and perform real-time fault diagnosis. Compare the fault results detected by the anomaly detection model with the known fault modes in the software operation anomaly database, and confirm whether they conform. According to the comparison results, determine the specific fault type.

[0059] Step 4: Risk assessment.

[0060] S41. Monitor the risk parameters in the risk area. When the risk parameters exceed the threshold, generate a trend curve related to the risk parameters that exceed the threshold. According to the obtained trend curve, calculate the fault probability change index, risk parameter sensitivity index coefficient and fault impact severity index corresponding to the trend curve in turn. The fault probability change index is expressed as the variance of the change in the probability of fault occurrence after the risk parameter changes; the risk parameter sensitivity index coefficient is expressed as obtaining the expected value variance corresponding to each risk parameter and the total variance corresponding to the risk parameter, and obtaining the risk parameter sensitivity index coefficient corresponding to the risk parameter. The calculation formula is as follows:

[0061]

[0062] in, represents the risk parameter sensitivity index coefficient; n represents the number of risk parameters; represents the expected value variance corresponding to the i-th risk parameter; It represents the total variance corresponding to the risk parameter. The calculation formula for the expected value variance corresponding to each risk parameter is:

[0063] Among them, X j represents the i-th risk parameter; P(X j ) represents the probability of failure; Y[X=X j ] is expressed as P(X j ) conditional expected value; the fault impact severity index is expressed as the coefficient value corresponding to the maximum value of the current fault occurrence probability selected from the software operation anomaly database. According to the fault probability change index, the risk parameter sensitivity index coefficient and the fault impact severity index, the comprehensive evaluation value is calculated. The comprehensive evaluation value is expressed as the weighted sum of the fault probability change index, the risk parameter sensitivity index coefficient and the fault impact severity index. The fault frequency corresponding to the comprehensive evaluation value is obtained. Based on the comprehensive evaluation value and the fault frequency, the formula Get the risk index and output the risk index as the risk situation of each risk area, where R is the risk index, represents the comprehensive evaluation value, f represents the fault frequency, ω1 represents the weight of the comprehensive evaluation value, and ω2 represents the weight of the fault frequency.

[0064] S42, establish a prediction model and evaluate performance: according to the determined fault type, establish a software operation fault prediction model through a classification algorithm. Input the initial sample set of software fault feature data under different operating conditions, classify all software fault feature data in the initial sample set, and obtain several sample subsets with different key features. Perform loop processing on each sample subset, resample the sample subset of software fault feature data using stratified sampling method, perform principal component analysis on the new sample subset, obtain the principal component coefficient vector related to the software fault type, repeat the execution to obtain several groups of principal component coefficient vectors, and merge the principal component coefficient vectors into a principal component coefficient matrix. Rearrange the coefficient matrix, and transform the initial sample set based on the coefficient matrix to generate a new sample set of software fault feature data. Train the software operation fault prediction model based on the new sample set of software fault feature data to obtain a sub-model of the software operation fault prediction model, and repeat multiple times to obtain an integrated software operation fault prediction model. Real-time software fault feature data during software operation is obtained, and the real-time software fault feature data is input into an integrated software operation fault prediction model. The real-time software fault feature data is analyzed and predicted through the software operation fault prediction model, and an evaluation of the future software operation performance is obtained.

[0065] Step 5: Execute maintenance measures.

[0066] According to the evaluation results of the prediction model, formulate and implement corresponding maintenance or optimization measures. Take preventive maintenance measures, including automatically restarting services, adjusting resource configuration, etc., to avoid potential failures.

[0067] The beneficial effects of this embodiment are: by establishing a comprehensive abnormality database, real-time monitoring of the operating status of functional modules, in-depth analysis of risks and faults, accurate assessment of risks and prediction of faults, intelligent management of software systems is achieved. Advanced algorithms (such as t-SNE, SVM, etc.) are used to improve the accuracy and efficiency of fault diagnosis, dynamically adjust maintenance strategies, and effectively prevent potential problems. In addition, the system supports automated execution of maintenance measures, which significantly improves the stability and operation efficiency of the system, reduces manual intervention, reduces operating costs, and enhances the ability to cope with complex environments.

[0068] Embodiment 2:

[0069] See also Figure 2 As shown, the software automation operation and maintenance system based on AI technology in this embodiment includes a software operation anomaly database module, a functional module operation monitoring module, a risk and fault analysis module, a risk assessment module and a maintenance measure execution module.

[0070] Software operation anomaly database module: responsible for collecting and storing historical data, including logs during normal operation, performance indicators (such as CPU usage, memory usage), known failure modes and their corresponding log modes, and changes in performance indicators.

[0071] Functional module operation monitoring module: used to identify the functional modules of the software and the corresponding operating parameters of each module, and generate a software operation flow chart.

[0072] Risk and Fault Analysis Module: It is used to extract risk parameters and functional problems from actual execution situations, determine key parameters using decision trees and clustering algorithms, and diagnose fault types through anomaly detection algorithms.

[0073] Risk assessment module: used to monitor risk parameters within the risk area, calculate comprehensive assessment values ​​and risk indicators, establish prediction models and evaluate their performance.

[0074] Maintenance measures execution module: formulates and executes corresponding maintenance or optimization measures according to the evaluation results of the prediction model.

[0075] The beneficial effects of this embodiment are as follows: by integrating five modules, namely, software operation abnormality database, function module monitoring, risk and fault analysis, risk assessment and maintenance measures execution, the whole process management from data collection to automatic maintenance is realized. The system uses advanced algorithms and models to improve the accuracy and efficiency of fault diagnosis, monitors and dynamically assesses risks in real time, and thus effectively prevents potential faults. In addition, it also supports the automated execution of maintenance measures, which significantly improves the stability and operation efficiency of the system, reduces the need for manual intervention, and reduces operating costs.

[0076] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

[0077] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. The software automation operation and maintenance method based on AI technology is characterized by: The method flow is as follows: Step 1: Collect historical data and set association rules between parameters; Step 2: Identify software function modules and their operation parameters, generate a software operation flow chart, verify the execution status of the function modules and the use of parameters, and output an alarm signal or an alarm signal corresponding to the abnormal duration as the actual execution status based on the comparison result between the module execution interval and the preset time interval; Step 3: Extract risk parameters and functional problems from the actual execution situation, use decision trees to implement risk partitioning, analyze software operation abnormality feature data to determine key parameters, and determine the fault type by comparing the anomaly detection algorithm with the software operation abnormality database; Step 4: Monitor the risk parameters in the risk area. When the risk parameters exceed the threshold, generate a trend curve and calculate the comprehensive evaluation value. Based on the comprehensive evaluation value and the failure frequency, obtain the risk index, and establish a prediction model to evaluate the performance, so as to predict the future performance of the software. Step 5: Based on the evaluation results of the prediction model, formulate and implement corresponding maintenance or optimization measures.

2. The software automation operation and maintenance method based on AI technology according to claim 1 is characterized in that: In step 1, the historical data includes logs, performance indicators during normal operation, and known failure modes and corresponding log modes and performance indicator changes. The collected performance indicators include CPU usage and memory usage.

3. The software automation operation and maintenance method based on AI technology according to claim 1 is characterized in that: In step 2, the specific method of verifying the execution time and parameter usage corresponding to the execution state of each software function module is as follows: Obtain the execution time and parameter usage of each software function module, determine the time difference between the current execution time and the previous execution time, calculate the module execution interval, and compare the module execution interval with the preset time interval; If the module time interval exceeds the preset time interval, the parameter usage in the current module execution interval is obtained to determine whether the current parameter usage meets expectations. If so, an alarm message related to the current module execution interval is generated. If not, the current parameter usage is recorded and an alarm signal about the parameter usage is generated. If the module time interval is less than the preset time interval, the start time and end time of the current software function module are verified, the corresponding start time and end time are set as the abnormal duration, and an alarm signal corresponding to the abnormal duration is generated, and the verified alarm signal is output as the actual execution status.

4. The software automation operation and maintenance method based on AI technology according to claim 1 is characterized in that: In the step three, the method of extracting risk parameters and functional problems from the actual execution situation is: mapping the obtained alarm signal to the functional problem, when one alarm signal is associated with multiple functional problems, the functional problem is associated with the alarm signal according to the different scenarios corresponding to the functional problem, obtaining the parameters of the associated functional problem and the operating parameters corresponding to the alarm signal, comparing the parameters of the functional problem with the operating parameters corresponding to the alarm signal, and obtaining the risk parameters existing in the operating parameters.

5. The software automation operation and maintenance method based on AI technology according to claim 1 is characterized in that: In step 3, the process of using decision tree to realize risk partition is as follows: The fault cause, component feature and fault feature are combined to obtain the combined feature. The child nodes of the decision tree represent the combined feature, the root node represents all functional problems to be evaluated, and the leaf nodes represent the output risk partitioning results. The Gini impurity of the fault feature of each child node of the decision tree is calculated in turn, and the decision tree is divided according to the Gini impurity of the fault feature, and sorted according to the eigenvalue size of the component feature. The leaf nodes of the decision tree that have been finally sorted and divided are used as the current risk partitioning results to obtain the corresponding risk areas.

6. The software automation operation and maintenance method based on AI technology according to claim 1, characterized in that: The step 4 also includes: the calculation formula of the risk parameter sensitivity index coefficient is: in, represents the risk parameter sensitivity index coefficient; n represents the number of risk parameters; represents the expected value variance corresponding to the i-th risk parameter; It represents the total variance corresponding to the risk parameter. The calculation formula for the expected value variance corresponding to each risk parameter is: Among them, X j represents the i-th risk parameter; P(X j ) represents the probability of failure; Y[X=X j ] is expressed as P(X j ) conditional expected value; the fault impact severity index is expressed as the coefficient value corresponding to the maximum value of the current fault occurrence probability selected from the software operation anomaly database; according to the fault probability change index, the risk parameter sensitivity index coefficient and the fault impact severity index, a comprehensive evaluation value is calculated, and the comprehensive evaluation value is expressed as the weighted sum of the fault probability change index, the risk parameter sensitivity index coefficient and the fault impact severity index; the fault frequency corresponding to the comprehensive evaluation value is obtained, and based on the comprehensive evaluation value and the fault frequency, the formula is used Get the risk index and output the risk index as the risk situation of each risk area, where R is the risk index, represents the comprehensive evaluation value, f represents the fault frequency, ω1 represents the weight of the comprehensive evaluation value, and ω2 represents the weight of the fault frequency.

7. The software automation operation and maintenance method based on AI technology according to claim 1 is characterized in that: The step four includes: inputting an initial sample set of software fault feature data under different operating conditions, classifying all software fault feature data in the initial sample set to obtain several sample subsets with different key features, resampling each sample subset using a stratified sampling method, performing principal component analysis to obtain a principal component coefficient vector, repeating the execution to obtain several groups of principal component coefficient vectors, and merging the principal component coefficient vectors into a principal component coefficient matrix, rearranging the coefficient matrix and transforming the initial sample set based on it, generating a new sample set of software fault feature data, training a software operation fault prediction model based on the new sample set of software fault feature data to obtain a sub-model of the software operation fault prediction model, repeating the cycle multiple times to obtain an integrated software operation fault prediction model, acquiring real-time software fault feature data during software operation, and inputting the real-time software fault feature data into the integrated software operation fault prediction model, analyzing and predicting the real-time software fault feature data through the software operation fault prediction model, and obtaining an evaluation of the future operation performance of the software.

8. The software automation operation and maintenance method based on AI technology according to claim 1 is characterized in that: The maintenance or optimization measures include automatically restarting services and adjusting resource configurations.

9. A software automation operation and maintenance system based on AI technology, used to implement the software automation operation and maintenance method based on AI technology as claimed in claim 1, characterized in that: The system comprises: Software operation abnormality database module: responsible for collecting and storing historical data, including logs, performance indicators, known failure modes and their corresponding log modes and performance indicator changes during normal operation; Functional module operation monitoring module: used to identify the functional modules of the software and the corresponding operating parameters of each module, and generate a software operation flow chart; Risk and Fault Analysis Module: used to extract risk parameters and functional problems from actual execution, determine key parameters using decision trees and clustering algorithms, and diagnose fault types through anomaly detection algorithms; Risk assessment module: used to monitor risk parameters in risk areas, calculate comprehensive assessment values ​​and risk indicators, establish prediction models and evaluate their performance; Maintenance measures execution module: formulates and executes corresponding maintenance or optimization measures according to the evaluation results of the prediction model.

Citation Information

Patent Citations

  • System operation and maintenance decision support system based on artificial intelligence

    CN118278778A

  • Simulator-based automobile operation fault identification method and system

    CN118656583A

  • Chemical production informatization management platform

    CN118761613A

  • Operation and maintenance system and method

    US20210271582A1

Cited By

  • Intelligent fault detection system and method for management machine

    CN120873921A

  • Variable pitch system fault diagnosis method, system and equipment for predictive maintenance and medium

    CN121452127A