Software operation and maintenance method and device, electronic equipment and storage medium
By collecting and filtering software operation status logs and using artificial intelligence models to generate operation and maintenance operation commands, the problem of low accuracy of artificial intelligence models in software operation and maintenance is solved, and the automation and efficiency of operation and maintenance operations are achieved.
Patent Information
- Application Number
- CN202510821459.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, artificial intelligence models lack effective screening and precise utilization of operation status logs in software operation and maintenance, resulting in poor accuracy of operation and maintenance operation commands, and the inability to dynamically adjust the importance of logs, affecting operation and maintenance efficiency and the accuracy of abnormal detection.
By collecting the software's operating status logs, filtering key logs with greater importance than preset values, using artificial intelligence models to generate operation and maintenance operation commands, adjusting the importance of logs in combination with the time attenuation model and semantic similarity, dynamically adjusting the importance of logs, and optimizing resource allocation.
The accuracy of the output operation and maintenance operation commands of the artificial intelligence model is improved, the automation of operation and maintenance operations is realized, manual intervention is reduced, and the operation and maintenance efficiency and system applicability and reliability are improved.
Smart Images

Figure CN120335850A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a software operation and maintenance method, apparatus, electronic device, and storage medium. Background Art
[0002] In the context of the continuous deepening of intelligent operation and maintenance technologies, the automated operation and maintenance system based on a knowledge base has achieved remarkable results. However, in related technologies, when an artificial intelligence model outputs an operation and maintenance operation command, due to the lack of effective screening and precise utilization of the running state logs, its accuracy is poor, making it difficult to meet the requirements of maintaining complex software systems.
[0003] Specifically, first, related technologies fail to dynamically adjust the importance level of logs and cannot perform differential processing based on the timeliness and log level of the logs, resulting in critical logs and ordinary logs being treated equally, reducing the operation and maintenance efficiency and the accuracy of anomaly detection. Second, related technologies do not fully consider the semantic association between logs and historical anomalies and the current load situation of the system, cannot quickly identify known anomalies based on semantic similarity, and cannot reasonably allocate resources to preferentially handle urgent anomalies under high load. These problems limit the applicability and reliability of artificial intelligence models in maintaining complex software systems.
[0004] Therefore, how to improve the accuracy of the operation and maintenance operation commands output by an artificial intelligence model is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] This application provides a software operation and maintenance method, apparatus, electronic device, and storage medium, which improves the accuracy of the operation and maintenance operation commands output by an artificial intelligence model.
[0006] This application provides a software operation and maintenance method, including: collecting the running state logs of software, and when it is detected that the software has an anomaly, collecting the running state logs of the software within a preset time period; screening critical logs with an importance level greater than a preset value from the running state logs within the preset time period; wherein, the importance level of the running state logs is positively correlated with the log level of the running state logs and negatively correlated with the generation time of the running state logs; inputting the critical logs into an artificial intelligence model to generate an operation and maintenance operation command corresponding to the critical logs by using the artificial intelligence model.
[0007] The present application also provides a software operation and maintenance device, including: a collection module, configured to collect the operation status logs of the software, and when it is detected that the software has an exception, collect the operation status logs of the software within a preset time period; a screening module, configured to screen out key logs with an importance level greater than a preset value from the operation status logs within the preset time period; wherein, the importance level of the operation status logs is positively correlated with the log level of the operation status logs and negatively correlated with the generation time of the operation status logs; a generation module, configured to input the key logs into an artificial intelligence model to generate operation and maintenance operation commands corresponding to the key logs by using the artificial intelligence model.
[0008] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above software operation and maintenance methods when executing the computer program.
[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program implements the steps of any of the above software operation and maintenance methods when executed by a processor.
[0010] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above software operation and maintenance methods when executed by a processor.
[0011] The software operation and maintenance method provided by the present application collects the operation status logs of the software, and when it is detected that the software has an exception, collects the operation status logs of the software within a preset time period. Further, key logs with an importance level greater than a preset value are screened out from the operation status logs within the preset time period, wherein the importance level of the operation status logs is positively correlated with the log level and negatively correlated with the generation time. This screening mechanism can accurately extract log information that is more valuable for exception analysis. Inputting the key logs into the artificial intelligence model enables the artificial intelligence model to generate corresponding operation and maintenance operation commands based on the operation status logs with a shorter generation time and a higher log level, thereby significantly improving the accuracy of the operation and maintenance operation commands output by the model. This method not only realizes the automation of operation and maintenance operations, reduces the dependence on manual intervention, and improves the operation and maintenance efficiency, but also further enhances the applicability and reliability of the artificial intelligence model in the operation and maintenance of complex software systems by accurately screening log information.
[0012] Furthermore, the present application dynamically adjusts the importance level of logs through a time decay model. As time goes by, the importance of logs gradually decreases, which conforms to the timeliness characteristics of log information. Newer logs can better reflect the current operating state of the system, while older logs may no longer be relevant. Through this dynamic adjustment, the system can more flexibly process logs at different times, avoid over - focusing on outdated logs, and thus improve the overall performance and resource utilization efficiency of the system.
[0013] Moreover, the present application designs a decay coefficient adjustment formula based on the semantic similarity between the running state logs and the historical exception logs and the current load. By considering the semantic similarity between the logs and the historical exception logs, it can accurately identify whether the current log indicates a known problem. If the semantic similarity between the current log and the historical exception log is high, it means that the log may be related to a known abnormal situation. At this time, the decay coefficient is reduced, so that the importance of the log can be retained for a longer time. This mechanism ensures that logs related to known exceptions will not be ignored due to the passage of time, thus improving the response speed and processing efficiency for known problems. Secondly, the introduced current load factor enables the evaluation of log importance to be adjusted according to the real - time operating state of the system. When the system load is high, resources are tense, and any exception may have a greater impact on the system performance. At this time, the decay coefficient is increased to accelerate the decay speed of log importance, so that the system can give priority to handling current urgent problems and avoid low - priority logs occupying too many resources. On the contrary, when the system load is low, the decay coefficient will be correspondingly reduced, and the system can retain more historical data for in - depth analysis, providing richer information for subsequent operation and maintenance decisions.
[0014] The present application also discloses a software operation and maintenance device, an electronic device, a computer - readable storage medium, and a computer program product, which can also achieve the above - mentioned technical effects.
[0015] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0017] Figure 1 It is a flowchart of a software operation and maintenance method shown according to an exemplary embodiment.
[0018] Figure 2A flowchart for training an LSTM model shown according to an exemplary embodiment.
[0019] Figure 3 A flowchart for another software operation and maintenance method shown according to an exemplary embodiment.
[0020] Figure 4 A flowchart for an automated operation and maintenance shown according to an exemplary embodiment.
[0021] Figure 5 A structural diagram of a software operation and maintenance device shown according to an exemplary embodiment.
[0022] Figure 6 A structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0024] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0025] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0026] Embodiments of the present application provide a software operation and maintenance method, and the method will be described in detail in combination with the execution process of the software operation and maintenance method. Refer to Figure 1 A flowchart for a software operation and maintenance method shown according to an exemplary embodiment.
[0027] S101: Collect the running state logs of the software. When it is detected that the software has an anomaly, collect the running state logs of the software within a preset time period.
[0028] During the software operation and maintenance process, it is first necessary to monitor the running status of the software in real time. The running status logs of the software can be collected in real time through a log collection tool. The running status logs refer to the log files generated during the running process of the software. These log files record the running status of the software and can include information such as timestamps, log levels, and log contents. The log contents can include CPU (Central Processing Unit) usage, memory occupancy, response time, etc.
[0029] In this step, when it is detected that the software has an anomaly, such as the software process crashing, freezing, running out of resources, etc., collect the running status logs of the software within a preset time period. These logs record various running conditions of the software during this time period and provide basic data support for subsequent anomaly analysis and generation of operation and maintenance operation commands.
[0030] S102: Screen out the critical logs with an importance level greater than a preset value from the running status logs within a preset time period; among them, the importance level of the running status logs is positively correlated with the log level of the running status logs and negatively correlated with the generation time of the running status logs.
[0031] In this step, screen out the critical logs with an importance level greater than a preset value from the running status logs within a preset time period. The importance level of the logs is positively correlated with their log levels, that is, the higher the log level, the higher its importance level. For example, the importance level of the running status log with a log level of ERROR is higher than that of the running status log with a log level of WARNING, and the importance level of the running status log with a log level of WARNING is higher than that of the running status log with a log level of INFO. The importance level of the running status logs is not only related to the log level but also related to the generation time of the logs. Specifically, the importance level of the logs decreases over time, that is, the earlier the generation time of the log, the lower its importance level. This design takes into account the timeliness of the log information because the newer logs can better reflect the current running status of the software. For example, if an error log was generated a few minutes ago, it can indicate the current existing problem better than a similar error log generated a few hours ago. In this way, the system can pay more attention to the most recent log information that may have a greater impact on the current operation and maintenance decisions.
[0032] As a feasible implementation method, after collecting the running status logs of the software, it further includes: setting the importance level of the running status logs based on the log levels of the running status logs; when the software is running normally, reducing the importance level of the running status logs based on the time decay model; when it is detected that the software has an anomaly, stop reducing the importance level of the running status logs.
[0033] In a specific implementation, the collected operation status logs are assigned an initial importance level according to their log levels. For example, logs at the error level may be assigned the highest importance level, such as 1.0, logs at the warning level may be assigned a lower importance level, such as 0.7, and logs at the information level may be assigned the lowest importance level, such as 0.3. Subsequently, when the software is running normally, the system gradually reduces the importance levels of these logs based on the time decay model to reflect the timeliness of the log information. However, once an anomaly in the software is detected, the system stops reducing the importance levels of the relevant logs to ensure that these logs that may be related to the anomaly can be given priority consideration and analysis. This approach helps to focus resources and attention on the most relevant logs, thereby improving the efficiency of fault diagnosis and repair.
[0034] As a feasible implementation, the time decay model is: ; where is the importance level set according to the log level of the operation status log, e is the natural constant, is the current decay coefficient, used to control the decay speed. The larger the current decay coefficient , the faster the decay speed. Conversely, the larger the current decay coefficient , the slower the decay speed. is the time difference between the current time and the generation time of the operation status log, , tcurrent is the current time, terror is the generation time of the operation status log, is the importance level of the operation status log at the current time t.
[0035] It can be seen that the importance level of the log is dynamically adjusted through the time decay model. As time goes by, the importance of the log gradually decreases, which conforms to the timeliness characteristics of the log information. Newer logs can better reflect the current running state of the system, while older logs may no longer be relevant. Through this dynamic adjustment, the system can handle logs at different times more flexibly, avoid over-focusing on outdated logs, and thus improve the overall performance and resource utilization efficiency of the system.
[0036] As a feasible implementation, determining the initial decay coefficient includes: determining the expected decay speed according to the current business requirements, and determining the initial decay coefficient according to the expected decay speed; where the initial decay coefficient is positively correlated with the expected decay speed.
[0037] In a specific implementation, the initial decay coefficient can be adjusted according to the current business requirements. For example, if it is desired that the importance decays to 50% within 1 hour, then (assuming the time unit is seconds), if it is desired that the importance decays to 10% within 24 hours, then .
[0038] As a feasible implementation manner, the process of determining the current decay coefficient includes: determining an initial decay coefficient, and adjusting the initial decay coefficient according to the semantic similarity degree between the operation status log and the historical exception log and / or the current load to determine the current decay coefficient; wherein, the current decay coefficient is negatively correlated with the similarity degree between the operation status log and the historical exception log, and / or, the current decay coefficient is positively correlated with the current load.
[0039] In a specific implementation, an initial decay rate can be preset, and then the initial decay coefficient is adjusted according to the semantic similarity degree between the current operation status log and the historical exception log and / or the current load, so as to determine the current decay coefficient. Specifically, the historical exception log can be the historical key log collected when the software has an exception. By comparing the semantic similarity degree between the current operation status log and the historical exception log, the relationship between the current operation status log and the known exception is evaluated. If the current operation status log has a high semantic similarity with the historical exception log, it indicates that the current operation status log may indicate a similar problem, so its importance should not decay quickly. On the contrary, if the similarity is low, the importance of the log can decay faster. The current decay coefficient also refers to the current load. A high load may mean that the system resources are tense, and any exception may have a greater impact on the system performance. Therefore, it is necessary to monitor the log information more closely. In this case, the decay coefficient will increase accordingly to slow down the decay rate of the log importance. In summary, the determination of the current decay coefficient is negatively correlated with the similarity degree between the operation status log and the historical exception log, that is, the higher the similarity, the smaller the decay coefficient, and the slower the decay of the log importance. At the same time, the current decay coefficient is positively correlated with the current load, that is, the higher the load, the larger the decay coefficient, and the faster the decay of the log importance. This dynamic adjustment mechanism enables the log management system to respond more flexibly to different operation and maintenance scenarios, improving the timeliness and accuracy of fault detection and handling.
[0040] As a feasible implementation manner, adjusting the initial decay coefficient according to the semantic similarity degree between the operation status log and the historical exception log and / or the current load to determine the current decay coefficient includes: adjusting the initial decay coefficient according to the decay coefficient adjustment formula to determine the current decay coefficient; wherein, the decay coefficient adjustment formula is: ; wherein, is the current decay coefficient, is the initial decay coefficient, is the semantic similarity influence coefficient, is the semantic similarity degree between the operation status log Lt and the historical exception log C, is the load impact coefficient, is the current load.
[0041] In a specific implementation, the initial attenuation coefficient is adjusted through the above attenuation coefficient adjustment formula to determine the current attenuation coefficient. Among them, is the semantic similarity impact coefficient, and its value range is between 0 and 1, which is used to adjust the correlation degree of historical log features. represents the semantic similarity degree between the running state log Lt and the historical abnormal log C, and can be obtained by calculating the cosine similarity through a pre-trained language model, such as BERT (Bidirectional Encoder Representations from Transformers, bidirectional encoder representation). is the load impact coefficient, and its value is also between 0 and 1, which is used to adjust the impact of real-time load on attenuation. is a real-time load function, defined as: , where CPU(t) and Memory(t) are the current CPU and memory utilization rates, and Threshold is a preset load threshold, such as 80%. When the system is under high load (such as >1), the attenuation rate is automatically increased, that is, is increased, and urgent exceptions are preferentially processed; when the load is low, the attenuation rate is decreased, that is, is decreased, and more historical data is retained for in-depth analysis.
[0042] The action mechanism of the above attenuation coefficient adjustment formula is as follows: when the current log is semantically similar to the known abnormal cluster, that is, the value is high, the system reduces the attenuation rate by reducing the coefficient, so as to extend the retention time of the log importance. On the contrary, if the log represents an unknown exception, that is, the semantic similarity is low, the system will maintain a high attenuation rate to avoid interference from redundant data. In addition, the system load will also affect the attenuation rate. When the system is in a high load state, that is, the value is greater than 1, the system increases the attenuation rate by increasing the coefficient, so as to preferentially process urgent exceptions. When the system load is low, the attenuation rate is decreased, so as to retain more historical data for in-depth analysis.
[0043] It can be seen that by considering the semantic similarity between the log and the historical abnormal logs, it is possible to accurately identify whether the current log indicates a known problem. If the semantic similarity between the current log and the historical abnormal logs is high, it means that the log may be related to known abnormal situations. At this time, the decay coefficient will be decreased, so that the importance of the log can be retained for a longer time. This mechanism ensures that the logs related to known abnormalities will not be ignored over time, thus improving the response speed and processing efficiency for known problems. Secondly, the introduced current load factor enables the evaluation of log importance to be adjusted according to the real-time operating state of the system. When the system load is high, resources are tense, and any abnormality may have a greater impact on the system performance. At this time, the decay coefficient will be increased to accelerate the decay speed of log importance, so that the system can give priority to handling current urgent problems and avoid low-priority logs occupying too many resources. On the contrary, when the system load is low, the decay coefficient will be correspondingly decreased, and the system can retain more historical data for in-depth analysis, providing richer information for subsequent operation and maintenance decisions.
[0044] For known abnormalities, through the above implementation methods, it is possible to accurately identify the logs that are semantically similar to historical abnormalities, such as "Kafka connection failure", automatically decrease the decay coefficient, and preferentially call historical repair solutions to improve the processing efficiency. For brand-new types of logs (such as component errors that have never been recorded), due to the low semantic similarity, the decay coefficient remains at a high level. However, if the system is simultaneously in a high-load state, it will trigger a "two-factor alarm", forcefully retain the log importance and trigger manual intervention to achieve sensitive capture of unknown abnormalities. At the same time, the decay strategy is adjusted in real time in combination with the system load to avoid low-priority logs occupying too much computing resources when resources are tense, and to achieve a balance between operation and maintenance efficiency and system performance through dynamic resource scheduling.
[0045] In summary, the above decay coefficient adjustment formula dynamically adjusts the initial decay coefficient by combining semantic similarity and system load, realizing the intelligence and self-adaptability of log importance evaluation. This mechanism not only improves the processing efficiency for known abnormalities, but also optimizes the resource allocation of the system under different load conditions, making the software operation and maintenance process more efficient and accurate.
[0046] S103: Input the key log into the artificial intelligence model to generate an operation and maintenance command corresponding to the key log by using the artificial intelligence model.
[0047] In this step, the key logs are input into an artificial intelligence model, such as an LSTM (Long Short-Term Memory) model. This artificial intelligence model is trained based on a large amount of historical log data and corresponding operation and maintenance operations, and can learn the mapping relationship between different abnormal situations and corresponding operation and maintenance operations. That is, the artificial intelligence model can analyze the key logs and generate corresponding operation and maintenance commands. These commands are designed to solve the abnormalities or problems recorded in the logs, such as restarting services, clearing caches, or modifying configurations. For example, if an error message of "insufficient memory" appears in the log, the model may generate an operation and maintenance command of "increase memory allocation", thus realizing the automated processing of abnormal situations. In this way, corresponding operation and maintenance commands can be quickly and accurately generated for the abnormalities in software operation, improving the operation and maintenance efficiency and system stability. In this way, the artificial intelligence model can automatically provide operation and maintenance solutions, reduce manual intervention, and improve the operation and maintenance efficiency and response speed.
[0048] In specific implementation, first, it is necessary to collect historical fault logs and their corresponding operation and maintenance records as the basis for constructing the training dataset. These data are sourced from system log files and operation and maintenance records. The system log files record fault information during software operation, while the operation and maintenance records record the repair operations of the operation and maintenance personnel for these faults. The collected data needs to be preprocessed, including steps such as data cleaning and word embedding. Data cleaning is to remove irrelevant information and ensure the quality of the data; word embedding is to map words into a low-dimensional dense vector space to capture the semantic relationships between words. Commonly used word embedding methods include pre-trained word vectors, such as Word2Vec (Word to Vector), GloVe (Global Vectors for Word Representation), FastText, etc. In addition, a dynamic training method can be adopted, where an Embedding layer is embedded in the model and the word vectors are continuously updated as the model is trained. The advantage of this method is that it can generate vectors with lower dimensions while capturing the semantic and context information of words, providing richer semantic features for model training. After that, the preprocessed data is divided into a training set, a validation set, and a test set. The training set is used to train the model, usually accounting for 70%-80% of the total data. The validation set is used to adjust the model hyperparameters and prevent overfitting, usually accounting for 10%-15% of the total data. The test set is used to evaluate the model performance, usually accounting for 10%-15% of the total data. After constructing the dataset, a suitable model architecture is selected, such as the LSTM network structure, which includes an input layer, an LSTM layer, a fully connected layer, and an output layer. Input layer: Receives the preprocessed data. The input layer receives the preprocessed data. The LSTM layer is responsible for capturing long-term dependencies in the time series. A single-layer or multi-layer LSTM structure can be designed according to the complexity of the task, and the number of neurons in each layer also needs to be adjusted according to the specific task. The fully connected layer maps the output of the LSTM layer to the target space, and the output layer is designed according to the task type. For example, the Softmax function is used in classification tasks, and a linear output is used in regression tasks. In addition, a loss function needs to be set. For classification tasks, cross-entropy loss is usually adopted. At the same time, parameters such as the learning rate and momentum also need to be set to optimize the model training process. The process of training the model is to input the training set data into the constructed model and continuously update the model parameters through the backpropagation algorithm. After each round of training (epoch), the validation set is used to evaluate the performance of the model to ensure that the model can gradually learn the patterns in the data during training and avoid overfitting.The parameters involved in the training process include the batch size, which is the number of samples used in each training; the number of epochs, which is the number of times the model traverses the entire training set; and the learning rate, which is used to control the step size of parameter updates. Finally, a series of evaluation metrics are used to measure the performance of the model. For classification tasks, common evaluation metrics include accuracy, precision, recall, and F1-score, etc. These metrics can reflect the accuracy and reliability of the model in processing classification tasks from different perspectives. Through these evaluation metrics, the performance of the model can be comprehensively understood, and the model can be further optimized and adjusted as needed. The entire training process is an iterative and optimization process. By continuously adjusting the model structure and parameters, an artificial intelligence model with excellent performance is finally obtained. This model can accurately analyze new fault logs and generate corresponding operation and maintenance commands, thus realizing the automation and intelligence of software operation and maintenance, and improving the operation and maintenance efficiency and system stability.
[0049] Taking the LSTM model as an example, the process of training the LSTM model is as Figure 2 shown and includes the following steps: collecting historical fault logs and repair operation records, data preprocessing, dividing the training set and the training set, constructing the LSTM model, training the LSTM model, and saving the trained model.
[0050] It can be seen that this step quickly identifies and responds to software anomalies in an automated manner, reducing manual intervention and improving operation and maintenance efficiency.
[0051] The software operation and maintenance method provided by the embodiments of the present application collects the running status logs of the software and, when detecting that the software has an anomaly, collects the running status logs of the software within a preset time period. Further, key logs with an importance level greater than a preset value are screened out from the running status logs within the preset time period, where the importance level of the running status logs is positively correlated with the log level and negatively correlated with the generation time. This screening mechanism can accurately extract log information that is more valuable for anomaly analysis. The key logs are input into the artificial intelligence model, enabling the artificial intelligence model to generate corresponding operation and maintenance commands based on the running status logs with a shorter generation time and a higher log level, thus significantly improving the accuracy of the operation and maintenance commands output by the model. This method not only realizes the automation of operation and maintenance operations, reduces the dependence on manual intervention, and improves operation and maintenance efficiency, but also further enhances the applicability and reliability of the artificial intelligence model in the operation and maintenance of complex software systems by accurately screening log information.
[0052] The embodiment of the present application discloses a software operation and maintenance method. Compared with the previous embodiment, the technical solution is further described and optimized in this embodiment. Refer to Figure 3 , which is a flowchart of another software operation and maintenance method shown according to an exemplary embodiment.
[0053] S201: Collect the running status logs of the software. When it is detected that the software has an anomaly, collect the running status logs of the software within a preset time period.
[0054] S202: Screen out the key logs in the running status logs within the preset time period whose importance level is greater than a preset value; wherein, the importance level of the running status logs is positively correlated with the log level of the running status logs and negatively correlated with the generation time of the running status logs.
[0055] S203: Input the key logs into the artificial intelligence model to generate the operation and maintenance operation commands corresponding to the key logs by using the artificial intelligence model.
[0056] S204: If the artificial intelligence model cannot output the operation and maintenance operation commands corresponding to the anomaly, receive and record the operation and maintenance operation commands for the anomaly through the input interface.
[0057] In specific implementation, in some cases, the artificial intelligence model may not be able to accurately identify or generate the operation and maintenance operation commands for a specific anomaly. At this time, receive and record the operation and maintenance operation commands manually input by the operation and maintenance personnel through the input interface. The input interface can be a GUI (Graphical User Interface) or a command line interface CLI (Command Line Interface). The operation and maintenance personnel can input the repair commands through these interfaces. In this step, record the operation and maintenance operation commands received through the input interface for subsequent model updates.
[0058] S205: After it is detected that the software has returned to normal, update the artificial intelligence model based on the key logs and the operation and maintenance operation commands received from the input interface.
[0059] In specific implementation, after the software returns to normal operation after performing the operation and maintenance operations, search backward from the current moment for the running status logs within the preset time period, and at the same time intercept the operation and maintenance operation commands from the current moment back to the start of the operation. Use these data to update the artificial intelligence model. By continuously learning new anomaly situations and the corresponding operation and maintenance operations, the model can be continuously optimized to improve the ability to handle unknown anomalies.
[0060] As a feasible implementation, updating the artificial intelligence model based on the key logs and the operation and maintenance operation commands received from the input interface includes: encapsulating the key logs and the operation and maintenance operation commands received from the input interface as training data; when the quantity of the training data reaches a preset value, updating the artificial intelligence model based on the training data.
[0061] In specific implementation, the key logs and the operation and maintenance operation commands are first encapsulated into training data. These training data contain the specific information of software abnormal situations and the corresponding operation and maintenance operations, providing opportunities for the artificial intelligence model to learn and optimize. When the collected training data reaches a certain preset value, the system will use these data to update the artificial intelligence model. This regular update mechanism ensures that the model can continuously learn new information, thereby improving the accuracy of its prediction and decision-making.
[0062] As a feasible implementation, after encapsulating the key logs and the operation and maintenance operation commands received from the input interface as training data, it further includes: determining the log category of the key logs and adding the log category to the training data.
[0063] In specific implementation, after encapsulating the key logs and the operation and maintenance operation commands as training data, the category of the key logs is determined, and this category information is added to the training data. The purpose of this step is to enable the artificial intelligence model to identify and learn different categories of log data, thereby improving the model's log classification and processing capabilities.
[0064] Specifically, a pre-classification model can be used to process the log data. The goal of this model is to classify the original log text, such as "Failed to connect to Kafka broker", into the corresponding technical fields or components, such as Kafka (a messaging system), Database (database), Network (network), etc. The classification categories are defined according to business requirements, for example: Kafka: logs related to Kafka; Database: logs related to the database; Network: logs related to the network; Security: logs related to security; Other: other unclassified logs.
[0065] In this way, the log data can be classified according to the technical fields to which they belong, and this category information is added as features to the training data, providing richer information for the training of the model.
[0066] As a feasible implementation, after encapsulating the key logs and the operation and maintenance operation commands received from the input interface as training data, it further includes: performing preprocessing operations on the training data to organize the training data into structured data.
[0067] In specific implementation, it is also necessary to perform preprocessing operations on the training data to organize the training data into structured data. The preprocessing operations include steps such as data cleaning, formatting, and vectorization. The purpose is to make the data more standardized and unified, facilitating model processing and analysis.
[0068] The specific steps are as follows: 1. Data collection: Collect a large amount of log data from log files to ensure coverage of all target categories. For example, "Failed to connect to Kafka broker" (unable to connect to the Kafka broker) → Kafka (message system); "Database query timeout" (database query timeout) → Database (database); "Networklatency detected" (network latency detected) → Network (network); "Unauthorized accessattempt" (unauthorized access attempt) → Security (security); "Application startedsuccessfully" (application started successfully) → Other (others).
[0069] 2. Text cleaning: Remove irrelevant information such as timestamps and special characters. For example, clean "2023-10-01 12:00:00[ERROR] Failed to connect to Kafka broker" to "Failed to connect to Kafkabroker".
[0070] 3. Word segmentation: Split the log text into words or phrases. For example, "Failed to connect to Kafkabroker" → ["Failed", "to", "connect", "to", "Kafka", "broker"].
[0071] 4. Vectorization: Convert the text into a numerical form for easy model processing. Common methods include: TF-IDF: A vectorization method based on term frequency-inverse document frequency; Word embedding: Such as Word2Vec, GloVe, or pre-trained BERT models.
[0072] Through these preprocessing operations, the training data is organized into structured data, providing high-quality input for model training.
[0073] As a feasible implementation, updating the artificial intelligence model based on training data includes: adding the training data to the training set to update the training set; retraining the artificial intelligence model using the updated training set to obtain a trained artificial intelligence model.
[0074] In a specific implementation, the pre-classified new data is added to the existing training set. If the newly added data reaches a preset quantity (such as 10 pieces), it triggers the retraining of the model, that is, retraining the artificial intelligence model using the updated training set. In this way, the model can continuously learn new information and adjust its parameters and structure according to this new information, thereby improving the accuracy and adaptability of the model. Through this implementation method of updating the model based on training data, the artificial intelligence model can continuously evolve to adapt to the changing operation and maintenance environment and improve the efficiency of fault detection and repair.
[0075] The automated operation and maintenance flow chart is as Figure 4 shown. The system logs are monitored in real time. During the monitoring process, potential fault logs are identified and the corresponding background operations are recorded. After completing an operation and maintenance task, the new logs and operation records are added to the training set. Check whether the newly added data volume reaches the threshold. If not, return to the step of monitoring the system logs in real time. If so, use the updated training set to retrain the artificial intelligence model.
[0076] Thus, when the artificial intelligence model cannot output the operation and maintenance operation command corresponding to the exception, the operation and maintenance operation command for the exception is received and recorded through the input interface, ensuring that in the case where the model cannot handle it, the operation and maintenance personnel can still perform effective intervention. In addition, when the software returns to normal, the artificial intelligence model is updated based on the running state logs of the software and the operation and maintenance operation commands received from the input interface, enabling the model to continuously learn new exception situations and corresponding operation and maintenance operations, thereby improving the accuracy and adaptability of the model and further enhancing the intelligent level of the operation and maintenance system.
[0077] Based on the above embodiments, as a preferred implementation manner, on the basis of collecting the running state logs of the acquisition software, context information related to the logs is further collected, such as user operation records, system configuration changes, external environment changes, etc. These context information can provide richer background knowledge for log analysis and help to more accurately understand the log content. Using an artificial intelligence model, combined with the collected context information, in-depth analysis of key logs is carried out. The model not only considers the information in the log itself, but also considers the related context information to identify potential abnormal patterns and correlations. When generating operation and maintenance operation commands, in addition to based on the log content, context information is also considered to formulate more accurate and effective operation and maintenance strategies. For example, if the log shows insufficient memory and the context information indicates that there have been a large number of user accesses recently, then the model may recommend increasing memory allocation or optimizing the cache strategy. When updating the artificial intelligence model, the context information is incorporated into the training data. In this way, the model can not only understand the log data during the learning process, but also learn the relationship between the log and the context information, thereby improving the generalization ability and prediction accuracy of the model. By continuously collecting new log data and context information, the model is continuously updated and optimized. This continuous learning mechanism enables the model to adapt to the constantly changing operation and maintenance environment and improve the intelligent level of automated operation and maintenance.
[0078] It can be seen that by considering the context information, the model can more accurately identify and understand abnormal logs, thereby improving the accuracy of fault diagnosis. The operation and maintenance operation commands generated in combination with the context information are more accurate and effective and can better solve practical problems. Context-aware model training enables the model to learn the relationship between the log and the context information, improving the generalization ability and prediction accuracy of the model.
[0079] The following introduces an application embodiment provided by the present application. For a software system that can normally print logs, the automated operation and maintenance includes the following: 1. Real-time log capture and dynamic evaluation: In the event trigger stage, the system continuously monitors the logs and captures relevant information. For example, at timestamp 14:30:00, the system records an error-level log: "Node resource allocation failed, insufficient available memory"; at timestamp 14:30:05, a warning-level log is recorded: "etcd storage engine write latency exceeds the standard"; at timestamp 14:30:06, an information-level log is recorded: "XXXXXXXX". The system performs dynamic attenuation processing on these logs. The initial importance is set as follows: The initial weight of ERROR-level logs is 1.0, WARNING-level is 0.7, and INFO-level is 0.3. An exponential decay model is adopted, and a 1-hour half-life parameter is configured. After 45 minutes, the importance of ERROR logs decays to 0.6, and WARNING logs decay to 0.42.
[0080] 2. Anomaly Detection and Log Retrospection: When it is found that the software cannot run properly, the system automatically retrospects the logs with an importance level greater than or equal to 0.5 in the past 2 hours. In this example, the high-priority log is a memory allocation error (current importance 0.6), while the etcd write latency (current importance 0.42, below the threshold) is a secondary-priority log.
[0081] 3. Automated Repair and State Recovery: In the stage of abnormal state determination, the system finds that the reason for the software's abnormal operation is not in the initial knowledge base, so the model cannot provide repair operations. At this time, manual intervention is required for troubleshooting. The operation and maintenance personnel find the following abnormal metrics in the system: the memory usage in the node resource utilization has exceeded 95% for 5 consecutive minutes, and the average time taken to create a Pod in the service scheduling latency has exceeded the baseline value by 300%. The operation and maintenance personnel start to execute the actions to repair the system, and the background records each operation instruction of the operation and maintenance personnel, such as executing commands like systemctl status mysqld. During the repair process by the operation and maintenance personnel, the system continuously detects whether the program is running properly. When it detects that the program is running properly, it enters the data encapsulation process. The generated structured event record contains the key error logs and their decayed importance, as well as the sequence of repair operations executed.
[0082] 4. Knowledge Base Evolution and Model Optimization: In the data standardization processing stage, the encapsulated event data is converted into 256-dimensional semantic vectors through a pre-trained language model, and the sequence of repair operations is encoded into a standardized operation instruction set. When the above situations accumulate to a certain number, it triggers the retraining of the model. For example, when the accumulated new event data reaches 10, it triggers the full-scale retraining of the LSTM model to achieve the evolution of the knowledge base and the optimization of the model. Through this series of automated operation and maintenance processes, the software system can achieve rapid response and repair of faults, while continuously optimizing its own operation and maintenance capabilities, improving the stability and reliability of the system.
[0083] Next, a software operation and maintenance device provided by an embodiment of the present application will be introduced. The software operation and maintenance device described below can be referred to mutually with the software operation and maintenance method described above. Refer to Figure 5 , the structural diagram of a software operation and maintenance device shown according to an exemplary embodiment.
[0084] The collection module 100 is used to collect the running state logs of the software. When it detects that the software has an anomaly, it collects the running state logs of the software within a preset time period.
[0085] The screening module 200 is used to screen the key logs with an importance level greater than a preset value from the running state logs within a preset time period; wherein, the importance level of the running state logs is positively correlated with the log level of the running state logs and negatively correlated with the generation time of the running state logs.
[0086] A generation module 300, configured to input critical logs into an artificial intelligence model to generate operation and maintenance operation commands corresponding to the critical logs by using the artificial intelligence model.
[0087] The software operation and maintenance device provided by the embodiments of the present application collects the running status logs of the software, and when it detects that the software has an anomaly, it collects the running status logs of the software within a preset time period. Further, critical logs with an importance level greater than a preset value are screened out from the running status logs within the preset time period, where the importance level of the running status logs is positively correlated with the log level and negatively correlated with the generation time. This screening mechanism can accurately extract log information that is more valuable for anomaly analysis. By inputting the critical logs into the artificial intelligence model, the artificial intelligence model can generate corresponding operation and maintenance operation commands based on the running status logs with a shorter generation time and a higher log level, thereby significantly improving the accuracy of the operation and maintenance operation commands output by the model. This method not only realizes the automation of operation and maintenance operations, reduces the dependence on manual intervention, and improves operation and maintenance efficiency, but also further improves the applicability and reliability of the artificial intelligence model in the operation and maintenance of complex software systems by accurately screening log information.
[0088] Based on the above embodiments, as a preferred implementation manner, it further includes: a setting module, configured to set the importance level of the running status logs based on the log level of the running status logs; an attenuation module, configured to, when the software is running normally, reduce the importance level of the running status logs based on a time decay model; and when it detects that the software has an anomaly, stop reducing the importance level of the running status logs.
[0089] Based on the above embodiments, as a preferred implementation manner, the time decay model is: ; where is the importance level set based on the log level of the running status logs, e is the natural constant, is the current attenuation coefficient, is the time difference between the current time and the generation time of the running status logs, is the importance level of the running status logs at the current time t.
[0090] Based on the above embodiments, as a preferred implementation manner, it further includes: a first determination module, configured to determine an initial attenuation coefficient; an adjustment module, configured to adjust the initial attenuation coefficient according to the semantic similarity degree between the running status logs and the historical anomaly logs and / or the current load to determine the current attenuation coefficient; where the current attenuation coefficient is negatively correlated with the similarity degree between the running status logs and the historical anomaly logs, and / or, the current attenuation coefficient is positively correlated with the current load.
[0091] Based on the above embodiments, as a preferred implementation, the first determination module is specifically configured to: determine an expected attenuation rate according to the current service requirement, and determine an initial attenuation coefficient according to the expected attenuation rate; wherein, the initial attenuation coefficient is positively correlated with the expected attenuation rate.
[0092] Based on the above embodiments, as a preferred implementation, the adjustment module is specifically configured to: adjust the initial attenuation coefficient according to the attenuation coefficient adjustment formula to determine the current attenuation coefficient; wherein, the attenuation coefficient adjustment formula is: ; wherein, is the current attenuation coefficient, is the initial attenuation coefficient, is the semantic similarity influence coefficient, is the semantic similarity degree between the running state log Lt and the historical exception log C, is the load influence coefficient, is the current load.
[0093] Based on the above embodiments, as a preferred implementation, it further includes: a construction module, configured to obtain historical fault logs and corresponding operation and maintenance operation records, and construct a training set based on the historical fault logs and corresponding operation and maintenance operation records; a training module, configured to train an artificial intelligence model using the training set to obtain a trained artificial intelligence model; wherein, the artificial intelligence model includes a long short-term memory neural network model.
[0094] Based on the above embodiments, as a preferred implementation, it further includes: a receiving module, configured to, when the artificial intelligence model cannot output an operation and maintenance operation command corresponding to an exception, receive and record the operation and maintenance operation command for the exception through the input interface; an updating module, configured to, when it is detected that the software has returned to normal, update the artificial intelligence model based on the key logs and the operation and maintenance operation commands received from the input interface.
[0095] Based on the above embodiments, as a preferred implementation, the updating module is specifically configured to: encapsulate the key logs and the operation and maintenance operation commands received from the input interface as training data; when the number of training data reaches a preset value, update the artificial intelligence model based on the training data.
[0096] Based on the above embodiments, as a preferred implementation, it further includes: a second determination module, configured to determine the log category of the key logs and add the log category to the training data.
[0097] Based on the above embodiments, as a preferred implementation, it further includes: a preprocessing module, configured to perform preprocessing operations on the training data to organize the training data into structured data.
[0098] Based on the above embodiments, as a preferred implementation, the update module is specifically configured to: add the training data to the training set to update the training set; and retrain the artificial intelligence model using the updated training set to obtain a trained artificial intelligence model.
[0099] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0100] Embodiments of the present application further provide an electronic device. Figure 6 As shown in the structural diagram of an electronic device according to an exemplary embodiment, Figure 6 the electronic device includes: a communication interface 1 capable of interacting with other devices such as network devices; a processor 2 connected to the communication interface 1 to enable information interaction with other devices, and when running a computer program, executing the software operation and maintenance method provided by one or more of the above technical solutions. The computer program is stored on a memory 3.
[0101] Of course, in actual application, the various components in the electronic device are coupled together through a bus system 4. It can be understood that the bus system 4 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 4 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 6 all the various buses are labeled as the bus system 4 in
[0102] The memory 3 in the embodiments of the present application is used to store various types of data to support the operation of the electronic device. Examples of these data include: any computer program for operating on the electronic device.
[0103] It can be understood that the memory 3 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 3 described in the embodiments of the present application is intended to include, but is not limited to, these and any other suitable types of memories.
[0104] The method disclosed in the embodiments of the present application above can be applied to the processor 2 or implemented by the processor 2. The processor 2 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 2 or the instructions in the form of software. The above-mentioned processor 2 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 2 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present application, it can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, which is located in the memory 3. The processor 2 reads the program in the memory 3 and combines its hardware to complete the steps of the foregoing method.
[0105] When the processor 2 executes the program, it implements the corresponding processes in the various methods of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.
[0106] The embodiments of the present application also provide a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is set to execute the steps in any of the above software operation and maintenance method embodiments when running.
[0107] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical disks, etc., all kinds of media that can store computer programs.
[0108] The embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by the processor 2, it implements the steps in any of the above software operation and maintenance method embodiments.
[0109] The embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by the processor 2, it implements the steps in any of the above software operation and maintenance method embodiments.
[0110] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0111] The above has introduced in detail a software operation and maintenance system, method, device, equipment, medium, and product provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A software operation and maintenance method, characterized in that, Including: Collecting the operation status log of the software, and when it is detected that the software has an exception, collecting the operation status log of the software within a preset time period; Screening key logs with an importance level greater than a preset value from the operation status logs within the preset time period; wherein, the importance level of the operation status log is positively correlated with the log level of the operation status log and negatively correlated with the generation time of the operation status log; Inputting the key logs into an artificial intelligence model to generate an operation and maintenance operation command corresponding to the key logs by using the artificial intelligence model.
2. The software operation and maintenance method according to claim 1, wherein After collecting the operation status log of the software, it further includes: Setting the importance level of the operation status log based on the log level of the operation status log; When the software is running normally, reducing the importance level of the operation status log based on a time decay model; When it is detected that the software has an exception, stopping reducing the importance level of the operation status log.
3. The software operation and maintenance method according to claim 2, characterized in that The time decay model is: ; Among them, is the importance level set based on the log level of the running status log, e is the natural constant, is the current attenuation coefficient, is the time difference between the current time and the generation time of the running status log, is the importance level of the running status log at the current time t.
4. The software operation and maintenance method according to claim 3, wherein It also includes: Determining an initial decay coefficient, and adjusting the initial decay coefficient according to the semantic similarity degree between the operation status log and the historical exception log and / or the current load to determine the current decay coefficient; wherein, the current decay coefficient is negatively correlated with the similarity degree between the operation status log and the historical exception log, and / or, the current decay coefficient is positively correlated with the current load.
5. The software operation and maintenance method according to claim 4, wherein Determining the initial decay coefficient includes: Determining an expected decay speed according to the current business requirements, and determining the initial decay coefficient according to the expected decay speed; wherein, the initial decay coefficient is positively correlated with the expected decay speed.
6. The software operation and maintenance method according to claim 4, wherein Adjusting the initial decay coefficient according to the semantic similarity degree between the operation status log and the historical exception log and / or the current load to determine the current decay coefficient includes: Adjusting the initial decay coefficient according to a decay coefficient adjustment formula to determine the current decay coefficient; wherein, the decay coefficient adjustment formula is: ; Among them, is the current attenuation coefficient, is the initial attenuation coefficient, is the semantic similarity influence coefficient, is the semantic similarity degree between the running state log Lt and the historical exception log C, is the load influence coefficient, is the current load.
7. The software operation and maintenance method according to claim 1, wherein It also includes: Obtaining historical fault logs and corresponding operation and maintenance operation records, and constructing a training set based on the historical fault logs and corresponding operation and maintenance operation records; Training an artificial intelligence model by using the training set to obtain a trained artificial intelligence model; wherein, the artificial intelligence model includes a long short-term memory neural network model.
8. The software operation and maintenance method according to claim 1, wherein After inputting the key logs into the artificial intelligence model, it further includes: If the artificial intelligence model cannot output an operation and maintenance operation command corresponding to the exception, receiving and recording an operation and maintenance operation command for the exception through an input interface; When it is detected that the software resumes normal operation, updating the artificial intelligence model based on the key logs and the operation and maintenance operation commands received from the input interface.
9. The software operation and maintenance method according to claim 8, characterized in that Updating the artificial intelligence model based on the key logs and the operation and maintenance operation commands received from the input interface includes: Encapsulating the key logs and the operation and maintenance operation commands received from the input interface into training data; When the number of the training data reaches a preset value, updating the artificial intelligence model based on the training data.
10. The software operation and maintenance method according to claim 9, wherein After encapsulating the critical log and the operation and maintenance operation command received from the input interface into training data, it further includes: Determine the log category of the critical log and add the log category to the training data.
11. The software operation and maintenance method according to claim 9, wherein After encapsulating the critical log and the operation and maintenance operation command received from the input interface into training data, it further includes: Perform a preprocessing operation on the training data to organize the training data into structured data.
12. The software operation and maintenance method according to claim 9, wherein Update the artificial intelligence model based on the training data, including: Add the training data to the training set to update the training set; Retrain the artificial intelligence model using the updated training set to obtain a trained artificial intelligence model.
13. A software operation and maintenance device, characterized in that, It includes: A collection module for collecting the running state logs of the software. When it detects that the software has an exception, it collects the running state logs of the software within a preset time period; A screening module for screening critical logs with an importance level greater than a preset value from the running state logs within the preset time period; wherein, the importance level of the running state logs is positively correlated with the log level of the running state logs and negatively correlated with the generation time of the running state logs; A generation module for inputting the critical log into the artificial intelligence model to generate an operation and maintenance operation command corresponding to the critical log by using the artificial intelligence model.
14. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps performed by the software operation and maintenance method according to any one of claims 1 to 12 when executing the computer program.
15. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed, the steps performed by the software operation and maintenance method according to any one of claims 1 to 12 are implemented.