Intelligent operation and maintenance method and device, computer equipment and storage medium

By building an operation and maintenance knowledge base and using a large language model to analyze fault characteristics, the problems of low efficiency and poor accuracy of traditional operation and maintenance methods have been solved, efficient and accurate operation and maintenance of software systems have been achieved, and the operation and maintenance effects in the medical and financial fields have been significantly improved.

CN120743323APending Publication Date: 2025-10-03PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510835570.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional operation and maintenance methods are inefficient and inaccurate in complex software systems, making it difficult to quickly locate and resolve faults, affecting the stability and security of software systems. This can lead to serious consequences, especially in the medical and financial fields.

Method used

Build an operation and maintenance knowledge base, analyze operation data in real time through pre-trained large language models, generate diagnostic results of fault characteristics, retrieve operation and maintenance solutions from the knowledge base, and automatically perform processing operations.

Benefits of technology

It achieves efficient and accurate operation and maintenance of software systems, quickly detects and repairs faults, reduces downtime, improves operation and maintenance efficiency and accuracy, and ensures system stability and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743323A_ABST
    Figure CN120743323A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to service system platforms of medical health, financial science and technology and the like, and discloses an intelligent operation and maintenance method and device, computer equipment and a storage medium. Acquiring operation data of a target software system in real time, and analyzing whether the operation data contains fault features or not through a pre-trained large language model; when the fault feature is detected, generating a diagnosis result corresponding to the fault feature through the large language model; searching a corresponding target operation and maintenance scheme from the operation and maintenance knowledge base according to the diagnosis result, and performing operation and maintenance processing on the target software system according to the target operation and maintenance scheme; therefore, efficient and accurate intelligent operation and maintenance of the software system can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent operation and maintenance method, apparatus, computer equipment, and computer-readable storage medium. Background Art

[0002] Currently, the operation and maintenance of modern software systems, especially complex enterprise applications and distributed systems, faces numerous challenges. Traditional O&M methods rely primarily on manual experience and technical documentation, requiring operators to manually monitor the operating status of the software system, troubleshoot the causes of faults, and implement appropriate remediation measures. This approach suffers from the following issues: 1. Low O&M efficiency: Manual troubleshooting is time-consuming and labor-intensive, making it difficult to quickly locate issues, especially in complex software systems. This leads to low O&M efficiency; 2. Poor O&M accuracy: Relying on personal experience can easily lead to misjudgments due to inexperience or oversights, compromising O&M effectiveness.

[0003] In the healthcare sector, the stability and reliability of software systems are directly related to the quality of medical services and patient safety. Software systems such as hospital information systems (HIS), electronic medical records (EMRs), and picture-assisted imaging systems (PACS) carry vast amounts of patient data and medical business processes. Traditional operations and maintenance methods for medical software systems suffer from low efficiency and accuracy, which can lead to data loss, business interruptions, and even endanger patient safety.

[0004] In the field of FinTech, the efficient operation of software systems is key to ensuring the security and efficiency of financial services. With the rapid development of technologies such as internet finance, mobile payments, and blockchain, FinTech software systems face higher operational and maintenance requirements. Traditional O&M methods for financial software systems suffer from low efficiency and accuracy, which can lead to transaction delays, data leaks, and even financial risks.

[0005] Based on this, how to provide an intelligent operation and maintenance method, device, computer equipment and computer-readable storage medium that can realize efficient and accurate intelligent operation and maintenance of software systems is an urgent problem to be solved by technical personnel in this field. Summary of the Invention

[0006] In view of the above-mentioned deficiencies in the prior art, the object of the present invention is to provide an intelligent operation and maintenance method, apparatus, computer equipment and computer-readable storage medium, aiming to solve the problem of how to achieve efficient and accurate intelligent operation and maintenance of software systems.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions:

[0008] In a first aspect, the present invention provides an intelligent operation and maintenance method, which includes:

[0009] Build an operation and maintenance knowledge base;

[0010] Acquire the target software system's operating data in real time and analyze whether the operating data contains fault characteristics using a pre-trained large language model;

[0011] When the fault feature is detected, a diagnosis result corresponding to the fault feature is generated by the large language model;

[0012] According to the diagnosis result, a corresponding target operation and maintenance plan is retrieved from the operation and maintenance knowledge base, and operation and maintenance processing is performed on the target software system according to the target operation and maintenance plan.

[0013] In a second aspect, the present invention provides an intelligent operation and maintenance device, comprising:

[0014] Construction module, used to build the operation and maintenance knowledge base;

[0015] An acquisition module is used to acquire the operating data of the target software system in real time and analyze whether the operating data contains fault characteristics through a pre-trained large language model;

[0016] a generating module, configured to generate a diagnosis result corresponding to the fault feature by using the large language model when the fault feature is detected;

[0017] The operation and maintenance module is used to retrieve a corresponding target operation and maintenance plan from the operation and maintenance knowledge base according to the diagnosis result, and perform operation and maintenance processing on the target software system according to the target operation and maintenance plan.

[0018] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the intelligent operation and maintenance method as described above when executing the computer program.

[0019] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program implements the intelligent operation and maintenance method as described above when executed by a processor.

[0020] Compared with the prior art, the present invention provides an intelligent operation and maintenance method, apparatus, computer equipment and computer-readable storage medium, wherein an operation and maintenance knowledge base is constructed; operation data of a target software system is acquired in real time, and a pre-trained large language model is used to analyze whether the operation data contains fault characteristics; when the fault characteristics are detected, a diagnosis result corresponding to the fault characteristics is generated by the large language model; based on the diagnosis result, a corresponding target operation and maintenance plan is retrieved from the operation and maintenance knowledge base, and operation and maintenance processing is performed on the target software system according to the target operation and maintenance plan; thus, the present invention can realize efficient and accurate intelligent operation and maintenance of software systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 A schematic diagram of an application environment of an intelligent operation and maintenance method provided by one embodiment of the present invention.

[0023] Figure 2 A flowchart of an intelligent operation and maintenance method provided by one embodiment of the present invention.

[0024] Figure 3 A schematic diagram of program modules of an intelligent operation and maintenance device provided by one embodiment of the present invention.

[0025] Figure 4 A schematic diagram of the structure of a computer device provided in one embodiment of the present invention.

[0026] Figure 5 Another structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0029] It will also be understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0030] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0031] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0032] References to "one embodiment" or "some embodiments" in the present specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0033] It should be understood that the order of execution of the steps in the following embodiments does not necessarily mean the order in which they are executed. The order in which each process is executed should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0034] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.

[0035] An intelligent operation and maintenance method provided by an embodiment of the present invention can be applied in Figure 1In the application environment shown, the client and server communicate via a network. The client includes, but is not limited to, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, personal digital assistants (PDAs), and other computer devices. The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0036] See also Figure 2 An embodiment of the present invention provides an intelligent operation and maintenance method, wherein the method comprises the following steps:

[0037] S100, build an operation and maintenance knowledge base;

[0038] S200, acquiring operating data of the target software system in real time, and analyzing whether the operating data contains fault characteristics using a pre-trained large language model;

[0039] S300: When the fault feature is detected, a diagnosis result corresponding to the fault feature is generated by using the large language model;

[0040] S400: Retrieve a corresponding target operation and maintenance solution from the operation and maintenance knowledge base according to the diagnosis result, and perform operation and maintenance processing on the target software system according to the target operation and maintenance solution.

[0041] Furthermore, the intelligent operation and maintenance method of this embodiment achieves efficient and accurate intelligent operation and maintenance of the software system through a series of collaborative steps, which is specifically reflected in the following aspects:

[0042] 1. Efficient fault detection

[0043] By acquiring real-time operational data from the target software system and analyzing it using a pre-trained large language model, this method can quickly identify whether the operational data contains fault signatures. The large language model's powerful natural language processing capabilities enable it to efficiently process and analyze large amounts of complex operational data, enabling immediate detection of potential fault signs and avoiding the inefficiency and lag of traditional manual monitoring. This real-time monitoring and rapid response mechanism significantly improves fault detection efficiency, ensuring that software system faults are detected and addressed at an early stage.

[0044] 2. Accurate fault diagnosis

[0045] When fault signatures are detected, the large language model generates detailed fault diagnosis results. The large language model not only identifies fault signatures but also combines contextual information with existing knowledge base content to generate accurate fault cause analysis and impact assessment. This deep learning-based diagnostic approach avoids the limitations of traditional manual diagnosis, which relies on personal experience and subjective judgment. It provides more objective and accurate diagnostic information, providing a reliable basis for subsequent operations and maintenance.

[0046] 3. Recommended intelligent operation and maintenance solutions

[0047] Based on the generated diagnostic results, this method retrieves a target O&M solution that matches the fault characteristics from the O&M knowledge base. This O&M knowledge base contains a wealth of O&M experience and solutions, built on historical data and expert knowledge, covering a wide range of fault scenarios. Through intelligent retrieval and matching mechanisms, this method quickly finds the most appropriate solution for the current fault, avoiding the blind actions often associated with lack of experience in traditional O&M and improving the accuracy and reliability of O&M decisions.

[0048] 4. Automated operation and maintenance

[0049] After determining the target O&M plan, this method uses automated O&M tools to perform O&M on the target software system. Automated tools can quickly execute repairs, optimizations, and other operations according to the pre-set O&M plan, reducing the need for manual intervention and improving the efficiency and accuracy of O&M operations. This automated mechanism not only quickly resolves issues but also prevents secondary failures caused by human error, ensuring that the software system can quickly return to normal operation.

[0050] 5. Self-learning and adaptive capabilities

[0051] By continuously accumulating new fault data and operational experience, this method optimizes the performance of the operational knowledge base and large language model. The large language model continuously updates its knowledge through continuous learning, improving its ability to identify new fault patterns. The operational knowledge base continuously enriches its solution library, improving its ability to handle complex faults. This self-learning and adaptive mechanism enables the method to continuously optimize its performance over time, better responding to changes in software systems and emerging fault types.

[0052] The intelligent O&M method of this embodiment achieves efficient and accurate O&M of software systems by organically combining efficient fault detection, accurate fault diagnosis, intelligent O&M solution recommendations, automated O&M operations, and self-learning and adaptive capabilities. This method not only improves O&M efficiency and reduces downtime, but also enhances O&M accuracy and reliability, reduces O&M costs, and provides a strong guarantee for the stable operation of software systems.

[0053] It is understandable that the intelligent operation and maintenance method provided in the embodiment of the present invention can be applied to intelligent operation and maintenance scenarios related to the medical and health field. The following is a specific example:

[0054] Scenario Background: A large hospital's hospital information system (HIS) and electronic medical record system (EMR) are core software systems for daily operations, carrying out key functions such as patient information management, medical process records, and medication management. The stability and data security of these systems are directly related to the quality of medical services and patient safety.

[0055] Specific Problem: One day, the hospital's EMR system experienced slow response times, causing frequent lags when medical staff accessed patient records, severely impacting medical efficiency. Traditional operations and maintenance (O&M) methods required manual troubleshooting of database performance, network connectivity, server load, and other aspects, which was time-consuming and difficult to quickly identify.

[0056] Apply the intelligent operation and maintenance method of the present invention:

[0057] 1. Build an operations knowledge base: The operations team pre-built an operations knowledge base that includes common failure modes, solutions, and operational procedures for medical software systems. For example, the knowledge base includes detailed solutions for database performance bottlenecks, network latency, server overloads, and other issues.

[0058] 2. Real-time acquisition and analysis of operational data: Monitoring devices deployed in the EMR system capture real-time system operational data, including server CPU and memory usage, database query response time, and network traffic. This data is transmitted to the intelligent operations and maintenance platform in real time and analyzed using a pre-trained large language model. The model quickly identifies abnormal database query response times and identifies fault characteristics that indicate performance bottlenecks.

[0059] 3. Generate diagnostic results: The large language model generates a detailed diagnostic report based on the analysis results, indicating that the cause of the failure is the failure of the database index, which leads to low query efficiency and affects the overall response speed of the system.

[0060] 4. Retrieve and execute the operation and maintenance plan: Based on the diagnosis results, the intelligent operation and maintenance platform retrieves the corresponding operation and maintenance plan from the operation and maintenance knowledge base and automatically invokes the automated operation and maintenance tool to perform database index optimization operations. Once the operation is completed, system performance returns to normal, medical staff can smoothly access patient medical records, and medical work can be carried out efficiently.

[0061] Technical effect: Through the intelligent operation and maintenance method of the present invention, the hospital's EMR system failures can be quickly detected and repaired, reducing system downtime, improving the efficiency and quality of medical services, and ensuring patient safety.

[0062] It is understandable that the intelligent operation and maintenance method provided in the embodiment of the present invention can also be applied to intelligent operation and maintenance scenarios related to the financial technology field. The following is a specific example:

[0063] Scenario Background: A fintech company provides online payment and wealth management services. Its software system needs to process large amounts of financial transaction data, ensuring real-time transactions and data security. With the rapid growth of its business, the system complexity has increased, placing increasingly stringent demands on operations and maintenance.

[0064] Specific issue: During a promotional event, the company's payment system experienced transaction delays. Users waited a long time after submitting payment requests to receive confirmation, leading to numerous complaints. Traditional operations and maintenance methods required personnel to individually troubleshoot multiple aspects, including payment interfaces, server performance, and network connectivity, making it difficult to quickly identify the issue.

[0065] Apply the intelligent operation and maintenance method of the present invention:

[0066] 1. Build an operations knowledge base: The operations team pre-built an operations knowledge base that includes common failure modes, solutions, and operational procedures for fintech systems. For example, the knowledge base includes detailed solutions for payment interface failures, server overloads, network delays, and other issues.

[0067] 2. Real-time acquisition and analysis of operational data: Monitoring devices deployed within the payment system capture real-time system operational data, including server CPU and memory usage, payment interface response time, and network traffic. This data is transmitted to the intelligent operations and maintenance platform in real time and analyzed using a pre-trained large language model. The model quickly identifies abnormal payment interface response times and identifies the fault characteristics of an interface failure.

[0068] 3. Generate diagnostic results: The large language model generates a detailed diagnostic report based on the analysis results, indicating that the cause of the failure is a temporary failure of the third-party service provider of the payment interface, which caused the interface call to fail and affected the normal completion of the transaction.

[0069] 4. Retrieve and execute an operation and maintenance plan: Based on the diagnostic results, the intelligent operation and maintenance platform retrieves the corresponding operation and maintenance plan from the operation and maintenance knowledge base, automatically invokes automated operation and maintenance tools, switches to a backup payment interface service provider, and optimizes the system configuration. Once the operation is completed, the payment system returns to normal, transaction delays are resolved, and user complaints are reduced.

[0070] Technical effect: Through the intelligent operation and maintenance method of the present invention, the payment system failures of financial technology companies can be quickly detected and repaired, reducing transaction delays, improving user satisfaction, and ensuring the stability and security of financial services.

[0071] Through the specific examples in the above two fields, it can be seen that the intelligent operation and maintenance method of the present invention can efficiently and accurately cope with the operation and maintenance challenges of complex software systems, significantly improve the operation and maintenance efficiency and quality, and has broad application prospects and practical value.

[0072] Furthermore, in one embodiment, the intelligent operation and maintenance method, wherein the construction of the operation and maintenance knowledge base specifically includes the steps of:

[0073] Collect operation and maintenance data from historical operation and maintenance information of several software systems;

[0074] Pre-processing the collected operation and maintenance data by cleaning, deduplication, and formatting;

[0075] The pre-processed operation and maintenance data is classified, and the classified operation and maintenance data is associated using knowledge graph technology to form an operation and maintenance knowledge base.

[0076] Furthermore, the specific implementation process of the steps in this embodiment is roughly as follows:

[0077] 1. Collect operation and maintenance data from the historical operation and maintenance information of several software systems

[0078] 1.1: Determine the source of data

[0079] Determine the scope of software systems for which operation and maintenance data needs to be collected, including different versions of the target software system, different deployment environments (such as development, testing, and production environments), and related subsystems.

[0080] Identify data source channels, such as system log files, monitoring tool records, historical operation records of the operation and maintenance team, user feedback reports, etc.

[0081] 1.2: Design a data collection plan

[0082] Design data collection tools or scripts based on the identified data sources to ensure efficient and complete collection of operation and maintenance data.

[0083] Determine the time range and frequency of data collection, for example, collect operations data for the past year, or collect data on a monthly basis.

[0084] 1.3: Perform data collection

[0085] Use designed data collection tools or scripts to collect operation and maintenance data from various data sources.

[0086] Ensure the stability and reliability of the data collection process to avoid data loss or corruption.

[0087] 2. Pre-process the collected operation and maintenance data by cleaning, deduplication and formatting

[0088] 2.1: Data Cleansing

[0089] Check the collected operation and maintenance data and remove invalid, erroneous, and noisy data. For example, delete duplicate error messages and malformed records in log files.

[0090] Handle missing values ​​in the data, such as filling in default values ​​or deleting records with missing values.

[0091] 2.2: Data Deduplication

[0092] Detect and remove duplicate data records to ensure data uniqueness. For example, duplicate entries caused by duplicate log records are removed.

[0093] 2.3: Data formatting

[0094] Convert operational data from different sources and formats into a unified format for easy processing. For example, unify the timestamp format in log data into the standard ISO8601 format.

[0095] Standardize the data, such as normalizing numerical data to the range of [0,1], to facilitate subsequent analysis and modeling.

[0096] 3. Classify the pre-processed operation and maintenance data, and use knowledge graph technology to associate the classified operation and maintenance data to form an operation and maintenance knowledge base

[0097] 3.1: Data Classification

[0098] Based on the content and characteristics of operation and maintenance data, it is classified into different categories, such as fault type, solution, operation, impact scope, etc.

[0099] Use machine learning algorithms (such as clustering algorithms) or manual labeling to classify the pre-processed operation and maintenance data.

[0100] 3.2: Building a Knowledge Graph

[0101] Using knowledge graph technology, the classified operation and maintenance data is associated to form a knowledge graph. For example, the fault type is associated with the corresponding solution, operation, and impact range.

[0102] Define the types of nodes and edges in the knowledge graph. For example, nodes represent fault types, solutions, etc., and edges represent the relationships between them.

[0103] 3.3: Optimization and improvement of knowledge graph

[0104] Based on historical operation and maintenance experience and expert knowledge, the knowledge graph is optimized and supplemented to ensure its completeness and accuracy.

[0105] The knowledge graph is updated regularly to reflect the latest operation and maintenance status and experience of the software system.

[0106] Through the above process, this embodiment can efficiently build a structured and associated operation and maintenance knowledge base, providing a solid foundation for subsequent intelligent operation and maintenance.

[0107] Furthermore, in one embodiment, the intelligent operation and maintenance method, wherein the real-time acquisition of the operating data of the target software system and the analysis of whether the operating data contains fault characteristics using a pre-trained large language model, specifically comprises the following steps:

[0108] Obtaining real-time operating data of the target software system through the API interface of the target software system;

[0109] Preprocessing the operating data by cleaning, standardizing and reducing its dimension;

[0110] Performing feature extraction on the preprocessed operating data to obtain a feature vector;

[0111] The extracted feature vector is input into a pre-trained large language model for semantic analysis to identify whether the operating data contains fault features.

[0112] Furthermore, the specific implementation process of the steps in this embodiment is roughly as follows:

[0113] 1. Obtain the target software system's operating data in real time through the target software system's API interface

[0114] 1.1: Determine the data collection points

[0115] Analyze the architecture of the target software system and determine the key nodes where operation data needs to be collected, such as server status interface, application log interface, database performance indicator interface, etc.

[0116] 1.2: Configure API interface

[0117] Configure the data collection tool according to the target software system's API documentation to ensure real-time access to operational data through the API. This configuration includes authentication information, request frequency, data format, and more.

[0118] 1.3: Deploy data collection devices

[0119] Deploy a data acquisition device on the server or monitoring platform of the target software system to ensure that the device can operate stably and transmit the collected operation data to the operation and maintenance monitoring platform in real time.

[0120] 1.4: Test data collection

[0121] Test the data acquisition device to verify whether it can correctly obtain operating data and ensure the integrity and accuracy of the data.

[0122] 2. Preprocessing of the operating data by cleaning, standardization and dimensionality reduction

[0123] 2.1: Data Cleansing

[0124] Check the collected operational data and remove invalid, erroneous, and noisy data. For example, delete duplicate records in the logs and filter out abnormal performance indicator values.

[0125] 2.2: Data Standardization

[0126] Convert operational data from different sources and formats into a unified format. For example, standardize timestamps to a standard format and normalize numerical data to the [0, 1] range.

[0127] 2.3: Data Dimensionality Reduction

[0128] Use dimensionality reduction techniques (such as principal component analysis PCA, linear discriminant analysis LDA, etc.) to reduce the dimension of running data, remove redundant information, and improve data processing efficiency.

[0129] 2.4: Data Validation

[0130] Verify the preprocessed data to ensure its integrity and consistency, and provide high-quality input for subsequent feature extraction and analysis.

[0131] 3. Extract features from the pre-processed running data to obtain feature vectors

[0132] 3.1: Define feature extraction rules

[0133] Based on the operating characteristics of the target software system, define feature extraction rules to determine which operating data indicators are related to fault characteristics. For example, CPU usage, memory usage, network latency, log error rate, etc.

[0134] 3.2: Applying feature extraction algorithms

[0135] Use machine learning or deep learning algorithms (such as convolutional neural network (CNN), recurrent neural network (RNN), etc.) to extract features from the preprocessed running data and generate feature vectors.

[0136] 3.3: Feature Vector Optimization

[0137] The extracted feature vectors are optimized to remove redundant features and retain the most representative features to improve the efficiency and accuracy of subsequent analysis.

[0138] 3.4: Feature Vector Verification

[0139] The extracted feature vectors are verified to ensure that they can accurately reflect the characteristic information of the running data and provide reliable input for subsequent semantic analysis.

[0140] 4. Input the extracted feature vector into the pre-trained large language model for semantic analysis to identify whether the operating data contains fault features

[0141] 4.1: Preparing a large language model

[0142] Load a pre-trained large language model to ensure that the model has been properly trained and optimized to understand the semantic information of the running data.

[0143] 4.2: Input feature vector

[0144] The extracted feature vector is input into the large language model, and the model will analyze it based on the semantic information of the feature vector.

[0145] 4.3: Semantic Analysis and Fault Identification

[0146] The large language model performs semantic analysis on feature vectors to identify whether the operational data contains fault signatures. For example, the model can use contextual understanding to identify whether an anomaly in a performance metric matches a known fault pattern.

[0147] 4.4: Generate analysis results

[0148] The large language model generates analysis results, including information such as whether fault features are detected, the type of fault features, and possible causes.

[0149] 4.5: Result Verification and Feedback

[0150] Verify the analysis results generated by the large language model to ensure their accuracy and reliability. Based on the verification results, make necessary adjustments and optimizations to the model to improve the accuracy of subsequent analysis.

[0151] Through the above process, this embodiment can realize real-time acquisition, preprocessing, feature extraction and semantic analysis of the target software system operation data, so as to efficiently and accurately detect whether the operation data contains fault characteristics, and provide strong support for subsequent fault diagnosis and operation and maintenance operations.

[0152] Furthermore, in one embodiment, the intelligent operation and maintenance method, wherein when the fault feature is detected, generating a diagnosis result corresponding to the fault feature by using the large language model, specifically comprises the steps of:

[0153] When the large language model detects that the operating data contains the fault feature, performing type recognition on the fault feature to determine the fault type of the fault feature;

[0154] According to the fault characteristics and fault types, and in combination with a context-aware algorithm, a diagnosis result corresponding to the fault characteristics is generated through the large language model.

[0155] Furthermore, the specific implementation process of the steps in this embodiment is roughly as follows:

[0156] 1. When the large language model detects that the operating data contains the fault feature, the fault feature is identified to determine the fault type of the fault feature.

[0157] 1.1: Fault Signature Detection

[0158] The large language model analyzes the input feature vector to determine whether there are fault characteristics in the operating data. If fault characteristics are detected, the subsequent type identification process is triggered.

[0159] 1.2: Fault feature extraction

[0160] Extract key information from detected fault features, such as abnormal values ​​of fault-related indicators, error codes in logs, system status changes, etc.

[0161] 1.3: Type Identification Model Call

[0162] Call the pre-trained fault type identification model (which can be a classification model based on machine learning or deep learning) and input the extracted key information into the model.

[0163] 1.4: Fault type determination

[0164] The fault type identification model outputs the specific type of fault characteristics based on the input key information, such as "database connection failure", "memory leak", "excessive network latency", etc.

[0165] 1.5: Type recognition result verification

[0166] Verify the identified fault type to ensure its accuracy and reliability. If the identification result is uncertain, it can be returned to the operation and maintenance personnel for manual review.

[0167] 2. Based on the fault characteristics and fault type, and in combination with the context-aware algorithm, the large language model generates the diagnostic results corresponding to the fault characteristics.

[0168] 2.1: Context Information Collection

[0169] Collect contextual information related to fault characteristics, including system operating status before and after the fault occurs, user operation records, historical fault data, etc.

[0170] 2.2: Application of context-aware algorithms

[0171] Apply context-aware algorithms (such as neural networks and attention mechanisms) to perform correlation analysis on fault features and their contextual information to explore potential causal relationships.

[0172] 2.3: Diagnosis result generation

[0173] The fault signature type and the analysis results of the context-aware algorithm are input into the large language model to generate detailed diagnostic results. The diagnostic results should include the fault type, possible cause, scope of impact, and recommended troubleshooting directions.

[0174] 2.4: Optimization of diagnostic results

[0175] Optimize the generated diagnostic results to ensure their accuracy and readability. For example, natural language processing technology can be used to polish the diagnostic results to make them easier for operations personnel to understand.

[0176] 2.5: Diagnosis result verification and feedback

[0177] The generated diagnostic results are verified to ensure they accurately reflect the fault situation. Based on the verification results, the large language model and context-aware algorithms are optimized to improve the accuracy of subsequent diagnosis.

[0178] Through the above process, this embodiment can realize the type identification of fault characteristics and the generation of detailed diagnostic results, providing accurate and comprehensive fault information to operation and maintenance personnel, so as to quickly locate the problem and take effective operation and maintenance measures.

[0179] Furthermore, in one embodiment, the intelligent operation and maintenance method, wherein, according to the diagnosis result, a corresponding target operation and maintenance plan is retrieved from the operation and maintenance knowledge base, and operation and maintenance processing is performed on the target software system according to the target operation and maintenance plan, specifically comprises the steps of:

[0180] Analyzing the diagnosis results to determine the impact of the fault characteristics;

[0181] When the impact of the fault feature exceeds a preset threshold, a target operation and maintenance solution matching the fault feature is retrieved from the operation and maintenance knowledge base;

[0182] According to the target operation and maintenance plan, the target software system is operated and maintained through automated operation and maintenance tools.

[0183] Furthermore, the intelligent operation and maintenance method, wherein the analyzing the diagnosis result to determine the impact degree of the fault characteristics, specifically comprises the steps of:

[0184] Constructing an impact assessment model;

[0185] Normalizing the diagnosis results so that the data format of the diagnosis results meets the input requirements of the impact assessment model;

[0186] The normalized diagnosis result is input into the impact assessment model, and the impact degree of the fault feature is determined according to the output result of the impact assessment model.

[0187] Furthermore, the intelligent operation and maintenance method, wherein, after the operation and maintenance processing of the target software system is performed by an automated operation and maintenance tool according to the target operation and maintenance plan, further specifically comprises the steps of:

[0188] During the operation and maintenance process of the target software system, real-time monitoring of the operation and maintenance effect of the target operation and maintenance solution;

[0189] Dynamically adjust the target operation and maintenance plan according to the monitored operation and maintenance effect;

[0190] Based on the adjusted target operation and maintenance plan, the target software system is operated and maintained by the automated operation and maintenance tool until the operation and maintenance work of the target software system is completed.

[0191] Furthermore, the specific implementation process of the steps in this embodiment is roughly as follows:

[0192] 1. Analyze the diagnostic results to determine the impact of the fault characteristics

[0193] 1.1 Constructing an impact assessment model

[0194] 1.1.1: Define evaluation metrics

[0195] Determine key indicators for evaluating the impact of fault characteristics, such as the extent of system performance degradation, the number of affected users, service interruption time, and fault repair difficulty.

[0196] 1.1.2: Designing the evaluation model architecture

[0197] Select an appropriate evaluation model architecture, such as a rule-based model, a machine learning model, or a deep learning model. For example, use a decision tree model or a neural network model to comprehensively assess the impact of fault characteristics.

[0198] 1.1.3: Model training and optimization

[0199] Use historical impact assessment data to train the assessment model and adjust model parameters to optimize the accuracy of the assessment results. This ensures that the model can accurately output the impact of fault characteristics based on the input diagnostic results.

[0200] 1.2 Normalizing the Diagnosis Results

[0201] 1.2.1: Data Preprocessing

[0202] Preprocess each indicator in the diagnostic results to remove outliers and noise data to ensure data accuracy.

[0203] 1.2.2: Normalization transformation

[0204] The indicator values ​​in the diagnosis results are converted to a unified scale range, such as [0,1] or [1,1], so as to meet the input requirements of the impact assessment model.

[0205] 1.2.3: Data format verification

[0206] Check whether the normalized data format meets the input requirements of the impact assessment model to ensure that the data can be smoothly input into the model for assessment.

[0207] 1.3 Input the normalized diagnostic results into the impact assessment model

[0208] 1.3.1: Input Data

[0209] The normalized diagnostic results were input into the impact assessment model.

[0210] 1.3.2: Model Inference

[0211] The model performs inference based on the input normalized data and calculates the impact of the fault characteristics.

[0212] 1.3.3: Output results

[0213] The model outputs an impact assessment result of the fault feature, such as a quantitative score or a classification label (such as "low impact", "medium impact", "high impact").

[0214] 2. When the impact of the fault feature exceeds a preset threshold, a target operation and maintenance solution matching the fault feature is retrieved from the operation and maintenance knowledge base.

[0215] 2.1 Determine whether the impact of the fault characteristics exceeds the threshold

[0216] 2.1.1: Setting Thresholds

[0217] Based on operation and maintenance requirements and historical data, set a preset threshold for the degree of impact, such as 0.8 (indicating high impact).

[0218] 2.1.2: Comparative evaluation results

[0219] Compare the impact assessment results output by the model with the preset threshold to determine whether it exceeds the threshold.

[0220] 2.1.3: Triggering the retrieval process

[0221] If the impact exceeds the threshold, the process of retrieving the operation and maintenance plan from the operation and maintenance knowledge base is triggered.

[0222] 2.2 Retrieve the target operation and maintenance plan from the operation and maintenance knowledge base

[0223] 2.2.1: Query the Knowledge Base

[0224] According to the type and impact of the fault characteristics, the matching operation and maintenance plan is searched from the operation and maintenance knowledge base.

[0225] 2.2.2: Scheme ranking

[0226] Sort the retrieved operation and maintenance solutions and give priority to the solution that best matches the fault characteristics.

[0227] 2.2.3: Solution Verification

[0228] Verify the retrieved operation and maintenance solutions to ensure their applicability and reliability.

[0229] 3. According to the target operation and maintenance plan, the target software system is operated and maintained through automated operation and maintenance tools

[0230] 3.1 Prepare automated operation and maintenance tools

[0231] 3.1.1: Select Tools

[0232] Select appropriate automated operation and maintenance tools based on the requirements of the target operation and maintenance plan.

[0233] 3.1.2: Configuration Tools

[0234] Configure automated operation and maintenance tools to ensure they can perform corresponding operations according to the operation and maintenance plan.

[0235] 3.1.3: Testing Tools

[0236] Test automated operation and maintenance tools to verify whether they can perform operation and maintenance operations correctly.

[0237] 3.2 Performing Operations and Maintenance

[0238] 3.2.1: Input operation and maintenance plan

[0239] Input the target operation and maintenance plan into the automated operation and maintenance tool.

[0240] 3.2.2: Execute Operation

[0241] Automated operation and maintenance tools perform operation and maintenance operations on the target software system according to the operation and maintenance plan, such as restarting services, repairing configurations, and updating software.

[0242] 3.2.3: Monitoring operation process

[0243] Monitor the execution process of operation and maintenance operations in real time to ensure smooth operations and promptly identify and address potential problems.

[0244] 4. During the operation and maintenance process of the target software system, real-time monitoring of the operation and maintenance effect of the target operation and maintenance solution

[0245] 4.1 Real-time monitoring of operation and maintenance effects

[0246] 4.1.1: Setting monitoring indicators

[0247] Determine key indicators for monitoring operation and maintenance effectiveness, such as system performance indicators, whether the fault is resolved, user feedback, etc.

[0248] 4.1.2: Deploy monitoring tools

[0249] Deploy real-time monitoring tools in the target software system, such as performance monitoring tools, log analysis tools, etc.

[0250] 4.1.3: Collect monitoring data

[0251] Collect monitoring data during operation and maintenance in real time to ensure the accuracy and completeness of the data.

[0252] 4.2 Analysis of monitoring data

[0253] 4.2.1: Data Preprocessing

[0254] The collected monitoring data is preprocessed to remove noise data and ensure data quality.

[0255] 4.2.2: Data Analysis

[0256] Analyze monitoring data, evaluate the effectiveness of operation and maintenance operations, determine whether the fault has been resolved, whether system performance has returned to normal, etc.

[0257] 4.2.3: Generate analysis report

[0258] Generate an operation and maintenance effect analysis report based on the analysis results to provide a basis for subsequent operation and maintenance plan adjustments.

[0259] 5. Dynamically adjust the target operation and maintenance plan based on the monitored operation and maintenance results

[0260] 5.1 Evaluate whether the operation and maintenance effect meets the requirements

[0261] 5.1.1: Setting effectiveness standards

[0262] Based on the operation and maintenance objectives, set evaluation criteria for operation and maintenance effects, such as system performance returning to normal levels and faults being completely resolved.

[0263] 5.1.2: Comparison of actual results

[0264] Compare the actual operation and maintenance results with the set standards to determine whether they meet the requirements.

[0265] 5.1.3: Triggering the Adjustment Process

[0266] If the operation and maintenance results do not meet expectations, the dynamic adjustment process of the operation and maintenance plan will be triggered.

[0267] 5.2 Dynamically adjust the operation and maintenance plan

[0268] 5.2.1: Analyze and adjust requirements

[0269] Based on the operation and maintenance effect analysis report, analyze the content that needs to be adjusted, such as optimizing configuration parameters, adding repairs, etc.

[0270] 5.2.2: Adjustment plan

[0271] Based on the analysis results, the target operation and maintenance plan is dynamically adjusted to generate a new operation and maintenance plan.

[0272] 5.2.3: Verify the adjustment plan

[0273] Verify the adjusted operation and maintenance plan to ensure its feasibility and effectiveness.

[0274] 6. Based on the adjusted target operation and maintenance plan, the target software system is operated and maintained through the automated operation and maintenance tool until the operation and maintenance of the target software system is completed.

[0275] 6.1 Performing Adjusted Operations

[0276] 6.1.1: Update Automation Tools

[0277] Input the adjusted operation and maintenance plan into the automated operation and maintenance tool and update the tool configuration.

[0278] 6.1.2: Re-execute the operation

[0279] The automated operation and maintenance tool re-executes operation and maintenance operations on the target software system according to the adjusted plan.

[0280] 6.1.3: Continuous Monitoring

[0281] During the execution of operation and maintenance operations, the operating status of the target software system is continuously monitored to ensure the effectiveness of the operation and maintenance operations.

[0282] 6.2 Confirmation of completion of operation and maintenance work

[0283] 6.2.1: Check the operation and maintenance effect

[0284] According to the set operation and maintenance effect evaluation criteria, check the operating status of the target software system to confirm whether the fault is completely resolved and whether the system performance has returned to normal.

[0285] 6.2.2: Recording Operation and Maintenance Results

[0286] The results of operation and maintenance operations are recorded in the operation and maintenance knowledge base to provide reference for subsequent operation and maintenance work.

[0287] 6.2.3: Close the operation and maintenance task

[0288] Confirm that the operation and maintenance work of the target software system is completed and close the related operation and maintenance tasks.

[0289] Through the above process, this embodiment can achieve efficient and accurate intelligent O&M of the target software system. From detecting and diagnosing fault characteristics, to searching and executing O&M solutions, to monitoring O&M results and dynamically adjusting solutions, the entire process forms a closed-loop intelligent O&M process, ensuring the stable and efficient operation of the software system.

[0290] As can be seen from the above method embodiments, the intelligent operation and maintenance method provided by the present invention includes: building an operation and maintenance knowledge base; acquiring the operating data of the target software system in real time, and analyzing the operating data using a pre-trained large language model to determine whether it contains fault characteristics; when the fault characteristics are detected, generating a diagnostic result corresponding to the fault characteristics using the large language model; based on the diagnostic result, retrieving the corresponding target operation and maintenance solution from the operation and maintenance knowledge base, and performing operation and maintenance processing on the target software system according to the target operation and maintenance solution. In this way, the method of the present invention can achieve efficient and accurate intelligent operation and maintenance of software systems.

[0291] It should be understood that although the present application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative work, and these operation steps are not necessarily performed in the order of the embodiments or flowcharts. The order of steps listed in the embodiments or flowcharts is only one way of executing the steps among many steps and does not represent the only execution order. It should be noted that there is not necessarily a certain order between the above steps. Those of ordinary skill in the art can understand from the description of the embodiments of the present invention that in different embodiments, the above steps may have different execution orders, that is, they may be executed in parallel, or they may be executed in an interchangeable manner, etc. Moreover, at least a portion of the steps in the embodiments or flowcharts may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but may be executed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be executed in turn, alternately or synchronously with other steps or at least a portion of the sub-steps or stages of other steps.

[0292] Based on the above method embodiment, please refer to Figure 3 Another embodiment of the present invention further provides an intelligent operation and maintenance device, wherein the device includes:

[0293] Construction module 11, used to build an operation and maintenance knowledge base;

[0294] An acquisition module 12 is configured to acquire operating data of a target software system in real time and analyze whether the operating data contains fault characteristics using a pre-trained large language model;

[0295] A generating module 13 is configured to generate a diagnosis result corresponding to the fault feature by using the large language model when the fault feature is detected;

[0296] The operation and maintenance module 14 is configured to retrieve a corresponding target operation and maintenance solution from the operation and maintenance knowledge base according to the diagnosis result, and perform operation and maintenance processing on the target software system according to the target operation and maintenance solution.

[0297] Furthermore, in one embodiment, the intelligent operation and maintenance device, wherein the building of the operation and maintenance knowledge base specifically includes:

[0298] Collect operation and maintenance data from historical operation and maintenance information of several software systems;

[0299] Pre-processing the collected operation and maintenance data by cleaning, deduplication, and formatting;

[0300] The pre-processed operation and maintenance data is classified, and the classified operation and maintenance data is associated using knowledge graph technology to form an operation and maintenance knowledge base.

[0301] Furthermore, in one embodiment, the intelligent operation and maintenance device, wherein the real-time acquisition of operating data of the target software system and the analysis of whether the operating data contains fault characteristics using a pre-trained large language model, specifically includes:

[0302] Obtaining real-time operating data of the target software system through the API interface of the target software system;

[0303] Preprocessing the operating data by cleaning, standardizing and reducing its dimension;

[0304] Performing feature extraction on the preprocessed operating data to obtain a feature vector;

[0305] The extracted feature vector is input into a pre-trained large language model for semantic analysis to identify whether the operating data contains fault features.

[0306] Furthermore, in one embodiment, the intelligent operation and maintenance device, wherein when the fault feature is detected, generating a diagnosis result corresponding to the fault feature by using the large language model, specifically includes:

[0307] When the large language model detects that the operating data contains the fault feature, performing type recognition on the fault feature to determine the fault type of the fault feature;

[0308] According to the fault characteristics and fault types, and in combination with a context-aware algorithm, a diagnosis result corresponding to the fault characteristics is generated through the large language model.

[0309] Furthermore, in one embodiment, the intelligent operation and maintenance device, wherein the step of retrieving a corresponding target operation and maintenance solution from the operation and maintenance knowledge base based on the diagnosis result, and performing operation and maintenance processing on the target software system according to the target operation and maintenance solution, specifically includes:

[0310] Analyzing the diagnosis results to determine the impact of the fault characteristics;

[0311] When the impact of the fault feature exceeds a preset threshold, a target operation and maintenance solution matching the fault feature is retrieved from the operation and maintenance knowledge base;

[0312] According to the target operation and maintenance plan, the target software system is operated and maintained through automated operation and maintenance tools.

[0313] Furthermore, in the intelligent operation and maintenance device, the step of analyzing the diagnosis result to determine the impact of the fault characteristics specifically includes:

[0314] Constructing an impact assessment model;

[0315] Normalizing the diagnosis results so that the data format of the diagnosis results meets the input requirements of the impact assessment model;

[0316] The normalized diagnosis result is input into the impact assessment model, and the impact degree of the fault feature is determined according to the output result of the impact assessment model.

[0317] Furthermore, the intelligent operation and maintenance device, after performing operation and maintenance processing on the target software system by an automated operation and maintenance tool according to the target operation and maintenance plan, further specifically includes:

[0318] During the operation and maintenance process of the target software system, real-time monitoring of the operation and maintenance effect of the target operation and maintenance solution;

[0319] Dynamically adjust the target operation and maintenance plan according to the monitored operation and maintenance effect;

[0320] Based on the adjusted target operation and maintenance plan, the target software system is operated and maintained by the automated operation and maintenance tool until the operation and maintenance work of the target software system is completed.

[0321] It should be noted that, in the embodiment of the device of the present invention, the information interaction, execution process and other contents between the above modules are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the aforementioned method embodiment part and will not be repeated here.

[0322] Based on the above method embodiment, another embodiment of the present invention further provides a computer device, which can be a server, and its internal structure diagram can be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the functions or steps on the server side of the intelligent operation and maintenance method in any of the above method embodiments are implemented.

[0323] Based on the above method embodiment, another embodiment of the present invention further provides a computer device, which can be a client, and its internal structure diagram can be as follows: Figure 5As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the functions or steps of the client side of the intelligent operation and maintenance method in any of the above method embodiments are implemented.

[0324] Those skilled in the art will understand that Figure 4 and Figure 5 The structural diagram shown in the figure is only a schematic diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more components than shown in the figure, or combine certain components, or have a different component arrangement.

[0325] The processor referred to herein may be a CPU, other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or any conventional processor, etc.

[0326] The memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of the operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be a hard disk of the computer device, and in other embodiments, it can also be an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device. Furthermore, the memory can also include both an internal storage unit of the computer device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program. The memory can also be used to temporarily store data that has been output or is about to be output.

[0327] Based on the above method embodiments, another embodiment of the present invention further provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the intelligent operation and maintenance method described in any of the above method embodiments. The computer-readable storage medium can be either non-volatile or volatile.

[0328] It should be noted that the above-mentioned functions or steps that can be implemented by computer-readable storage media or computer devices, and the technical effects brought about by the functions / steps, can be found in the relevant descriptions in the aforementioned method embodiments. To avoid repetition, they will not be described one by one here.

[0329] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM). The disclosed memory components or memories of the operating environments described herein are intended to comprise one or more of these and / or any other suitable types of memory.

[0330] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, in the embodiment of the device of the present invention, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual application, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the above-mentioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0331] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0332] In the embodiments provided by the present invention, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0333] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0334] It should be noted that if software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. An intelligent operation and maintenance method, characterized in that: include: Build an operation and maintenance knowledge base; Acquire the target software system's operating data in real time and analyze whether the operating data contains fault characteristics using a pre-trained large language model; When the fault feature is detected, a diagnosis result corresponding to the fault feature is generated by the large language model; According to the diagnosis result, a corresponding target operation and maintenance plan is retrieved from the operation and maintenance knowledge base, and operation and maintenance processing is performed on the target software system according to the target operation and maintenance plan.

2. The intelligent operation and maintenance method according to claim 1, characterized in that: The construction of the operation and maintenance knowledge base includes: Collect operation and maintenance data from historical operation and maintenance information of several software systems; Pre-processing the collected operation and maintenance data by cleaning, deduplication, and formatting; The pre-processed operation and maintenance data is classified, and the classified operation and maintenance data is associated using knowledge graph technology to form an operation and maintenance knowledge base.

3. The intelligent operation and maintenance method according to claim 1, characterized in that: The real-time acquisition of the target software system's operating data and the analysis of whether the operating data contains fault characteristics using a pre-trained large language model include: Obtaining real-time operating data of the target software system through the API interface of the target software system; Preprocessing the operating data by cleaning, standardizing and reducing its dimension; Performing feature extraction on the preprocessed operating data to obtain a feature vector; The extracted feature vector is input into a pre-trained large language model for semantic analysis to identify whether the operating data contains fault features.

4. The intelligent operation and maintenance method according to claim 1, characterized in that: When the fault feature is detected, generating a diagnosis result corresponding to the fault feature by using the large language model includes: When the large language model detects that the operating data contains the fault feature, performing type recognition on the fault feature to determine the fault type of the fault feature; According to the fault characteristics and fault types, and in combination with a context-aware algorithm, a diagnosis result corresponding to the fault characteristics is generated through the large language model.

5. The intelligent operation and maintenance method according to claim 1, characterized in that: The step of retrieving a corresponding target operation and maintenance solution from the operation and maintenance knowledge base based on the diagnosis result, and performing operation and maintenance processing on the target software system according to the target operation and maintenance solution, includes: Analyzing the diagnosis results to determine the impact of the fault characteristics; When the impact of the fault feature exceeds a preset threshold, a target operation and maintenance solution matching the fault feature is retrieved from the operation and maintenance knowledge base; According to the target operation and maintenance plan, the target software system is operated and maintained through automated operation and maintenance tools.

6. The intelligent operation and maintenance method according to claim 5, characterized in that: The analyzing the diagnosis result to determine the impact degree of the fault characteristics includes: Constructing an impact assessment model; Normalizing the diagnosis results so that the data format of the diagnosis results meets the input requirements of the impact assessment model; The normalized diagnosis result is input into the impact assessment model, and the impact degree of the fault feature is determined according to the output result of the impact assessment model.

7. The intelligent operation and maintenance method according to claim 5, characterized in that: After performing operation and maintenance processing on the target software system by using an automated operation and maintenance tool according to the target operation and maintenance plan, the method further includes: During the operation and maintenance process of the target software system, real-time monitoring of the operation and maintenance effect of the target operation and maintenance solution; Dynamically adjust the target operation and maintenance plan according to the monitored operation and maintenance effect; Based on the adjusted target operation and maintenance plan, the target software system is operated and maintained by the automated operation and maintenance tool until the operation and maintenance work of the target software system is completed.

8. An intelligent operation and maintenance device, characterized in that: include: Construction module, used to build the operation and maintenance knowledge base; An acquisition module is used to acquire the operating data of the target software system in real time and analyze whether the operating data contains fault characteristics through a pre-trained large language model; a generating module, configured to generate a diagnosis result corresponding to the fault feature by using the large language model when the fault feature is detected; The operation and maintenance module is used to retrieve a corresponding target operation and maintenance plan from the operation and maintenance knowledge base according to the diagnosis result, and perform operation and maintenance processing on the target software system according to the target operation and maintenance plan.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the intelligent operation and maintenance method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the intelligent operation and maintenance method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Distributed database operation and maintenance method and device and related equipment

    CN121350004A