Method and device for operating and maintaining web service and related equipment
By using a large language model to perform semantic understanding of multi-source data and generate operation and maintenance strategies, the problem of existing tools relying on manual operations is solved, intelligent operation and maintenance is realized, and efficiency and stability are improved.
Patent Information
- Application Number
- CN202510829207.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-03
AI Technical Summary
Existing operation and maintenance automation tools are highly dependent on manual judgment and operation in key links such as exception handling, problem diagnosis and emergency response. They lack flexibility and intelligence and are difficult to adapt to the dynamics and unpredictability of complex systems, resulting in low efficiency and high error rates.
By collecting multi-source operation data and using pre-trained large language models for semantic understanding, we can perform anomaly detection, root cause analysis, and trend prediction, generate operation and maintenance strategies, and optimize model parameters through machine learning to achieve intelligent operation and maintenance management.
It improves operation and maintenance efficiency, reduces labor costs, reduces error rates, and enhances the stability and adaptability of web services.
Smart Images

Figure CN120743598A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and in particular to a method, apparatus, and related equipment for operating and maintaining web services. Background Art
[0002] In the modern information technology sector, operations and maintenance (O&M) are crucial for ensuring system stability and high availability. As the IT infrastructure of enterprises and organizations grows increasingly complex, the O&M workload becomes immense, demanding, and repetitive, placing increasing demands on the professional skills and experience of O&M personnel. While existing O&M automation tools can partially alleviate O&M pressure, they still rely heavily on manual judgment and manipulation in key areas such as exception handling, problem diagnosis, and emergency response. Furthermore, traditional O&M automation tools often rely on fixed rules and preset scripts, which lacks flexibility and intelligence, making them difficult to adapt to rapidly changing O&M scenarios and unknown failure modes. This rigid automation model is unable to cope with the inherent dynamism and unpredictability of modern complex systems, resulting in low efficiency, high costs, and difficulty in reducing error rates. Therefore, developing an O&M approach that overcomes these limitations of existing O&M automation tools and achieves a truly intelligent, adaptive, efficient, and low-error O&M approach, thereby significantly improving the stability and adaptability of web services, has become a critical and pressing issue within the industry. Summary of the Invention
[0003] In view of this, embodiments of the present application provide a method, apparatus, and related equipment for operating and maintaining a web service to at least or partially solve the above-mentioned problems.
[0004] In a first aspect, an embodiment of the present application provides a method for operating and maintaining a web service, comprising:
[0005] Collecting multi-source operation data from a web service system and preprocessing the multi-source data, the multi-source operation data including at least one of log data, performance indicators, and alarm information of the web service system;
[0006] Based on the pre-trained large language model, semantic understanding is performed on the pre-processed data to perform at least one data analysis service of anomaly detection, root cause analysis, and trend prediction according to the results of the semantic understanding;
[0007] Generate corresponding operation and maintenance strategy information based on the result of the data analysis service, wherein the operation and maintenance strategy information includes at least one of suggested operation information and automated execution operation information;
[0008] Collect operation and maintenance result data of the operation and maintenance management of the web service according to the operation and maintenance operation strategy information, so as to adjust the parameters of the pre-trained large language model or generate a new response strategy according to the operation and maintenance result data for subsequent operation and maintenance management of the web service.
[0009] Optionally, in an embodiment of the present application, the anomaly detection process includes:
[0010] Utilizing the natural language understanding capability of the large language model, extracting key operational features from the preprocessed multi-source operational data, wherein the key operational features include one or more of abnormal keywords, abnormal state description information, abnormal operational mode information, and operational time series feature information;
[0011] Determining the result of the anomaly detection according to the key features of the anomaly;
[0012] Optionally, in one embodiment of the present application, the root cause analysis process includes:
[0013] Parsing error information and error type information from the preprocessed multi-source data using the large language model; determining an error frequency and an error time corresponding to the error information;
[0014] Based on the error information, error type information, error frequency and error time, combined with the system log context relationship corresponding to the error information, the potential causal relationship of multiple events in the system operation is determined; based on the potential causal relationship and system historical data, the cause of the system operation and maintenance abnormality is determined.
[0015] Optionally, in one embodiment of the present application, the trend prediction process includes:
[0016] Analyzing the operation time series feature information and system event-driven data extracted from the pre-processed multi-source operation data using the large language model and machine learning algorithm;
[0017] Based on the analysis results, external factors are integrated to predict system resource usage trends and potential load change information, and early warning is issued based on the resource usage trends and potential load change information;
[0018] The external factors include at least one of seasonal traffic fluctuation information and traffic fluctuation information caused by activities on a specific date.
[0019] Optionally, in an embodiment of the present application, the method further includes:
[0020] Obtain historical traffic operation mode information and system resource consumption information of web services;
[0021] performing correlation analysis on the historical traffic pattern information and the system resource consumption information using the large language model;
[0022] According to the result of the correlation analysis, the change of the system resource demand is predicted, and an early warning is issued according to the result of the change prediction.
[0023] Optionally, in one embodiment of the present application, the collecting, by a machine learning method, operation and maintenance result data of the operation and maintenance management of the web service according to the operation and maintenance operation strategy information, so as to adjust the parameters of the pre-trained large language model or generate a new response strategy according to the validity of the operation and maintenance result data, includes:
[0024] Collecting the operation and maintenance result data through machine learning methods;
[0025] Analyzing the success rate and execution effect of the operation and maintenance operations contained in the operation and maintenance result data;
[0026] According to the success rate and execution effect of the operation and maintenance operation, the parameters of the large language model are iteratively adjusted or a new response strategy is generated.
[0027] Optionally, in an embodiment of the present application, the method further comprises: receiving a marking operation of the operation and maintenance result by a user, so as to determine the validity of the operation and maintenance result data according to the marking operation;
[0028] According to the marking operation, the parameters of the large language model are optimized and adjusted or a new response strategy is generated.
[0029] Optionally, in an embodiment of the present application, the method further includes:
[0030] Build a natural language interactive web interface using a large language model to respond to natural language instructions input by users and generate corresponding operation and maintenance response instructions for response execution.
[0031] In a second aspect, based on the method for operating and maintaining a web service described in the first aspect of the present application, an embodiment of the present application further provides an apparatus for operating and maintaining a web service, comprising:
[0032] a collection module for collecting multi-source operation data from the web service system and preprocessing the multi-source data, wherein the multi-source operation data includes at least one of log data, performance indicators and alarm information of the web service system;
[0033] an analysis module for performing semantic understanding on the preprocessed data based on a pretrained large language model, and performing at least one data analysis service of anomaly detection, root cause analysis, and trend prediction based on the results of the semantic understanding;
[0034] An operation and maintenance module, configured to generate corresponding operation and maintenance strategy information based on the result of the data analysis service, for performing operation and maintenance management on the web service; wherein the operation and maintenance strategy information includes at least one of suggested operation information and automated execution operation information;
[0035] An adjustment module is used to collect operation and maintenance result data of the operation and maintenance management of the web service according to the operation and maintenance operation strategy information, so as to adjust the parameters of the pre-trained large language model or generate a new response strategy according to the operation and maintenance result data for subsequent operation and maintenance management of the web service.
[0036] In a third aspect, an embodiment of the present application further provides a computer storage medium having computer executable instructions stored thereon, which, when executed, can perform any one of the methods for operating and maintaining a web service as described in the first aspect of the embodiment of the present application.
[0037] In a fourth aspect, an embodiment of the present application further provides an electronic device, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;
[0038] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to execute any one of the methods for operating and maintaining a web service as described in the first aspect of the embodiment of the present application.
[0039] The present application provides a method, apparatus, and related equipment for operating and maintaining a web service. The method collects multi-source operating data from a web service system and pre-processes the multi-source data, wherein the multi-source operating data includes at least one of log data, performance indicators, and alarm information of the web service system. Based on a pre-trained large language model, the pre-processed data is semantically understood to perform at least one data analysis service of anomaly detection, root cause analysis, and trend prediction based on the semantic understanding results. Based on the results of the data analysis services, corresponding operation and maintenance operation strategy information is generated, wherein the operation and maintenance operation strategy information includes at least one of recommended operation information and automated execution operation information. Operation and maintenance result data of the web service operation and maintenance management performed based on the operation and maintenance operation strategy information is collected, and parameters of the pre-trained large language model are adjusted or new response strategies are generated based on the operation and maintenance result data for subsequent operation and maintenance management of the web service. The method fully utilizes the powerful text understanding and generation capabilities of the large language model to achieve intelligent operation and maintenance automation, thereby improving operation and maintenance efficiency, reducing labor costs, reducing error rates, and enhancing the stability and adaptability of web services. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0041] Figure 1 A schematic diagram of a workflow of a method for operating and maintaining a web service provided in an embodiment of the present application;
[0042] Figure 2 A schematic diagram of the structure of a device for operating and maintaining web services provided in an embodiment of the present application.
[0043] Figure 3 A structural diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.
[0045] It should be understood that the various steps described in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.
[0046] Example 1
[0047] The present invention provides a method for operating and maintaining a web service. Figure 1 As shown, Figure 1 A workflow diagram of a method for operating and maintaining a web service provided in an embodiment of the present application includes:
[0048] Step S101: Collect multi-source operation data from a web service system and pre-process the multi-source data, wherein the multi-source operation data includes at least one of log data, performance indicators, and alarm information of the web service system.
[0049] In this stage of the embodiment of the present application, multi-source data is collected as the basis for operation and maintenance to ensure the accuracy of the generated targeted operation and maintenance strategy. The collection process includes real-time collection of raw operation and maintenance data from various web systems. The collected multi-source data includes but is not limited to diversified data such as system logs, application logs, network traffic data, CPU usage, memory usage, disk I / O, response time, error codes, and various alarm information, so that the collected data has a high abundance. Preprocessing the collected data is the key to determining the quality of subsequent large language model analysis. Specifically, the preprocessing includes data cleaning of these massive amounts of heterogeneous data to remove noise and redundant information, normalization processing, such as unifying data format, time error information, and structuring processing to convert unstructured data into structured data that can be quickly understood. For example, the collected raw log data may contain a large amount of unstructured text. Through preprocessing, key events, timestamps, error levels, and other information are extracted and structured to ensure the integrity, consistency, and high quality of the data, providing a reliable foundation for the subsequent in-depth analysis of the large language model. This strict control of data quality is a prerequisite for the large language model to exert its powerful semantic understanding and pattern recognition capabilities. It can also ensure the comprehensiveness and robustness of information processing during the operation and maintenance of the system designed according to this solution.
[0050] Step S102: Based on the pre-trained large language model, semantic understanding is performed on the pre-processed data to perform at least one data analysis service of anomaly detection, root cause analysis, and trend prediction according to the results of the semantic understanding.
[0051] In an embodiment of the present application, a pre-trained large language model is used to perform semantic understanding on the pre-processed data, which can efficiently extract key information for operation and maintenance management from the collected multi-source unstructured text, and then quickly and accurately perform analysis services such as anomaly detection, root cause analysis of faults, or working trend prediction of system working software and hardware based on this key information. This analysis capability far exceeds the traditional analysis capability based on preset rules or regular expression matching, realizes efficient and intelligent perception of operation and maintenance events, and can better understand the implicit meaning and contextual relationship in the collected multi-source data such as log entries, rather than just the technical effect that can be achieved by literal matching of a certain data.
[0052] Optionally, in one embodiment of the present application, the anomaly detection process includes: utilizing the natural language understanding capabilities of the large language model to extract key operational features from the preprocessed multi-source operational data, wherein the key operational features include one or more of abnormal keywords, abnormal state description information, abnormal operational mode information, and operational time series feature information; and determining the anomaly detection result based on the key abnormal features. This embodiment of the present application is illustrative, using extracted log data as an example. The large language model can accurately extract keywords or abnormal state descriptions, such as "error," "failure," "timeout," and "unable to connect," from the text of the log data. The large language model can understand the context of these keywords and determine whether they truly represent an anomaly, thereby avoiding false positives. Identifying abnormal operational mode information refers to utilizing the large language model to identify abnormal patterns that occur in the system within a specific time period based on historical data and known operational anomalies. For example, when the key operational feature of "high frequency of database connection failures" is identified in the log data, it can be determined that the anomaly is caused by an abnormal operational mode, such as a full connection pool or an overloaded service system database service. The running time series feature information refers to the feature information of a certain hardware or service being used intensively in a certain period of time in the system. For example, if the CPU processing resource utilization rate is higher than 95% for more than a certain time (such as 10 minutes), it is determined that the system has a high service occupancy anomaly. Of course, the embodiment of the present application is only illustrative here to illustrate it, and does not mean that the present application is limited to this. For example, a large language model can also be used to combine historical log data for advanced pattern matching. For example, in the collected historical data, a certain type of database connection timeout error occurs frequently under high load. The large language model can apply this pattern to the new log to determine whether there is a similar problem in the current log and perform association analysis, thereby achieving more accurate abnormal pattern recognition. The abnormal detection method defined in the embodiment of the present application is different from the traditional fixed rule or threshold detection method. The large language model can combine contextual information, historical background and semantic association for intelligent judgment, thereby identifying hidden, potential or new abnormalities, including those abnormal events that do not conform to the preset rules. Thus, it is possible to discover subtle or complex abnormalities that are difficult to capture by traditional methods, significantly improving the accuracy and coverage of early warning and targeted operation and maintenance management for abnormal events.
[0053] Optionally, in one embodiment of the present application, the root cause analysis process includes: using the large language model to parse error information and error type information from the preprocessed multi-source data; determining the error frequency and error time corresponding to the error information; based on the error information, error type information, error frequency and error time, combined with the system log context relationship corresponding to the error information, determining the potential causal relationship of multiple events in the system operation, and based on the potential causal relationship and system historical data, determining the cause of the system operation and maintenance anomaly. Specifically, in an embodiment of the present application, the large language model can friendly parse the error information in the log, accurately determine the error type (such as connection timeout, service crash, memory overflow, etc.), and quickly analyze the frequency of the error, the precise time and the load status information of the system when the error occurs. The large language model can deeply analyze the contextual relationship of multiple log item information, events and monitoring indicators in the preprocessed multi-source data based on the parsed error information, thereby accurately identifying the potential causal relationship between multiple system operation events. For example, when a service crashes, the large language model can analyze whether there is a significant fluctuation in the CPU usage or memory consumption of the system during the same period, helping to determine whether the problem is related to excessive resource consumption. In the embodiment of the present application, the purpose of root cause analysis is to find the root cause of the anomaly. When dealing with this problem, the large language model can determine the correlation between multi-source data based on the multiple error attribute information about the anomaly parsed from the multi-source data. Thereby inferring the root cause of the potential anomaly. Especially in the complex and changeable environment of system logs, the model can help identify the causal chain hidden behind the complex data and difficult to discover through simple rules. This capability enables the system to perform in-depth fault diagnosis, thereby significantly shortening the troubleshooting time.
[0054] In addition, in an optional implementation of an embodiment of the present application, the root cause analysis process may also include: using a large language model to learn error information related to anomalies in historical data, and identifying common abnormal patterns and their corresponding causes. For example, when a certain fault in the system often occurs in the following situations: under high load conditions, the database connection pool is exhausted; when memory consumption is too high, the application crashes. The large language model is based on learning error information of similar historical working modes. When it again identifies similar high-load and application crash logs from the extracted key information, it can quickly and efficiently identify the root cause of the anomaly, thereby greatly improving the efficiency and robustness of the operation and maintenance management of the system during the implementation of the method described in this solution.
[0055] Optionally, in one embodiment of the present application, the trend prediction process includes: using the large language model and machine learning algorithm to analyze the operating time series feature information and system event-driven data extracted from the pre-processed multi-source operating data; based on the analysis results, integrating external factors to predict system resource usage trends and potential load change information, and providing early warnings based on the resource usage trends and potential load change information; wherein the external factors include at least one of seasonal traffic fluctuation information and traffic fluctuation information caused by activities on specific dates. In the embodiment of the present application, the operating time series feature information includes the usage of key resources such as CPU, memory, disk, and network bandwidth, and is predicted in combination with the time dimension. For example, using a machine learning algorithm, by extracting rules and patterns from a large amount of data, the model can make predictions or decisions without explicit programming. This method enables the large language model to quickly and accurately learn about the changing trends of the CPU usage and memory usage of the deployed system, thereby predicting under what circumstances the system will reach a high load critical point or exhaustion in the future. Analysis based on system event-driven data refers to making predictions based on specific events (such as service restarts, traffic peaks, deployment changes, and other runtimes). For example, analyze whether the response time of certain services has increased dramatically during peak traffic times in the past, and predict the impact of similar events in the future. At the same time, the large language model also integrates external factors for prediction at this stage to obtain more accurate and intelligent trend analysis results. For example, during holidays or certain promotional activities, the traffic of the web service system may surge. The use of a large language model can predict the growth trend of this service demand in advance and provide suggestions for early expansion. Through the correlation analysis of historical traffic patterns and resource consumption, the model can predict changes in the system's demand for software and hardware resources in different time dimensions, and remind whether additional resources need to be configured in advance to avoid performance bottlenecks in the system. The present application here limits the use of a large language model to integrate multi-source data for trend prediction, which can make the results of trend prediction more comprehensive and reliable, and can also transform the operation and maintenance method from traditional passive response to active prevention.
[0056] Optionally, in one implementation of the embodiment of the present application, the trend prediction process further includes: obtaining historical traffic operation pattern information and system resource consumption information of the web service, performing correlation analysis on the historical traffic pattern information and the system resource consumption information using the large language model, predicting changes in system resource demand based on the results of the correlation analysis, and issuing early warnings based on the results of the change prediction. This enables efficient and accurate anomaly identification and trend prediction.
[0057] Step S103: Generate corresponding operation and maintenance strategy information based on the results of the data analysis service, and the operation and maintenance strategy information includes at least one of advisory operation information and automated execution operation information. In the embodiment of the present application, a large language model is used to generate highly targeted operation and maintenance strategy information as a response operation based on the anomaly detection results, root cause analysis results, and trend prediction results of the analysis service, and the response specifically includes advisory operation information and automated execution operation information, so that the method described in the embodiment of the present application has a better balance of efficiency and security when applied. Specifically, for example, by distinguishing abnormal faults corresponding to different risk levels and generating corresponding operation and maintenance strategy information, an intelligent hierarchical response is achieved, so that the method described in the embodiment of the present application has better operation and maintenance flexibility.
[0058] Specifically, in the embodiment of the present application, the recommended operation refers to an operation involving key system functions, high risks, or operations that require final manual confirmation, which can be notified to the operation and maintenance personnel in the form of suggestions through notification systems such as web interfaces, emails, or instant messaging tools (such as WeChat, DingTalk). The operation and maintenance personnel can decide whether to execute based on the suggestions and their own judgment. For example, it is recommended to expand capacity, upgrade the software version, adjust the size of the database connection pool, etc. This "human-computer collaboration" model ensures that the final control of humans is retained in key decisions, thereby improving the security of the system and the trust of the operation and maintenance personnel. The automated execution operation is for low-risk, highly repeatable and verified safe operations, which can be directly automated according to preset trigger conditions and execution logic. For example. Restarting unresponsive services, cleaning up disk space, adjusting non-critical configuration parameters, automatic expansion and contraction (within the scope of security policies), etc. The automated execution of these operations significantly improves the efficiency of operation and maintenance and reduces potential errors caused by manual intervention.
[0059] Furthermore, in an optional implementation of the embodiments of the present application, the operation and maintenance operation policy information used to respond to identified anomalies is configured with one or more of its specific trigger conditions, execution logic, execution priority, and rollback strategy, which can be customized by the operation and maintenance personnel. This high degree of configurability ensures that the method described in the embodiments can flexibly adapt to the various business requirements, security policies, and operation and maintenance practices of different systems, so that the response method of the generated operation and maintenance operation policy information can be dynamically adjusted as the system environment and business requirements change, thereby maintaining its long-term effectiveness.
[0060] Step S104: Collecting operation and maintenance result data for the web service based on the operation and maintenance strategy information, and adjusting the parameters of the pre-trained large language model or generating a new response strategy based on the operation and maintenance result data for subsequent operation and maintenance management of the web service. This application optimizes the response strategy and model performance of the large language model in this manner, thereby forming a data-driven closed-loop optimization, further improving the accuracy of methods such as anomaly detection, root cause analysis, and trend prediction.
[0061] Optionally, in an embodiment of the present application, the operation and maintenance result data of the operation and maintenance management of the web service according to the operation and maintenance operation strategy information is collected by a machine learning method to adjust the parameters of the pre-trained large language model or generate a new response strategy according to the validity of the operation and maintenance result data, including: collecting the operation and maintenance result data by a machine learning method; analyzing the operation and maintenance operation success rate and execution effect contained in the operation and maintenance result data; iteratively adjusting the parameters of the large language model or generating a new response strategy according to the operation and maintenance operation success rate and execution effect. The embodiment of the present application further introduces a machine learning method at this stage, so that the solution can achieve the intelligence and adaptability of the system after implementation. By integrating machine learning methods (such as reinforcement learning, supervised learning, etc.), the response strategy and model performance of the large language model can be optimized efficiently and continuously. The large language model can automatically generate or optimize new response strategies and strategies based on the learned new patterns and best practices, so that the prevention described in the embodiment of the present application can continuously evolve and adapt to the ever-changing IT environment and new operation and maintenance conditions, effectively improving the ability of the large language model to face various new abnormal problems that may arise in a dynamic environment, and ensuring the long-term stability of the system.
[0062] Optionally, in one embodiment of the present application, the method further includes: receiving a marking operation of the user on the operation and maintenance result to determine the validity of the operation and maintenance result data based on the marking operation; and optimizing and adjusting the parameters of the large language model or generating a new response strategy based on the marking operation. The embodiment of the present application defines that the method also supports a feedback mechanism for operation and maintenance personnel to mark the effectiveness of the operation, so that the method described in this solution can be further adjusted and optimized accordingly. It also improves the flexibility of the method described in the embodiment of the present application.
[0063] Optionally, in one embodiment of the present application, the method further includes: constructing a natural language interactive web interface using a large language model, which is used to respond to natural language instructions input by the user and generate corresponding operation and maintenance response instructions for response execution. The embodiment of the present application is limited to human-computer interaction based on a large prediction model, thereby significantly improving the efficiency, convenience and intuitiveness of the interaction between operation and maintenance personnel and the system, and lowering the operational threshold. Moreover, this interactive method is very intuitive, making complex operation and maintenance operations simple and easy to understand, and improving the decision-making speed and efficiency of operation and maintenance personnel.
[0064] The present application provides a method for operating and maintaining a web service, comprising collecting multi-source operational data from a web service system and preprocessing the multi-source data, the multi-source operational data including at least one of log data, performance indicators, and alarm information of the web service system; performing semantic understanding on the preprocessed data based on a pre-trained large language model, and performing at least one data analysis service of anomaly detection, root cause analysis, and trend prediction based on the results of the semantic understanding; generating corresponding operational strategy information based on the results of the data analysis service, the operational strategy information including at least one of recommended operational information and automated execution operational information; collecting operational results data of the web service operation and maintenance management based on the operational strategy information, and adjusting parameters of the pre-trained large language model or generating a new response strategy based on the operational results data for subsequent operational management of the web service. The method can fully utilize the powerful text understanding and generation capabilities of the large language model to achieve intelligent operational automation, thereby improving operational efficiency, reducing labor costs, reducing error rates, and enhancing the stability and adaptability of the web service.
[0065] Example 2
[0066] Based on the method for operating and maintaining a web service provided in the first embodiment of the present application, the embodiment of the present application further provides a device for operating and maintaining a web service, such as Figure 2 As shown, Figure 2 This is a schematic diagram of the structure of an apparatus 20 for operating and maintaining a web service provided in an embodiment of the present application. The apparatus 20 for operating and maintaining a web service includes:
[0067] A collection module 201 is configured to collect multi-source operation data from a web service system and pre-process the multi-source data, wherein the multi-source operation data includes at least one of log data, performance indicators, and alarm information of the web service system;
[0068] An analysis module 202 is configured to perform semantic understanding on the preprocessed data based on a pretrained large language model, and to perform at least one data analysis service of anomaly detection, root cause analysis, and trend prediction based on the results of the semantic understanding;
[0069] The operation and maintenance module 203 is configured to generate corresponding operation and maintenance operation strategy information based on the result of the data analysis service, for performing operation and maintenance management on the web service; wherein the operation and maintenance operation strategy information includes at least one of recommended operation information and automated execution operation information;
[0070] The adjustment module 204 is used to collect the operation and maintenance result data of the operation and maintenance management of the web service according to the operation and maintenance operation strategy information, so as to adjust the parameters of the pre-trained large language model or generate a new response strategy according to the operation and maintenance result data for subsequent operation and maintenance management of the web service.
[0071] Optionally, in an embodiment of the present application, the anomaly detection process includes:
[0072] Utilizing the natural language understanding capability of the large language model, extracting key operational features from the preprocessed multi-source operational data, wherein the key operational features include one or more of abnormal keywords, abnormal state description information, abnormal operational mode information, and operational time series feature information;
[0073] Determining the result of the anomaly detection according to the key features of the anomaly;
[0074] Optionally, in one embodiment of the present application, the root cause analysis process includes:
[0075] Parsing error information and error type information from the preprocessed multi-source data using the large language model; determining an error frequency and an error time corresponding to the error information;
[0076] Based on the error information, error type information, error frequency and error time, combined with the system log context relationship corresponding to the error information, the potential causal relationship of multiple events in the system operation is determined; based on the potential causal relationship and system historical data, the cause of the system operation and maintenance abnormality is determined.
[0077] Optionally, in one embodiment of the present application, the trend prediction process includes:
[0078] Analyzing the operation time series feature information and system event-driven data extracted from the pre-processed multi-source operation data using the large language model and machine learning algorithm;
[0079] Based on the analysis results, external factors are integrated to predict system resource usage trends and potential load change information, and early warning is issued based on the resource usage trends and potential load change information;
[0080] The external factors include at least one of seasonal traffic fluctuation information and traffic fluctuation information caused by activities on a specific date.
[0081] Optionally, in one embodiment of the present application, the device 20 further includes a prediction module (not shown in the accompanying drawings), which is used to obtain historical traffic operation mode information and system resource consumption information of the web service; use the large language model to perform correlation analysis on the historical traffic pattern information and the system resource consumption information; predict changes in system resource requirements based on the results of the correlation analysis, and issue early warnings based on the results of the change predictions.
[0082] Optionally, in one embodiment of the present application, the adjustment module 204 is also used to: collect the operation and maintenance result data through a machine learning method; analyze the operation and maintenance operation success rate and execution effect contained in the operation and maintenance result data; and iteratively adjust the parameters of the large language model or generate a new response strategy based on the operation and maintenance operation success rate and execution effect.
[0083] Optionally, in one embodiment of the present application, the device 20 further includes an optimization module (not shown in the drawings), which is used to receive a user's marking operation on the operation and maintenance result, so as to determine the validity of the operation and maintenance result data based on the marking operation; and according to the marking operation, optimize and adjust the parameters of the large language model or generate a new response strategy.
[0084] Optionally, in one embodiment of the present application, the device 20 also includes an interaction module (not shown in the drawings), which is used to build a natural language interactive web interface using a large language model, and is used to respond to natural language instructions input by the user and generate corresponding operation and maintenance response instructions for response execution.
[0085] The present application provides a device for operating and maintaining a web service, comprising a collection module for collecting multi-source operation data from a web service system and pre-processing the multi-source data, wherein the multi-source operation data includes at least one of log data, performance indicators, and alarm information of the web service system; an analysis module for performing semantic understanding on the pre-processed data based on a pre-trained large language model, and performing at least one data analysis service of anomaly detection, root cause analysis, and trend prediction based on the results of the semantic understanding; an operation and maintenance module for generating corresponding operation and maintenance operation strategy information based on the results of the data analysis service, wherein the operation and maintenance operation strategy information includes at least one of recommended operation information and automated execution operation information; and an adjustment module for collecting operation and maintenance result data of the operation and maintenance management of the web service based on the operation and maintenance operation strategy information, and adjusting the parameters of the pre-trained large language model or generating a new response strategy based on the operation and maintenance result data for subsequent operation and maintenance management of the web service. The device has a simple structure and can fully utilize the powerful text understanding and generation capabilities of the large language model to achieve intelligent operation and maintenance automation, thereby improving operation and maintenance efficiency, reducing labor costs, reducing error rates, and enhancing the stability and adaptability of the web service.
[0086] Example 3:
[0087] An embodiment of the present application further provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements any one of the methods for operating and maintaining a web service as described in the first embodiment of the present application.
[0088] Example 4:
[0089] The present application also provides an electronic device, such as Figure 3 As shown, Figure 3 An embodiment of the present application provides a schematic structural diagram of an electronic device 30, which includes:
[0090] One or more processors 301, a communication interface 302, a memory 303 and a communication bus 304, wherein the processor 301, the memory 303 and the communication interface 302 communicate with each other via the communication bus 304;
[0091] Memory 303, used to store one or more programs;
[0092] When the one or more programs are executed by the one or more processors 301, the one or more processors 301 implement any one of the methods for operating and maintaining a web service as described in the first embodiment of the present application.
[0093] Thus far, this application has described specific embodiments of the present subject matter. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing can be advantageous.
[0094] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system layer onto a PLD through their own programming, eliminating the need for chip manufacturers to design and manufacture dedicated integrated circuit chips. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0095] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0096] The system layers, devices, modules, or units described in the above embodiments may be implemented by computer chips or physical devices, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0097] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0098] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0099] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, system layers, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0100] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0101] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system-level embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0102] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for operating and maintaining a web service, characterized in that: include: Collecting multi-source operation data from a web service system and preprocessing the multi-source data, the multi-source operation data including at least one of log data, performance indicators, and alarm information of the web service system; Based on the pre-trained large language model, semantic understanding is performed on the pre-processed data to perform at least one data analysis service of anomaly detection, root cause analysis, and trend prediction according to the results of the semantic understanding; Generate corresponding operation and maintenance strategy information based on the result of the data analysis service, wherein the operation and maintenance strategy information includes at least one of suggested operation information and automated execution operation information; Collect operation and maintenance result data of the operation and maintenance management of the web service according to the operation and maintenance operation strategy information, so as to adjust the parameters of the pre-trained large language model or generate a new response strategy according to the operation and maintenance result data for subsequent operation and maintenance management of the web service.
2. The method for operating and maintaining a web service according to claim 1, wherein: The anomaly detection process includes: Utilizing the natural language understanding capability of the large language model, extracting key operational features from the preprocessed multi-source operational data, wherein the key operational features include one or more of abnormal keywords, abnormal state description information, abnormal operational mode information, and operational time series feature information; The result of the anomaly detection is determined according to the key features of the anomaly.
3. The method for operating and maintaining a web service according to claim 1, wherein: The root cause analysis process includes: Parsing error information and error type information from the preprocessed multi-source data using the large language model; determining an error frequency and an error time corresponding to the error information; Based on the error information, error type information, error frequency and error time, combined with the system log context relationship corresponding to the error information, the potential causal relationship of multiple events in the system operation is determined; based on the potential causal relationship and system historical data, the cause of the system operation and maintenance abnormality is determined.
4. The method for operating and maintaining a web service according to claim 1, wherein: The process of trend prediction includes: Analyzing the operation time series feature information and system event-driven data extracted from the pre-processed multi-source operation data using the large language model and machine learning algorithm; Based on the analysis results, external factors are integrated to predict system resource usage trends and potential load change information, and early warning is issued based on the resource usage trends and potential load change information; The external factors include at least one of seasonal traffic fluctuation information and traffic fluctuation information caused by activities on a specific date.
5. The method for operating and maintaining a web service according to claim 4, wherein: The method further comprises: Obtain historical traffic operation mode information and system resource consumption information of web services; performing correlation analysis on the historical traffic pattern information and the system resource consumption information using the large language model; According to the result of the correlation analysis, the change of the system resource demand is predicted, and an early warning is issued according to the result of the change prediction.
6. The method for operating and maintaining a web service according to claim 1, wherein: The collecting of operation and maintenance result data of the operation and maintenance management of the web service according to the operation and maintenance operation strategy information, and adjusting the parameters of the pre-trained large language model or generating a new response strategy according to the operation and maintenance result data for subsequent operation and maintenance management of the web service, includes: Collecting the operation and maintenance result data through machine learning methods; Analyzing the success rate and execution effect of the operation and maintenance operations contained in the operation and maintenance result data; According to the success rate and execution effect of the operation and maintenance operation, the parameters of the large language model are iteratively adjusted or a new response strategy is generated.
7. The method for operating and maintaining a web service according to claim 6, wherein: The method further includes: receiving a marking operation of the operation and maintenance result by a user, and determining the validity of the operation and maintenance result data according to the marking operation; According to the marking operation, the parameters of the large language model are optimized and adjusted or a new response strategy is generated.
8. The method for operating and maintaining a web service according to claim 1, wherein: The method further comprises: Build a natural language interactive web interface using a large language model to respond to natural language instructions input by users and generate corresponding operation and maintenance response instructions for response execution.
9. A device for operating and maintaining a web service, characterized in that: include: a collection module for collecting multi-source operation data from the web service system and preprocessing the multi-source data, wherein the multi-source operation data includes at least one of log data, performance indicators and alarm information of the web service system; an analysis module for performing semantic understanding on the preprocessed data based on a pretrained large language model, and performing at least one data analysis service of anomaly detection, root cause analysis, and trend prediction based on the results of the semantic understanding; An operation and maintenance module, configured to generate corresponding operation and maintenance strategy information based on the result of the data analysis service, for performing operation and maintenance management on the web service; wherein the operation and maintenance strategy information includes at least one of suggested operation information and automated execution operation information; An adjustment module is used to collect operation and maintenance result data of the operation and maintenance management of the web service according to the operation and maintenance operation strategy information, so as to adjust the parameters of the pre-trained large language model or generate a new response strategy according to the operation and maintenance result data for subsequent operation and maintenance management of the web service.
10. A computer storage medium, characterized in that The computer storage medium stores computer-executable instructions, which, when executed, execute the method for operating and maintaining a web service according to any one of claims 1 to 8.
Citation Information
Cited By
Heterogeneous gateway management method, device and equipment based on large model agent and medium
CN121462436A