Service management method, device, equipment, medium and product

Through unified monitoring platform and chaos engineering model training, the monitoring data of the virtual cloud platform is integrated and the fault processing model is generated, which solves the problem of monitoring data integration in the virtual cloud platform, realizes efficient fault discovery and processing, and improves management efficiency.

CN119938448APending Publication Date: 2025-05-06EVERSEC BEIJING TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510032705.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively integrate monitoring data at all levels in the virtual cloud platform, resulting in inefficient management and difficulty in time to detect fault services and troubleshooting.

Method used

By introducing a unified monitoring platform, integrating monitoring data at all levels, and model training is carried out based on chaos engineering, candidate fault code models and self-healing function models are generated, and stored in the fault knowledge base, so as to achieve unified monitoring and fault handling of target application services of virtual cloud platform.

Benefits of technology

It realizes unified monitoring and data integration of target application services of the virtual cloud platform, timely discovers fault services and performs fault handling, and improves the management efficiency of application services in the virtual cloud platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938448A_ABST
    Figure CN119938448A_ABST
Patent Text Reader

Abstract

The invention discloses a service management method, device and equipment, a medium and a product. The method comprises the following steps: in response to a service monitoring request for a virtual cloud platform, performing monitoring and data acquisition on a target application service running on the virtual cloud platform to obtain target monitoring data; according to the target monitoring data, performing model training based on chaos engineering to obtain a candidate fault code model and a candidate self-healing function model, and storing the candidate fault code model and the candidate self-healing function model to a fault knowledge base; and under the condition of determining that the target application service has a fault according to the target monitoring data, performing fault processing on the target application service according to the fault type corresponding to the target monitoring data and the fault knowledge base. According to the technical scheme, monitoring and data integration can be carried out on the target application services on the virtual cloud platform in a unified mode, so that fault services are found in time, fault processing is carried out, and the management efficiency of the application services in the virtual cloud platform is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of virtual cloud platforms, and in particular to a service management method, device, equipment, medium and product. Background Art

[0002] With the continuous development of virtual cloud platform technology, a multi-level monitoring solution has been adopted to manage various application services running in the virtual cloud platform. However, the multi-level monitoring solution uses a variety of monitoring tools, which makes the centralized management and analysis of monitoring data complicated, and the monitoring data at each level is difficult to effectively integrate, which reduces the management efficiency of application services in the virtual cloud platform.

[0003] Therefore, how to uniformly monitor and integrate the target application services on the virtual cloud platform, so as to timely discover faulty services and handle them, and improve the management efficiency of application services in the virtual cloud platform, is a problem that needs to be solved urgently. Summary of the invention

[0004] The present invention provides a service management method, device, equipment, medium and product to uniformly monitor and integrate data of target application services on a virtual cloud platform, so as to timely discover faulty services and handle faults, thereby improving the management efficiency of application services in the virtual cloud platform.

[0005] According to one aspect of the present invention, there is provided a service management method, comprising:

[0006] In response to a service monitoring request for the virtual cloud platform, monitoring and data collection are performed on a target application service running on the virtual cloud platform to obtain target monitoring data;

[0007] According to the target monitoring data, model training is performed based on chaos engineering to obtain candidate fault code models and candidate self-healing function models and store them in the fault knowledge base;

[0008] When it is determined according to the target monitoring data that a target application service fails, the target application service is fault processed according to the fault type and the fault knowledge base corresponding to the target monitoring data.

[0009] According to another aspect of the present invention, there is provided a service management device, comprising:

[0010] A monitoring module, for responding to a service monitoring request for the virtual cloud platform, monitoring and collecting data on a target application service running on the virtual cloud platform to obtain target monitoring data;

[0011] A module is obtained, which is used to perform model training based on chaos engineering according to target monitoring data to obtain candidate fault code models and candidate self-healing function models and store them in a fault knowledge base;

[0012] The processing module is used to perform fault processing on the target application service according to the fault type and fault knowledge base corresponding to the target monitoring data when it is determined that the target application service has a fault according to the target monitoring data.

[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the service management method described in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the service management method described in any embodiment of the present invention when executed.

[0018] According to another aspect of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the service management method of any embodiment of the present invention is implemented.

[0019] The technical solution of the embodiment of the present invention monitors and collects data of the target application service running on the virtual cloud platform in response to the service monitoring request of the virtual cloud platform to obtain the target monitoring data; performs model training based on chaos engineering according to the target monitoring data to obtain the candidate fault code model and the candidate self-healing function model and store them in the fault knowledge base; when it is determined that the target application service fails according to the target monitoring data, the target application service is fault-handled according to the fault type and fault knowledge base corresponding to the target monitoring data. By uniformly monitoring and integrating the data of the target application service on the virtual cloud platform, the faulty service can be discovered and handled in time, thereby improving the management efficiency of the application service in the virtual cloud platform.

[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 is a flow chart of a service management method provided by Embodiment 1 of the present invention;

[0023] Figure 2 is a flow chart of a service management method provided by Embodiment 2 of the present invention;

[0024] Figure 3 is a structural block diagram of a service management device provided in Embodiment 3 of the present invention;

[0025] Figure 4 It is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first", "second", "target", "candidate", "alternative", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.

[0028] It should be noted that in the relevant technologies, a multi-level monitoring solution is often used to monitor and operate the virtual cloud platform, namely the infrastructure layer, the virtual layer and the application layer. However, there are the following shortcomings: 1. Management complexity caused by the diversification of monitoring tools: Due to the use of multiple monitoring tools, the centralized management and analysis of monitoring data becomes complicated, and relevant personnel need to spend a lot of time switching between different tools. 2. Insufficient data integration and analysis: It is difficult to effectively integrate monitoring data at all levels, and the lack of a unified data analysis platform makes it impossible to fully understand the system operation status and difficult to achieve early warning and prediction of potential problems. 3. Redundancy and false alarms of alarm information: Due to unreasonable alarm rule settings, it is easy to cause redundancy and false alarms of alarm information, which increases the workload of relevant personnel and may miss really important problems. 4. Limited degree of automation: Although automated operation and maintenance tools have been introduced, in actual applications, the degree of automation is still limited, and many tasks still require manual intervention, making it difficult to give full play to the advantages of automated operation and maintenance. 5. Lack of intelligent operation and maintenance: The current operation and maintenance solutions mainly rely on static rules and experience, lack the application of intelligent means (such as machine learning and artificial intelligence), and cannot achieve adaptive and self-healing intelligent operation and maintenance.

[0029] In response to the above problems, the present invention makes improvements in the following aspects: 1. Unified monitoring platform: Introduce a unified monitoring platform to integrate monitoring data at all levels, provide centralized data management and analysis capabilities, and reduce the complexity of operation and maintenance. 2. Intelligent alarm management: Optimize alarm rules, use machine learning algorithms to implement intelligent alarms, reduce false alarms and redundant alarms, and improve the accuracy and timeliness of alarms. 3. Improve the level of automation: Further promote and apply automated operation and maintenance tools to achieve a higher level of automated configuration, deployment and troubleshooting, and reduce manual intervention. 4. Introduce intelligent operation and maintenance technology: Use artificial intelligence and big data analysis technology to improve the intelligence level of operation and maintenance, achieve comprehensive perception and intelligent decision-making of the system's operating status, and gradually achieve the goal of adaptive and self-healing operation and maintenance. Through these improvement measures, the monitoring and operation and maintenance efficiency of the virtual cloud platform can be comprehensively improved, the operation and maintenance costs can be reduced, and the stability and reliability of the system can be improved. The specific implementation method will be introduced in detail in the subsequent embodiments.

[0030] Embodiment 1

[0031] Figure 1 is a flow chart of a service management method provided in the first embodiment of the present invention; this embodiment can be applied to the case where application services running in a virtual cloud platform are monitored and fault handling is performed when a fault occurs. The method can be executed by a service management device, which can be implemented in the form of hardware and / or software. The service management device can be configured in an electronic device and executed by a virtual cloud platform, such as Figure 1 As shown, the service management method includes:

[0032] S101. In response to a service monitoring request for a virtual cloud platform, monitor and collect data on a target application service running on the virtual cloud platform to obtain target monitoring data.

[0033] Among them, the virtual cloud platform is used to monitor and coordinate the application services and application systems it manages. A service monitoring request refers to a request to monitor the operation of application services managed by the virtual cloud platform. Optionally, the target application service includes a cluster container service and a virtual machine service; the cluster container service includes middleware services, and the virtual machine service includes bare metal system services. Container service is a general term. Products such as mail systems and OA (Office Automation) systems can all be deployed based on containers, which can be called container services at the technical level. Cluster container services can also include production line services.

[0034] Optionally, based on the service discovery interface pre-configured on the virtual cloud platform, application services running on the virtual cloud platform can be automatically discovered according to information such as resource scheduling, network allocation, and service operation status of the virtualization cluster, that is, the running application services can be determined. Furthermore, relevant personnel can pre-configure whether to enable monitoring and self-healing functions for different management application services managed on the virtual cloud platform, generate configuration information, and after responding to a service monitoring request for the virtual cloud platform and determining the running application services running on the virtual cloud platform, determine the application services with monitoring and self-healing functions enabled in the running application services as target application services according to the configuration information.

[0035] Optionally, the target application service running on the virtual cloud platform is monitored and data is collected to obtain target monitoring data, including: determining the target application service running on the virtual cloud platform, and using at least one monitoring process to perform offload monitoring on the target application service; in the process of performing offload monitoring on the target application service, data is collected based on the calling interface of each target application service to obtain target monitoring data.

[0036] One target application service corresponds to one monitoring process, and the number of target application services and monitoring processes is the same.

[0037] Exemplarily, the monitoring interval corresponding to each target application service can be determined, and based on the monitoring interval, the CPU (Central Processing Unit), memory, network, storage, and operating status of the target application service can be periodically detected and data collected by calling the interface to obtain target monitoring data.

[0038] Optionally, after obtaining the target monitoring data, the target monitoring data corresponding to the target application service may be stored in a local database.

[0039] S102. Perform model training based on chaos engineering according to target monitoring data to obtain candidate fault code models and candidate self-healing function models and store them in a fault knowledge base.

[0040] Chaos engineering is the science of deliberately injecting faults into a system to measure resilience. Like any scientific method, chaos engineering focuses on experiments / hypotheses and then compares the results to the control (steady state). A candidate fault code model refers to a model that outputs fault resolution solutions corresponding to different fault types, and a candidate self-healing function model refers to a model used to handle faults. The fault knowledge base can store candidate fault code models and candidate self-healing function models corresponding to different fault types. Candidate self-healing function models can be self-healing function models such as config, resource, net, state, and hardware.

[0041] Exemplarily, if the fault type is service configuration parameter error information, the corresponding fault code model may be config-error. Through the fault code model, the corresponding self-healing function model may be automatically called for processing, wherein the corresponding self-healing function model may be a config self-healing function model.

[0042] Optionally, the core of chaos engineering is feedback and adaptation. Model training is performed by injecting faults or anomalies during the training of AI (Artificial Intelligence) models. Specifically, the chaos engineering fault model automates model training and updating through the CI / CD (Continuous Integration / Continuous Deployment) process. The AI ​​big model update is based on the fact that after the chaos engineering simulation failure, the model can absorb new training data in real time through online learning and gradually update the weights. By retraining the model, new data (such as robustness data) containing the chaos engineering simulation results are input in batches to update the model. By inputting the target monitoring data into the trained AI big model, the candidate fault code model and candidate self-healing function model corresponding to the service failure in the target monitoring data can be obtained, so as to facilitate subsequent fault handling when a fault is detected.

[0043] S103: When it is determined according to the target monitoring data that a target application service fails, the target application service is fault processed according to the fault type and fault knowledge base corresponding to the target monitoring data.

[0044] Optionally, when it is determined that a target application service has failed based on the target monitoring data, fault handling is performed on the target application service based on the fault type and fault knowledge base corresponding to the target monitoring data, including: determining the service operating status of the target application service based on the target monitoring data, and determining whether the target application service has failed based on the service operating status; if so, determining a target fault handling script based on the fault type and fault knowledge base corresponding to the target monitoring data to perform fault handling on the target application service.

[0045] The fault type may be, for example, a service configuration parameter error or memory overflow. The fault handling script may be a service restart script, a configuration modification script, a resource adjustment script, or a master-slave switching script. The service operation status may be normal operation or abnormal operation.

[0046] Optionally, based on the operating parameters in the target monitoring data, when the operating parameters are detected as running, the service operating state of the target application service is determined to be normal operation, otherwise the service operating state is determined to be abnormal. If the service operating state of the target application service is normal operation, it can be determined that the target application service has failed. If the service operating state of the target application service is abnormal, it can be determined that the target application service has not failed.

[0047] Optionally, based on the fault type and fault knowledge base corresponding to the target monitoring data, a target fault handling script is determined to perform fault handling on the target application service, including: determining the fault type corresponding to the target monitoring data, and matching it in the fault knowledge base according to the fault type to obtain a target fault code model and a target self-healing function model; determining the target fault handling script based on the target fault code model and the target self-healing function model, and performing fault handling on the target application service by calling the fault handling script.

[0048] Optionally, the parameter value of each indicator in the target monitoring data may be compared with a preset normal range of the parameters to determine whether each indicator of the target application service is normal, and based on the judgment result, determine the fault type of the target monitoring data.

[0049] Optionally, after the target application service is fault-handled, it also includes: determining whether the fault handling is successful based on the service self-healing status, and if so, updating the service status information; if not, updating the fault knowledge base based on the target monitoring data, fault type, fault code model, self-healing function model and corresponding fault handling script.

[0050] Optionally, you can use the status field in the database to confirm whether the self-healing is successful. For example, the service internal error return field information is 500; the service status is normal and the return field information is 200; if the service return field is not 200, all other fields will display warning information as prompted.

[0051] Optionally, if the service self-healing of the target application service is successful, the data will be sent to the service information module and the database alarm information will be compared to change the alarm status, return the service repair success status information and cancel the alarm. If the self-healing is not successful, the target monitoring data, fault type, fault code model, self-healing function model and corresponding fault handling script can be written into the fault knowledge base to update the fault knowledge base.

[0052] For example, if the fault type is memory overflow, it cannot be self-healed through the resource self-healing function, and the service name and fault error information will be written into the fault knowledge base in a form format.

[0053] The technical solution of the embodiment of the present invention monitors and collects data of the target application service running on the virtual cloud platform in response to the service monitoring request of the virtual cloud platform to obtain the target monitoring data; performs model training based on chaos engineering according to the target monitoring data to obtain the candidate fault code model and the candidate self-healing function model and store them in the fault knowledge base; when it is determined that the target application service fails according to the target monitoring data, the target application service is fault-handled according to the fault type and fault knowledge base corresponding to the target monitoring data. By uniformly monitoring and integrating the data of the target application service on the virtual cloud platform, the faulty service can be discovered and handled in time, thereby improving the management efficiency of the application service in the virtual cloud platform.

[0054] Embodiment 2

[0055] Figure 2 is a flow chart of a service management method provided in Embodiment 2 of the present invention; based on the above embodiments, this embodiment provides a preferred example of managing target application services in a virtual cloud platform, specifically, Figure 2 As shown, the method includes the following process:

[0056] S201. In response to a service monitoring request for a virtual cloud platform, determine a target application service running on the virtual cloud platform, and use at least one monitoring process to perform offload monitoring on the target application service.

[0057] S202: In the process of performing offload monitoring on target application services, data collection is performed based on the calling interface of each target application service to obtain target monitoring data.

[0058] S203. Perform model training based on chaos engineering according to target monitoring data to obtain candidate fault code models and candidate self-healing function models and store them in a fault knowledge base.

[0059] S204: Determine a service running state of the target application service according to the target monitoring data, and determine whether the target application service is faulty according to the service running state.

[0060] S205: If yes, determine the fault type corresponding to the target monitoring data, and match it in the fault knowledge base according to the fault type to obtain a target fault code model and a target self-healing function model.

[0061] S206: Determine a target fault handling script according to the target fault code model and the target self-healing function model, and perform fault handling on the target application service by calling the fault handling script.

[0062] S207: Determine whether the fault handling is successful according to the service self-healing status. If the fault handling is successful, update the service status information.

[0063] S208. If the fault handling fails, the fault knowledge base is updated according to the target monitoring data, fault type, fault code model, self-healing function model and corresponding fault handling script.

[0064] The technical solution of the present invention can support automated operation and maintenance monitoring of various virtualized cloud platforms, and can automatically monitor, detect and self-heal faults for all service resources of physical servers, virtual machines and containers. It has strong scalability and supports customized detection and self-healing templates. It has the advantages of high concurrency based on the multi-coroutine mechanism and occupies less resources. The monitoring is more real-time, and service failure problems can be automatically solved through the fault knowledge base and self-healing model. There is no need for manual operation and maintenance processing, which greatly reduces personnel investment and responds to service alarm problems more quickly, thereby more efficiently ensuring the stability and continuity of services.

[0065] Embodiment 3

[0066] Figure 3 : is a structural block diagram of a service management device provided in the third embodiment of the present invention; this embodiment can be applied to the situation where application services running in the virtual cloud platform are monitored and fault handling is performed when a fault occurs. The service management device provided in the embodiment of the present invention can execute the service management method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method; the service management device can be implemented in the form of hardware and / or software, and configured in an electronic device with service management function, and executed by the virtual cloud platform, such as Figure 3 As shown, the service management device specifically includes:

[0067] The monitoring module 301 is used to monitor and collect data of the target application service running on the virtual cloud platform in response to the service monitoring request of the virtual cloud platform to obtain target monitoring data;

[0068] Obtaining module 302, for performing model training based on chaos engineering according to target monitoring data to obtain candidate fault code models and candidate self-healing function models and store them in a fault knowledge base;

[0069] The processing module 303 is used to perform fault processing on the target application service according to the fault type and fault knowledge base corresponding to the target monitoring data when it is determined that the target application service fails according to the target monitoring data.

[0070] The technical solution of the embodiment of the present invention monitors and collects data of the target application service running on the virtual cloud platform in response to the service monitoring request of the virtual cloud platform to obtain the target monitoring data; performs model training based on chaos engineering according to the target monitoring data to obtain the candidate fault code model and the candidate self-healing function model and store them in the fault knowledge base; when it is determined that the target application service fails according to the target monitoring data, the target application service is fault-handled according to the fault type and fault knowledge base corresponding to the target monitoring data. By uniformly monitoring and integrating the data of the target application service on the virtual cloud platform, the faulty service can be discovered and handled in time, thereby improving the management efficiency of the application service in the virtual cloud platform.

[0071] Furthermore, the processing module 303 may include:

[0072] A judgment unit, used to determine a service running state of a target application service according to target monitoring data, and determine whether the target application service is faulty according to the service running state;

[0073] The processing unit is used to, if yes, determine a target fault handling script according to the fault type and fault knowledge base corresponding to the target monitoring data to perform fault handling on the target application service.

[0074] Furthermore, the processing unit is specifically used for:

[0075] Determine the fault type corresponding to the target monitoring data, and match it in the fault knowledge base according to the fault type to obtain the target fault code model and the target self-healing function model;

[0076] According to the target fault code model and the target self-healing function model, the target fault handling script is determined, and the fault handling of the target application service is performed by calling the fault handling script.

[0077] Furthermore, the monitoring module 301 is specifically used for:

[0078] Determine a target application service running on the virtual cloud platform, and use at least one monitoring process to perform offload monitoring on the target application service;

[0079] In the process of offloading and monitoring the target application services, data collection is performed based on the calling interface of each target application service to obtain the target monitoring data.

[0080] Furthermore, the above device is also used for:

[0081] Determine whether the fault handling is successful based on the service self-healing status, and if so, update the service status information;

[0082] If not, the fault knowledge base is updated according to the target monitoring data, fault type, fault code model, self-healing function model and corresponding fault handling script.

[0083] Furthermore, the target application service includes a cluster container service and a virtual machine service; the cluster container service includes a middleware service, and the virtual machine service includes a bare metal system service.

[0084] Embodiment 4

[0085] Figure 4 It is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0086] like Figure 4 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0087] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0088] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the service management method.

[0089] In some embodiments, the service management method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the service management method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the service management method in any other appropriate manner (e.g., by means of firmware).

[0090] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0091] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0092] In the context of the present invention, a computer readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or equipment. A computer readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, a computer readable storage medium may be a machine readable signal medium. A more specific example of a machine readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0093] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0094] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0095] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0096] In one embodiment, the embodiment of the present invention further includes a computer program product, the computer program product includes a computer program, and the computer program implements the service management method of any embodiment of the present invention when executed by a processor.

[0097] In the process of implementation, the computer program product can be written in one or more programming languages ​​or a combination thereof to perform the computer program code of the present invention, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet).

[0098] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0099] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A service management method, characterized in that: include: In response to a service monitoring request for the virtual cloud platform, monitoring and data collection are performed on a target application service running on the virtual cloud platform to obtain target monitoring data; According to the target monitoring data, model training is performed based on chaos engineering to obtain candidate fault code models and candidate self-healing function models and store them in the fault knowledge base; When it is determined according to the target monitoring data that a target application service fails, the target application service is fault processed according to the fault type and the fault knowledge base corresponding to the target monitoring data.

2. The method according to claim 1, characterized in that When it is determined that the target application service fails according to the target monitoring data, the target application service is fault-handled according to the fault type and fault knowledge base corresponding to the target monitoring data, including: Determine the service operation status of the target application service according to the target monitoring data, and determine whether the target application service is faulty according to the service operation status; If so, a target fault handling script is determined according to the fault type and fault knowledge base corresponding to the target monitoring data to handle the fault of the target application service.

3. The method according to claim 2, characterized in that According to the fault type and fault knowledge base corresponding to the target monitoring data, determine the target fault handling script to handle the fault of the target application service, including: Determine the fault type corresponding to the target monitoring data, and match it in the fault knowledge base according to the fault type to obtain the target fault code model and the target self-healing function model; According to the target fault code model and the target self-healing function model, the target fault handling script is determined, and the fault handling of the target application service is performed by calling the fault handling script.

4. The method according to claim 1, characterized in that: Monitor and collect data on the target application services running on the virtual cloud platform to obtain target monitoring data, including: Determine a target application service running on the virtual cloud platform, and use at least one monitoring process to perform offload monitoring on the target application service; In the process of offloading and monitoring the target application services, data collection is performed based on the calling interface of each target application service to obtain the target monitoring data.

5. The method according to claim 1, characterized in that After troubleshooting the target application service, it also includes: Determine whether the fault handling is successful based on the service self-healing status, and if so, update the service status information; If not, the fault knowledge base is updated according to the target monitoring data, fault type, fault code model, self-healing function model and corresponding fault handling script.

6. The method according to claim 1, characterized in that in, The target application services include cluster container services and virtual machine services; cluster container services include middleware services, and virtual machine services include bare metal system services.

7. A service management device, characterized in that: include: A monitoring module, for responding to a service monitoring request for the virtual cloud platform, monitoring and collecting data on a target application service running on the virtual cloud platform to obtain target monitoring data; A module is obtained, which is used to perform model training based on chaos engineering according to target monitoring data to obtain candidate fault code models and candidate self-healing function models and store them in a fault knowledge base; The processing module is used to perform fault processing on the target application service according to the fault type and fault knowledge base corresponding to the target monitoring data when it is determined that the target application service has a fault according to the target monitoring data.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the service management method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the service management method according to any one of claims 1 to 6 when executed.

10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the service management method according to any one of claims 1 to 6.