Service availability monitoring method and device, medium and product
By adaptively generating the target service availability threshold, the problem that manual experience value in the prior art cannot adapt to business changes is solved, the adaptability of service availability monitoring and timely handling of faults is achieved, and the stability of the system and the accuracy of business processing is improved.
Patent Information
- Application Number
- CN202510627781.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-07-22
AI Technical Summary
The existing service availability monitoring methods rely on manual experience values and cannot adapt to business development and changes, resulting in the inability to detect and deal with system failures in a timely and accurate manner.
By adaptively generating the target service availability threshold based on the real-time data of the system service, and performing fault isolation or recovery processing when the availability rate is below the threshold, the adaptively generated target service availability threshold is used instead of the traditional empirical fixed value.
It realizes the service availability monitoring adaptive business development and changes, timely discovers and accurately deals with faulty system services, and improves system stability and correctness of business processing.
Smart Images

Figure CN120358170A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fintech, and particularly to a method, device, medium and product for monitoring service availability. Background Art
[0002] At present, the banking industry has gradually expanded from the traditional financial field to the full scenarios of financial and non-financial fields, and the business scenarios are becoming more and more complex. The stability of system operation and the correctness of business processing are crucial.
[0003] The general methods for traditional monitoring systems to monitor service availability usually adopt means such as heartbeat detection, fixed service success rate, and fixed service failure rate. However, the services within the system are gradually developing, and the existing monitoring indicators change with the business development. Relying on manual experience values to set fixed service availability evaluation values can no longer meet the requirements, and there is an urgent need for a dynamic monitoring method for service availability that can adapt to the changes in business development. Summary of the Invention
[0004] The present invention provides a method, device, medium and product for monitoring service availability to solve the problem that the existing service availability monitoring method relies on manual experience values and cannot adapt to the changes in business development.
[0005] According to one aspect of the present invention, there is provided a method for monitoring service availability, including:
[0006] When the current service monitoring index meets the index monitoring condition, determine the current service availability rate according to the real-time data of the system service, and obtain a target service availability rate threshold that is adaptively generated and matches the current service monitoring index;
[0007] When the current service availability rate is less than the target service availability rate threshold, perform fault isolation or service recovery processing on the current service according to the current service index definition information.
[0008] According to another aspect of the present invention, there is provided a device for monitoring service availability, including:
[0009] A data acquisition module, configured to determine the current service availability rate according to the real-time data of the system service and obtain a target service availability rate threshold that is adaptively generated and matches the current service monitoring index when the current service monitoring index meets the index monitoring condition;
[0010] A service isolation and recovery module, configured to perform fault isolation or service recovery processing on the current service according to the current service index definition information when the current service availability rate is less than the target service availability rate threshold.
[0011] According to another aspect of the present invention, there is provided an electronic device, where the electronic device includes:
[0012] At least one processor; and
[0013] A memory communicatively connected to the at least one processor; wherein
[0014] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the service availability monitoring method according to any embodiment of the present invention.
[0015] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the service availability monitoring method according to any embodiment of the present invention when executed.
[0016] According to another aspect of the present invention, there is provided a computer program product including a computer program that implements the service availability monitoring method according to any embodiment of the present invention when executed by a processor.
[0017] The technical solution of the embodiment of the present invention, when the current service monitoring index meets the index monitoring condition, determines the current service availability rate according to the real-time data of the system service, and obtains a target service availability rate threshold adaptively generated and matching the current service monitoring index, so that when the current service availability rate is less than the target service availability rate threshold, the current service is subjected to fault isolation or service recovery processing according to the current service index definition information. In this solution, the target service availability rate threshold is adaptively generated instead of using a traditional empirical fixed value, which can better adapt to the changes in the development of the system service, and can timely detect and handle the faulty system service, solving the problem that the existing service availability monitoring method depends on manual empirical values and cannot adapt to the changes in business development, enabling the service availability monitoring to adapt to the changes in business development and timely and accurately detect and handle the faulty system service.
[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1Flowchart of a service availability monitoring method provided in Embodiment 1 of the present invention;
[0021] Figure 2 Flowchart of a service availability monitoring method provided in Embodiment 2 of the present invention;
[0022] Figure 3 Schematic structural diagram of a service availability monitoring device provided in Embodiment 4 of the present invention;
[0023] Figure 4 Schematic structural diagram of an electronic device that can be used to implement the embodiments of the present invention is shown. Detailed implementation manners
[0024] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] Embodiment 1
[0027] Figure 1 Flowchart of a service availability monitoring method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of monitoring and promptly disposing of system business failures based on service availability. This method can be executed by a service availability monitoring device, which can be implemented in the form of hardware and / or software, and the service availability monitoring device can be configured in an electronic device.
[0028] As Figure 1 shown, the method includes:
[0029] Step 110: When the current service monitoring metrics meet the metric monitoring conditions, determine the current service availability rate based on the real-time system service data, and obtain the target service availability rate threshold that is adaptively generated and matches the current service monitoring metrics.
[0030] Among them, the current service monitoring metrics can be the monitoring metrics of the current system service. Exemplarily, the current service monitoring metrics can include, but are not limited to, information such as metric types (such as business services, application services, etc.), metric names, metric monitoring time intervals, metric fault isolation thresholds, metric effective status, and associated service lists. The metric monitoring conditions can be the conditions for monitoring the metrics of the current service. Exemplarily, the metric monitoring conditions can include, but are not limited to, the metric effective status of the current service being effective, and the running time of the current service being within the metric monitoring time interval. The real-time system service data can be the data associated with system service transactions. Exemplarily, the real-time system service data can include, but are not limited to, business transaction details, business transaction logs, and system operation logs.
[0031] Among them, the current service availability rate can be used to describe the actual success rate of the current service. Optionally, when the current service is a business service, the metric type can be specifically divided into transaction success rate, user login success rate, transaction coverage rate, and business closed-loop rate, etc. When the current service is a business service, the metric type can be specifically divided into service response time, background process health, and monitoring task running status, etc. The target service availability rate threshold can be a pre-determined service availability rate threshold that matches the current service. Optionally, the target service availability rate threshold can be determined based on any big data analysis algorithm or artificial intelligence algorithm.
[0032] It should be noted that the collection of business transaction details is information and data that have been authorized by the user or fully authorized by all parties, and the processing of related data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or reject.
[0033] In the embodiment of the present invention, it can be determined whether the current service monitoring metrics meet the metric monitoring conditions. If the current service monitoring metrics meet the metric monitoring conditions, then obtain the real-time system service data, calculate the current service availability rate according to the service success definition of the current service, and adaptively generate a target service availability rate threshold that matches the current service monitoring metrics before monitoring the metrics of the current service.
[0034] Among them, the calculation logic of the current service availability is determined by the definition of the current service success, that is, the ratio of the actual number of successes of the current service to the total number of the current service. For example, when the current service is a business service, the success or failure of the business transaction result is not a consideration in the calculation of the current service availability (i.e., the business service success rate). Under the transaction dimension, the current service availability is the ratio of the number of successful business service transactions to the total number of transactions. When the current service is an application service, the existence of both the request and the response can be understood as a successful application service.
[0035] Step 120: When the current service availability is less than the target service availability threshold, perform fault isolation or service recovery processing on the current service according to the current service indicator definition information.
[0036] The current service indicator definition information may be a rule definition of a service monitoring indicator of the current service. Exemplarily, the current service indicator definition information may include but is not limited to the current service monitoring indicator, target service availability threshold, and logic definition of service success.
[0037] In an embodiment of the present invention, the current service availability rate can be compared with the target service availability rate threshold. If the current service availability rate is less than the target service availability rate threshold, it indicates that the current service may have a fault. Therefore, based on the current service indicator definition information, it is determined whether the current service needs to directly isolate the fault or perform service recovery processing.
[0038] The technical solution of the embodiment of the present invention determines the current service availability rate according to the real-time data of the system service when the current service monitoring indicator meets the indicator monitoring conditions, and obtains the adaptively generated target service availability rate threshold that matches the current service monitoring indicator, so that when the current service availability rate is less than the target service availability rate threshold, the current service is fault-isolated or service recovery is processed according to the current service indicator definition information. In this solution, the target service availability rate threshold is adaptively generated instead of using the traditional empirical fixed value, which can better adapt to changes in the development of system services, and can timely discover and handle faulty system services, solving the problem that the existing service availability monitoring method relies on artificial experience values and cannot adapt to changes in business development, and can make service availability monitoring adaptive to changes in business development, and timely and accurately discover and handle faulty system services.
[0039] Embodiment 2
[0040] Figure 2 This is a flow chart of a service availability monitoring method provided in the second embodiment of the present invention. This embodiment is specific based on the above embodiment and provides a specific optional implementation method for determining the current service availability rate based on the real-time data of the system service. Figure 2 As shown, the method includes:
[0041] Step 210: When the current service monitoring metrics meet the metric monitoring conditions, filter the real-time data of the system service to obtain the data to be analyzed for the current service.
[0042] Among them, the data to be analyzed for the current service can be the data related to the calculation of the current service availability rate in the real-time data of the system service.
[0043] Specifically, when the current service monitoring metrics meet the metric monitoring conditions, based on the type of the current service, the data to be analyzed for the current service associated with the calculation of the current service availability rate can be filtered out from the real-time data of the system service.
[0044] Exemplarily, when the current service is a business service, the data to be analyzed for the current service can be the business transaction details and business transaction logs in the real-time data of the system service. When the current service is an application service, the data to be analyzed for the current service can be the business transaction details, business transaction logs, and system operation logs in the real-time data of the system service.
[0045] Step 220: According to the service success determination condition matched by the current service and the data to be analyzed for the current service, calculate the current service availability rate, and obtain the target service availability rate threshold adaptively generated and matched with the current service monitoring metrics.
[0046] Among them, the service success determination condition can be the definition of service success.
[0047] In the embodiment of the present invention, the service success determination condition matched by the current service can be obtained, so as to identify the data to be analyzed for the current service according to the service success determination condition matched by the current service, obtain the actual number of successful times of the current service, and determine the total number of times of the current service according to the data to be analyzed for the current service, so as to use the actual number of successful times of the current service and the total number of times of the current service to determine the current service availability rate, and adaptively generate the target service availability rate threshold matched with the current service monitoring metrics.
[0048] In an optional embodiment of the present invention, calculating the current service availability rate according to the service success determination condition matched by the current service and the data to be analyzed for the current service may include: when the current service is a business service, the service success determination condition matched by the current service is the first service success determination condition; calculating the current service availability rate according to the first service success determination condition and the data to be analyzed for the current service; when the current service is an application service, the service success determination condition matched by the current service is the second service success determination condition; calculating the current service availability rate according to the second service success determination condition and the data to be analyzed for the current service.
[0049] Among them, the first service success determination condition may be a service success determination condition matching the business service. The second service success determination condition may be a service success determination condition matching the application service.
[0050] In an embodiment of the present invention, when the current service is a business service, the first service success determination condition can be obtained, and then the current service availability rate can be calculated according to the first service success determination condition and the current service data to be analyzed. When the current service is an application service, the second service success determination condition can be obtained, and then the current service availability rate can be calculated according to the second service success determination condition and the current service data to be analyzed, that is, the service availability rate is calculated separately for different types of system services to better fit different service success rate evaluation methods.
[0051] In an optional embodiment of the present invention, obtaining a target service availability rate threshold that is adaptively generated and matches the current service monitoring metrics may include: obtaining a target service availability rate threshold that is adaptively generated and matches the current service monitoring metrics through a service availability rate threshold prediction model and / or an expert comprehensive calculation rule.
[0052] Among them, the service availability rate threshold prediction model may be a regression model for predicting the success rate of system services. The expert comprehensive calculation rule may be an expert rule system for predicting the success rate of system services.
[0053] In an embodiment of the present invention, a target service availability rate threshold that is obtained by separately analyzing the historical data of system services through a service availability rate threshold prediction model and matches the current service monitoring metrics can be obtained, or a target service availability rate threshold that is adaptively generated by an expert comprehensive calculation rule and matches the current service monitoring metrics can be obtained. It is also possible to correct the analysis result based on the expert comprehensive calculation rule after analyzing the historical data of system services based on the service availability rate threshold prediction model, and use the final corrected result as the target service availability rate threshold that matches the current service monitoring metrics. The target service availability rate threshold determined based on this solution can adapt to business development changes and does not rely on empirical fixed values.
[0054] Step 230, when the current service availability rate is less than the target service availability rate threshold, perform fault isolation or service recovery processing on the current service according to the current service metric definition information.
[0055] In an optional embodiment of the present invention, performing fault isolation or service recovery processing on the current service according to the current service metric definition information may include: determining a metric fault isolation threshold according to the current service metric definition information; isolating the current service when the metric monitoring value of the current service is lower than the metric fault isolation threshold; or, calling the current service recovery script to recover the current service when the metric monitoring value of the current service is greater than or equal to the metric fault isolation threshold.
[0056] Among them, the index monitoring value can be used to describe the unavailability degree of the current service. The index fault isolation threshold can be a preset threshold for evaluating whether the service is isolated.
[0057] In the embodiment of the present invention, the previous service index definition information can be parsed to determine the index fault isolation threshold, and the difference between 1 and the current service availability rate is used as the index monitoring value. Then, the index monitoring value of the current service is compared with the index fault isolation threshold. If the index monitoring value of the current service is lower than the index fault isolation threshold, the current service is isolated. When the index monitoring value of the current service is greater than or equal to the index fault isolation threshold, the current service recovery script is called to attempt to recover the current service, which can isolate the abnormal system services that need to be isolated for timely operation and maintenance, and attempt to recover the abnormal system services that do not need to be directly isolated, avoiding the long-term offline of system services.
[0058] In an optional embodiment of the present invention, after calling the current service recovery script to recover the current service, it may further include: if the recovery of the current service fails, the current service is fault-isolated.
[0059] In the embodiment of the present invention, if the current service recovery fails after calling the current service recovery script to attempt to recover the current service, the current service is further fault-isolated to prevent entering a dead loop of recovering the current service.
[0060] In an optional embodiment of the present invention, before obtaining the target service availability rate threshold that is adaptively generated and matches the current service monitoring index, it may further include: cleaning the system service historical data in the target time window to obtain service availability rate threshold evaluation sample data; inputting the service availability rate threshold evaluation sample data into a preset regression algorithm model for service availability rate threshold analysis training to obtain a service availability rate threshold prediction model.
[0061] Among them, the target time window can be a preset time window. The target time window can be the previous day or the previous week, etc., which can be set by oneself. The system service historical data can be the data associated with the system service transactions in the target time window. The service availability rate threshold prediction model is used to predict the available rate threshold of the system service. The service availability rate threshold evaluation sample data can be the data obtained after cleaning the system service historical data in the target time window. The preset regression algorithm model can be any regression algorithm model.
[0062] In the embodiments of the present invention, the historical data of system services in a target time window can be cleaned to obtain sample data for evaluating the service availability threshold. Then, the sample data for evaluating the service availability threshold is input into a preset regression algorithm model for analysis and training of the service availability threshold, and the trained preset regression algorithm model is used as the service availability threshold prediction model. By data cleaning, the quality of the sample data for evaluating the service availability threshold is ensured, providing a reliable data basis for adaptively generating the service availability threshold of the system service.
[0063] The technical solution of the embodiments of the present invention screens the real-time data of the system service when the current service monitoring index meets the index monitoring condition to obtain the data to be analyzed for the current service. Then, according to the service success determination condition matched by the current service and the data to be analyzed for the current service, the availability rate of the current service is calculated, and the target service availability threshold matched with the current service monitoring index generated adaptively is obtained. Thus, when the availability rate of the current service is less than the target service availability threshold, the current service is subjected to fault isolation or service recovery processing according to the current service index definition information. In this solution, through the screening operation, the data required for calculating the availability rate of the current service is quickly located from the massive data, and the actual number of successful times of the current service is effectively and accurately identified through the service success determination condition. Therefore, the availability rate of the current service is accurately calculated based on the corresponding service success determination condition and the data to be analyzed for the current service, and the target service availability threshold is generated adaptively instead of using a traditional empirical fixed value, which can better adapt to the changes in the development of the system service, timely discover and handle the faulty system service, solve the problem that the existing service availability monitoring method relies on manual empirical values and cannot adapt to the changes in business development, and enable the service availability monitoring to adapt to the changes in business development and timely and accurately discover and handle the faulty system service.
[0064] Embodiment III
[0065] Embodiment III of the present invention provides an alternative embodiment of a service availability monitoring system, and its specific implementation manner can be referred to the following embodiments. Among them, the same or corresponding technical terms as those in the above embodiments will not be elaborated herein.
[0066] The service availability monitoring system includes a data collection component, a policy configuration component, a data cleaning and service availability intelligent model component, a service availability monitoring component, a service availability calculation and update component, and a service monitoring alarm component. Each component cooperates with each other to jointly realize the real-time monitoring of service availability and the dynamic adjustment of the service availability threshold.
[0067] The data collection component is used to collect and store the system service transaction data (such as business transaction details, business transaction logs, and system operation logs).
[0068] The policy configuration component includes a policy configuration maintenance unit, an expert rule unit, and a policy release unit. The policy configuration maintenance unit supports users to configure the service availability threshold prediction model and the expert comprehensive calculation rule. The expert rule unit is used to set the adjustment rule for the results of the service availability threshold prediction model based on the experience of domain experts. The expert comprehensive calculation rule in the expert rule unit can be set or not. Similar to expert rules, when the trading volume is in interval 1, the allowable fluctuation value of the service availability threshold is aa; when the trading volume is in interval 2, the allowable fluctuation value of the service availability threshold is bb, etc. The policy release unit is used to officially release the policies maintained by the configuration. The existing policies will automatically become invalid, and only one valid policy configuration is allowed for one service.
[0069] The data cleaning and service availability intelligent model component includes a data cleaning unit and a service availability intelligent model unit. The data cleaning unit is used to perform legality checks on the basic information such as the types, dictionary values, and lengths of the system service transaction data collected, including logical legality checks such as whether there are duplicates and data errors. It screens and processes the business transaction details and business transaction logs, and classifies the record status into three categories: success, failure, and exception according to the rules. It performs desensitization processing on sensitive fields according to the field configuration rules. The order of the foregoing specific data cleaning operations is not restricted.
[0070] The service availability intelligent model unit inputs the cleaned business transaction details, business transaction logs, and system operation logs (i.e., the service availability threshold evaluation sample data) into the service availability threshold prediction model to calculate the service availability threshold. The service availability threshold prediction model can adopt the ridge regression algorithm. First, it uses the correlation analysis method to perform correlation analysis on the data processed by the data cleaning unit to obtain the correlation scores of all factors (transaction time, transaction type, trading volume, customer type, bank card type, partner type, etc.). Then, it uses the ridge regression cross-validation algorithm to perform mean regression processing on the service monitoring indicators. Further, it uses the polynomial fitting method to calculate the weights of the service availability threshold evaluation sample data based on the data integrity analysis and...
[0071] The service availability monitoring component includes a monitoring metric definition unit, a log analysis unit, a business service monitoring unit, an application service monitoring unit, and a fault isolation and recovery unit. The monitoring metric definition unit is used to configure the service metric definitions for business service monitoring metrics and application service monitoring metrics. The log analysis unit is used to schedule the metric logic tasks (i.e., service availability calculation logic) corresponding to the current service monitoring metrics by the task invocation center. Specifically, it performs logical statistical analysis on business details, business logs, etc. through the metric logic tasks, calculates and stores the relevant result information. The business service monitoring unit, through the task scheduling center, schedules the business services in the monitoring metric definition unit according to the task configuration. If the current service monitoring metrics meet the metric monitoring conditions, it calls the log analysis unit to obtain the current service availability, judges the relationship between the current service availability and the corresponding service availability threshold. If the current service availability is greater than or equal to the corresponding service availability threshold, it indicates that the current business service metrics are normal, otherwise it indicates that the metrics are abnormal. When the current business service is abnormal, it transfers the current service metric definition information to the fault isolation and recovery unit for processing. The application service monitoring unit has a similar operation logic to the business service, except that the definition of service success is different and the service availability calculation logic will be different. The fault isolation and recovery unit receives the fault events of the business service monitoring unit and the application service monitoring unit. If the metric monitoring value of the current service is lower than the metric fault isolation threshold, the service needs to be isolated and the alarm information is transferred to the service monitoring alarm component for processing. If the metric monitoring value of the current service is greater than or equal to the metric fault isolation threshold, it can call the service recovery logic in the current service metric definition information to attempt to recover the service. If it cannot be recovered, the current service is then isolated.
[0072] The service availability calculation and update component is used to call the service availability intelligent model unit after the end-of-day business is completed to determine the latest service availability threshold, and determine the final service availability threshold according to the policy configuration component, and then update the final service availability threshold to the current service metric definition information and release the latest definition.
[0073] The service monitoring alarm component includes a monitoring warning information generation unit, a monitoring warning notification unit, and a monitoring warning storage unit. The monitoring alarm generation unit is used to generate monitoring warning information (information such as warning name, warning level, notification personnel list, and notification channels). The monitoring warning notification unit notifies the relevant personnel of the alarm information through mobile phones, text messages, emails, intelligent outbound calls, etc. The monitoring warning storage unit is used to store the monitoring warning event information for post-event problem analysis and auditing, etc.
[0074] The service availability monitoring of this solution adopts an adaptive dynamic calculation mode instead of the traditional fixed value based on experience. By combining comprehensive factors such as historical transaction volume, date, and major industry adjustments, and using a regression algorithm, the availability rates of all services within a week are estimated, and the estimation results are more accurate. Based on the model calculation results, combined with expert rules, it is allowed to adjust the estimated value based on the configuration strategy on the basis of the estimated value calculated by the model algorithm, and the result is more dynamic and forward-looking.
[0075] Embodiment 4
[0076] Figure 3 It is a schematic structural diagram of a service availability monitoring device provided in Embodiment 4 of the present invention. As Figure 3 shown, the device includes:
[0077] A data acquisition module 310, configured to determine the current service availability rate according to the real-time data of the system service when the current service monitoring index meets the index monitoring condition, and obtain a target service availability rate threshold that is adaptively generated and matches the current service monitoring index;
[0078] A service isolation and recovery module 320, configured to perform fault isolation or service recovery processing on the current service according to the current service index definition information when the current service availability rate is less than the target service availability rate threshold.
[0079] The technical solution of the embodiment of the present invention determines the current service availability rate according to the real-time data of the system service when the current service monitoring index meets the index monitoring condition, and obtains a target service availability rate threshold that is adaptively generated and matches the current service monitoring index, so that when the current service availability rate is less than the target service availability rate threshold, fault isolation or service recovery processing is performed on the current service according to the current service index definition information. In this solution, the target service availability rate threshold is adaptively generated instead of using the traditional fixed value based on experience, which can better adapt to the changes in the development of the system service, and can timely discover and handle the faulty system service, solving the problem that the existing service availability monitoring method depends on manual experience values and cannot adapt to the changes in business development, enabling the service availability monitoring to adapt to the changes in business development and timely and accurately discover and handle the faulty system service.
[0080] Optionally, the data acquisition module 310 includes a first data acquisition unit and a second data acquisition unit. The first data acquisition unit is configured to screen the real-time data of the system service to obtain the current service data to be analyzed; calculate the current service availability rate according to the service success discrimination condition matched by the current service and the current service data to be analyzed.
[0081] Optionally, the first data acquisition unit is specifically configured to, when the current service is a business service, the service success determination condition matched by the current service is the first service success determination condition; calculate the availability rate of the current service according to the first service success determination condition and the data to be analyzed of the current service; when the current service is an application service, the service success determination condition matched by the current service is the second service success determination condition; calculate the availability rate of the current service according to the second service success determination condition and the data to be analyzed of the current service.
[0082] Optionally, the second data acquisition unit is configured to obtain a target service availability threshold that is adaptively generated to match the monitoring metrics of the current service through a service availability threshold prediction model and / or an expert comprehensive calculation rule.
[0083] Optionally, the service isolation and recovery module 320 is specifically configured to determine a metric fault isolation threshold according to the current service metric definition information; isolate the current service when the metric monitoring value of the current service is lower than the metric fault isolation threshold; or, when the metric monitoring value of the current service is greater than or equal to the metric fault isolation threshold, call the current service recovery script to recover the current service.
[0084] Optionally, the service availability monitoring device further includes a service recovery failure handling module, configured to isolate the current service if the recovery of the current service fails.
[0085] Optionally, the service availability monitoring device further includes a data cleaning and model training module, configured to clean the system service historical data in a target time window to obtain service availability threshold evaluation sample data; input the service availability threshold evaluation sample data into a preset regression algorithm model for service availability threshold analysis training to obtain a service availability threshold prediction model; wherein, the service availability threshold prediction model is used to predict the availability threshold of the system service.
[0086] The service availability monitoring device provided by the embodiments of the present invention can execute the service availability monitoring method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0087] Embodiment 5
[0088] Figure 4The figure shows a schematic structural diagram of an electronic device that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0089] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as ROM 12, RAM 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the ROM 12 or the computer program loaded from the storage unit 18 into the RAM 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The I / O interface 15 is also connected to the bus 14. The ROM 12 is a read-only memory, the RAM 13 is a random access memory, and the I / O interface 15 is an input / output interface.
[0090] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0091] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the service availability monitoring method.
[0092] In some embodiments, the service availability monitoring method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the service availability monitoring method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the service availability monitoring method by any other suitable means (e.g., by means of firmware).
[0093] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0094] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0095] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0096] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0097] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0098] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of traditional physical hosts and VPS servers, such as high management difficulty and weak business scalability.
[0099] The embodiments of the present application also disclose a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the service availability monitoring method provided in any embodiment of the present application. This program product and the service availability monitoring methods disclosed in the embodiments of the present application belong to the same inventive concept, and thus will not be elaborated herein.
[0100] It should be understood that various forms of the processes shown above may be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0101] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A service availability monitoring method, characterized in that, Including: When the current service monitoring metrics meet the metric monitoring conditions, determine the current service availability rate according to the real-time system service data, and obtain the target service availability rate threshold that is adaptively generated and matches the current service monitoring metrics; When the current service availability rate is less than the target service availability rate threshold, perform fault isolation or service recovery processing on the current service according to the current service metric definition information.
2. The service availability monitoring method according to claim 1, wherein Determining the current service availability rate according to the real-time system service data includes: Filter the real-time system service data to obtain the data to be analyzed for the current service; Calculate the current service availability rate according to the service success determination condition matched by the current service and the data to be analyzed for the current service.
3. The service availability monitoring method according to claim 2, wherein Calculating the current service availability rate according to the service success determination condition matched by the current service and the data to be analyzed for the current service includes: When the current service is a business service, the service success determination condition matched by the current service is the first service success determination condition; calculate the current service availability rate according to the first service success determination condition and the data to be analyzed for the current service; When the current service is an application service, the service success determination condition matched by the current service is the second service success determination condition; calculate the current service availability rate according to the second service success determination condition and the data to be analyzed for the current service.
4. The service availability monitoring method according to claim 1, wherein Obtaining the target service availability rate threshold that is adaptively generated and matches the current service monitoring metrics includes: Obtain the target service availability rate threshold that is adaptively generated and matches the current service monitoring metrics through a service availability rate threshold prediction model and / or an expert comprehensive calculation rule.
5. The service availability monitoring method according to claim 1, wherein, Performing fault isolation or service recovery processing on the current service according to the current service metric definition information includes: Determine the metric fault isolation threshold according to the current service metric definition information; When the metric monitoring value of the current service is lower than the metric fault isolation threshold, isolate the current service; or, When the metric monitoring value of the current service is greater than or equal to the metric fault isolation threshold, call the current service recovery script to recover the current service.
6. The service availability monitoring method according to claim 5, wherein After calling the current service recovery script to recover the current service, it further includes: If the recovery of the current service fails, perform fault isolation on the current service.
7. The service availability monitoring method according to claim 4, characterized in that Before obtaining the target service availability rate threshold that is adaptively generated and matches the current service monitoring metrics, it further includes: Clean the historical system service data in the target time window to obtain the sample data for service availability rate threshold evaluation; Input the sample data for service availability rate threshold evaluation into a preset regression algorithm model for service availability rate threshold analysis training to obtain a service availability rate threshold prediction model; Wherein, the service availability rate threshold prediction model is used to predict the availability rate threshold of the system service.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, enables the at least one processor to execute the service availability monitoring method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for implementing the service availability monitoring method according to any one of claims 1-7 when the computer instructions are executed by a processor.
10. A computer program product, characterized in that, The computer program product includes a computer program which, when executed by a processor, implements the service availability monitoring method according to any one of claims 1-7.