A method and apparatus for service resource monitoring

By establishing a distributed monitoring platform and utilizing peer-to-peer centers and distribution nodes to process monitoring data, the problem of cloud platform monitoring systems being unable to uniformly monitor multiple business resources has been solved, achieving full coverage monitoring and anomaly alarms for massive business resources.

CN113986677BActive Publication Date: 2025-10-21JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111302291.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-04
Publication Date
2025-10-21
Estimated Expiration
2041-11-04

AI Technical Summary

Technical Problem

Existing cloud platform monitoring technologies cannot monitor multiple service resources in a unified manner, causing the monitoring system to fail when the cloud platform malfunctions.

Method used

Establish a distributed monitoring platform to receive monitoring data from business resource domains through peer centers, determine whether there are any anomalies, and convert them into anomaly alarm data when anomalies occur. Support the addition of new business resource domains and information dissemination, and use distribution nodes and parallel data processing nodes for data processing and analysis.

Benefits of technology

It achieves full coverage monitoring of multiple service resources, ensuring effective monitoring even when the cloud platform fails, thus improving the coverage and accuracy of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113986677B_ABST
    Figure CN113986677B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and device for service resource monitoring, which is applied to a distributed monitoring platform. The distributed monitoring platform comprises at least one distributed monitoring platform peer center, and each distributed monitoring platform peer center corresponds to at least one service resource domain. The method comprises the following steps: determining a service resource domain to be monitored; receiving monitoring data of a service resource sent by the service resource domain through the distributed monitoring platform peer center, and judging whether the service resource domain is abnormal according to the monitoring data; and in the case that it is determined that the service resource domain is abnormal, calling the distributed monitoring platform peer center to convert the monitoring data of the abnormal service resource domain into abnormal alarm data. Through the establishment of a distributed monitoring platform, full coverage monitoring of multiple service resources is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method and device for monitoring business resources. Background Art

[0002] Existing monitoring technologies are used to monitor resources at all levels of cloud platforms, primarily targeting independent business resources or systems by deploying agents on key nodes. However, there is no unified approach to monitoring these independent business resources. Consequently, when the cloud platform itself fails, such as when the control plane fails, the monitoring system becomes ineffective. Summary of the Invention

[0003] The present disclosure provides a method and apparatus for monitoring business resources, which achieve full coverage monitoring of multiple business resources by establishing a distributed monitoring platform.

[0004] In a first aspect, the present disclosure provides a method for business resource monitoring, which is applied to a distributed monitoring platform, wherein the distributed monitoring platform includes at least one distributed monitoring platform peer center, and each of the distributed monitoring platform peer centers corresponds to at least one business resource domain;

[0005] The method specifically includes:

[0006] Determine the business resource domain to be monitored;

[0007] Receiving, through the distributed monitoring platform peer center, monitoring data of the service resources sent by the service resource domain, and determining whether an abnormality occurs in the service resource domain based on the monitoring data;

[0008] When it is determined that an abnormality occurs in the business resource domain, the distributed monitoring platform peer center is called to convert the abnormal monitoring data into abnormal alarm data.

[0009] According to the method for monitoring business resources provided by the present disclosure, the method further includes:

[0010] In the case of a newly added business resource domain, the corresponding relationship between the newly added business resource domain and the distributed monitoring platform peer center is determined, and the information of the newly added business resource domain is published to the distributed monitoring platform.

[0011] According to the method for monitoring service resources provided by the present disclosure, before determining the service resource domain to be monitored, the method includes:

[0012] Generate a random traffic reservation table based on the data reception index of the distributed monitoring platform peer center, wherein the random traffic reservation table includes the amount of data sent in the time and service resource domains;

[0013] The random traffic reservation table is sent to the business resource domain, and the business resource domain is called to divide the business resources in the business resource domain into levels according to the time of the random traffic reservation table and the amount of data sent by the business resource domain, and determine the various levels of the business resource domain; wherein the various levels of the business resource domain have a priority order.

[0014] According to the method for business resource monitoring provided by the present disclosure, the distributed monitoring platform peer center includes a first-level diversion node, a second-level diversion node and a parallel data processing node;

[0015] The receiving, through the distributed monitoring platform peer center, the monitoring data of the service resources sent by the service resource domain, and determining whether an abnormality occurs in the service resource domain according to the monitoring data, includes:

[0016] Calling the first-level shunt node to receive the monitoring data, and adding the monitoring data to a high-speed processing queue in the first-level shunt node;

[0017] Calling the secondary shunt node to perform position hash processing on the monitoring data in the high-speed processing queue to determine the corresponding first parallel data processing node;

[0018] The secondary diversion node is called to send the monitoring data to the first parallel data processing node, and the first parallel data processing node determines whether an abnormality occurs in the business resource domain based on the monitoring data.

[0019] According to the method for business resource monitoring provided by the present disclosure, the distributed monitoring platform peer center further includes a fault aggregation node;

[0020] The determining, by the first parallel data processing node based on the monitoring data, whether an abnormality occurs in the service resource domain includes:

[0021] Analyzing the monitoring data by the first parallel data processing node to determine whether an abnormality occurs in the monitoring data;

[0022] If the monitoring data is abnormal, the monitoring data is reported to the fault aggregation node;

[0023] An abnormality inquiry message is sent to the service resource domain corresponding to the monitoring data through the fault aggregation node, and whether an abnormality occurs in the service resource domain is determined according to the returned response message.

[0024] According to the service resource monitoring method provided by the present disclosure, analyzing the monitoring data by the first parallel data processing node to determine whether an abnormality occurs in the monitoring data includes:

[0025] Performing a unified analysis of the monitoring data sent by each level of the service resource domain through the first parallel processing node to determine whether an abnormality occurs in the monitoring data at each level;

[0026] The determining whether an abnormality occurs in the service resource domain according to the returned response message includes:

[0027] It is determined whether an abnormality occurs at the level of the service resource domain according to the returned response message.

[0028] According to the method for monitoring business resources provided by the present disclosure, the distributed monitoring platform peer center receives the business resource monitoring data sent by the business resource domain, and determines whether an abnormality occurs in the business resource domain based on the monitoring data, further comprising:

[0029] Sending the monitoring data corresponding to the high-priority service resources directly to the secondary diversion node;

[0030] Calling the secondary offload node to perform location hash processing on the monitoring data corresponding to the high-priority service resource to obtain a second parallel data processing node;

[0031] The secondary diversion node is called to send the monitoring data to the second parallel data processing node, and the second parallel data processing node determines whether an abnormality occurs in the business resource domain based on the monitoring data.

[0032] According to the service resource monitoring method provided by the present disclosure, judging whether an abnormality occurs in the service resource domain according to the returned response message includes:

[0033] If the response message does not obtain feedback data, it is confirmed that an exception has occurred;

[0034] If the response message obtains feedback data, it is confirmed that no abnormality has occurred.

[0035] In a second aspect, the present disclosure provides a device for monitoring business resources, which is provided on a distributed monitoring platform, wherein the distributed monitoring platform includes at least one distributed monitoring platform peer center, and each of the distributed monitoring platform peer centers corresponds to at least one business resource domain;

[0036] The device specifically includes:

[0037] A determination module, used to determine the business resource domain to be monitored;

[0038] a receiving module, configured to receive monitoring data of the service resources sent by the service resource domain through the peer center of the distributed monitoring platform, and determine whether an abnormality occurs in the service resource domain based on the monitoring data;

[0039] The conversion module is used to call the distributed monitoring platform peer center to convert the monitoring data of the abnormality into abnormal alarm data when it is determined that the business resource domain has an abnormality.

[0040] In a third aspect, the present disclosure provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for monitoring business resources as described in any one of the above items are implemented.

[0041] In a fourth aspect, the present disclosure provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for business resource monitoring as described in any one of the above items.

[0042] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program, which implements the steps of the method for business resource monitoring as described in any one of the above items when executed by a processor.

[0043] The present disclosure provides a method and apparatus for business resource monitoring, which first determines a business resource domain to be monitored, wherein the business resource domain includes all business resources, receives monitoring data of the business resources sent by the business resource domain through a distributed monitoring platform peer center in a distributed monitoring platform, and calls the distributed monitoring platform peer center to determine whether an abnormality has occurred in the business resource domain based on the monitoring data. Each distributed monitoring platform peer center corresponds to at least one business resource domain, so by using the distributed monitoring platform peer center to determine whether an abnormality has occurred in the business resource domain, it is possible to monitor a large number of business resource domains; when it is determined that an abnormality has occurred in the business resource domain, the distributed monitoring platform peer center is called to convert the abnormal monitoring data into abnormal alarm data. The present disclosure achieves full coverage monitoring of multiple business resources by establishing a distributed monitoring platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the present disclosure or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0045] Figure 1 This is an overall layout diagram of a distributed monitoring platform provided by an embodiment of the present disclosure;

[0046] Figure 2 This is a flow chart of a method for monitoring a service resource domain provided by an embodiment of the present disclosure;

[0047] Figure 3 It is a block diagram of the various levels of the financial resource domain A provided by the embodiment of the present disclosure;

[0048] Figure 4 It is a level block diagram of each level of the financial resource domain B provided by the embodiment of the present disclosure;

[0049] Figure 5 This is one of the flow charts provided in the embodiment of the present disclosure for determining whether an abnormality occurs in the business resource domain;

[0050] Figure 6 This is the second flowchart of determining whether an abnormality occurs in the business resource domain provided by the embodiment of the present disclosure;

[0051] Figure 7 This is a schematic diagram of the overall process of the method for monitoring service resource domains provided by an embodiment of the present disclosure;

[0052] Figure 8 This is a schematic diagram of the overall process under a special case of the method for monitoring the service resource domain provided by an embodiment of the present disclosure;

[0053] Figure 9 Schematic diagram of the structure of a device for monitoring service resources provided by an embodiment of the present disclosure;

[0054] Figure 10 It is a structural diagram of the electronic device provided by the present disclosure. DETAILED DESCRIPTION

[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the embodiments of the present disclosure.

[0056] The distributed monitoring platform includes at least one distributed monitoring platform peer center. If each distributed monitoring platform peer center is regarded as a node, the distributed monitoring platform can be understood as a monitoring platform composed of nodes that communicate through the network and coordinate work to complete common tasks. The purpose of establishing a distributed monitoring platform is to utilize more distributed monitoring platform peer centers to monitor more business resource domains.

[0057] Correspondingly, each distributed monitoring platform peer center may correspond to one business resource domain or multiple business resource domains.

[0058] Reference Figure 1As shown, it is an overall layout diagram of a distributed monitoring platform provided by an embodiment of the present disclosure. Figure 1 The layout in the figure is based on the example that the distributed monitoring platform includes three distributed monitoring platform peer centers, and each distributed monitoring platform peer center corresponds to three business resource domains.

[0059] In the case of multiple distributed monitoring platform peer centers, the distributed monitoring platform peer centers use a ring network structure for data intercommunication. When an abnormality occurs in one of the distributed monitoring platform peer centers, the business resource domain corresponding to the abnormal distributed monitoring platform peer center can be diverted to the adjacent upstream and downstream distributed monitoring platform peer centers through network switching. This achieves the effect of monitoring the business resource domain corresponding to one of the distributed monitoring platform peer centers even if an abnormality occurs in the distributed monitoring platform peer center.

[0060] It is understandable that the settings of the peer center and business resource domain of the distributed monitoring platform can be independently set by those skilled in the art according to actual needs or application scenarios, and this disclosure does not limit this.

[0061] Reference Figure 2 FIG. 1 is a flow chart of a method for monitoring a service resource domain according to an embodiment of the present disclosure, the method comprising:

[0062] 210. Determine the business resource domain to be monitored.

[0063] In this step, the business resource domain can be understood as a general term for all resources in any business system being monitored. The business resource domain to be monitored can be any, such as a financial resource domain.

[0064] 220 , receiving the monitoring data of the service resources sent by the service resource domain through the distributed monitoring platform peer center, and determining whether an abnormality occurs in the service resource domain based on the monitoring data.

[0065] In this step, the financial resource domain is taken as an example. The financial resource domain includes all financial resources in the monitored financial system, where the financial resources can be one or more of server resources, virtual machine resources, network resources or storage system resources.

[0066] Anomalies can be understood as abnormal events that occur during the operation of the financial resource domain, such as a downtime event.

[0067] Specifically, the distributed monitoring platform peer center receives the monitoring data of financial resources sent by the financial resource domain, and determines whether the financial resource domain has an abnormality based on the monitoring data.

[0068] 230. When it is determined that an abnormality occurs in the business resource domain, the distributed monitoring platform peer center is called to convert the abnormal monitoring data into abnormal alarm data.

[0069] In this step, taking the financial resource domain as an example, when it is determined that an abnormal event occurs during the operation of the financial resource domain, the distributed monitoring platform peer center is called to convert the monitoring data corresponding to the abnormal event into abnormal alarm data.

[0070] The present disclosure provides a method for business resource monitoring, which first determines the business resource domain to be monitored, where all business resources are included in the business resource domain, receives monitoring data of the business resources sent by the business resource domain through the distributed monitoring platform peer center in the distributed monitoring platform, and calls the distributed monitoring platform peer center to determine whether an abnormality has occurred in the business resource domain based on the monitoring data. Each distributed monitoring platform peer center corresponds to at least one business resource domain, so by using the distributed monitoring platform peer center to determine whether an abnormality has occurred in the business resource domain, it is possible to monitor a large number of business resource domains; when it is determined that an abnormality has occurred in the business resource domain, the distributed monitoring platform peer center is called to convert the abnormal monitoring data into abnormal alarm data. The present disclosure achieves full coverage monitoring of multiple business resources by establishing a distributed monitoring platform.

[0071] The method provided in the embodiment of the present disclosure further includes:

[0072] In the case of a newly added business resource domain, the corresponding relationship between the newly added business resource domain and the distributed monitoring platform peer center is determined, and the information of the newly added business resource domain is published to the distributed monitoring platform.

[0073] In this step, when there are multiple distributed monitoring platform peer centers, different distributed monitoring platform peer centers can correspond to different business resource domains. Therefore, in the case of a new business resource domain, the correspondence between the new business resource domain and the distributed monitoring platform peer center is determined, and the information of the new business resource domain is published to the distributed monitoring platform so that the distributed monitoring platform can monitor the new business resource domain.

[0074] The method provided in the embodiment of the present disclosure includes the following steps 211 to 212 before step 210:

[0075] Step 211: Generate a random traffic reservation table based on the data reception index of the distributed monitoring platform peer center, wherein the random traffic reservation table includes the amount of data sent in the time and service resource domains.

[0076] In this step, generating a random traffic reservation table based on the data reception indicators of the distributed monitoring platform peer center can be understood as calculating and processing the free space of business resources in the business resource domain within a certain time window T, and generating random and uniform data reception indicators in time within the next time window T. The distributed monitoring platform peer center will collect the generated data reception indicators to generate a random traffic reservation table.

[0077] Step 212: Send the random traffic reservation table to the business resource domain, and call the business resource domain to divide the business resources in the business resource domain into levels according to the time of the random traffic reservation table and the amount of data sent by the business resource domain, and determine the various levels of the business resource domain; wherein, the various levels of the business resource domain have a priority order.

[0078] In this step, the random traffic reservation table may be sent to the service resource domain via a network, and the random traffic reservation table may be used to achieve temporal balance of monitoring data.

[0079] Each level of the business resource domain has a priority order, and the priority order can be set according to the specific application scenario. This disclosure does not limit the priority order.

[0080] Correspondingly, the business resources in the business resource domain are divided into levels to determine the levels of the business resource domain. For example, the business resource domain is a financial resource domain.

[0081] Reference Figure 3 As shown in FIG, a block diagram of each level of the financial resource domain A provided by the embodiment of the present disclosure is provided. For the financial resource domain A, a virtualization or cloud platform architecture can be adopted to divide the financial resources A into: available zone collectors, fault domain collectors, cabinet collectors, server collectors, and virtual machine collectors; refer to FIG. Figure 4 As shown, a block diagram of the various levels of the financial resource domain B provided by the embodiment of the present disclosure is provided. A distributed physical node architecture can be adopted for financial resource B, and financial resource B can be divided into: available zone collectors, fault domain collectors, cabinet collectors, and server collectors. Among them, the collectors of monitoring data at each level are designed for the financial resources at that level. For example, the available zone collectors can include the power supply status, temperature and humidity status, or personnel on-the-job status of the data center. It is understandable that the architecture adopted by each financial resource domain and the content contained in each collector can be independently set by those skilled in the art based on actual needs or application scenarios, and this disclosure does not limit this.

[0082] In the method provided by the embodiment of the present disclosure, the distributed monitoring platform peer center includes a first-level diversion node, a second-level diversion node and a parallel data processing node.

[0083] Specifically, the distributed monitoring platform peer center can be implemented in a hierarchical manner using cloud platform technology, and the distributed monitoring platform peer center can be divided into first-level diversion nodes, second-level diversion nodes and parallel data processing nodes accordingly.

[0084] Reference Figure 5 FIG. 1 is a flowchart of determining whether an abnormality occurs in the service resource domain according to an embodiment of the present disclosure, including:

[0085] 510 , calling the first-level diversion node to receive the monitoring data, and adding the monitoring data to the high-speed processing queue in the first-level diversion node.

[0086] In this step, a processing system for receiving monitoring data and a large-capacity memory cache system are provided in the first-level diversion node. The processing system is used to receive monitoring data, and the large-capacity memory cache system is used to add the monitoring data to the high-speed processing queue in the first-level diversion node.

[0087] Setting up a processing system for receiving monitoring data and a large-capacity memory cache system in the first-level diversion node can effectively prevent the system network traffic peak from causing excessive pressure and damage to the distributed monitoring platform peer center.

[0088] 520 , calling the secondary offload node to perform position hash processing on the monitoring data in the high-speed processing queue to determine the corresponding first parallel data processing node.

[0089] In this step, the secondary diversion node is called to receive the monitoring data from the primary diversion node, and the monitoring data is balanced in time through the random traffic reservation table. Therefore, the monitoring data in the secondary diversion node is balanced monitoring data.

[0090] Position hashing refers to storing monitoring data in a hash table, establishing a mapping relationship between the monitoring data and the storage location of the monitoring data in the hash table, so that each monitoring data corresponds to a unique location in the hash table. When a certain monitoring data needs to be obtained, the monitoring data to be obtained is mapped to the hash table through a hash function, and the corresponding storage location is the position hash corresponding to the monitoring data to be obtained.

[0091] Performing location hashing on monitoring data can achieve spatial balance of monitoring data.

[0092] 530 , calling the secondary offload node to send the monitoring data to the first parallel data processing node, and having the first parallel data processing node determine whether an abnormality occurs in the service resource domain based on the monitoring data.

[0093] In this step, the first parallel data processing node may be a running virtual machine. The first parallel processing node is responsible for processing the received monitoring data, performing unified processing, and determining whether an abnormality occurs in the service resource domain corresponding to the monitoring data.

[0094] In the method provided by the embodiment of the present disclosure, the distributed monitoring platform peer center further includes a fault aggregation node.

[0095] In this step, the fault aggregation node is responsible for collecting the fault analysis results of the parallel processing nodes.

[0096] Specifically, step 530 includes the following steps 531 to 533:

[0097] Step 531: Analyze the monitoring data through the first parallel data processing node to determine whether an abnormality occurs in the monitoring data.

[0098] Step 532: If the monitoring data is abnormal, report the monitoring data to the fault aggregation node.

[0099] Step 533: Send an abnormality inquiry message to the service resource domain corresponding to the monitoring data through the fault aggregation node, and determine whether an abnormality occurs in the service resource domain based on the returned response message.

[0100] Steps 531 to 533 are further described with the following examples:

[0101] Taking the monitoring data of financial resource domain B as an example, the monitoring data of financial resource domain B is analyzed by the first parallel data processing node to determine whether there is any abnormality in the monitoring data of financial resource domain B. If the monitoring data corresponding to the cabinet collector in financial resource domain B is down, the monitoring data corresponding to the cabinet collector needs to be reported to the fault aggregation node; the fault aggregation node sends an abnormal query message to financial resource domain B. The abnormal query message sent can be a network connection probe data message. Based on the returned response message, it is determined whether the cabinet collector in the financial resource domain B is really down.

[0102] Step 531 specifically includes:

[0103] The first parallel processing node uniformly analyzes the monitoring data sent from each level of the service resource domain to determine whether an abnormality occurs in the monitoring data at each level.

[0104] In this step, financial resource B is used as an example. Financial resource B is divided into: availability zone collectors, fault domain collectors, cabinet collectors, and server collectors. That is, the monitoring data corresponding to the availability zone collectors, fault domain collectors, cabinet collectors, and server collectors are uniformly analyzed to determine whether the monitoring data corresponding to each level is abnormal.

[0105] By determining whether the monitoring data at each level of the business resource domain is abnormal, accurate positioning can be achieved.

[0106] Step 533 specifically includes:

[0107] It is determined whether an abnormality occurs at the level of the service resource domain according to the returned response message.

[0108] In this step, the response message can be understood as a message in which the financial resource domain answers or responds to the abnormal inquiry message sent.

[0109] Reference Figure 6 As shown, the second flowchart of determining whether an abnormality occurs in the business resource domain provided by the embodiment of the present disclosure further includes:

[0110] 610 : Send the monitoring data corresponding to the high-priority service resources directly to the secondary diversion node.

[0111] In this step, taking financial resource domain A as an example, the priority order of financial resource A is: availability zone collector, fault domain collector, cabinet collector, server collector, virtual machine collector, among which availability zone collector has a high priority and virtual machine collector has a low priority.

[0112] The monitoring data corresponding to the high-priority availability zone collector in financial resource domain A is sent directly to the second-level diversion node, skipping the first-level diversion node.

[0113] 620 , calling the secondary offload node to perform location hash processing on the monitoring data corresponding to the high-priority service resource to obtain a second parallel data processing node.

[0114] In this step, the secondary offload node performs location hashing on the received monitoring data corresponding to the high-priority available zone collector to obtain a second parallel data processing node.

[0115] 630 , calling the secondary offload node to send the monitoring data to the second parallel data processing node, and having the second parallel data processing node determine whether an abnormality occurs in the service resource domain based on the monitoring data.

[0116] In this step, the second parallel data processing node receives the monitoring data corresponding to the high-priority available zone collector, and determines whether an abnormality occurs in the corresponding business resource domain.

[0117] The method provided by the embodiment of the present disclosure, step 533 further includes:

[0118] If the response message does not obtain feedback data, it is confirmed that an exception has occurred;

[0119] If the response message obtains feedback data, it is confirmed that no abnormality has occurred.

[0120] Further, the implementation of this disclosure is further explained: Figure 7 FIG. 7 is a schematic diagram of the overall process of the method for monitoring the service resource domain provided by the embodiment of the present disclosure, which specifically includes steps 710 to 780:

[0121] The method for business resource monitoring proposed in the embodiment of the present disclosure is applied to a distributed monitoring platform, and is explained by taking an example in which the distributed monitoring platform includes a distributed monitoring platform peer center, and the distributed monitoring platform peer center corresponds to a business resource domain, wherein the distributed monitoring platform peer center includes a first-level diversion node, a second-level diversion node, a parallel data processing node and a fault aggregation node.

[0122] Before monitoring a business resource domain, it is necessary to collect monitoring data for the business resource domain. The collection of monitoring data is implemented based on a hierarchical structure. Taking financial resource domain A as an example, a virtualized architecture is adopted for financial resource A, which is divided into: availability zone collector, fault domain collector, cabinet collector, server collector, and virtual machine collector. Among them, the availability zone collector has a high priority, and the virtual machine collector has a low priority.

[0123] 710 , determining financial resource domain A, and monitoring the monitoring data of the business resources in financial resource A.

[0124] 720 , calling the first-level diversion node to receive the monitoring data, and adding the monitoring data to the high-speed processing queue of the first-level diversion node.

[0125] 730 , calling the secondary offload node to perform position hash processing on the monitoring data in the high-speed processing queue to determine the corresponding first parallel data processing node.

[0126] 740 , performing unified analysis on the monitoring data sent from each level of the service resource domain A through the first parallel processing node to determine whether the monitoring data in each level is abnormal, so as to achieve accurate positioning.

[0127] 750. If the monitoring data is abnormal, the monitoring data will be reported to the fault aggregation node.

[0128] 760 , calling the fault aggregation node to send an exception query message to the business resource domain A corresponding to the monitoring data, and judging whether an exception occurs in the business resource domain A based on the returned response message.

[0129] 770. If the response message does not obtain feedback data, it is determined that an abnormality occurs in the service resource domain A.

[0130] 780. If the response message obtains feedback data, confirm that the business resource domain A has no abnormality.

[0131] Reference Figure 8 The figure shows the overall process diagram of the method for business resource domain monitoring provided by the embodiment of the present disclosure under special circumstances. There is a special case, that is, if during the business resource monitoring process, the monitoring data has been marked as high priority, such as the availability zone collector in the financial resource domain A is high priority, then steps 810 to 870 are executed:

[0132] 810. Determine the monitoring data corresponding to the collector in the high-priority availability zone in the financial resource domain A.

[0133] 820 , calling the secondary diversion node to perform location hash processing on the monitoring data corresponding to the high-priority available zone collector, and determining the corresponding second parallel data processing node.

[0134] 830 , calling the secondary offload node to send the monitoring data to the second parallel data processing node, and the second parallel data processing node determines whether an abnormality occurs in the service resource domain A based on the monitoring data.

[0135] 840. If the monitoring data is abnormal, the monitoring data will be reported to the fault aggregation node.

[0136] 850 , calling the fault aggregation node to send an exception query message to the business resource domain A corresponding to the monitoring data, and judging whether an exception occurs in the business resource domain A according to the returned response message.

[0137] 860. If the response message does not obtain feedback data, it is determined that an abnormality occurs in the service resource domain A.

[0138] 870. If the response message obtains feedback data, confirm that the business resource domain A has no abnormality.

[0139] Based on any of the above embodiments, Figure 9 A structural diagram of the business resource monitoring device provided in an embodiment of the present disclosure is provided on a distributed monitoring platform, wherein the distributed monitoring platform includes at least one distributed monitoring platform peer center, and each of the distributed monitoring platform peer centers corresponds to at least one business resource domain.

[0140] The device specifically includes:

[0141] The determination module 910 is configured to determine a service resource domain to be monitored.

[0142] The receiving module 920 is configured to receive the monitoring data of the service resources sent by the service resource domain through the distributed monitoring platform peer center, and determine whether an abnormality occurs in the service resource domain based on the monitoring data.

[0143] The conversion module 930 is used to call the distributed monitoring platform peer center to convert the monitoring data of the abnormality into abnormal alarm data when it is determined that the business resource domain has an abnormality.

[0144] The present disclosure provides a device for monitoring business resources. The device first determines the business resource domain to be monitored, which includes all business resources. The distributed monitoring platform peer center in the distributed monitoring platform receives the monitoring data of the business resources sent by the business resource domain, and calls the distributed monitoring platform peer center to determine whether an abnormality occurs in the business resource domain based on the monitoring data. Each distributed monitoring platform peer center corresponds to at least one business resource domain. Therefore, by using the distributed monitoring platform peer center to determine whether an abnormality occurs in the business resource domain, it is possible to monitor a large number of business resource domains. When it is determined that an abnormality occurs in the business resource domain, the distributed monitoring platform peer center is called to convert the abnormal monitoring data into abnormal alarm data. The present disclosure achieves full coverage monitoring of multiple business resources by establishing a distributed monitoring platform.

[0145] Based on any of the above embodiments, the apparatus further includes a new module configured to:

[0146] In the case of a newly added business resource domain, the corresponding relationship between the newly added business resource domain and the distributed monitoring platform peer center is determined, and the information of the newly added business resource domain is published to the distributed monitoring platform.

[0147] Based on any of the above embodiments, the device further includes:

[0148] A generation module is used to generate a random traffic reservation table based on the data reception index of the distributed monitoring platform peer center, wherein the random traffic reservation table includes the amount of data sent in the time and business resource domains.

[0149] A division module is used to send the random traffic reservation table to the business resource domain, and call the business resource domain to divide the business resources in the business resource domain into levels according to the time of the random traffic reservation table and the amount of data sent by the business resource domain, and determine the various levels of the business resource domain; wherein the various levels of the business resource domain have a priority order.

[0150] Based on any of the above embodiments, the distributed monitoring platform peer center includes a first-level diversion node, a second-level diversion node and a parallel data processing node.

[0151] The receiving module 920 specifically includes:

[0152] The adding unit is used to call the first-level diversion node to receive the monitoring data, and add the monitoring data to the high-speed processing queue in the first-level diversion node.

[0153] The processing unit is used to call the secondary diversion node to perform position hash processing on the monitoring data in the high-speed processing queue to determine the corresponding first parallel data processing node.

[0154] The judgment unit is configured to call the secondary diversion node to send the monitoring data to the first parallel data processing node, and judge whether an abnormality occurs in the business resource domain based on the monitoring data through the first parallel data processing node.

[0155] Based on any of the above embodiments, the distributed monitoring platform peer center further includes a fault aggregation node.

[0156] The judging unit includes:

[0157] The analyzing subunit is configured to analyze the monitoring data through the first parallel data processing node to determine whether an abnormality occurs in the monitoring data.

[0158] The reporting subunit is used to report the monitoring data to the fault aggregation node if an abnormality occurs in the monitoring data.

[0159] The judgment subunit is configured to send an abnormality inquiry message to the service resource domain corresponding to the monitoring data through the fault aggregation node, and judge whether an abnormality occurs in the service resource domain according to the returned response message.

[0160] Based on any of the above embodiments, the analysis subunit is specifically configured to:

[0161] Performing a unified analysis of the monitoring data sent by each level of the service resource domain through the first parallel processing node to determine whether an abnormality occurs in the monitoring data at each level;

[0162] The judgment subunit is specifically used to:

[0163] It is determined whether an abnormality occurs at the level of the service resource domain according to the returned response message.

[0164] Based on any of the above embodiments, the receiving module 920 further includes:

[0165] The sending unit is used to send the monitoring data corresponding to the high-priority business resources directly to the secondary diversion node.

[0166] An acquisition unit is configured to call the secondary diversion node to perform location hash processing on the monitoring data corresponding to the high-priority business resource, and acquire a second parallel data processing node.

[0167] The calling unit is configured to call the secondary diversion node to send the monitoring data to the second parallel data processing node, and determine whether an abnormality occurs in the business resource domain based on the monitoring data by the second parallel data processing node.

[0168] Based on any of the above embodiments, judging whether an abnormality occurs at the level of the service resource domain according to the returned response message is specifically used for:

[0169] If the response message does not obtain feedback data, it is confirmed that an exception has occurred.

[0170] If the response message obtains feedback data, it is confirmed that no abnormality has occurred.

[0171] Figure 10 An example of a physical structure diagram of an electronic device is shown below. Figure 10 As shown, the electronic device may include: a processor 1001, a communications interface 1002, a memory 1003, and a communication bus 1004, wherein the processor 1001, the communications interface 1002, and the memory 1003 communicate with each other via the communication bus 1004. The processor 1001 may call the logic instructions in the memory 1003 to execute a method for monitoring business resources, the method comprising: determining a business resource domain to be monitored; receiving, through the distributed monitoring platform peer center, monitoring data of business resources sent by the business resource domain, and determining whether an abnormality has occurred in the business resource domain based on the monitoring data; and, if it is determined that an abnormality has occurred in the business resource domain, calling the distributed monitoring platform peer center to convert the abnormal monitoring data into abnormality alarm data.

[0172] In addition, the logic instructions in the above-mentioned memory 1003 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the embodiment of the present disclosure is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0173] On the other hand, the present disclosure also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the business resource monitoring method provided by the above methods, which method includes: determining the business resource domain to be monitored; receiving monitoring data of the business resources sent by the business resource domain through the distributed monitoring platform peer center, and judging whether an abnormality occurs in the business resource domain based on the monitoring data; when it is determined that an abnormality occurs in the business resource domain, calling the distributed monitoring platform peer center to convert the abnormal monitoring data into abnormal alarm data.

[0174] On the other hand, the present disclosure also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the above-mentioned business resource monitoring methods, the methods comprising: determining the business resource domain to be monitored; receiving monitoring data of the business resources sent by the business resource domain through the distributed monitoring platform peer center, and judging whether an abnormality occurs in the business resource domain based on the monitoring data; and when it is determined that an abnormality occurs in the business resource domain, calling the distributed monitoring platform peer center to convert the abnormal monitoring data into abnormal alarm data.

[0175] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0176] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.

Claims

1. A method for monitoring business resources, characterized in that: Applied to a distributed monitoring platform, wherein the distributed monitoring platform includes at least one distributed monitoring platform peer center, each of which corresponds to at least one business resource domain; the distributed monitoring platform peer centers communicate with each other via a network and coordinate their work to complete a common task; The method specifically includes: Determine the business resource domain to be monitored; Receiving, through the distributed monitoring platform peer center, monitoring data of the service resources sent by the service resource domain, and determining whether an abnormality occurs in the service resource domain based on the monitoring data; When it is determined that an abnormality occurs in the business resource domain, calling the distributed monitoring platform peer center to convert the abnormal monitoring data into abnormal alarm data; Before determining the service resource domain to be monitored, the following steps are included: Generate a random traffic reservation table based on the data reception index of the peer center of the distributed monitoring platform; the random traffic reservation table includes the time and the amount of data sent in the service resource domain; The random traffic reservation table is sent to the business resource domain, and the business resource domain is called to divide the business resources in the business resource domain into levels according to the time of the random traffic reservation table and the amount of data sent by the business resource domain, so as to determine the various levels of the business resource domain; wherein the various levels of the business resource domain have a priority order; the random traffic reservation table enables the monitoring data to be balanced in time.

2. The method for monitoring business resources according to claim 1, wherein: The method further comprises: In the case of a newly added business resource domain, the corresponding relationship between the newly added business resource domain and the distributed monitoring platform peer center is determined, and the information of the newly added business resource domain is published to the distributed monitoring platform.

3. The method for monitoring business resources according to claim 1, wherein: The distributed monitoring platform peer center includes a first-level diversion node, a second-level diversion node and a parallel data processing node; The receiving, through the distributed monitoring platform peer center, the monitoring data of the service resources sent by the service resource domain, and determining whether an abnormality occurs in the service resource domain according to the monitoring data, includes: Calling the first-level shunt node to receive the monitoring data, and adding the monitoring data to a high-speed processing queue in the first-level shunt node; Calling the secondary shunt node to perform position hash processing on the monitoring data in the high-speed processing queue to determine the corresponding first parallel data processing node; The secondary diversion node is called to send the monitoring data to the first parallel data processing node, and the first parallel data processing node determines whether an abnormality occurs in the business resource domain based on the monitoring data.

4. The method for monitoring business resources according to claim 3, wherein: The distributed monitoring platform peer center also includes a fault aggregation node; The determining, by the first parallel data processing node based on the monitoring data, whether an abnormality occurs in the service resource domain includes: Analyzing the monitoring data by the first parallel data processing node to determine whether an abnormality occurs in the monitoring data; If the monitoring data is abnormal, the monitoring data is reported to the fault aggregation node; An abnormality inquiry message is sent to the service resource domain corresponding to the monitoring data through the fault aggregation node, and whether an abnormality occurs in the service resource domain is determined according to the returned response message.

5. The method for monitoring business resources according to claim 4, characterized in that: The analyzing the monitoring data by the first parallel data processing node to determine whether an abnormality occurs in the monitoring data includes: Performing a unified analysis of the monitoring data sent by each level of the service resource domain through the first parallel processing node to determine whether an abnormality occurs in the monitoring data at each level; The determining whether an abnormality occurs in the service resource domain according to the returned response message includes: It is determined whether an abnormality occurs at the level of the service resource domain according to the returned response message.

6. The method for monitoring business resources according to claim 3, wherein: The receiving, through the distributed monitoring platform peer center, the monitoring data of the service resources sent by the service resource domain, and determining whether an abnormality occurs in the service resource domain according to the monitoring data, further includes: Sending the monitoring data corresponding to the high-priority service resources directly to the secondary diversion node; Calling the secondary offload node to perform location hash processing on the monitoring data corresponding to the high-priority service resource to obtain a second parallel data processing node; The secondary diversion node is called to send the monitoring data to the second parallel data processing node, and the second parallel data processing node determines whether an abnormality occurs in the business resource domain based on the monitoring data.

7. The method for monitoring service resources according to claim 4, wherein: The determining whether an abnormality occurs in the service resource domain according to the returned response message includes: If the response message does not obtain feedback data, it is confirmed that an exception has occurred; If the response message obtains feedback data, it is confirmed that no abnormality has occurred.

8. A device for monitoring business resources, characterized in that: Set up on a distributed monitoring platform, wherein the distributed monitoring platform includes at least one distributed monitoring platform peer center, each of the distributed monitoring platform peer centers corresponds to at least one business resource domain; the distributed monitoring platform peer centers communicate with each other through a network and coordinate work to complete a common task; The device specifically includes: A determination module, used to determine the business resource domain to be monitored; a receiving module, configured to receive monitoring data of the service resources sent by the service resource domain through the peer center of the distributed monitoring platform, and determine whether an abnormality occurs in the service resource domain based on the monitoring data; A conversion module, configured to, when determining that an abnormality occurs in the business resource domain, call the distributed monitoring platform peer center to convert the abnormal monitoring data into abnormal alarm data; A generation module is used to generate a random traffic reservation table based on the data reception index of the distributed monitoring platform peer center; the random traffic reservation table includes the time and the amount of data sent in the business resource domain; the random traffic reservation table is sent to the business resource domain, and the business resource domain is called to hierarchically divide the business resources in the business resource domain according to the time of the random traffic reservation table and the amount of data sent in the business resource domain, and determine the various levels of the business resource domain; wherein the various levels of the business resource domain have a priority order; the random traffic reservation table enables the monitoring data to be balanced in time.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for monitoring service resources according to any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the service resource monitoring method according to any one of claims 1 to 7 are implemented.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the service resource monitoring method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method for monitoring grade classification of virtual machine under cloud data center environment

    CN104360924A

  • Large-scale electric power equipment monitoring and alarm data real-time processing method and system

    CN107968840A

  • Competition-based resource reservation method and equipment

    CN111836370A

  • Service system fault analysis method and device

    CN112269718A