Problem analysis early warning method and device

By dividing multi-layer structures in CDN scenarios and monitoring indicator factors, the problem of low troubleshooting efficiency in CDN scenarios is solved, rapid positioning and handling of faults is achieved, and troubleshooting efficiency is improved.

CN120066900APending Publication Date: 2025-05-30SHANGHAI BILIBILI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510218346.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In CDN scenarios, it is difficult to check the faults of business scenarios in a timely manner, and it is impossible to alert the data interfaces, and the alarms of multiple basic indicators cannot be associated, resulting in low fault processing efficiency.

Method used

Through a multi-layer structure based on business scenario division, the indicator factors and their association relationships of each layer are determined, the indicator factors of non-business layers are monitored, whether they meet the early warning conditions, and problem positioning and alarming is carried out based on the association relationship.

Benefits of technology

It realizes the rapid positioning and handling of faults in business scenarios, improves the efficiency of troubleshooting, promptly warns of business appearance problems, and reduces the time for manual inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066900A_ABST
    Figure CN120066900A_ABST
Patent Text Reader

Abstract

The invention discloses a problem analysis and early warning method and device. The method comprises the steps that a multi-layer structure is obtained based on business scene division; wherein the multi-layer structure has an up-and-down hierarchical relationship and comprises a business layer and a non-business layer; determining an index factor of each layer in the multi-layer structure and an association relationship between the index factors of each layer; monitoring an index factor of a non-service layer, and judging whether the index factor accords with an early warning condition or not; and if the pre-warning condition is met, performing problem positioning alarm according to the met pre-warning condition and the incidence relation between the index factors. According to the multi-layer structure and the incidence relation of the index factors, by monitoring the basic index factors (such as basic data, equipment and the like) of each non-business layer, the abnormity existing in the basic index factors can be found in time, the business presentation problem can be early warned, and the troubleshooting of the problem is accelerated through the incidence relation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technologies, and in particular, to a problem analysis and warning method and apparatus. Background Art

[0002] In the CDN scenario, a business scenario is split into multiple microservices, such as dozens of microservices. Each microservice has its own responsibilities and is responsible for its own related work. The overall business process is executed after being connected in series by the business itself. When a failure occurs during the execution of a business scenario, the failures mainly come from the following aspects: basic environment failures and business failures. Basic environment failures refer to failures of the hardware (such as CPU, memory, disk, etc.) for deploying services, the operating system (such as file descriptors, sockets, file systems, etc.), the network, etc.; business failures refer to failures caused by the business itself, such as business changes, bugs existing in the business itself, etc. However, when a failure occurs, it is impossible to quickly troubleshoot problems, impossible to alarm data class interfaces, and impossible to associate the alarms of multiple basic metrics. Therefore, there is an urgent need for a problem analysis and warning method. Summary of the Invention

[0003] In view of the above problems, embodiments of the present application are proposed to provide a problem analysis and warning method and apparatus that overcome the above problems or at least partially solve the above problems.

[0004] According to a first aspect of the embodiments of the present application, a problem analysis and warning method is provided, which includes:

[0005] Obtaining a multi-layer structure based on the division of the business scenario; wherein, there is an upper and lower hierarchical relationship between the multi-layer structures, including a business layer and a non-business layer;

[0006] Determining the index factors of each layer in the multi-layer structure, and the association relationship between the index factors of each layer;

[0007] Monitoring the index factors of the non-business layer to determine whether the warning conditions are met;

[0008] If the warning conditions are met, perform problem location and warning according to the met warning conditions and the association relationship between the index factors.

[0009] Optionally, determining the index factors of each layer in the multi-layer structure, and the association relationship between the index factors of each layer further includes:

[0010] According to the multi-layer structure, respectively determining the index factors of each layer; wherein, the index factors include custom factors;

[0011] According to the upper and lower hierarchical relationship of the multi-layer structure, constructing the association relationship between the index factors of the lower layer and the index factors of the upper layer.

[0012] Optionally, the association relationships between the metric factors include one-to-one, one-to-many, and / or many-to-one association relationships.

[0013] Optionally, monitoring the metric factors of the non-business layer and determining whether the warning conditions are met further includes:

[0014] According to the metric factors of each non-business layer, pre-construct warning conditions corresponding to the metric factors;

[0015] Based on the metric factors of each layer, monitor the monitoring objects corresponding to the metric factors in the business scenario in real time, and determine whether the warning conditions are met according to the real-time execution situation of the monitoring objects.

[0016] Optionally, performing problem location and warning according to the met warning conditions and the association relationships between the metric factors further includes:

[0017] Determine the target metric factors of the corresponding target level according to the met warning conditions;

[0018] According to the association relationships between the metric factors, determine the metric factors of the business layer associated with the target metric factors, and / or other associated metric factors at the non-business level;

[0019] Determine business warning information according to the metric factors of the business layer, and perform problem location according to the target metric factors and / or associated metric factors.

[0020] Optionally, the multi-layer structure includes: a data collection layer, a basic metric layer, a data processing layer, a perception layer, and a business layer.

[0021] Optionally, the method further includes:

[0022] According to the problem location warning, perform warning correction processing according to the preset processing rules.

[0023] According to the second aspect of the embodiments of the present application, a problem analysis and warning device is provided, which includes:

[0024] A division module, adapted to divide a multi-layer structure based on a business scenario; wherein, there is an upper and lower hierarchical relationship between the multi-layer structures, including a business layer and a non-business layer;

[0025] A relationship module, adapted to determine the metric factors of each layer in the multi-layer structure, and the association relationships between the metric factors of each layer;

[0026] A monitoring module, adapted to monitor the metric factors of the non-business layer and determine whether the warning conditions are met;

[0027] An alarm module, adapted to perform problem location and alarm according to the met warning conditions and the association relationships between the metric factors if the warning conditions are met.

[0028] According to a third aspect of the embodiments of the present application, a computing device is provided, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0029] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above problem analysis and warning method.

[0030] According to a fourth aspect of the embodiments of the present application, a computer storage medium is provided, and at least one executable instruction is stored in the storage medium, and the executable instruction causes the processor to perform the operations corresponding to the above problem analysis and warning method.

[0031] According to a fifth aspect of the embodiments of the present application, a computer program product is provided, including at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above problem analysis and warning method.

[0032] According to the problem analysis and warning method and device provided by the present application, according to the multi-layer structure and the correlation relationship of the index factors, the anomalies existing in the basic index factors (such as basic data, devices, etc.) of each non-business layer can be detected in time, the problems of business appearance can be warned, and the investigation of problems can be accelerated through the correlation relationship.

[0033] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically exemplified below. Description of the Drawings

[0034] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0035] Figure 1 A flowchart of a problem analysis and warning method according to an embodiment of the present application is shown;

[0036] Figure 2 A flowchart of a problem analysis and warning method according to another embodiment of the present application is shown;

[0037] Figure 3 A multi-layer structure diagram is shown;

[0038] Figure 4Shows a schematic diagram of the index factors of a multi-layer structure;

[0039] Figure 5 Shows a schematic diagram of the structure of a problem analysis and warning device according to an embodiment of the present application;

[0040] Figure 6 Shows a schematic diagram of the structure of a computing device according to an embodiment of the present application. Detailed implementation manners

[0041] Hereinafter, exemplary embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be completely conveyed to those skilled in the art.

[0042] First, the noun terms involved in one or more embodiments of the present application are explained.

[0043] Patrol inspection: In the CDN scenario, daily interface call checks are performed on the basic services of the CDN, and the content returned by the interface is verified based on experience. If problems are found, alarms are issued.

[0044] Basic metrics: Metrics that are not related to the business, common metrics in computer software, such as various metrics of CPU, memory, disk, file descriptors, network, etc.

[0045] Business metrics: Metrics that are related to the business and are strongly related to the business.

[0046] Figure 1 Shows a flowchart of a problem analysis and warning method according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:

[0047] Step S101, obtaining a multi-layer structure based on business scenario division.

[0048] Before performing problem analysis and warning on an overall business scenario, the structure can be divided based on the business scenario. When dividing, it can be divided according to the functional services provided by the business in the business scenario, or different levels can be determined according to the granularity of demand splitting, etc. In the CDN scenario, the business includes, for example, live broadcast services, on-demand services, etc. For specific business implementations, different microservices can include, for example, data acquisition, processing, generation of business data, business data push, business implementation, etc. Based on different business processes and functions provided by microservices in the business scenario, the business scenario is divided, and the specific division rules are not limited here.

[0049] In this embodiment, the number of layers of the multi-layer structure is not limited. When there are more layers, more subsequent associated index factors will be involved, the business division will be more precise, and it will also be convenient for problem location and alarm. Specifically, it is divided according to the actual implementation situation and is not limited here.

[0050] When splitting, there is an upper and lower hierarchical relationship between the multi-layer structures. That is, in the multi-layer structure, the output of any lower layer serves as the input of the upper layer, and there is an upper and lower hierarchical relationship between the two layers. The multi-layer structure includes a business layer and a non-business layer. The non-business layer is the lower level of the business layer. The business layer is the specific functional layer of the business, which corresponds to the display of the surface problems of the business and is strongly associated with the specific business. The non-business layer can include, for example, data-related layers (such as basic data collection, processing, business data processing, etc.), device-related layers (such as the resource usage of devices, etc.). The non-business layer is more about collecting and processing basic data, business data, etc., and device resource occupancy, which has no or weak association with the business. Regular inspections can be carried out on the non-business layer to facilitate timely problem discovery.

[0051] Step S102: Determine the index factors of each layer in the multi-layer structure and the association relationship between the index factors of each layer.

[0052] After obtaining the multi-layer structure through division, for each level in the multi-layer structure, index factors can be defined for each level. When determining the index factors, it can be based on being easily obtainable in the business scenario and reflecting the important characteristics of the business scenario. For example, in the business scenario of a video website, play data, resource occupancy, bit rate, etc. can all be used as index factors.

[0053] In the multi-layer structure, corresponding index factors need to be determined for each layer according to the different layers divided. The index factors of the business layer can be determined according to the requirements of the business scenario, such as business indicators, which can include custom business-related index factors. In the non-business layer, for example, the index factors of the device-related layer can be set as device status, CPU, memory, disk, network, etc., and the index factors of the data-related layer can be set as data storage-related index factors, data processing-related index factors, etc. Data storage-related index factors such as databases, message storage queues, logs, caches, etc., and data processing-related index factors such as data classification, sorting, etc. The index factors of the non-business layer can also be custom index factors, which can include basic indicators, etc. The above are just examples and are specifically set according to the actual implementation situation and are not limited here.

[0054] Between the upper and lower levels in the multi-layer structure, the output of the index factors at the lower level is used as the input of the index factors at the upper level. When establishing the correlation relationship between the index factors of each layer, one or more index factors at the upper level can be selected from multiple index factors at the lower level to construct the correlation relationship. It is also possible to construct the correlation relationship between multiple index factors at the upper level and one index factor at the lower level. That is, the correlation relationship between the index factors includes one-to-one, one-to-many, or many-to-one correlation relationships. The specific correlation relationship is constructed according to the actual situation and is not limited here.

[0055] Step S103: Monitor the index factors of the non-business layer and determine whether they meet the warning conditions.

[0056] The index factors of the business layer correspond to the functions of the business and present the superficial problems during the execution of the business scenario. In this embodiment, by monitoring the index factors of the non-business layer, based on the daily inspection of data and equipment, the abnormal monitoring of the index factors of the non-business layer can be gradually lifted and aggregated to the business layer, more clearly pointing to the specific superficial problems in the business layer. That is, the basic index factors of the non-business layer are pre-associated with the superficial problems in the business layer, and there is no need for developers to check problems one by one. It is possible to directly associate and locate problems between the superficial problems and the index factors of the non-business layer based on the monitoring of the index factors of the non-business layer.

[0057] Specifically, based on the multi-layer structure, the index factors in the non-business layer can be monitored, and the monitoring can be continuously performed as part of the daily inspection. By monitoring the index factors of the non-business layer, such as monitoring the device status, monitoring the resource occupancy of CPU, memory, disk, network, etc., monitoring the data storage situation, monitoring the data processing status, monitoring the data acquisition situation, etc., based on the monitoring, it is determined whether the warning conditions are met. The warning conditions can be set in advance according to the business scenario, and the warning conditions can include one or more. For example, the resource occupancy is higher than the preset occupancy threshold, such as 80%, and the packet loss rate calculated from the data is greater than the preset packet loss threshold, such as 20%. The above is for illustrative purposes, and the specific settings are based on the actual situation and are not limited here. When making the judgment, it can be determined according to the business scenario. For example, when one warning condition is met, it is determined that the warning conditions are met, or when multiple combined warning conditions are met, it is determined that the warning conditions are met. The specific settings are based on the actual situation and are not limited here.

[0058] When it is determined that the warning conditions are met, step S104 is executed; otherwise, no processing is required and the monitoring continues.

[0059] Step S104: Locate and alarm the problem according to the met warning conditions and the correlation relationship between the index factors.

[0060] When it is determined that the warning conditions are met, problem location can be performed according to the met warning conditions. For example, according to the met warning conditions, the target index factors at the corresponding target level are determined. According to the association relationship, it can be gradually ascended to the surface problems corresponding to the business layer, or it can be determined whether there are problems with other associated index factors in each non-business layer. When there are problems with the target index factors at the target level in the non-business layer, there may also be problems with other associated index factors at non-target levels, and further investigation is required.

[0061] Based on the determined target index factors and other associated index factors at non-target levels, problem location can be performed. According to the index factors of the business layer associated with the target index factors, the surface problems are determined, so that problem location warnings can be issued. The content of the warning can be determined according to the surface problems, so that the abnormal situation can be clearly understood according to the content of the warning. Problem location can directly locate to the abnormal occurrence point without checking one by one, so as to realize the monitoring of the basic index factors in the non-business layer, warn the surface problems of the business, and realize rapid problem location and investigation through the association relationship.

[0062] According to the problem analysis and warning method provided by the present application, according to the multi-layer structure and the association relationship of the index factors, the anomalies existing in the basic index factors (such as basic data, equipment, etc.) of each non-business layer can be detected in time, the surface problems of the business can be warned, and the investigation of the problems can be accelerated through the association relationship.

[0063] Figure 2 The flowchart of the problem analysis and warning method according to an embodiment of the present application is shown, as Figure 2 shown, the method includes the following steps:

[0064] Step S201, a multi-layer structure is obtained based on the division of the business scenario.

[0065] When dividing it into a multi-layer structure based on the business scenario, it can be divided into a business layer and a non-business layer. Among them, the business layer is the upper-layer structure, and the non-business layer is the lower-layer structure. The business layer corresponds to the surface problems in the process of executing the business scenario, and the non-business layer corresponds to the basic data, equipment, etc. When dividing, it can be divided into different levels according to the requirements. The more levels there are, the more index factors are set, the finer the association is, and the more accurate the business directivity is. It is specifically set according to the actual situation and is not limited here.

[0066] The obtained multi-layer structure can be as Figure 3 shown, including: a data acquisition layer, a basic index layer, a data processing layer, a perception layer, and a business layer. The upper and lower layer relationships between the layers are as Figure 3As shown in the figure, the business layer is at the topmost level, the data collection layer is at the bottommost level, and in the middle are the basic index layer, the data processing layer, and the perception layer in sequence. The above-described multi-layer structure is for illustrative purposes, and can be divided according to the actual implementation situation, which is not limited here.

[0067] Step S202: According to the multi-layer structure, determine the index factors of each layer respectively, and construct the association relationship between the index factors of the lower layer and the index factors of the upper layer according to the upper and lower layer relationship of the multi-layer structure.

[0068] In the multi-layer structure, the bottommost layer is the data collection layer. The data collection layer collects data according to data sources such as databases, message queues, logs, caches, and monitoring related to the business and basic services, and can store it in the data collection layer table for subsequent business processing. For the data collection layer, the determined index factors it contains can be as Figure 4 shown, including index factors such as databases, message queues, caches, and logs, and other index factors of other data sources can also be customized, which is not limited here. The basic index layer is related to the device's own state and the usage of device resources. Above the data collection layer, the basic index layer is set, and the index factors it contains are as Figure 4 shown, including device status, CPU, memory, disk, network, etc. The device status includes the status of each device (such as servers, databases, etc.), and the index factors such as CPU, memory, and disk are used to determine the resource usage; the network resources include the total network bandwidth, network usage resources, etc. The basic index layer can customize other index factors of devices, resources, etc., which is not limited here. The data processing layer includes various calculations on the data, and its index factors include processing such as normalization, weighted average, time filtering, sorting, and classification. Here is for illustrative purposes, and corresponding index factors can be set according to the actual implementation situation for the calculation and processing of data, which is not limited here. The perception layer is used to perceive data related to the business. Taking video live broadcast and on-demand as examples, the index factors it contains include index factors concerned about live broadcast lags such as the bit rate is greater than 6M, the change history data blocks of the live broadcast service in the past 5 minutes, the CPU (corresponding to the live broadcast server) is greater than 80%, the network packet loss rate (corresponding to the live broadcast server) is greater than 20%, etc., and, the disk (on-demand server) is greater than 95%, the change history data blocks of the on-demand service in the past 5 minutes, the network packet loss rate (corresponding to the on-demand server) is greater than 20%, the CPU (corresponding to the on-demand server) is greater than 80%, etc. index factors related to on-demand, all of which belong to the index factors concerned about on-demand lags. As the topmost layer, the business layer's concerned index factors are related to the surface problems. Taking live broadcast and on-demand as examples, the index factors of the business layer include live broadcast lags and on-demand lags. For other business scenarios, the perception layer and the business layer can determine the index factors of the perception layer and the business layer according to the specific business, which is not limited here.

[0069] When determining the index factors of each layer, the correlation relationship between the index factors of the upper and lower levels can be constructed based on the already determined index factors of the lower level. For example, after the index factors of the basic index layer and the data collection layer are determined, the correlation relationship between the index factors of the basic index layer and the data collection layer can be constructed. As shown by the line segments with inclusion arrows between the layers in Figure 4 , there are line segments with inclusion arrows between the index factor database of the data collection layer and the index factor network and CPU of the basic index layer, indicating that there is a correlation relationship between the index factor database of the data collection layer and the index factor network and CPU of the basic index layer. For example, data is read from the database through the network, and CPU resources are used during the reading process; there is a line segment with an inclusion arrow between the index factor log of the data collection layer and the index factor disk of the basic index layer, indicating that there is a correlation relationship between the index factor log of the data collection layer and the index factor disk of the basic index layer, such as reading the log from the disk. There is a correlation relationship between the index factor CPU of the basic index layer and the index factor normalization of the data processing layer, a correlation relationship between the index factor normalization of the data processing layer and the index factor CPU (live server) of the perception layer being greater than 80%, a correlation relationship between the index factor CPU (live server) of the perception layer being greater than 80% and the index factor live freeze of the service layer, a correlation relationship between the index factor disk of the basic index layer and the index factor sorting of the data processing layer, a correlation relationship between the index factor sorting of the data processing layer and the index factor disk (VOD server) of the perception layer being greater than 95%, a correlation relationship between the index factor disk (VOD server) of the perception layer being greater than 95% and the index factor VOD freeze of the service layer, a correlation relationship between the index factor network of the basic index layer and the index factor time filtering of the data processing layer, a correlation relationship between the index factor normalization of the data processing layer and the index factor network packet loss rate (live server) of the perception layer being greater than 20%, a correlation relationship between the index factor network packet loss rate (live server) of the perception layer being greater than 20% and the index factor live freeze of the service layer, a correlation relationship between the index factor bit rate of the perception layer being greater than 6M, the change history data block of the live service in the past 5 minutes, the CPU (live corresponding to the live server) being greater than 80%, and the network packet loss rate (live corresponding to the live server) being greater than 20% and the index factor live freeze of the service layer, a correlation relationship between the index factor disk (VOD server) of the perception layer being greater than 95%, the change history data block of the VOD service in the past 5 minutes, the network packet loss rate (VOD corresponding to the VOD server) being greater than 20%, and the CPU (VOD corresponding to the VOD server) being greater than 80% and the index factor VOD freeze of the service layer. There is a correlation relationship between the index factors of adjacent upper and lower levels, so that when monitoring the index factors of the non-service layer and finding abnormalities in the index factors of the non-service layer, the reasons for the corresponding surface problems can be gradually lifted layer by layer, and the surface problems can be solved in a timely manner.

[0070] Figure 4Only some of the association relationships are shown, and not all association relationships are shown. In specific implementation, association relationships are established for the index factors between adjacent layers according to the implementation situation, which is not limited here. When establishing the association relationship between the index factors of adjacent two levels, an index factor at the lower level can be associated with an index factor at the upper level. For example, the index factor log at the data collection layer is associated with the index factor disk at the basic index layer. An index factor at the lower level can also be associated with multiple index factors at the upper level. For example, the index factor database at the data collection layer is associated with the index factors network and CPU at the basic index layer. Multiple index factors at the lower level can also be associated with an index factor at the upper level. For example, the index factors disk (VOD server) > 95%, change history data block of VOD service in the past 5 minutes, network packet loss rate (VOD corresponding to VOD server) > 20%, and CPU (VOD corresponding to VOD server) > 80% at the perception layer are associated with the index factor VOD lag at the service layer. It is specifically set according to the implementation situation and is not limited here.

[0071] Step S203: According to the index factors of each non-service layer, pre-construct warning conditions corresponding to the index factors.

[0072] For each index factor of each non-service layer, corresponding warning conditions can be pre-constructed respectively. The warning conditions can include setting a threshold or a threshold range for the index factor. When the threshold or the threshold range is exceeded, a warning can be issued, etc. Such as the threshold of database connection duration, the threshold of remaining network bandwidth, the threshold of CPU occupancy rate, the threshold of normalization processing time, etc. The threshold or the threshold range can be set according to historical experience data and can also be adjusted in combination with real-time situations. It is specifically set according to the implementation situation in combination with specific services and is not limited here.

[0073] Furthermore, when pre-constructing the warning conditions, relevant warning conditions can be constructed for one index factor, such as setting the occupancy rate threshold of the CPU, or relevant warning conditions can be jointly constructed for multiple index factors, such as when multiple index factors all exceed their respective thresholds, it is a warning condition, etc., which is not limited here.

[0074] Step S204: Based on the index factors of each layer, monitor in real time the monitoring objects corresponding to the index factors in the business scenario, and judge whether it meets the warning conditions according to the real-time execution situation of the monitoring objects.

[0075] According to the index factors of each layer, the monitoring objects corresponding to each index factor in the real-time monitoring business scenario can be monitored through daily inspections, etc., such as monitoring the usage of resources such as the CPU, memory, and disk of the device, monitoring the network usage bandwidth, remaining bandwidth, network lag situation, etc., monitoring database connections, database access, etc., monitoring the processing duration of various data, etc., which will not be listed one by one here.

[0076] Monitoring can adopt real-time monitoring to obtain the real-time execution situation of each monitoring object, and compare it with the pre-constructed warning conditions to determine whether it meets the warning conditions. By judging whether it meets the warning conditions, it is determined whether to give a warning. If it meets the warning conditions, step S205 is executed; otherwise, no processing is required and the monitoring continues.

[0077] Step S205: Locate and alarm the problem according to the met warning conditions and the correlation relationship between the index factors.

[0078] After determining that the warning conditions are met, according to the met warning conditions, the target index factors of the target layer corresponding to the warning conditions can be determined first. For example, if the met warning condition is that the network remaining bandwidth is less than the threshold, the target index factor corresponding to the target layer can be determined as the network in the basic index layer.

[0079] According to the correlation relationship between the index factors, the index factors of the business layer associated with the target index factor, as well as other non-business-level associated index factors, etc. can be determined. For example Figure 4 in, according to the correlation relationship, floating up layer by layer, the index factor of the business layer such as live broadcast lag can be determined. If the target index factor is that the packet loss rate > 20%, according to the correlation relationship, the non-business layer associated index factors are found, including time filtering, network, database, etc. According to the index factors of the business layer, business alarm information can be determined. For example, the alarm information is "There is a risk of live broadcast lag". According to the target index factor and the associated index factors, the problem is located. When locating the problem, it can be located according to the packet loss rate, time filtering, network, database, etc. It is possible that abnormalities occur in all non-business layers, or it is possible that the index factors of a certain layer are abnormal. One or more of the target index factor and the associated index factors can be targeted for investigation to reduce large-scale investigation and improve the problem location efficiency.

[0080] Step S206: Perform alarm correction processing according to the problem location alarm according to the preset processing rules.

[0081] After the problem location warning, based on the accuracy of the problem location warning, it is possible to automatically perform warning correction processing on the target index factor and the associated index factor, so as to automatically troubleshoot and solve problems. For example, if the problem location warning determines that the high CPU load in the basic index layer causes live broadcast lag and the specific device can be located, the corresponding device can be automatically taken offline, and the user traffic on the device can be scheduled to other devices, thus solving the root cause of the live broadcast lag problem.

[0082] Furthermore, during the warning correction processing, it is possible to perform automatic warning correction processing on different problem causes according to the preset processing rules, which is convenient for automatically solving the problem causes and improving the execution efficiency of the business scenario.

[0083] According to the problem analysis and warning method provided by this application, the business scenario is divided into a multi-layer structure of a business layer and a non-business layer, and the index factors of each layer in the multi-layer structure and the association relationship of the index factors are set, so as to pre-define the association between the surface problems in the business layer and the causes of the problems. By monitoring the index factors of each non-business layer, it is possible to intuitively obtain the possible anomalies of each index factor in the non-business layer, so as to timely warn of the possible surface problems in the business, and then based on the warning, directly locate the cause and fundamentally solve the problem. Based on the warning, the cause of the problem can be solved through automatic correction processing, realizing the efficient and normal execution of the business scenario.

[0084] Figure 5 The structure diagram of a problem analysis and warning device provided by an embodiment of this application is shown. As Figure 5 shown, the device includes:

[0085] A division module 510, adapted to obtain a multi-layer structure based on the division of the business scenario; among them, there is an upper and lower hierarchical relationship between the multi-layer structures, including a business layer and a non-business layer;

[0086] A relationship module 520, adapted to determine the index factors of each layer in the multi-layer structure, and the association relationship between the index factors of each layer;

[0087] A monitoring module 530, adapted to monitor the index factors of the non-business layer and determine whether the warning conditions are met;

[0088] An alarm module 540, adapted to perform problem location warning according to the met warning conditions and the association relationship between the index factors if the warning conditions are met.

[0089] Optionally, the relationship module 520 is further adapted to:

[0090] According to the multi-layer structure, determine the index factors of each layer respectively; among them, the index factors include custom factors;

[0091] Construct the association relationship between the metric factors of the next layer and those of the upper layer according to the hierarchical relationship of the multi-layer structure.

[0092] Optionally, the association relationships between metric factors include one-to-one, one-to-many, and / or many-to-one association relationships.

[0093] Optionally, the monitoring module 530 is further adapted to:

[0094] Pre-construct warning conditions corresponding to the metric factors according to the metric factors of each non-business layer;

[0095] Based on the metric factors of each layer, monitor in real time the monitored objects corresponding to the metric factors in the business scenario, and determine whether the warning conditions are met according to the real-time execution situation of the monitored objects.

[0096] Optionally, the alarm module 540 is further adapted to:

[0097] Determine the target metric factors of the corresponding target layer according to the met warning conditions;

[0098] Determine the metric factors of the business layer associated with the target metric factors, and / or other associated metric factors at non-business levels, according to the association relationships between metric factors;

[0099] Determine business alarm information according to the metric factors of the business layer, and perform problem location according to the target metric factors and / or associated metric factors.

[0100] Optionally, the multi-layer structure includes: a data collection layer, a basic metric layer, a data processing layer, a perception layer, and a business layer.

[0101] Optionally, the device further includes: a correction processing module 550, which is adapted to perform alarm correction processing according to the problem location alarm according to preset processing rules.

[0102] The descriptions of the above modules refer to the corresponding descriptions in the method embodiments and will not be elaborated here.

[0103] According to the problem analysis and warning device provided by the present application, according to the multi-layer structure and the association relationships of metric factors, it is possible to monitor the basic metric factors (such as basic data, devices, etc.) of each non-business layer in a timely manner, discover the anomalies existing in the basic metric factors, warn of business appearance problems, and accelerate the troubleshooting of problems through the association relationships.

[0104] The present application also provides a non-volatile computer storage medium, which stores at least one executable instruction, and the executable instruction can execute the operations corresponding to the problem analysis and warning method in any of the above method embodiments.

[0105] The present application also provides a computer program product, which includes at least one executable instruction or computer program. The executable instruction or computer program enables the processor to perform operations corresponding to the problem analysis and warning method in any of the above method embodiments.

[0106] Figure 6 FIG. 4 shows a schematic structural diagram of a computing device according to an embodiment of the present application. The specific implementation of the computing device is not limited in the specific embodiments of the present application.

[0107] As Figure 6 shown, the computing device may include: a processor 602, a communication interface 604, a memory 606, and a communication bus 608.

[0108] Among them:

[0109] The processor 602, the communication interface 604, and the memory 606 communicate with each other through the communication bus 608.

[0110] The communication interface 604 is used to communicate with network elements of other devices such as clients or other servers.

[0111] The processor 602 is used to execute the program 610, and specifically can execute relevant steps in the above problem analysis and warning method embodiments.

[0112] Specifically, the program 610 may include program code, and the program code includes computer operation instructions.

[0113] The processor 602 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the present application. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0114] The memory 606 is used to store the program 610. The memory 606 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.

[0115] The program 610 can specifically be used to cause the processor 602 to execute the problem analysis and early warning method in any of the above method embodiments. For the specific implementation of each step in the program 610, reference can be made to the corresponding steps and descriptions in the corresponding units in the above problem analysis and early warning embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated here.

[0116] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings provided herein. The structure required to construct such systems will be apparent from the above description. In addition, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the descriptions made above for specific languages are for the purpose of disclosing the preferred embodiments of the present application.

[0117] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0118] Similarly, it should be understood that, for the purpose of streamlining the present application and assisting in understanding one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present application.

[0119] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be adopted to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise explicitly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature that provides the same, equivalent or similar purpose.

[0120] In addition, those skilled in the art can understand that although some of the embodiments herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of this application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0121] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components according to the present application. The present application can also be implemented as a device or apparatus program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0122] It should be noted that the above embodiments are illustrative of the present application rather than restrictive thereof, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A problem analysis and early warning method, comprising: A multi-layer structure is obtained based on the business scenario division; wherein, there is an upper and lower hierarchical relationship between the multi-layer structure, including a business layer and a non-business layer; Determining the index factors of each layer in the multi-layer structure, and the correlation relationship between the index factors of each layer; Monitor the indicator factors of the non-business layer to determine whether they meet the early warning conditions; If the warning conditions are met, the problem is located and an alarm is issued based on the warning conditions and the correlation between the indicator factors.

2. The method according to claim 1, wherein: The determining of the index factors of each layer in the multi-layer structure and the correlation between the index factors of each layer further includes: According to the multi-layer structure, the index factors of each layer are determined respectively; wherein the index factors include custom factors; According to the upper and lower hierarchical relationship of the multi-layer structure, a correlation relationship between the index factors of the next layer and the index factors of the previous layer is constructed.

3. The method according to claim 1 or 2, wherein: The association relationship between the indicator factors includes a one-to-one, one-to-many and / or many-to-one association relationship.

4. The method according to any one of claims 1 to 3, wherein: The monitoring of the indicator factors of the non-business layer and judging whether the early warning conditions are met further comprises: According to the indicator factors of each non-business layer, pre-constructing the warning conditions corresponding to the indicator factors; Based on the indicator factors of each layer, the monitoring objects corresponding to the indicator factors in the business scenario are monitored in real time, and whether the early warning conditions are met is determined according to the real-time execution status of the monitoring objects.

5. The method according to claim 1, wherein: The problem location alarm according to the corresponding early warning conditions and the correlation between the indicator factors further includes: Determine the target indicator factors of the corresponding target level according to the early warning conditions that meet the requirements; Determine the business-level indicator factors associated with the target indicator factors and / or other non-business-level associated indicator factors according to the association relationship between the indicator factors; The service alarm information is determined according to the indicator factors of the service layer, and the problem is located according to the target indicator factors and / or associated indicator factors.

6. The method according to any one of claims 1 to 5, wherein: The multi-layer structure includes: a data collection layer, a basic indicator layer, a data processing layer, a perception layer and a business layer.

7. The method according to any one of claims 1 to 6, wherein: The method further comprises: Locate the alarm based on the problem and perform alarm correction processing according to preset processing rules.

8. A problem analysis and early warning device, comprising: A partitioning module is suitable for obtaining a multi-layer structure based on business scenarios; wherein there is an upper and lower hierarchical relationship between the multi-layer structures, including a business layer and a non-business layer; A relationship module, adapted to determine the index factors of each layer in the multi-layer structure, and the correlation relationship between the index factors of each layer; A monitoring module, adapted to monitor the indicator factors of the non-business layer to determine whether the warning conditions are met; The alarm module is suitable for locating the problem and issuing an alarm if the warning conditions are met based on the correlation between the warning conditions and indicator factors.

9. A computing device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the problem analysis and early warning method according to any one of claims 1-7.

10. A computer storage medium, wherein at least one executable instruction is stored in the storage medium, and the executable instruction enables a processor to execute operations corresponding to the problem analysis and early warning method according to any one of claims 1 to 7.

11. A computer program product, comprising at least one executable instruction, wherein the executable instruction enables a processor to execute operations corresponding to the problem analysis and early warning method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent operation and maintenance method and system of multi-level IT system

    CN121255527A

  • Intelligent operation and maintenance method and system for multi-level IT system

    CN121255527B