An attribution method and device based on full-link monitoring of data
By setting up a monitoring module between the functional module and the service module, and automatically locate faults with topological relationships, the problem of difficult manual positioning is solved, and the effect of fast and accurate positioning of fault modules is achieved.
Patent Information
- Application Number
- CN202111249451.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-26
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-10-26
AI Technical Summary
In the prior art, it is difficult, untimely and inefficient to manually locate online problems, making it difficult to accurately locate the cause of the fault in complex systems.
By setting up a monitoring module between the functional module and the service module, using the topological relationship to automatically locate the fault module, combined with the preset fault cause display, automated fault positioning is achieved.
实现了在故障发生后快速、准确定位故障功能模块和服务模块,降低了人工分析的工作量,提高了定位效率和准确性,减少了系统停机时间。
Smart Images

Figure CN114064335B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software testing, and in particular, to an attribution method and device based on full-link data monitoring. Background Art
[0002] Online problems usually refer to large-scale problems or events that affect the availability of online services, mainly including three parts: problem discovery, problem localization, and problem solving. Problem discovery can be achieved through multiple means such as user complaints and monitoring alarms, while problem localization, that is, attribution, requires finding the root cause of the problem so as to give the most appropriate solution. For complex systems, problem localization is a very crucial but difficult task. In the prior art, through the monitoring of core business metrics, online problems of abnormal metric types can be discovered in a timely manner, and technicians view alarm logs and system logs for manual analysis and problem localization. For complex systems composed of multiple single points, more attention is paid to the monitoring and alarm of various metric result data, and it is difficult to locate the root cause.
[0003] In the process of implementing the present invention, the applicant found that there are at least the following problems in the prior art:
[0004] In the prior art, it is difficult, untimely, and inefficient to manually locate problems. Summary of the Invention
[0005] Embodiments of the present invention provide an attribution method and device based on full-link data monitoring, which solve the problems of difficult, untimely, and inefficient manual problem localization.
[0006] To achieve the above object, on the one hand, embodiments of the present invention provide an attribution method based on full-link data monitoring, including:
[0007] Monitoring corresponding functional modules through respective monitoring modules;
[0008] When an alarm occurs in the monitoring module corresponding to a specific functional module, traversing and checking the alarm information of the monitoring modules corresponding to the functional modules in each functional module dependency branch of the specific functional module according to the functional module dependency topology relationship, and respectively determining for each functional module dependency branch that if an alarm is found in the functional module dependency branch, then the functional module corresponding to the last alarm-occurring monitoring module in the functional module dependency branch is used as the faulty functional module that causes the alarm in the monitoring module corresponding to the specific functional module;
[0009] Wherein, corresponding monitoring modules are added to each functional module in advance.
[0010] Further, it further includes: monitoring each service module in the corresponding functional module through each monitoring module, where a functional module is composed of service modules;
[0011] Traverse and check the alarm information of the monitoring modules corresponding to each service module in each service module dependency branch of the faulty functional module according to the service module dependency topology relationship within the faulty functional module, and for each service module dependency branch, determine that if an alarm is found in this service module dependency branch, then use the service module corresponding to the last monitoring module that alarms in this service module dependency branch as the faulty service module;
[0012] Add corresponding monitoring modules for each service module within each functional module in advance.
[0013] Further, it also includes: presenting the preset fault causes corresponding to the faulty functional module and / or the faulty service module;
[0014] Among them, the preset fault causes are pre-bound to each functional module and / or service module.
[0015] Further, the traversing and checking of the alarm information of the monitoring modules corresponding to each functional module in each functional module dependency branch of the specific functional module according to the functional module dependency topology relationship, and for each functional module dependency branch, determining that if an alarm is found in this functional module dependency branch, then using the functional module corresponding to the last monitoring module that alarms in this functional module dependency branch as the faulty functional module that causes the monitoring module of the specific functional module to alarm, specifically:
[0016] Check the alarm information of the monitoring modules corresponding to each functional module on each search path according to the search path defined in the preset configuration information. For each search path, determine that if an alarm is found in this search path, then use the functional module corresponding to the last monitoring module that alarms on this search path as the faulty functional module, and give the corresponding fault cause of this faulty functional module according to the preset configuration information;
[0017] The traversing and checking of the alarm information of the monitoring modules corresponding to each service module in each service module dependency branch of the faulty functional module according to the service module dependency topology relationship within the faulty functional module, and for each service module dependency branch, determining that if an alarm is found in this service module dependency branch, then using the service module corresponding to the last monitoring module that alarms in this service module dependency branch as the faulty service module, specifically:
[0018] Check the alarm information of the monitoring modules corresponding to each service module on each search path according to the search paths defined in the preset configuration information. For each search path, determine that if an alarm is found in the search path, the service module corresponding to the last monitoring module that alarms on the search path is used as the faulty service module, and the corresponding cause of the fault of the faulty service module is given according to the preset configuration information.
[0019] Among them, the configuration information is established according to the functional module dependency topology relationship and / or the service module dependency topology relationship, and records the corresponding causes of faults of each functional module and / or each service module on the search path.
[0020] Further, the monitoring module is a kafka module;
[0021] The monitoring of the corresponding functional modules by each monitoring module includes:
[0022] Monitor the exposure volume of the exposure logs of the corresponding functional modules through the kafka module, and alarm for the corresponding functional modules when the exposure volume is greater than or equal to the specified exposure threshold;
[0023] Among them, the kafka module is pre-connected between each functional module, and the exposure logs generated by the pre-functional modules in each functional module are input to the post-functional modules through the kafka module.
[0024] On the other hand, an attribution device based on full-link data monitoring according to an embodiment of the present invention includes:
[0025] A function monitoring unit for monitoring corresponding functional modules through each monitoring module;
[0026] A function fault discovery unit for, when an alarm occurs in the monitoring module corresponding to a specific functional module, traversing and checking the alarm information of the monitoring modules corresponding to the functional modules in each functional module dependency branch of the specific functional module according to the functional module dependency topology relationship, and respectively determining for each functional module dependency branch that if an alarm is found in the functional module dependency branch, the functional module corresponding to the last monitoring module that alarms in the functional module dependency branch is used as the faulty functional module that causes the alarm of the monitoring module of the specific functional module;
[0027] Among them, corresponding monitoring modules are pre-added to each functional module.
[0028] Further, it further includes: a service monitoring unit for monitoring each service module in the corresponding functional module through each monitoring module, where the functional module is composed of service modules;
[0029] A service fault discovery unit, configured to traverse and check alarm information of monitoring modules corresponding to each service module in each service module dependency branch of the faulty function module according to the service module dependency topology relationship within the faulty function module, and respectively determine for each service module dependency branch that if an alarm is found in the service module dependency branch, the service module corresponding to the last alarm - occurring monitoring module in the service module dependency branch is used as the faulty service module;
[0030] Correspondingly, monitoring modules are added to each service module within each function module in advance.
[0031] Furthermore, it further includes: a fault display unit, configured to display preset fault reasons corresponding to the faulty function module and / or the faulty service module;
[0032] Wherein, the preset fault reasons are pre - bound to each function module and / or service module.
[0033] Furthermore, the function fault discovery unit is specifically configured to: according to the search paths defined in the preset configuration information, check alarm information of monitoring modules corresponding to each function module on each search path, respectively determine for each search path that if an alarm is found in the search path, the function module corresponding to the last alarm - occurring monitoring module on the search path is used as the faulty function module, and give the corresponding fault reason of the faulty function module according to the preset configuration information;
[0034] The service fault discovery unit is specifically configured to: according to the search paths defined in the preset configuration information, check alarm information of monitoring modules corresponding to each service module on each search path, respectively determine for each search path that if an alarm is found in the search path, the service module corresponding to the last alarm - occurring monitoring module on the search path is used as the faulty service module, and give the corresponding fault reason of the faulty service module according to the preset configuration information;
[0035] Wherein, the configuration information is established according to the function module dependency topology relationship and / or the service module dependency topology relationship, and records the respective fault reasons corresponding to each function module and / or each service module on the search path.
[0036] Furthermore, the monitoring module is a kafka module;
[0037] The function monitoring unit is specifically configured to:
[0038] Monitor the exposure volume of the exposure log of the corresponding function module through the kafka module, and alarm for the corresponding function module when the exposure volume is greater than or equal to the specified exposure threshold;
[0039] Among them, the function modules are pre-connected through the Kafka module, and the exposure logs generated by the pre-function modules in each function module are input to the post-function modules through the Kafka module.
[0040] The above technical solution has the following beneficial effects: By respectively setting corresponding monitoring modules for the function modules of interest or prone to failure or for all function modules, and when a failure alarm occurs in the monitoring module, automatically searching for the most fundamental faulty function module according to the dependency relationship between the function modules, so that after the failure alarm occurs, the most fundamental faulty function module can be automatically located in the shortest time, achieving the effect of quickly and efficiently locating the faulty module. When a failure alarm is detected, the failure location is automatically started, solving the problem of difficult manual location; Further, the service modules within the function modules are also monitored through the corresponding monitoring modules of each service module. For the service modules within the most fundamental faulty function module found, continue to locate the most fundamental faulty service module according to the dependency relationship between the service modules, so as to achieve the effect of quickly and efficiently locating the faulty service module; Further, by binding each function module and each service module to their respective common failure causes, when the faulty function module and the faulty service module are determined, the corresponding common failure causes are displayed, further reducing the difficulty of locating problems and achieving the effect of assisting maintenance personnel to quickly solve the failure problems. Further, according to the dependency relationship and common failure causes and other information of the function modules of interest or all and the service modules within the function modules, a search path configuration is established, and the faulty module is automatically searched and located according to the specified path according to the search path configuration, and the corresponding common failure causes are provided, which can significantly reduce the traversal of the dependency relationship paths of each function module and service module, and only traverse the dependency relationship paths prone to failure, significantly improving the traversal efficiency and further improving the efficiency of locating faults. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of an attribution method based on full-link data monitoring according to one of the embodiments of the present invention;
[0043] Figure 2 It is an architecture diagram of an attribution device based on full-link data monitoring according to one of the embodiments of the present invention;
[0044] Figure 3It is a schematic diagram of the traversal path for a functional module in one embodiment of the attribution method based on data full-link monitoring of the present invention;
[0045] Figure 4 It is a schematic diagram of the traversal path for a service module in one embodiment of the attribution method based on data full-link monitoring of the present invention. Detailed implementation manners
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0047] On the one hand, as Figure 1 shown, an embodiment of the present invention provides an attribution method based on data full-link monitoring, including:
[0048] Step S100, monitoring corresponding functional modules through respective monitoring modules;
[0049] Step S101, when an alarm occurs in the monitoring module corresponding to a specific functional module, traversing and checking the alarm information of the monitoring modules corresponding to the functional modules in each functional module dependency branch of the specific functional module according to the functional module dependency topology relationship, and respectively judging for each functional module dependency branch that if an alarm is found in the functional module dependency branch, then the functional module corresponding to the last alarmed monitoring module in the functional module dependency branch is used as the faulty functional module that causes the monitoring module corresponding to the specific functional module to alarm;
[0050] Among them, corresponding monitoring modules are added to each functional module in advance.
[0051] In some embodiments, the overall system function is realized by multiple functional modules calling each other according to a certain dependency topology relationship; corresponding monitoring modules are added to each functional module in advance. Specifically, corresponding monitoring modules can be added to some functional modules of interest, or corresponding monitoring modules can be added to all functional modules, which can be specifically determined according to requirements. The monitoring module will collect some key business metrics of the functional module, such as the number of requests per second, memory usage, exposure volume, etc., and analyze the key business metrics. If it is determined that the change in the key business metric triggers a preset alarm rule, the corresponding monitoring module will issue an alarm. When the monitoring module corresponding to a specific functional module issues an alarm, the corresponding specific functional module can be determined according to the alarmed monitoring module, and then starting from the specific functional module, according to the functional module dependency topology relationship of the specific functional module, traverse and check the monitoring modules corresponding to the functional modules in each functional module dependency branch of the specific functional module, and find the functional module corresponding to the last alarmed monitoring module in each functional module dependency branch as the faulty functional module. Specifically, there may be only one functional module dependency branch in which the faulty functional module is found, or there may be faulty functional modules in multiple functional module dependency branches, that is, there may be one or more faulty functional modules. When traversing and checking, each monitoring module corresponding to each functional module in each functional module dependency branch can be checked, but in this way, the traversal efficiency will be low; preferably, when a functional module does not issue an alarm, there is no alarm in the functional module dependency branch on which the functional module depends, so there is no need to continue checking the functional module dependency branch of the functional module, thereby improving the efficiency of the check. As Figure 3 shown Figure 3It is the functional module dependency topology of a simple functional module. Module 3 depends on the output of Module 2, and Module 2 depends on the output of Module 1. Data is transmitted between each module through Kafka. Each Kafka is used as a corresponding monitoring module to monitor the PV volume in each Kafka. When the PV volume in Kafka3 is abnormal, an alarm is triggered. The system will sequentially check Kafka2 of Module 2 and Kafka1 of Module 1 according to the functional module dependency topology of Module 3, and determine the final faulty functional module. For example, if there is no alarm in Kafka1, then the module 2 corresponding to Kafka2 is the last functional module with a monitoring module alarm in this dependency branch, that is, Module 2 is the faulty functional module in this dependency branch. It should be noted that in specific situations, there may be no monitoring module corresponding to any functional module in some functional module dependency branches with an alarm, indicating that the functional modules in these functional module dependency branches are normal; there may also be one or more faulty functional modules located in one or more functional module dependency branches; it is also possible that each functional module dependency branch is normal, and in this case, the faulty functional module is a specific functional module; since the located faulty functional module is the specific functional module itself or the module on which the specific functional module depends, the root cause of the monitoring module alarm triggered by the specific functional module lies in the one or more faulty functional modules determined. Technicians can continue to conduct a detailed analysis of the faulty functional module to ultimately solve the fault problem.
[0052] The embodiment of the present invention has the following technical effects: By monitoring each functional module through each monitoring module, when a specific functional module is abnormal, it will trigger a monitoring module alarm. According to the functional module dependency topology of the specific functional module, the faulty functional module is automatically located, avoiding manually collecting information, manually analyzing information, and manually inferring to locate the faulty functional module, realizing automatic fault location, improving the efficiency of fault location and fault resolution, significantly reducing the downtime of the system caused by fault resolution, and thus increasing the stable operation time of the system.
[0053] Furthermore, it also includes: monitoring each service module in the corresponding functional module through each monitoring module, where the functional module is composed of service modules;
[0054] Traverse and check the alarm information of the monitoring modules corresponding to each service module in each service module dependency branch of the faulty functional module according to the service module dependency topology in the faulty functional module, and respectively judge for each service module dependency branch that if an alarm is found in this service module dependency branch, then take the service module corresponding to the last monitoring module that alarms in this service module dependency branch as the faulty service module;
[0055] Add corresponding monitoring modules for each service module in each functional module in advance.
[0056] In some embodiments, each functional module is first monitored by each monitoring module. When a specific functional module is abnormal, the monitoring module will be triggered to alarm. According to the functional module dependency topology relationship of the specific functional module, the faulty functional module is automatically located, realizing the location of faults at the functional module level. However, so far, the more detailed fault location is still unclear. In this embodiment, corresponding monitoring modules are also set for each service module inside each functional module, and according to the service module dependency topology relationship of each service module inside the faulty functional module, each service module dependency branch in the service module dependency topology relationship is traversed and checked, and the service module corresponding to the last monitoring module that alarms in each service module dependency branch of the faulty functional module is found as the faulty service module; when traversing and checking, each monitoring module corresponding to each functional module in each functional module dependency branch can be checked, but in this way, the traversal efficiency will be low; preferably, when a certain functional module does not alarm, there is no alarm in the functional module dependency branch on which the functional module depends, so there is no need to continue checking the functional module dependency branch of this functional module, thereby improving the efficiency of the check. Similarly, it should be noted that in specific situations, there may be no alarm for the monitoring module corresponding to any service module in some dependency branches, indicating that the service modules in these service module dependency branches are normal; there may also be one or more service module dependency branches that locate the faulty service module; it is also possible that each service module dependency branch is normal. At this time, the fault may come from some implementation details in the faulty functional module or service modules not included in the monitoring, and technical personnel need to further analyze according to other information such as the log of the automatically located faulty functional module. After the faulty service module is located, technical personnel can continue to conduct a detailed analysis on the faulty service module in order to finally solve the fault problem.
[0057] The embodiment of the present invention has the following technical effects: By traversing the alarm situations of the monitoring modules of each functional module and each service module according to the functional module dependency relationship and the service module dependency relationship, the faulty functional module and the faulty service module inside each faulty functional module are finally determined, realizing the automatic location of faults, and the faults can be located in a smaller range of service modules, further reducing the workload of manual analysis and division of labor for location, improving the efficiency and accuracy of fault location, further reducing the downtime caused by fault resolution, and improving the stable operation time of the system.
[0058] Further, it further includes: displaying the preset fault causes corresponding to the faulty functional module and / or the faulty service module;
[0059] Wherein, the preset fault causes are pre-bound to each functional module and / or service module.
[0060] In some embodiments, by traversing the alarm situations of the monitoring modules of each functional module and / or service module according to the functional module dependency relationship and / or service module dependency relationship, the faulty functional module and / or the faulty service module within each faulty functional module are finally determined. At this time, the effect of automatically locating the fault location has been achieved. Considering that the faults occurring in any system are usually repetitive, limited in type, and limited in fault causes; in order to solve the faults more efficiently, by analyzing and statistically processing historical faults, the common fault causes of each functional module and / or service module can be obtained. Bind the obtained fault causes to the corresponding functional modules and / or service modules, and store the binding relationship. When the faulty functional module and / or faulty service module are determined, obtain the corresponding fault causes of each faulty functional module and / or each faulty service module through this binding relationship, and display the obtained fault causes to the technicians, which plays a role in guiding the technicians to solve the faults and achieves the effect of making full use of the historical experience of the technicians to solve the fault problems.
[0061] Further, traverse and check the alarm information of the monitoring modules corresponding to the functional modules in each functional module dependency branch of the specific functional module according to the functional module dependency topology relationship, and respectively judge for each functional module dependency branch that if an alarm is found in this functional module dependency branch, then use the functional module corresponding to the last alarm - occurring monitoring module in this functional module dependency branch as the faulty functional module that causes the monitoring module of the specific functional module to alarm. Specifically:
[0062] Check the alarm information of the monitoring modules corresponding to the functional modules on each search path according to the search path defined in the preset configuration information. Respectively judge for each search path that if an alarm is found in this search path, then use the functional module corresponding to the last alarm - occurring monitoring module on this search path as the faulty functional module, and give the corresponding fault cause of this faulty functional module according to the preset configuration information;
[0063] Traverse and check the alarm information of the monitoring modules corresponding to the service modules in each service module dependency branch of the faulty functional module according to the service module dependency topology relationship within the faulty functional module, and respectively judge for each service module dependency branch that if an alarm is found in this service module dependency branch, then use the service module corresponding to the last alarm - occurring monitoring module in this service module dependency branch as the faulty service module. Specifically:
[0064] According to the search paths defined in the preset configuration information, check the alarm information of the monitoring modules corresponding to each service module on each search path. For each search path, determine that if an alarm is found in the search path, the service module corresponding to the last monitoring module that alarms on the search path is used as the faulty service module, and the corresponding fault cause of the faulty service module is given according to the preset configuration information;
[0065] Among them, the configuration information is established according to the functional module dependency topology relationship and / or the service module dependency topology relationship, and records the corresponding fault causes of each functional module and / or each service module on the search path.
[0066] In some real-time examples, the configuration information is established in advance for the functional modules and / or service modules that are of interest or prone to failure or all according to the dependency topology relationship. The search paths are defined in the configuration information, and the corresponding fault causes of each functional module and / or each service module on the search path are recorded. When an alarm occurs in the monitoring module corresponding to a certain functional module or service module existing in the configuration information, each functional module and / or service module is checked in turn according to the order defined by the search path to determine the faulty functional module and / or faulty service module, and the corresponding fault cause is given. The implementation form of the configuration information may include but is not limited to a table; for example, as shown in Table 1, the upstream_name column is the upstream and downstream relationship of horizontal search, that is, the dependency relationship between functional modules or service modules, and the father_name column is the parent-child relationship of the vertical search multi-fork tree, that is, the dependency relationship from the functional module to the service module, and the reason is the empirical value, including the reasons for historical problems and some tracing ideas. For example, the functional module corresponding to kafka3 depends on the functional module corresponding to kafka2, and the functional module corresponding to kafka2 depends on the functional module corresponding to kafka1; the sub-modules, that is, service modules, depended on by the functional module corresponding to kafka2 include kafka4 and kafka6. Configuration information can be established for the functional modules and / or service modules that are prone to failure, skipping the functional modules and / or service modules in the middle part of the dependency relationship that are not prone to failure, so as to further improve the efficiency of fault location, and give the fault cause after the fault is located. If a new node is added, only the database needs to be modified, with high code reusability and strong maintainability.
[0067] Further, the monitoring module is a kafka module;
[0068] In some embodiments, the monitoring of the corresponding functional modules by each monitoring module includes:
[0069] Monitor the exposure volume of the exposure logs of the corresponding functional modules through the Kafka module, and alarm for the corresponding functional modules when the exposure volume is greater than or equal to the specified exposure threshold;
[0070] Among them, the Kafka module is pre-connected between each functional module, and the exposure logs generated by the pre-functional modules in each functional module are input to the post-functional modules through the Kafka module;
[0071] In some other embodiments, the monitoring of each service module in the corresponding functional module by each monitoring module includes: monitoring the exposure volume of the exposure logs of each service module inside each functional module through the Kafka module, and alarming for the corresponding service module when the exposure volume is greater than or equal to the specified exposure threshold;
[0072] Among them, the Kafka module is pre-connected between each service module inside the functional module, and the exposure logs generated by the pre-service modules in each service module are input to the post-service modules through the Kafka module.
[0073] In some embodiments, an embodiment of the present invention is applied to the field of fault location based on Kafka monitoring. The Kafka module is connected in series between the functional modules and / or service modules of interest, so that the exposure logs of the functional modules and / or service modules are transmitted through the Kafka module to the next functional module and / or service module; the specified exposure threshold can be determined according to specific requirements or through statistical analysis of historical data.
[0074] The above technical solution has the following beneficial effects: By separately setting corresponding monitoring modules for the function modules of interest or prone to failure, or for all function modules, and when a failure alarm occurs in the monitoring module, automatically finding the most fundamental faulty function module according to the dependency relationship between the function modules, so that after the failure alarm occurs, the most fundamental faulty function module can be automatically located in the shortest time, achieving the effect of timely and efficient location of the faulty module. When a failure alarm is detected, the failure location is automatically started, solving the problem of difficult manual location; further, the service modules within the function modules are also monitored through the corresponding monitoring modules of each service module. For the service modules within the found most fundamental faulty function module, continue to locate the most fundamental faulty service module according to the dependency relationship between the service modules, so as to achieve the effect of timely and efficient location of the faulty service module; further, by binding each function module and each service module to their respective common failure causes, when the faulty function module and the faulty service module are determined, the corresponding common failure causes are displayed, further reducing the difficulty of locating the problem and achieving the effect of assisting maintenance personnel to quickly solve the failure problem. Further, according to the dependency relationship and common failure causes and other information of each function module of interest or all function modules and the service modules within the function modules, a search path configuration is established. According to the search path configuration, the faulty module is automatically searched and located according to the specified path, and the corresponding common failure causes are provided, which can significantly reduce the traversal of the dependency relationship paths of each function module and service module, and only traverse the dependency relationship paths prone to failure, significantly improving the traversal efficiency and further improving the efficiency of locating the failure.
[0075] On the other hand, as Figure 2 shown, an attribution device based on full-link data monitoring provided by an embodiment of the present invention includes:
[0076] A function monitoring unit 200, configured to monitor corresponding function modules through each monitoring module;
[0077] A function failure discovery unit 201, configured to, when an alarm occurs in the monitoring module corresponding to a specific function module, traverse and check the alarm information of the monitoring modules corresponding to the function modules in each function module dependency branch of the specific function module according to the function module dependency topology relationship, and respectively determine for each function module dependency branch that if an alarm is found in the function module dependency branch, the function module corresponding to the last monitoring module that alarms in the function module dependency branch is used as the faulty function module that causes the monitoring module of the specific function module to alarm;
[0078] Wherein, corresponding monitoring modules are added to each function module in advance.
[0079] Further, it further includes:
[0080] A service monitoring unit, which is used to monitor each service module in the corresponding function module through each monitoring module, where the function module is composed of service modules;
[0081] A service fault discovery unit, which is used to traverse and check the alarm information of the monitoring modules corresponding to each service module in each service module dependency branch of the faulty function module according to the service module dependency topology relationship in the faulty function module, and respectively judge for each service module dependency branch that if an alarm is found in the service module dependency branch, then use the service module corresponding to the last monitoring module that generates an alarm in the service module dependency branch as the faulty service module;
[0082] Correspondingly, monitoring modules are added to each service module in each function module in advance.
[0083] Furthermore, it further includes:
[0084] A fault display unit, which is used to display the preset fault causes corresponding to the faulty function module and / or the faulty service module;
[0085] Wherein, the preset fault causes are bound to each function module and / or service module in advance.
[0086] Furthermore, the function fault discovery unit is specifically configured to: check the alarm information of the monitoring modules corresponding to each function module on each search path according to the search paths defined in the preset configuration information, and respectively judge for each search path that if an alarm is found in the search path, then use the function module corresponding to the last monitoring module that generates an alarm in the search path as the faulty function module, and give the corresponding fault cause of the faulty function module according to the preset configuration information;
[0087] The service fault discovery unit is specifically configured to: check the alarm information of the monitoring modules corresponding to each service module on each search path according to the search paths defined in the preset configuration information, and respectively judge for each search path that if an alarm is found in the search path, then use the service module corresponding to the last monitoring module that generates an alarm in the search path as the faulty service module, and give the corresponding fault cause of the faulty service module according to the preset configuration information;
[0088] Wherein, the configuration information is established according to the function module dependency topology relationship and / or the service module dependency topology relationship, and records the respective fault causes corresponding to each function module and / or each service module on the search path.
[0089] Furthermore, the monitoring module is a kafka module;
[0090] The function monitoring unit 200 is specifically configured to:
[0091] Monitor the exposure volume of the exposure logs of the corresponding function modules through the kafka module, and alarm for the corresponding function modules when the exposure volume is greater than or equal to the specified exposure threshold;
[0092] Among them, the kafka module is pre-connected between each function module, and the exposure logs generated by the pre-function module in each function module are input to the post-function module through the kafka module;
[0093] The service monitoring unit is specifically configured to: monitor the exposure volume of the exposure logs of each service module inside each function module through the kafka module, and alarm for the corresponding service module when the exposure volume is greater than or equal to the specified exposure threshold;
[0094] Among them, the kafka module is pre-connected between each service module inside the function module, and the exposure logs generated by the pre-service module in each service module are input to the post-service module through the kafka module.
[0095] The above technical solution has the following beneficial effects: By separately setting their corresponding monitoring modules for the function modules that are of interest or prone to failure or for all function modules, and when a failure alarm occurs in the monitoring module, automatically finding the most fundamental faulty function module according to the dependency relationship between each function module, so that after a failure alarm occurs, the most fundamental faulty function module can be automatically located in the shortest time, achieving the effect of timely and efficient location of the faulty module. When a failure alarm is detected, automatic fault location starts, solving the problem of difficult manual fault location; Further, the service modules within the function modules are also monitored through the corresponding monitoring modules of each service module. For the service modules within the most fundamental faulty function module found, continue to locate the most fundamental faulty service module according to the dependency relationship between the service modules, so as to achieve the effect of timely and efficient location of the faulty service module; Further, by binding each function module and each service module to their respective common fault causes, when the faulty function module and the faulty service module are determined, the corresponding common fault causes are displayed, further reducing the difficulty of problem location and achieving the effect of assisting maintenance personnel to quickly solve the fault problem. Further, according to the dependency relationship and common fault cause information of each function module and each service module that are of interest or all, a search path configuration is established. According to the search path configuration, automatically search and locate the faulty module according to the specified path, and provide the corresponding common fault causes, which can significantly reduce the traversal of the dependency relationship paths of each function module and service module, and only traverse the dependency relationship paths that are prone to failure, significantly improving the traversal efficiency and further improving the efficiency of fault location.
[0096] The above technical solutions of the embodiments of the present invention will be described in detail below in conjunction with specific application examples. For technical details not introduced during the implementation process, reference can be made to the relevant descriptions in the previous text.
[0097] Based on the full-link monitoring of data, the technical solution of the present invention finds the root fault module and the root cause of the problem through the upstream and downstream association relationships of the function module and the service module, improves the attribution efficiency, thereby shortening the processing time of online problems and reducing more losses of the system.
[0098] The following are the explanations of the terms involved:
[0099] Kafka is a high-throughput distributed publish-subscribe messaging system used to process streaming data;
[0100] PV is the exposure volume. Each time a user sees an advertisement or an event, it is recorded as an exposure.
[0101] The following takes the exposure system as an example to illustrate the technical solution of the present invention. As Figure 3 shown, Figure 3 is the main link of the system, including 3 function modules. The function modules are connected through Kafka. Function module 1 generates the original exposure log and writes it into Kafka1. Function module 2 reads the exposure log from Kafka1, and after some policy processing, writes the processed exposure log into Kafka2. Similarly, module 3 reads the exposure log from Kafka2, processes the exposure log, and then writes it into Kafka3. Finally, the index calculation uses the data in Kafka3. Here, the PV volume is taken as an example.
[0102] The first step: Monitoring and alarming. Each of the 3 Kafkas (i.e., Kafka1, Kafka2, and Kafka3) counts its own PV volume, sets a change rate threshold, and alarms when the threshold is exceeded;
[0103] The second step: Defining the fault area. The specific method is as follows. Horizontal search is performed according to the upstream and downstream relationships. As Figure 3, viewed horizontally along the arrow direction, it is from kafka1 to kafka2 to kafka3. The upstream of kafka3 is kafka2, and the upstream of kafka2 is kafka1. For example, if kafka3 alarms, we need to check its upstream kafka2. If kafka2 also alarms, we can infer that the abnormal alarm of kafka3 is caused by the abnormality of kafka2, and we can rule out the problem of functional module 3. Then we continue to look for the upstream kafka1. If kafka1 has no abnormality, we can basically conclude that it is a problem with functional module 2, which makes the data of kafka2 abnormal, and the chain reaction affects the downstream kafka3 to be abnormal, and then leads to the abnormal pv index data. Therefore, the fault area is functional module 2.
[0104] Step 3: Find the root module of the fault. The specific method is as follows. Through step 2, it can be determined that the fault area is functional module 2. Then, functional module 2 is further vertically divided as Figure 4 shown, Figure 4 Module 2 in it is functional module 2, Figure 4 Service A in it is service module A, service B is service module B, service A1 is service module A1, service B1 is service module B1, and service B2 is service module B2; according to Figure 4 the service module dependency topology relationship in it, traverse each branch of functional module 2 one by one. For example, for service A and service A1, check whether there are abnormal alarms. The tracing idea is the same as in step 2. If service A has an abnormality, continue to check service A1. If service A1 has a problem, it can be inferred that the abnormality of service A1 causes the abnormality of service A and module 2, and finally leads to the abnormal index calculation. Then service A1 is the root module where the fault occurs.
[0105] Step 4: Locate the root module where the fault occurs. Combining historical experience, possible causes and some tracing ideas can also be given. In actual problem troubleshooting, it can be found that students with troubleshooting experience can locate problems quickly, especially for complex systems, and the value of experience is higher. Therefore, it is very necessary to implement it. Refer to the high-frequency reasons or recent reasons (timeliness) and trace in a targeted manner to improve the efficiency of locating problems.
[0106] Finally, to implement the above process, a database configuration management method is adopted, including upstream and downstream relationships, multi-tree relationships, and historical experience, as shown in Table 1. upstream_name is the upstream and downstream relationship for horizontal search, father_name is the parent-child relationship of the multi-tree for vertical search, and reason is the experience value, including historical problem reasons and some tracing ideas. If a new node is added, only the database needs to be modified, with high code reusability and strong maintainability.
[0107]
[0108] Table 1 Lookup Path Configuration Table
[0109] The embodiments of the present invention have the following technical effects: A full - link data monitoring system is established. Through the upstream - downstream relationship in the horizontal direction and the multi - fork tree relationship in the vertical direction, the root failure module with problems is found and historical tracing experience is given, thereby improving the efficiency of online problem location and further shortening the processing time of online problems.
[0110] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of the present disclosure. The appended method claims present the elements of the various steps in an exemplary order and are not limited to the specific order or hierarchy recited.
[0111] In the above detailed description, various features are combined in a single embodiment to simplify the present disclosure. This disclosure method should not be construed as reflecting an intention that the embodiments of the claimed subject matter require more features than are clearly recited in each claim. On the contrary, as reflected by the appended claims, the present invention exists in a state with fewer features than all the features of the disclosed single embodiment. Therefore, the appended claims are hereby expressly incorporated into the detailed description, where each claim stands alone as a separate preferred embodiment of the present invention.
[0112] To enable any person skilled in the art to implement or use the present invention, the above - described disclosed embodiments have been described. For those skilled in the art, various modification ways of these embodiments are obvious, and the general principles defined herein can also be applied to other embodiments without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.
[0113] The above description includes examples of one or more embodiments. Of course, it is impossible to describe all possible combinations of components or methods for the purpose of describing the above - mentioned embodiments, but those of ordinary skill in the art should recognize that each embodiment can be further combined and arranged. Therefore, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. In addition, regarding the term "comprising" used in the specification or claims, the coverage of this term is similar to the term "including", as explained when "including:" is used as a transitional term in the claims. In addition, any term "or" used in the claims or the specification is to mean "non - exclusive or".
[0114] Those skilled in the art can also understand that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention can be implemented by electronic hardware, computer software, or a combination of both. To clearly show the interchangeability of hardware and software, the above-mentioned various illustrative components, units, and steps have generally described their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art can use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope protected by the embodiments of the present invention.
[0115] In the embodiments of the present invention, the various illustrative logical blocks or units can be implemented or operated to perform the described functions by a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of the above designs. The general-purpose processor can be a microprocessor. Optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.
[0116] The steps of the methods or algorithms described in the embodiments of the present invention can be directly embedded in hardware, software modules executed by a processor, or a combination of both. The software modules can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be disposed in an ASIC, and the ASIC can be disposed in a user terminal. Optionally, the processor and the storage medium can also be disposed in different components of the user terminal.
[0117] In one or more exemplary designs, the functions described in embodiments of the present invention may be implemented in hardware, software, firmware, or any combination of the three. If implemented in software, these functions may be stored on a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. A computer-readable medium includes both computer storage media and communication media that facilitate transfer of a computer program from one place to another. The storage media may be any available media that can be accessed by a general or special purpose computer. For example, such computer-readable media may include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions or data structures and that can be read by a general or special purpose computer or a general or special purpose processor. In addition, any connection can be properly defined as a computer-readable medium, for example, if software is transmitted from a website, server, or other remote source via a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wirelessly such as infrared, wireless, and microwave, it is also included within the defined computer-readable medium. The disks and discs include compact disks, laser disks, optical disks, DVDs, floppy disks, and Blu-ray disks, disks typically reproduce data magnetically, while discs typically reproduce data optically with a laser. Combinations of the above may also be included within the computer-readable medium.
[0118] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An attribution method based on full-link monitoring of data, characterized in that Including: Monitoring corresponding functional modules through respective monitoring modules; When an alarm occurs in the monitoring module corresponding to a specific functional module, traverse and check the alarm information of the monitoring modules corresponding to the functional modules in each functional module dependency branch of the specific functional module according to the functional module dependency topology relationship, and respectively determine for each functional module dependency branch that if an alarm is found in the functional module dependency branch, the functional module corresponding to the last monitoring module that alarms in the functional module dependency branch is used as the faulty functional module causing the alarm in the monitoring module of the specific functional module. Specifically: According to the search paths defined in the preset configuration information, check the alarm information of the monitoring modules corresponding to the functional modules on each search path, and respectively determine for each search path that if an alarm is found in the search path, the functional module corresponding to the last monitoring module that alarms in the search path is used as the faulty functional module, and give the corresponding fault reason of the faulty functional module according to the preset configuration information; wherein, the configuration information is established according to the functional module dependency topology relationship, and records the corresponding fault reasons of each functional module of interest or prone to failure on the search path; Showing the preset fault reason corresponding to the faulty functional module; wherein, the preset fault reason is pre-bound to each functional module, and the preset fault reason is the fault reason that repeatedly appears when each functional module fails through the analysis and statistics of historical faults; Wherein, corresponding monitoring modules are added to each functional module in advance, and the monitoring module is a kafka module; The monitoring of corresponding functional modules through respective monitoring modules includes: Monitoring the exposure volume of the exposure logs of the corresponding functional modules through the kafka module, and alarming for the corresponding functional modules when the exposure volume is greater than or equal to the specified exposure threshold; wherein, the kafka module is connected between each functional module in advance, and the exposure logs generated by the pre-functional modules in each functional module are input to the post-functional modules through the kafka module.
2. The attribution method based on full data link monitoring according to claim 1, wherein, Also including: Monitoring each service module in the corresponding functional module through respective monitoring modules, wherein the functional module is composed of service modules; Traversing and checking the alarm information of the monitoring modules corresponding to the service modules in each service module dependency branch of the faulty functional module according to the service module dependency topology relationship in the faulty functional module, and respectively determining for each service module dependency branch that if an alarm is found in the service module dependency branch, the service module corresponding to the last monitoring module that alarms in the service module dependency branch is used as the faulty service module; Adding corresponding monitoring modules to each service module in each functional module in advance.
3. The attribution method based on full data link monitoring according to claim 2, wherein Also including: Showing the preset fault reason corresponding to the faulty service module; Wherein, the preset fault reason is pre-bound to each service module.
4. The attribution method based on full data link monitoring according to claim 2, characterized in that Traverse and check the alarm information of the monitoring modules corresponding to each service module in each service module dependency branch of the faulty function module according to the service module dependency topology relationship in the faulty function module, and respectively determine for each service module dependency branch that if an alarm is found in the service module dependency branch, the service module corresponding to the last monitoring module that alarms in the service module dependency branch is used as the faulty service module. Specifically: Check the alarm information of the monitoring modules corresponding to each service module on each search path according to the search path defined in the preset configuration information. Respectively determine for each search path that if an alarm is found in the search path, the service module corresponding to the last monitoring module that alarms in the search path is used as the faulty service module, and give the corresponding cause of the fault for the faulty service module according to the preset configuration information; Among them, the configuration information is established according to the service module dependency topology relationship and records the respective causes of the faults corresponding to each service module on the search path.
5. An attribution device based on full-link monitoring of data, characterized in that, Including: A function monitoring unit for monitoring the corresponding function module through each monitoring module; A function fault discovery unit for, when an alarm occurs in the monitoring module corresponding to a specific function module, traversing and checking the alarm information of the monitoring modules corresponding to each function module in each function module dependency branch of the specific function module according to the function module dependency topology relationship, and respectively determining for each function module dependency branch that if an alarm is found in the function module dependency branch, the function module corresponding to the last monitoring module that alarms in the function module dependency branch is used as the faulty function module that causes the alarm of the monitoring module of the specific function module; The function fault discovery unit is specifically configured to: check the alarm information of the monitoring modules corresponding to each function module on each search path according to the search path defined in the preset configuration information, respectively determine for each search path that if an alarm is found in the search path, the function module corresponding to the last monitoring module that alarms in the search path is used as the faulty function module, and give the corresponding cause of the fault for the faulty function module according to the preset configuration information; among them, the configuration information is established according to the function module dependency topology relationship and records the respective causes of the faults corresponding to each function module that is of interest or prone to failure on the search path; A fault display unit for displaying the preset cause of the fault corresponding to the faulty function module; among them, the preset cause of the fault is pre-bound to each function module; the preset cause of the fault is the cause of the fault that repeatedly appears when each function module fails obtained through the analysis and statistics of historical faults; Among them, corresponding monitoring modules are added to each function module in advance, and the monitoring module is a kafka module; The function monitoring unit is specifically configured to: monitor the exposure volume of the exposure logs of the corresponding function modules through the Kafka module, and alarm for the corresponding function modules when the exposure volume is greater than or equal to the specified exposure threshold; wherein, the Kafka module is pre-connected between the function modules, and the exposure logs generated by the pre-function modules in each function module are input to the post-function modules through the Kafka module.
6. The attribution device based on data full-link monitoring according to claim 5, wherein It further includes: A service monitoring unit for monitoring each service module in the corresponding function module through each monitoring module, wherein the function module is composed of service modules; A service fault discovery unit for traversing and checking the alarm information of the monitoring modules corresponding to the service modules in each service module dependency branch of the faulty function module according to the service module dependency topology relationship in the faulty function module, and respectively judging for each service module dependency branch that if an alarm is found in the service module dependency branch, the service module corresponding to the last alarmed monitoring module in the service module dependency branch is used as the faulty service module; Monitoring modules corresponding to the service modules in each function module are added in advance.
7. The attribution device based on full data link monitoring according to claim 6, wherein, It further includes: A fault display unit for displaying the preset fault reasons corresponding to the faulty service modules; Wherein, the preset fault reasons are pre-bound to each service module.
8. The attribution device based on full-link data monitoring according to claim 6, wherein The service fault discovery unit is specifically configured to: check the alarm information of the monitoring modules corresponding to the service modules on each search path according to the search path defined in the preset configuration information, respectively judge for each search path that if an alarm is found in the search path, the service module corresponding to the last alarmed monitoring module on the search path is used as the faulty service module, and give the corresponding fault reason of the faulty service module according to the preset configuration information; Wherein, the configuration information is established according to the service module dependency topology relationship and records the respective corresponding fault reasons of the service modules on the search path.
Citation Information
Patent Citations
Alarm root cause analysis method, device and equipment and storage medium
CN109684181A
Micro-service monitoring method and device
CN111124830A