Operation and maintenance processing method and device, electronic equipment and storage medium

By acquiring and analyzing the values ​​of multiple operational metrics and using various rules for anomaly detection, the problem of low accuracy in anomaly analysis in traditional operation and maintenance models is solved, thereby improving the effectiveness of operation and maintenance.

CN120909819APending Publication Date: 2025-11-07SHENZHEN SIYUAN ELECTRONICS TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510785415.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In traditional operation and maintenance models, the accuracy of anomaly analysis is not high, resulting in poor operation and maintenance results.

Method used

By acquiring the values ​​of multiple operational metrics, the target operational metrics are determined, and anomaly analysis is performed using pre-defined basic rules, composite rules, and scenario rules, and preset operational and maintenance procedures are executed.

Benefits of technology

This improves the accuracy of anomaly analysis, thereby enhancing the effectiveness of operation and maintenance, and enabling more efficient fault detection and handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909819A_ABST
    Figure CN120909819A_ABST
Patent Text Reader

Abstract

The invention discloses an operation and maintenance processing method and device, electronic equipment and a storage medium, and relates to the technical field of computers. Determining a target operation index in the plurality of operation indexes based on the respective index values of the plurality of operation indexes; if the number of the target operation indexes is multiple, first target rules corresponding to the multiple target operation indexes are obtained from preset rules, the first target rules comprise basic rules, composite rules and scene rules corresponding to the target operation indexes, and the index values of the multiple target operation indexes are analyzed through the first target rules in sequence; and under the condition that at least part of the target operation indexes are analyzed to be abnormal, preset operation and maintenance processing is executed, so that the accuracy of abnormality analysis can be improved, and the effect of operation and maintenance processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, and more particularly, to an operation and maintenance processing method and device, an electronic device, and a storage medium. BACKGROUND

[0002] The traditional manual operation and maintenance mode cannot meet the requirements of stability, security, and efficiency of operation and maintenance. Therefore, some operation and maintenance processing methods through electronic devices have emerged. In the related art, abnormality analysis can be performed by setting rules.

[0003] However, in the related art, abnormality analysis is performed by setting rules, which results in low accuracy of abnormality analysis and further results in poor effect of operation and maintenance processing. SUMMARY

[0004] In view of the above problems, the embodiments of the present application provide an operation and maintenance processing method and device, an electronic device, and a storage medium to solve the problem of low accuracy of abnormality analysis in the related art and further poor effect of operation and maintenance processing, thereby improving the accuracy of abnormality analysis and further improving the effect of operation and maintenance processing.

[0005] In a first aspect, the embodiments of the present application provide an operation and maintenance processing method, including: obtaining index values of a plurality of running indexes, the plurality of running indexes including running indexes of at least one of a target device, a server, or a database, the server and the database being used to implement a business scenario of the target device; determining a target running index in the plurality of running indexes based on the index values of the plurality of running indexes; if the target running index is multiple, obtaining first target rules corresponding to the multiple target running indexes from pre-set rules, the first target rules including a basic rule, a composite rule, and a scenario rule corresponding to each target running index, the basic rule being used to indicate a first abnormal condition of the target running index, the composite rule being used to indicate a second abnormal condition composed of at least two first abnormal conditions, and the scenario rule being used to indicate a third abnormal condition of the business scenario of the target device; sequentially analyzing the index values of the multiple target running indexes through the first target rules, and executing a pre-set operation and maintenance processing in a case where at least part of the target running indexes is abnormal.

[0006] In a possible implementation manner, the method further includes: if the target running index is one, obtaining a second target rule corresponding to the target running index, the second target rule being used to indicate a fourth abnormal condition of the target running index; analyzing a feature value of the target running index through the second target rule, and executing the pre-set operation and maintenance processing in a case where the target running index is abnormal.

[0007] In a possible implementation, the acquiring the respective index value of each of the plurality of operation indexes comprises: acquiring the respective index value of each of the plurality of operation indexes at a first time; and determining a target operation index from the plurality of operation indexes based on the respective index value of each of the plurality of operation indexes, comprising: for any operation index, if the index value of the operation index at the first time is different from the index value of the operation index at a second time, determining the operation index as the target operation index, the second time being a time before the first time.

[0008] In a possible implementation, the plurality of operation indexes comprises a first part of operation indexes, a second part of operation indexes, and a third part of operation indexes, and the acquiring the respective index value of each of the plurality of operation indexes comprises: acquiring a running log of the target device, and acquiring the respective index value of each of the first part of operation indexes from the running log of the target device; detecting the operation index of the server to obtain the respective index value of each of the second part of operation indexes; and detecting the operation index of the database to obtain the respective index value of each of the third part of operation indexes.

[0009] In a possible implementation, the analyzing the respective index value of each of the plurality of target operation indexes by using the first target rule in sequence, and performing a preset operation and maintenance processing in a case where at least part of the target operation indexes is determined to be abnormal, comprises: determining whether the respective index value of each of at least part of the target operation indexes meets an abnormal condition indicated by the first target rule; if the respective index value of each of at least part of the target operation indexes meets at least one target abnormal condition indicated by the first target rule, determining that the respective index value of each of at least part of the target operation indexes that meets the target abnormal condition is abnormal; determining a severity of the abnormality, and performing an alarm based on the severity, and / or performing a pre-arranged target abnormality solving action to make at least part of the target operation indexes that is abnormal return to normal.

[0010] In a possible implementation, the performing the pre-arranged target abnormality solving action comprises: performing a pre-arranged target abnormality solving action corresponding to the target abnormal condition.

[0011] In a possible implementation, the method further comprises: performing an optimization processing on the pre-set rules, the optimization processing comprising at least one of the following: identifying a similarity between any two rules in the pre-set rules, and merging any two rules with a similarity greater than a similarity threshold; or detecting at least two conflicting rules in the pre-set rules, and retaining one of the at least two conflicting rules; or adjusting a priority of each rule in the pre-set rules according to an operation and maintenance effect, the priority being used to indicate an order of analyzing the index value.

[0012] In a second aspect, an embodiment of the present application provides an operation and maintenance processing apparatus, comprising: an acquisition module configured to acquire respective index values of a plurality of running indexes, the plurality of running indexes comprising running indexes of at least one of a target device, a server or a database, the server and the database being configured to implement a business scenario of the target device; an operation and maintenance processing module configured to determine a target running index in the plurality of running indexes based on the respective index values of the plurality of running indexes; if the target running index is a plurality, acquire a first target rule corresponding to the plurality of target running indexes from a pre-set rule, the first target rule comprising a basic rule, a composite rule and a scenario rule corresponding to each target running index, the basic rule being configured to indicate a first abnormal condition of the target running index, the composite rule being configured to indicate a second abnormal condition composed of at least two first abnormal conditions, and the scenario rule being configured to indicate a third abnormal condition of the business scenario of the target device; analyze the respective index values of the plurality of target running indexes through the first target rule in sequence, and execute a pre-set operation and maintenance processing in a case where at least part of the target running indexes are abnormal.

[0013] In a third aspect, an electronic device is provided, comprising: a processor; a memory, the memory storing computer readable instructions, the computer readable instructions being executed by the processor to implement the operation and maintenance processing method as above.

[0014] In a fourth aspect, a computer readable storage medium is provided, the computer readable storage medium storing computer readable instructions, the computer readable instructions being executed by a processor to implement the operation and maintenance processing method as above.

[0015] In a fifth aspect, a computer program product is provided, comprising computer instructions, the computer instructions being executed by a processor to implement the operation and maintenance processing method as above.

[0016] In the embodiment, by acquiring index values of a plurality of running indexes, the plurality of running indexes comprise running indexes of at least one of a target device, a server or a database, the server and the database being used to implement a business scenario of the target device; based on the index values of the plurality of running indexes, a target running index in the plurality of running indexes is determined; if the target running index is multiple, a first target rule corresponding to the multiple target running indexes is acquired from a pre-set rule, the first target rule comprising a basic rule, a composite rule and a scenario rule corresponding to each target running index, the basic rule being used to indicate a first abnormal condition of the target running index, the composite rule being used to indicate a second abnormal condition composed of at least two first abnormal conditions, and the scenario rule being used to indicate a third abnormal condition of the business scenario of the target device; the index values of the multiple target running indexes are analyzed in sequence through the first target rule, and in a case where at least part of the target running indexes is abnormal, a pre-set operation and maintenance processing is executed, so that the multiple rules can continue the abnormal analysis and the pre-set operation and maintenance processing can be executed in a case where an abnormality occurs, thereby improving the accuracy of the abnormal analysis and further improving the effect of the operation and maintenance processing. BRIEF DESCRIPTION OF DRAWINGS

[0017] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application. It is to be expressly understood that the drawings are included solely for purposes of illustration and that they are not to be construed as limiting the application.

[0018] Figure 1 is a schematic diagram of an application scenario of the present application according to an embodiment of the present application.

[0019] Figure 2 is a flowchart of an operation and maintenance processing method according to an embodiment of the present application.

[0020] Figure 3 is a schematic diagram of a framework of an operation and maintenance system according to an embodiment of the present application.

[0021] Figure 4 is a block diagram of an operation and maintenance processing apparatus according to an embodiment of the present application.

[0022] Figure 5 is a structural schematic diagram of a computer system of an electronic device suitable for implementing an embodiment of the present application. DETAILED DESCRIPTION

[0023] The embodiments of the present application will be described in detail below, examples of the embodiments are shown in the drawings, wherein the same or similar notations represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.

[0024] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below by referring to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0025] In the following description, the terms "first\second" and the like are only to distinguish similar objects, and do not represent a specific order of the objects. Understandably, "first\second" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0026] "Multiple" referred to herein means two or more. "And / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the front and rear associated objects are in an "or" relationship. In the following description, "some embodiments or some embodiment modes" describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict.

[0027] With the continuous expansion of enterprise information technology (IT) infrastructure and the continuous increase of business complexity, the traditional manual operation and maintenance mode has been unable to meet the requirements of modern enterprises on system stability, security and efficiency. The current enterprise operation and maintenance is facing the following pain points: difficulty in collecting multi-source heterogeneous data: the enterprise IT environment usually contains server clusters, business databases, middleware and other components, and the data format and interface are not unified. Abnormal detection lag: the traditional threshold-based alarm mechanism cannot timely discover potential problems. Low efficiency of fault handling: relying on manual experience for fault diagnosis and recovery, slow response speed. It is difficult to deposit operation and maintenance knowledge: a large amount of operation and maintenance experience exists in the individual brains of engineers, and cannot form standardized processes.

[0028] The existing operation and maintenance solutions in the market have the following limitations: open source monitoring tools (such as Zabbix and Prometheus) only provide basic monitoring functions and lack intelligent analysis and scene processing capabilities. Commercial APM products (such as Dynatrace and New Relic) are expensive and have low customization. Traditional operation and maintenance scripts lack unified management and scheduling platforms and have high maintenance costs. Existing solutions mostly focus on a single aspect (such as infrastructure or application performance) and lack an end-to-end panoramic view.

[0029] Although in the related art, rules are set for anomaly analysis, the rules used in the related art are relatively simple, such as setting a threshold in the rules, which results in low accuracy of anomaly analysis, and thus poor effect of operation and maintenance processing.

[0030] Therefore, embodiments of the present application propose an operation and maintenance processing method and device, an electronic device and a storage medium to solve the problem of low accuracy of anomaly analysis in the related art, and thus poor effect of operation and maintenance processing, thereby improving the accuracy of anomaly analysis and the effect of operation and maintenance processing.

[0031] Figure 1 FIG. 1 is a schematic diagram of an application scenario according to an embodiment of the present application. As shown in FIG. 1, the application scenario can include a target device 102, a server 104 and a database 106. The target device 102 can be a device performing a certain function, and the server 104 and the database 106 are used to implement a business scenario of the target device 102, which is related to the function of the target device. Figure 1

[0032] The target device 102 can be a hunting camera, a smartphone, a tablet computer, a notebook computer, a desktop computer, a smart television, a smart home device, a vehicle terminal, etc., which is not limited here.

[0033] The server 104 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, etc. basic cloud computing services.

[0034] The database 106 can be integrated in the server 104 or independent of the server 104, which is not limited here.

[0035] Taking the target device 102 as a hunting camera as an example.

[0036] ​A hunting camera is an automatic shooting device designed for outdoor environments, triggered by infrared sensing or motion detection to record wildlife activities, monitor outdoor environments, or assist hunting observations. The hunting camera can send the recorded video data to the server 104, and then the server 104 can store the recorded video data in the database 106, and then when it is needed to view the video data, a request can be sent to the server 104, and then the server 104 can obtain the video data from the database 106 in response to the request and send the video data to the party initiating the request. In this example, the business scenario can be a scenario in which video data is sent to the server 104, and the server 104 writes the video data to the database 106.

[0037] In this embodiment, the electronic device can obtain the index values of a plurality of running indexes, the plurality of running indexes including the running indexes of at least one of the target device 102, the server 104, or the database 106, the server 104 and the database 106 being used to implement the business scenario of the target device 102; based on the index values of the plurality of running indexes, determining a target running index in the plurality of running indexes; if there are multiple target running indexes, obtaining a first target rule corresponding to the multiple target running indexes from a pre-set rule, the first target rule including a basic rule, a composite rule, and a scenario rule corresponding to each target running index, the basic rule being used to indicate a first abnormal condition of the target running index, the composite rule being used to indicate a second abnormal condition composed of at least two first abnormal conditions, and the scenario rule being used to indicate a third abnormal condition of the business scenario of the target device 102; and sequentially analyzing the index values of the plurality of target running indexes through the first target rule, and executing a pre-set operation and maintenance processing in a case where at least part of the target running indexes are abnormal.

[0038] It should be noted that the electronic device of the present embodiment can be the target device 102 or the server 104, and can also be a separately arranged device, which is not limited herein.

[0039] The implementation details of the technical solutions of the embodiments of the present application are described in detail as follows: Figure 2 is a flowchart of an operation and maintenance processing method according to an embodiment of the present application, which can be executed by an electronic device, and with reference to Figure 2 as shown, the method at least includes steps S210 to S240, which are described in detail as follows: S210, obtaining the index values of a plurality of running indexes, the plurality of running indexes including the running indexes of at least one of the target device, the server, or the database, the server and the database being used to implement the business scenario of the target device.

[0040] In this embodiment, multiple operational metrics can be at least one of the following: operational metrics of the target device, operational metrics of the server, or operational metrics of the database. Optionally, the operational metrics of the target device may include, but are not limited to, at least one of the following: number of backup logins, number of device connections to servers, number of device binding failures, device traffic data, device network registration logs, number of devices that have not logged in for more than 24 hours, or number of devices that have not taken pictures for more than 24 hours. The operational metrics of the server may include, but are not limited to, at least one of the following: Central Processing Unit (CPU), memory, disk, network, or service operational status. Service operational status refers to the real-time running status of processes, applications, or middleware deployed on the server, such as indicating whether the service is running normally. In this embodiment, the target device can perform specific functions to generate data, and then send the data to the server. After receiving the data, the server writes the data into the database for storage, that is, the business scenario of the target device is realized through the server and the database.

[0041] In one possible implementation, the values ​​of multiple operational metrics are obtained, including: Obtain the operation logs of the target device, and extract the values ​​of the first part of the operation indicators from the operation logs of the target device; detect the operation indicators of the server to obtain the values ​​of the second part of the operation indicators; detect the operation indicators of the database to obtain the values ​​of the third part of the operation indicators.

[0042] In this embodiment, the metric values ​​of the target device, server, and database's respective operating metrics can be obtained.

[0043] It should be understood that the more operational metrics obtained, the higher the accuracy of anomaly analysis; conversely, the fewer operational metrics obtained, the higher the efficiency of anomaly analysis.

[0044] S220. Based on the individual values ​​of multiple operating indicators, determine the target operating indicator among the multiple operating indicators.

[0045] In this embodiment, operational metrics with non-zero values ​​can be used as target operational metrics. This reduces the analysis of invalid operational metrics, thereby reducing resource waste in operation and maintenance. In another possible implementation, operational metrics with corresponding values ​​greater than a first preset threshold can also be used as target operational metrics; this is not a limitation.

[0046] In another possible approach, the individual values ​​of multiple operational indicators are obtained at the first moment. Based on these individual values, the target operational indicator among the multiple operational indicators can be determined, which may include: For any operation index, if the index value of the operation index at the first time is different from the index value of the operation index at the second time, the operation index is determined as a target operation index, and the second time is a time before the first time.

[0047] In the embodiment, the index value of the operation index at the first time is different from the index value of the operation index at the second time, and the second time is a time before the first time, that is, the index value of the operation index at the current time changes relative to the index value at the historical time. Therefore, the operation index can be taken as the target operation index, and the index value changes, which may cause a risk to be abnormal.

[0048] The embodiment determines the operation index as the target operation index if the index value of the operation index at the first time is different from the index value of the operation index at the second time, and the second time is a time before the first time, that is, the index value of the operation index changes, and then uses the index value of the target operation index to determine whether there is an abnormality. In this way, resource waste during operation and maintenance processing can be reduced.

[0049] In the embodiment, the index value of the operation index at the first time is different from the index value of the operation index at the second time, and the second time is a time before the first time, that is, the index value of the operation index at the current time changes relative to the index value at the historical time. Therefore, the operation index can be taken as the target operation index, and the index value changes, which may cause a risk to be abnormal.

[0050] In this embodiment, the basic rule can be understood as a first abnormal condition corresponding to a single target running index. Optionally, the first abnormal condition can be that the index value of the target running index is greater than a second preset threshold. The second preset threshold is greater than the first preset threshold. The composite rule can be understood as a second abnormal condition corresponding to at least two target running indexes, which is composed on the basis of at least two first abnormal conditions corresponding to the at least two target running indexes. It should be noted that in this embodiment, the relationship between the at least two first abnormal conditions in the second abnormal condition can be an AND relationship, an OR relationship, or a NOT relationship. In this embodiment, if the at least two first abnormal conditions in the second abnormal condition have an AND relationship, it means that the at least two first abnormal conditions in the second abnormal condition are satisfied at the same time, and the second abnormal condition is satisfied. If the at least two first abnormal conditions in the second abnormal condition have an OR relationship, it means that one of the at least two first abnormal conditions in the second abnormal condition is satisfied, and the second abnormal condition is satisfied. If the at least two first abnormal conditions in the second abnormal condition have a NOT relationship, it means that the NOT indicated first abnormal condition is not satisfied, and other first abnormal conditions except the NOT indicated first abnormal condition are satisfied, and the second abnormal condition is considered to be satisfied. The third abnormal condition can mean that the index value of the target running index is greater than a third preset threshold, or the index value of the target running index is a preset index value, which can indicate a fault, for example, indicating that the Remote Dictionary Server (Redis) fails, or indicating that the service running state fails, etc.

[0051] For example, the following illustrates the basic rule, the composite rule and the scene rule.

[0052] Basic rule: preset index threshold rule (such as CPU utilization > 90% or disk utilization > 85%).

[0053] Composite rule: support AND / OR / NOT logic combination. An example of one kind of AND logic combination can be: high CPU usage and high memory usage and connection number surge (connection number increase greater than threshold). An example of one kind of OR logic combination can be: CPU utilization > 90% or high memory usage > 85%. An example of one kind of NOT logic combination can be: CPU utilization > 90% NOT network delay (Ping) value greater than 1000 milliseconds. It should be noted that the thresholds corresponding to different running indexes can be different.

[0054] Scenario rule: rule package specific to a certain business scenario (e.g., insufficient database connection pool, Redis middleware failure, etc.). The scenario rule in this embodiment can be a rule specific to the business scenario of the target device.

[0055] S240, sequentially analyze the respective index values of the plurality of target operation indicators through the first target rule, and in the case that at least part of the target operation indicators are found to be abnormal, perform the preset operation and maintenance processing.

[0056] In this embodiment, when at least part of the respective index values of the target operation indicators meet one of the abnormal conditions indicated by the first target rule, it is determined that the at least part of the target operation indicators are abnormal, and the analysis of the index values is stopped. In another possible implementation, it can also be determined whether each abnormal condition indicated by the first target rule is met. Optionally, each rule in the pre-set rules can be configured with a priority, and the priority is used to indicate the order of analyzing the index values. The higher the priority, the earlier the order of analyzing the index values using the rule. Optionally, the priority of the rule can be represented by a weight, and the greater the weight represents the higher priority.

[0057] In this embodiment, by obtaining the respective index values of the plurality of operation indicators, the plurality of operation indicators include the operation indicators of at least one of the target device, the server, or the database, and the server and the database are used to implement the business scenario of the target device; based on the respective index values of the plurality of operation indicators, the target operation indicators in the plurality of operation indicators are determined; if the target operation indicators are multiple, from the pre-set rules, the first target rules corresponding to the plurality of target operation indicators are obtained, the first target rules include the basic rules, the composite rules, and the scenario rules corresponding to each target operation indicator, the basic rules are used to indicate the first abnormal condition of the target operation indicators, the composite rules are used to indicate the second abnormal condition composed of at least two first abnormal conditions, and the scenario rules are used to indicate the third abnormal condition of the business scenario of the target device; sequentially analyze the respective index values of the plurality of target operation indicators through the first target rule, and in the case that at least part of the target operation indicators are found to be abnormal, perform the preset operation and maintenance processing. In this way, the abnormal analysis can be continued through the plurality of rules, and the preset operation and maintenance processing can be performed in the case of abnormality, thereby improving the accuracy of the abnormal analysis and further improving the effect of the operation and maintenance processing.

[0058] In one possible implementation, the method further includes: If the target operation index is one, a second target rule corresponding to the target operation index is acquired, the second target rule being used to indicate a fourth abnormal condition in which the target operation index is abnormal; the characteristic value of the target operation index is analyzed through the second target rule, and in a case where it is analyzed that the target operation index is abnormal, a preset operation and maintenance processing is executed.

[0059] The fourth abnormal condition in this embodiment can refer to the description of the first abnormal condition, which is not repeated here.

[0060] In this embodiment, if the target operation index is one, a second target rule corresponding to the target operation index is acquired, the second target rule being used to indicate a fourth abnormal condition in which the target operation index is abnormal; the characteristic value of the target operation index is analyzed through the second target rule, and in a case where it is analyzed that the target operation index is abnormal, a preset operation and maintenance processing is executed. In this way, even if the target operation index is one, abnormal analysis can be performed and operation and maintenance processing can be performed when an abnormality occurs, which can improve the comprehensiveness of operation and maintenance processing.

[0061] In a possible implementation, the index values of the plurality of target operation indexes are analyzed through the first target rule in sequence, and in a case where it is analyzed that at least part of the target operation indexes are abnormal, a preset operation and maintenance processing is executed, including: It is determined whether the index values of at least part of the target operation indexes satisfy the abnormal condition indicated by the first target rule; if the index values of at least part of the target operation indexes satisfy at least one target abnormal condition indicated by the first target rule, it is determined that the index values of at least part of the target operation indexes satisfying the target abnormal condition are abnormal; the severity of the abnormality is determined, and an alarm is given based on the severity, and / or a target abnormality solving action previously programmed is executed to make at least part of the target operation indexes that are abnormal return to normal.

[0062] In this embodiment, the severity of the abnormality can be determined by the difference between the index value and the threshold in the abnormal condition. The greater the difference, the higher the severity. The severity of the abnormality is different, and the way of the alarm is also different. The higher the severity, the more channels of the alarm. The target abnormality solving action can include but is not limited to restarting the service, expanding the node, expanding the disk, switching the traffic, which is not limited here.

[0063] In the embodiment, whether the index value of each of the at least part of the target operation indexes meets the abnormal condition indicated by the first target rule is determined, if the index value of each of the at least part of the target operation indexes meets at least one target abnormal condition indicated by the first target rule, it is determined that the index value of each of the at least part of the target operation indexes meeting the target abnormal condition is abnormal, the severity of the abnormality is determined, and the severity is used for alarm and / or performing the pre-arranged target abnormality solving action to make the at least part of the target operation indexes meeting the abnormality normal, so that the alarm and / or the pre-arranged target abnormality solving action can be automatically performed to make the at least part of the target operation indexes meeting the abnormality normal, thereby improving the intelligent degree of operation and maintenance.

[0064] In a possible implementation, the pre-arranged target abnormality solving action is performed, including: The pre-arranged target abnormality solving action corresponding to the target abnormal condition is performed.

[0065] In the embodiment, the mapping relationship between the abnormal condition and the abnormality solving action can be determined in advance, and the target abnormality solving action can be determined according to the mapping relationship after the target abnormal condition is determined, so that the pre-arranged target abnormality solving action corresponding to the target abnormal condition is performed. The action of the embodiment can be an action for at least one of the target device, the server or the database.

[0066] For example, if the service running state is abnormal, the service can be restarted, if the CPU utilization is greater than a threshold, the node can be expanded, if the disk utilization is greater than a threshold, the disk can be expanded, and if the network Ping value is high, the traffic can be switched, which is not limited herein.

[0067] In the embodiment, the pre-arranged target abnormality solving action corresponding to the target abnormal condition is performed, so that the abnormality related solving action can be performed, thereby improving the utilization rate of the resources for solving the abnormality.

[0068] In a possible implementation, the method further includes: The pre-set rule is optimized, and the optimization includes at least one of the following: Similarity between any two rules in the pre-set rule is identified, and any two rules with similarity greater than a similarity threshold are merged; or, At least two conflicting rules in the pre-set rule are detected, and one of the at least two conflicting rules is reserved; or, The priority of each rule in the pre-set rule is adjusted according to the operation and maintenance effect, and the priority is used for indicating the order of analyzing the index value.

[0069] When identifying the similarity between any two pre-defined rules, the similarity can be represented by the number of times both rules are simultaneously satisfied. The more times they are simultaneously satisfied, the higher the similarity. This allows any two rules with a similarity greater than a similarity threshold to be merged. During merging, these two rules can be combined using an OR relationship. This improves the efficiency of operation and maintenance processes.

[0070] When detecting at least two conflicting rules in a pre-defined set of rules, a directed graph can be used to construct a rule-event model. For newly added or modified rules, a path search is performed on the graph to search for similar rules and detect whether there are duplicates or conflicts (for example, if there is a CPU>80% alarm and a CPU>90% alarm is added, it is considered a duplicate; a restart for CPU>90% conflicts with a capacity expansion for CPU>90%). This embodiment improves the efficiency of operation and maintenance by retaining one of the at least two conflicting rules.

[0071] When adjusting the priority of pre-defined rules based on operational performance, the higher the operational performance, the higher the priority; conversely, the lower the priority, the lower the performance. This embodiment improves operational efficiency by dynamically adjusting the priority of each rule.

[0072] To facilitate understanding, the following examples illustrate the operation and maintenance process through an operation and maintenance system, based on the above examples.

[0073] Please see Figure 3 , Figure 3 This is a schematic diagram of the framework of an operation and maintenance system provided in an embodiment of this application. Figure 3The operation and maintenance system shown can include a data collector 310, a big data warehouse 320, an intelligent rule engine 330, and a task executor 340. Among them, the data collector 310 can collect data from Elasticsearch through a log inspection script, such as logs of a mobile target device, collect data from a server through a server inspection script, such as obtaining running indicators of the server, and collect data from a business database (also referred to as a database) through a database inspection script, such as obtaining running indicators of the business database. Then, the collected data is stored in the big data warehouse 320. Then, when operation and maintenance processing is needed, the corresponding target rule (such as the first target rule or the second target rule of the embodiment) can be obtained through the intelligent rule engine 330, and then the task executor 340 performs abnormality analysis, such as abnormality identification through the target rule (such as abnormality analysis through basic rules, composite rules, scene rules, etc.), and if an abnormality is found in the indicators, an alarm is given according to the degree of abnormality, such as WeChat alarm, email alarm, telephone alarm, etc., and event processing is performed (such as performing a pre-arranged action).

[0074] In general, the system can be divided into a data collection layer, a data processing layer, an intelligent rule engine layer, an execution control layer, and an alarm notification layer.

[0075] 2.1.1 Data collection layer: Elasticsearch data collection module: Collect business exception logs and other indicators on Elasticsearch through a log inspection script. Server monitoring module: Collect CPU, memory, disk, network, service running state, and other basic indicators through a lightweight Agent. Database monitoring module: Collect SQL execution efficiency, lock waiting, cache hit rate, and other database key indicators through a dedicated probe. 2.1.2 Data processing layer: Data collector: Realize data cleaning, format conversion, and standardization processing. Big data warehouse: Use Apache Doris high-performance data warehouse.

[0076] 2.1.3 Intelligent rule engine layer: Multi-level rule library: Basic rules: pre-set indicator threshold rules (such as CPU > 90%, disk > 85%, etc.), composite rules: support AND / OR / NOT logic combination (such as "CPU usage is high and memory usage is high and connection number suddenly increases"), and scene rules: rules package for specific business scenarios (such as insufficient database connection pool and redis middleware failure). In data analysis, the indicator values of multiple running indicators can be obtained, and then the target running indicator is determined.

[0077] Rule optimizer: automatic rule deduplication: identify and merge similar rules (similarity > 80%), conflict detection: find mutually contradictory rule combinations, weight adjustment: dynamically adjust rule priority based on historical effects.

[0078] 2.1.4 Execution control layer: Rule trigger engine: event-driven architecture: single metric change only triggers related rule calculation. Sliding window calculation: support 5 minutes / 1 hour time window aggregation.

[0079] Action orchestrator: pre-set action library: restart service, expand node, expand disk, switch traffic, etc. Workflow engine: support complex process orchestration (such as first try to restart, and automatically roll back after failure).

[0080] 2.1.5 Alarm notification layer: Multi-level alarm mechanism: hierarchical processing according to event severity. Low-level alarm: WeChat / DingDing notification. Medium-level alarm: email notification + WeChat / DingDing notification + ticket creation. High-level alarm: phone call + emergency response team linkage.

[0081] This embodiment constructs a three-in-one automatic operation and maintenance system of rule intelligent optimization, multi-source data fusion and closed-loop self-healing, and the system architecture is as Figure 3 The following technologies are used to improve operation and maintenance efficiency: data standardization adapter: unified collection and cleaning of data from different data sources, conversion to unified standard data, and provision to rule engine, dynamic rule weight algorithm: real-time calculation of current event weight according to pre-set rule engine, intelligent hierarchical processing of events, intelligent rule conflict detection: graph algorithm is used to detect rule conflicts, and when multiple events occur at the same time, over-processing is avoided, multi-level execution engine: workflow method is used to process events, realizing complex event process integration, closed-loop self-healing executor: action orchestration method is used to realize "detection-decision-execution-verification", completing the whole process from fault discovery to repair without manual intervention.

[0082] It should be understood that the technical effects of this embodiment include but are not limited to the following: Efficiency improvement: automatic data collection and analysis reduces more than 80% of manual operation.

[0083] Fault prevention: more than 90% of potential problems are discovered in advance through predictive analysis.

[0084] Cost reduction: reduce more than 50% of the average fault repair time (MTTR).

[0085] Knowledge sedimentation: form reusable operation and maintenance knowledge graph, reduce personnel flow risk.

[0086] Scalability: Modular design supports horizontal scaling, accommodating businesses of different sizes.

[0087] In addition, traditional rule engines can also be used as an alternative solution, such as using Drools or other rule engine models. Applicable scenario: environments with simple operation and maintenance scenarios and clear rules. Advantages: simple implementation and strong interpretability.

[0088] In addition, edge computing can also be set as an alternative solution by deploying lightweight analysis modules on edge nodes. Applicable scenario: industrial Internet of Things environments with extremely high real-time requirements. Advantages: reduces network latency and improves response speed.

[0089] In addition, a blockchain-based alternative solution can be set up to use smart contracts to achieve tamper-proof records of operation and maintenance. Applicable scenario: financial industry with strict audit and compliance requirements. Advantages: operation records are traceable and meet compliance requirements.

[0090] The following describes an apparatus embodiment of the present application, which can be used to execute the methods in the above-mentioned embodiments of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the above-mentioned method embodiments of the present application.

[0091] Figure 4 is a block diagram of an operation and maintenance processing apparatus according to an embodiment of the present application, as shown in Figure 4 The operation and maintenance processing apparatus includes an acquisition module 410 and an operation and maintenance processing module 420, wherein: The acquisition module 410 is configured to acquire index values of a plurality of running indexes, the plurality of running indexes including running indexes of at least one of a target device, a server, or a database, the server and the database being configured to implement a business scenario of the target device; and the operation and maintenance processing module 420 is configured to determine a target running index from the plurality of running indexes based on the index values of the plurality of running indexes, and if the target running index is multiple, acquire a first target rule corresponding to the multiple target running indexes from a pre-set rule, the first target rule including a basic rule, a composite rule, and a scenario rule corresponding to each target running index, the basic rule being configured to indicate a first abnormal condition of the target running index, the composite rule being configured to indicate a second abnormal condition composed of at least two first abnormal conditions, and the scenario rule being configured to indicate a third abnormal condition of an abnormal business scenario of the target device; analyze the index values of the multiple target running indexes through the first target rule in sequence, and execute a pre-set operation and maintenance processing when at least part of the target running indexes are found to be abnormal.

[0092] In a possible implementation, the operation and maintenance processing module 420 is further configured to: if the target operation index is one, acquire a second target rule corresponding to the target operation index, the second target rule being used to indicate a fourth abnormal condition in which the target operation index is abnormal; analyze the characteristic value of the target operation index through the second target rule, and perform a preset operation and maintenance processing in a case where it is analyzed that the target operation index is abnormal.

[0093] In a possible implementation, when the operation and maintenance processing module 420 acquires the index values of the plurality of operation indexes, the operation and maintenance processing module 420 is configured to: acquire the index values of the plurality of operation indexes at a first time; and when the operation and maintenance processing module 420 determines the target operation index in the plurality of operation indexes based on the index values of the plurality of operation indexes, the operation and maintenance processing module 420 is configured to: for any operation index, if the index value of the operation index at the first time is different from the index value of the operation index at a second time, determine the operation index as the target operation index, the second time being a time before the first time.

[0094] In a possible implementation, the plurality of operation indexes includes a first part of operation indexes, a second part of operation indexes, and a third part of operation indexes, and when the acquiring module 410 acquires the index values of the plurality of operation indexes, the acquiring module 410 is configured to: acquire a running log of the target device, and acquire the index values of the first part of operation indexes from the running log of the target device; detect the operation indexes of the server to obtain the index values of the second part of operation indexes; and detect the operation indexes of the database to obtain the index values of the third part of operation indexes.

[0095] In a possible implementation, when the operation and maintenance processing module 420 analyzes the index values of the plurality of target operation indexes through the first target rule in sequence, and performs a preset operation and maintenance processing in a case where it is analyzed that at least part of the target operation indexes is abnormal, the operation and maintenance processing module 420 is configured to: determine whether the index values of the at least part of the target operation indexes in the plurality of target operation indexes satisfy the abnormal condition indicated by the first target rule; determine a severity of the abnormality, and perform an alarm based on the severity, and / or perform a target abnormality solving action that is pre-arranged, so as to restore the at least part of the target operation indexes that is abnormal to normal.

[0096] In a possible implementation, when the operation and maintenance processing module 420 performs the target abnormality solving action that is pre-arranged, the operation and maintenance processing module 420 is configured to: perform the target abnormality solving action that is pre-arranged and corresponds to the target abnormal condition.

[0097] In a possible implementation, the operation and maintenance processing module 420 is further configured to perform optimization processing on the preset rules, the optimization processing including at least one of the following: identifying the similarity between any two rules in the preset rules, and merging any two rules with a similarity greater than a similarity threshold; or detecting at least two mutually conflicting rules in the preset rules, and retaining one of the at least two mutually conflicting rules; or adjusting the priority of each rule in the preset rules according to the operation and maintenance effect, the priority being used to indicate the order of analyzing the index value.

[0098] Figure 5 A structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. It should be noted that, Figure 5 The computer system 500 of the electronic device shown is only an example and should not impose any limitation on the functions and use range of the embodiments of the present application. The electronic device can be used to execute the operation and maintenance processing method provided by the application.

[0099] As Figure 5 shown, the computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or programs loaded from a storage portion 508 to a random access memory (RAM) 503, such as performing the methods in the above embodiments. In the RAM 503, various programs and data required for system operation are also stored. The CPU 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0100] The following components are connected to the I / O interface 505: an input section 506 including input devices such as a keyboard and mouse; an output section 507 including output devices such as a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), and a speaker; a storage section 508 including a hard disk; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as necessary. A removable medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 510 as necessary, so that a computer program read therefrom is installed into the storage section 508 as necessary.

[0101] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present application. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, various functions defined in the system of the present application are executed.

[0102] It should be noted that the computer-readable medium in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disc Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present application, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, transmit, propagate or transport a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, or the like, or any suitable combination thereof.

[0103] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In the flowcharts or block diagrams, each block can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than that shown in the drawings. For example, two blocks that are shown in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams or flowcharts, and the combination of blocks in the block diagrams or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0104] The units described in the embodiments of the present application can be implemented by software, or by hardware, or by a combination of software and hardware. The units described can also be located in a single processor. In some cases, the names of the units do not limit the units themselves.

[0105] As another aspect, the present application also provides a computer readable storage medium, which can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device. The computer readable storage medium carries computer readable instructions, which, when executed by a processor, implement the method in any of the above embodiments.

[0106] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory), or a combination thereof, and similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a whole module or unit of the function of the module or unit, or a part of the module or unit.

[0107] According to an aspect of the embodiments of the present application, a computer program product is provided, which includes computer instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the electronic device to perform the method in any of the above embodiments.

[0108] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. Indeed, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units.

[0109] Those skilled in the art can easily understand, through the above description of the embodiments, that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, or the like) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to perform the methods according to the embodiments of the present application.

[0110] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the embodiments disclosed herein. It is intended that the present application cover any and all variations of the present application that come within the scope of the claims and that the terms describe and of the specification be interpreted to cover such variations.

[0111] It should be understood that the present application is not limited to the precise construction that has been described above and illustrated in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the present application. The scope of the present application is limited only by the appended claims.

Claims

1. An operation and maintenance processing method, characterized by, The method comprises the following steps: obtaining index values of a plurality of running indexes, the plurality of running indexes comprising running indexes of at least one of a target device, a server or a database, the server and the database being used to implement a business scenario of the target device; determining a target running index in the plurality of running indexes based on the index values of the plurality of running indexes; if the target running index is a plurality, obtaining a first target rule corresponding to the plurality of target running indexes from a pre-set rule, the first target rule comprising a basic rule, a composite rule and a scenario rule corresponding to each target running index, the basic rule being used to indicate a first abnormal condition of the target running index, the composite rule being used to indicate a second abnormal condition composed of at least two first abnormal conditions, and the scenario rule being used to indicate a third abnormal condition of the business scenario of the target device; sequentially analyzing the index values of the plurality of target running indexes through the first target rule, and executing a pre-set operation and maintenance processing in a case where at least part of the target running indexes is abnormal.

2. The method of claim 1, wherein, The method further comprises: if the target running index is one, obtaining a second target rule corresponding to the target running index, the second target rule being used to indicate a fourth abnormal condition of the target running index; analyzing the characteristic value of the target running index through the second target rule, and executing the pre-set operation and maintenance processing in a case where the target running index is abnormal.

3. The method of claim 1, wherein, The method further comprises: obtaining index values of a plurality of running indexes, the plurality of running indexes comprising running indexes of at least one of a target device, a server or a database, the server and the database being used to implement a business scenario of the target device; determining a target running index in the plurality of running indexes based on the index values of the plurality of running indexes; The method further comprises:

4. The method of claim 1, wherein, if the target running index is one, obtaining a second target rule corresponding to the target running index, the second target rule being used to indicate a fourth abnormal condition of the target running index; analyzing the characteristic value of the target running index through the second target rule, and executing the pre-set operation and maintenance processing in a case where the target running index is abnormal. The method further comprises: obtaining index values of a plurality of running indexes, the plurality of running indexes comprising running indexes of at least one of a target device, a server or a database, the server and the database being used to implement a business scenario of the target device; 5. The method of claim 1, wherein, determining a target running index in the plurality of running indexes based on the index values of the plurality of running indexes; The method further comprises: if the target running index is one, obtaining a second target rule corresponding to the target running index, the second target rule being used to indicate a fourth abnormal condition of the target running index; analyzing the characteristic value of the target running index through the second target rule, and executing the pre-set operation and maintenance processing in a case where the target running index is abnormal. The method further comprises: obtaining index values of a plurality of running indexes, the plurality of running indexes comprising running indexes of at least one of a target device, a server or a database, the server and the database being used to implement a business scenario of the target device; determining a target running index in the plurality of running indexes based on the index values of the plurality of running indexes; The method further comprises: if the target running index is one, obtaining a second target rule corresponding to the target running index, the second target rule being used to indicate a fourth abnormal condition of the target running index; analyzing the characteristic value of the target running index through the second target rule, and executing the pre-set operation and maintenance processing in a case where the target running index is abnormal. If the index values of the at least part of the target running indicators satisfy at least one target abnormal condition indicated by the first target rule, it is determined that the index values of the at least part of the target running indicators satisfying the target abnormal condition are abnormal; determining the severity of the abnormality, and based on the severity, performing an alarm and / or executing a pre-arranged target abnormality solving action to make the at least part of the target running indicators abnormal return to normal.

6. The method of claim 5, wherein, The execution of the pre-arranged target abnormality solving action includes: executing the pre-arranged target abnormality solving action corresponding to the target abnormal condition.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: optimizing the pre-set rules, the optimization including at least one of the following: identifying the similarity between any two rules in the pre-set rules, and merging the any two rules with a similarity greater than a similarity threshold; or, detecting at least two conflicting rules in the pre-set rules, and retaining one of the at least two conflicting rules; or, adjusting the priority of each rule in the pre-set rules according to the operation and maintenance effect, the priority being used to represent the order of analyzing the index values.

8. An operation and maintenance processing apparatus characterized by comprising: including: an acquisition module, configured to acquire index values of a plurality of running indicators, the plurality of running indicators including running indicators of at least one of a target device, a server or a database, the server and the database being used to implement a business scenario of the target device; an operation and maintenance processing module, configured to determine target running indicators in the plurality of running indicators based on the index values of the plurality of running indicators; if the target running indicators are multiple, acquiring first target rules corresponding to the multiple target running indicators from pre-set rules, the first target rules including a basic rule, a composite rule and a scenario rule corresponding to each target running indicator, the basic rule being used to indicate a first abnormal condition of the target running indicator, the composite rule being used to indicate a second abnormal condition composed of at least two first abnormal conditions, and the scenario rule being used to indicate a third abnormal condition of a business scenario of the target device; sequentially analyzing the index values of the multiple target running indicators through the first target rules, and in the case that at least part of the target running indicators are abnormal, executing a pre-set operation and maintenance processing.

9. An electronic device, comprising: including: a processor; a memory, the memory having computer readable instructions stored thereon, the computer readable instructions being executed by the processor to implement the method of any one of claims 1-7.

10. A computer-readable storage medium having stored thereon computer-readable instructions, the computer-readable instructions comprising: When the computer readable instructions are executed by the processor, the method of any one of claims 1-7 is implemented.

Citation Information

Cited By

  • Resource level detection method and device of third-party script, equipment and storage medium

    CN121637188A