Intelligent operation and maintenance monitoring method and device, electronic equipment and storage medium
By building a global resource topology diagram of the IT system and calculating the global impact value, the quantitative evaluation problem of changing operations in IT operations and maintenance is solved, improving evaluation efficiency and reducing risks.
Patent Information
- Application Number
- CN202510243263.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the lack of pre-quantitative evaluation of IT operation and maintenance changes, resulting in low efficiency and limited effectiveness of manual review, and high risks in post-event monitoring.
Build a global resource topology diagram of the IT system, group resources according to the preset type and set weights, determine the associated resources through the global resource topology diagram, and calculate the global impact value of operation and maintenance change operations.
It realizes a quantitative assessment of the scope of impact of IT operation and maintenance changes, reducing the time cost of manual review and the risks of post-event monitoring.
Smart Images

Figure CN120353660A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of IT operation and maintenance, and in particular, to an intelligent operation and maintenance monitoring method, device, electronic device, and storage medium. Background Art
[0002] In the actual IT operation and maintenance scenario, operation and maintenance related personnel often need to perform change operations on basic resources, application resources, etc., and thus need to evaluate the impact scope of the change operation on the overall resources, minimize the impact scope as much as possible, and ensure the high efficiency and availability of the IT system.
[0003] In the related art, either manual pre-review is adopted, that is, through the operation and maintenance group meeting review, manually sorting out the resource topology relationship and analyzing the impact of the change on the upstream and downstream resources; or post-monitoring analysis is adopted, that is, using CMDB to monitor, using APM software to monitor applications, using skywalking tools to monitor the system interaction link, using SNMP protocol to monitor server metrics, etc., to perform abnormal alarm and emergency handling of the abnormality and then execute the change again.
[0004] However, the manual pre-review lacks general quantitative evaluation indicators, and the result depends on the review team and review quality, with high time cost and limited effect; although the post-monitoring analysis can accurately evaluate the impact scope, it is often a post-remedial measure, which may lead to irreparable business risks and cannot effectively support the availability, security, and stability of the system. Summary of the Invention
[0005] The present invention provides an intelligent operation and maintenance monitoring method, device, electronic device, and storage medium to solve the technical problem of being unable to pre-quantitatively calculate the impact scope of change operations in IT operation and maintenance.
[0006] In a first aspect, the present invention provides an intelligent operation and maintenance monitoring method, including: constructing a global resource topology relationship graph of an IT system; grouping each resource of the IT system according to a preset type, setting a first weight for each resource, and setting a second weight for each grouping label; before performing an operation and maintenance change operation on a target resource in the IT system, determining associated resources of the target resource according to the global resource topology relationship graph; and determining a global impact value of the operation and maintenance change operation according to the associated resources and the corresponding first weights, the grouping labels to which the associated resources belong, and the corresponding second weights.
[0007] In some embodiments, constructing a global resource topology relationship diagram of the IT system includes: sorting out the basic resources and application resources of the IT system, where the basic resources include at least one of servers and cabinets; determining the physical location of the cabinets, and associating the physical location of the cabinets with the management IP address and network IP address of the servers to form a basic resource topology relationship diagram; associating the network IP address of the servers with the application resources to form an application resource topology relationship diagram; and associating the basic resource topology relationship diagram and the application resource topology relationship diagram through the network IP address of the servers to form a global resource topology relationship diagram.
[0008] In some embodiments, determining the associated resources of the target resource according to the global resource topology relationship diagram includes: performing a depth-first search based on the global resource topology relationship diagram to determine the first resources directly affected by the target resource, where the first resources include the target resource itself; performing a breadth-first search based on the global resource topology relationship diagram to determine the second resources indirectly associated with the target resource, and the first resources and the second resources constitute the associated resources.
[0009] In some embodiments, the preset types include at least one of the following: cloud service type, cluster type, service priority type; the grouping tags corresponding to the cloud service type include at least one of the following: infrastructure as a service, platform as a service, software as a service; the grouping tags corresponding to the cluster type include at least one of the following: application cluster, data cluster, message cluster; the grouping tags corresponding to the service priority type include at least one of the following: basic service, critical business service, non-critical business service.
[0010] In some embodiments, the calculation formula of the global impact value is as follows:
[0011]
[0012] where f(x) represents the global impact value, n represents the total number of resources of the associated resources, m represents the total number of grouping tags of the grouping tags to which the associated resources belong, R i represents the i-th associated resource, Y i represents the first weight corresponding to the i-th associated resource, T j represents the j-th grouping tag, W j represents the second weight corresponding to the j-th grouping tag, C1 and C2 respectively represent preset constants, and Δx represents the manual correction offset.
[0013] In some embodiments, the method further includes at least one of the following: determining at least one preset quantitative evaluation index value according to the global impact value; and outputting a risk warning message when the global impact value is greater than a preset threshold.
[0014] Second aspect, the present invention provides an intelligent operation and maintenance monitoring device, including: a resource topology construction module, configured to construct a global resource topology relationship diagram of an IT system; a grouping and weighting module, configured to group each resource of the IT system according to a preset type, and set a first weight for each resource and a second weight for each grouping label; an associated resource loading module, configured to determine the associated resources of a target resource in the IT system according to the global resource topology relationship diagram before performing an operation and maintenance change operation on the target resource; an operation and maintenance impact calculation module, configured to determine a global impact value of the operation and maintenance change operation according to the associated resources and the corresponding first weights, the grouping labels to which the associated resources belong, and the corresponding second weights.
[0015] In some embodiments, the resource topology construction module is specifically configured to: sort out the basic resources and application resources of the IT system, where the basic resources include at least one of a server and a cabinet; determine the physical location of the cabinet, and associate the physical location of the cabinet with the management IP address and network IP address of the server to form a basic resource topology relationship diagram; associate the network IP address of the server with the application resources to form an application resource topology relationship diagram; associate the basic resource topology relationship diagram and the application resource topology relationship diagram through the network IP address of the server to form a global resource topology relationship diagram.
[0016] Third aspect, the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; the memory is used to store a computer program; the processor is configured to implement the steps of the intelligent operation and maintenance monitoring method according to any one of the first aspect when executing the program stored on the memory.
[0017] Fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and is characterized in that the computer program implements the steps of the intelligent operation and maintenance monitoring method according to any one of the first aspect when executed by a processor.
[0018] An intelligent operation and maintenance monitoring method, device, electronic device, and storage medium provided by an embodiment of the present invention determine its associated resources according to the global resource topology relationship before performing an operation and maintenance change on a certain resource, and calculate according to the weights of each preset resource and each resource grouping label, obtaining the global impact value of the operation and maintenance change operation, realizing the quantitative evaluation of the impact range of operation and maintenance changes in IT operation and maintenance, and solving the problems of low efficiency and high risk in manual review and post-event monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a schematic flowchart of an intelligent operation and maintenance monitoring method provided by an embodiment of the present invention;
[0022] Figure 2 It is a global resource topology relationship diagram provided by an embodiment of the present invention;
[0023] Figure 3 It is a schematic diagram of cloud service type grouping provided by an embodiment of the present invention;
[0024] Figure 4 It is a schematic diagram of service priority type grouping provided by an embodiment of the present invention;
[0025] Figure 5 For Figure 1 It is a detailed flowchart of step S101 in the illustrated embodiment;
[0026] Figure 6 For Figure 1 It is a detailed flowchart of step S103 in the illustrated embodiment;
[0027] Figure 7 It is a schematic structural diagram of an intelligent operation and maintenance monitoring device provided by an embodiment of the present invention;
[0028] Figure 8 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. Specific Embodiments
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0030] First, the nouns involved in the present invention are explained:
[0031] Operation and maintenance: It refers to a series of management, maintenance, monitoring, optimization, etc. work carried out to ensure and maintain various information technology infrastructures such as information systems, network devices, and software applications.
[0032] CMDB: Refers to the Configuration Management Database, which stores and manages various configuration information of devices in the IT architecture.
[0033] Cabinet: Refers to a metal cabinet for storing and installing electronic devices, such as server cabinets, network cabinets, etc.
[0034] APM: Refers to an IT application performance management tool that maintains measurable application performance metrics.
[0035] Rule engine: Refers to a certain judgment program maintained and executed in a rule set.
[0036] Figure 1 It is a schematic flowchart of an intelligent operation and maintenance monitoring method provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps:
[0037] Step S101, construct a global resource topology relationship diagram of the IT system.
[0038] Specifically, the IT system includes basic resources and application resources. Basic resources such as cabinets, servers, network devices, storage devices, various terminal devices, etc., and application resources such as application services, system processes, resource directories, etc. Build CMDB resource management, uniformly label and associate with IP addresses, and establish a global resource topology relationship diagram of the IT system. As Figure 2 It is a global resource topology relationship diagram provided by an embodiment of the present invention.
[0039] Step S102, group the resources of the IT system according to a preset type, and set a first weight for each resource and a second weight for each group label.
[0040] Specifically, there are multiple grouping bases. The CMDB resources can be effectively grouped according to the corresponding grouping bases, determine the group labels to which each resource belongs, and set weights (Y1, Y2... Y n ) for each resource, and set weights (W1, W2.... W m ) for each group label. The weights are obtained by comprehensively considering implementation costs, hardware costs, and software costs.
[0041] In some embodiments, the preset type includes at least one of the following: cloud service type, cluster type, service priority type; the group labels corresponding to the cloud service type include at least one of the following: Infrastructure as a Service, Platform as a Service, Software as a Service; the group labels corresponding to the cluster type include at least one of the following: application cluster, data cluster, message cluster; the group labels corresponding to the service priority type include at least one of the following: basic service, critical business service, non-critical business service.
[0042] Specifically, the grouping basis includes cloud service types and cluster types. For example, according to cloud service types, each resource is divided into Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). According to cluster types, it is divided into application clusters, data clusters, message clusters, etc. The grouping tags are named Tag1, Tag2....Tagm (abbreviated as T1, T2...T m ), and the specific grouping content is named Group1, Group2....Groupn (abbreviated as G1, G2…G n ), such as Figure 3 is a schematic diagram of cloud service type grouping provided by an embodiment of the present invention.
[0043] The grouping also includes service priority types. In the conventional change application scenario, grouping and division within the service group will also be performed to distinguish basic services, critical business services, non-critical business services, etc., such as Figure 4 is a schematic diagram of service priority type grouping provided by an embodiment of the present invention.
[0044] Step S103: Before performing an operation and maintenance change operation on the target resource in the IT system, determine the associated resources of the target resource according to the global resource topology diagram.
[0045] Specifically, before performing an operation and maintenance change operation on the target resource in the IT system, load the global resource topology diagram and retrieve the resources associated with the target resource. It should be noted that the associated resources include the target resource itself and other resources associated with the target resource.
[0046] Step S104: Determine the global impact value of the operation and maintenance change operation according to the associated resources and the corresponding first weight, the grouping tags to which the associated resources belong, and the corresponding second weight.
[0047] Specifically, input the associated resource identifier, the grouping tag identifier to which the associated resource belongs, and the corresponding first weight and second weight into the rule engine to calculate the impact value of the target resource on the overall resources in the case of an operation and maintenance change operation.
[0048] In some embodiments, the calculation formula of the global impact value is as follows:
[0049]
[0050] Among them, f(x) represents the global impact value, n represents the total number of associated resources, m represents the total number of grouping tags of the grouping tags to which the associated resources belong, R i represents the i-th associated resource, Y i represents the first weight corresponding to the i-th associated resource, Tj represents the j-th grouping label, W j represents the second weight corresponding to the j-th grouping label, C1 and C2 respectively represent preset constants, and Δx represents the manual correction offset.
[0051] It should be noted that C1 and C2 are constants used to regulate the different impacts of grouping labels and specific grouping contents on the stability of the IT system. By defining C1 and C2, the corresponding impact ratio can be adjusted to correct the data magnitude. In addition, in the calculation of certain specific resources, in addition to considering the weight, manual parameter correction Δx is also required. For example, in Figure 4 In the scenario where the report center service is regarded as a non-critical service, during non-high-load business periods, its impact degree may be reduced, so offset correction is required.
[0052] In some embodiments, the method further includes: determining at least one preset quantitative evaluation index value according to the global impact value; and outputting a risk warning message when the global impact value is greater than a preset threshold.
[0053] Specifically, after calculating the global impact value f(x), quantitative evaluation indexes such as risk level and impact range ratio can be output according to f(x); or when f(x) is relatively large, a risk warning can be triggered to assist the operation and maintenance personnel in judging whether to perform operation and maintenance change operations.
[0054] The intelligent operation and maintenance monitoring method provided in this embodiment, before performing operation and maintenance changes on a certain resource, determines its associated resources according to the global resource topology relationship, and calculates according to the weights of the preset resources and the grouping labels of each resource to obtain the global impact value of the operation and maintenance change operation, realizing the quantitative evaluation of the impact range of operation and maintenance changes in IT operation and maintenance, and solving the problems of low efficiency and high risk in manual review and post-event monitoring.
[0055] On the basis of the foregoing embodiments, Figure 5 For Figure 1 a detailed flowchart of step S101 in the illustrated embodiment, as Figure 5 shown, this step S101 includes the following steps:
[0056] Step S1011: Sort out the basic resources and application resources of the IT system, and the basic resources include at least one of servers and cabinets.
[0057] Specifically, sort out the basic resources, which include physical machines, cabinets, network devices, storage devices, various terminal devices, etc.; sort out the application resources, including application services, system processes, resource directories, etc.; each resource can be identified, such as abbreviated as R1, R2...Rn.
[0058] Step S1012: Determine the physical location of the cabinet, and associate the physical location of the cabinet with the management IP address and network IP address of the server to form a basic resource topology diagram.
[0059] Specifically, register the physical environment and the cabinet, associate the cabinet location with the IP address managed by the server, associate the IP address managed by the server with the network IP address of the server. At the same time, auxiliary information such as the brand and accessories of the server can also be uniformly associated with the IP address managed by the server to form a basic resource topology.
[0060] Step S1013: Associate the network IP address of the server with application resources to form an application resource topology diagram.
[0061] Specifically, use the server network IP to associate with application services, system processes, resource directories, etc. to form an application resource topology.
[0062] Step S1014: Associate the basic resource topology diagram and the application resource topology diagram through the network IP address of the server to form a global resource topology diagram.
[0063] Specifically, associate the basic resource topology with the application resource topology through the network IP address of the server to form a global resource topology diagram, as Figure 2 shown.
[0064] Based on the foregoing embodiments, by sorting out the basic resources and application resources of the IT system, the basic resources include at least one of servers and cabinets; determine the physical location of the cabinet, and associate the physical location of the cabinet with the management IP address and network IP address of the server to form a basic resource topology diagram; associate the network IP address of the server with application resources to form an application resource topology diagram; associate the basic resource topology diagram and the application resource topology diagram through the network IP address of the server to form a global resource topology diagram; that is, construct a global resource topology diagram through IP association to realize dynamic resource management.
[0065] Based on the foregoing embodiments, Figure 6 For Figure 1 a detailed flowchart of step S103 in the shown embodiment, as Figure 6 described, step S103 includes the following steps:
[0066] Step S1031: Perform a depth-first search based on the global resource topology diagram to determine the first resources directly affected by the target resource, and the first resources include the target resource itself.
[0067] Step S1032: Perform a breadth-first search based on the global resource topology relationship graph to determine the second resources indirectly associated with the target resource. The first resource and the second resources constitute the associated resources.
[0068] Specifically, before the operation and maintenance change operation, load the resource topology relationship and perform the following two types of searches. One is a depth-first search, that is, vertically traverse the upstream and downstream resources to lock the directly affected resource scope. The directly affected resources include the target resource itself. The other is a breadth-first search, that is, a horizontal expansion search to identify indirectly associated resources.
[0069] Take Figure 4 the scenario as an example. To meet the business requirements and upgrade a data center service, the upgrade of the data center service will affect the order inventory approval in the order center and the display of order volume-related data in the report center. Therefore, when calculating the impact value, it is necessary to perform a depth-first vertical search to lock that both the order center and the report center will be affected, and then perform a breadth-first horizontal search to determine whether other applications are affected by the association of the above three applications. Then, load the application resources, application labels, and corresponding weights into the rule engine for calculation to obtain the global impact value. That is, substituting into formula (1), R1 is the data center service, R2 is the order center service, R3 is the report center service, T1 is the basic application layer, T2 is the key application layer, and T3 is the non-critical application layer.
[0070] Based on the foregoing embodiments, perform a depth-first search based on the global resource topology relationship graph to determine the first resources directly affected by the target resource. The first resources include the target resource itself; perform a breadth-first search based on the global resource topology relationship graph to determine the second resources indirectly associated with the target resource. The first resources and the second resources constitute the associated resources; that is, based on the dual search mechanism of depth-first and breadth-first, automatically and comprehensively determine the associated resources.
[0071] Figure 7 The following is a schematic structural diagram of an intelligent operation and maintenance monitoring device provided by an embodiment of the present invention. As Figure 7 shown, the device includes:
[0072] A resource topology construction module 701, configured to construct a global resource topology relationship graph of an IT system;
[0073] A grouping and weighting module 702, configured to group the resources of the IT system according to a preset type, and set a first weight for each resource and a second weight for each group label;
[0074] An associated resource loading module 703, configured to determine the associated resources of the target resource according to the global resource topology relationship graph before performing an operation and maintenance change operation on the target resource in the IT system;
[0075] An operation and maintenance impact calculation module 704, configured to determine a global impact value of an operation and maintenance change operation according to the associated resources and corresponding first weights, grouping tags to which the associated resources belong, and corresponding second weights.
[0076] In some embodiments, the resource topology construction module 701 is specifically configured to:
[0077] Sort out the basic resources and application resources of the IT system, where the basic resources include at least one of servers and cabinets;
[0078] Determine the physical location of the cabinet, and associate the physical location of the cabinet with the management IP address and network IP address of the server to form a basic resource topology relationship diagram;
[0079] Associate the network IP address of the server with the application resources to form an application resource topology relationship diagram;
[0080] Associate the basic resource topology relationship diagram and the application resource topology relationship diagram through the network IP address of the server to form a global resource topology relationship diagram.
[0081] In some embodiments, the associated resource loading module 703 is specifically configured to:
[0082] Perform a depth-first search based on the global resource topology relationship diagram to determine the first resources directly affected by the target resource, where the first resources include the target resource itself;
[0083] Perform a breadth-first search based on the global resource topology relationship diagram to determine the second resources indirectly associated with the target resource, and the first resources and the second resources constitute the associated resources.
[0084] In some embodiments, the preset types include at least one of the following: cloud service type, cluster type, service priority type;
[0085] The grouping tags corresponding to the cloud service type include at least one of the following: infrastructure as a service, platform as a service, software as a service;
[0086] The grouping tags corresponding to the cluster type include at least one of the following: application cluster, data cluster, message cluster;
[0087] The grouping tags corresponding to the service priority type include at least one of the following: basic service, critical business service, non-critical business service.
[0088] In some embodiments, the calculation formula of the global impact value is as follows:
[0089]
[0090] Among them, f(x) represents the global influence value, n represents the total number of associated resources, m represents the total number of grouping labels of the grouping labels to which the associated resources belong, and R i represents the i-th associated resource, and Y i represents the first weight corresponding to the i-th associated resource, and T j represents the j-th grouping label, and W j represents the second weight corresponding to the j-th grouping label, C1 and C2 respectively represent preset constants, and Δx represents the manual correction offset.
[0091] In some embodiments, the operation and maintenance impact calculation module 704 is further configured to perform at least one of the following:
[0092] Determine at least one preset quantitative evaluation index value according to the global influence value;
[0093] Output a risk warning message when the global influence value is greater than a preset threshold.
[0094] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process and corresponding beneficial effects of the intelligent operation and maintenance monitoring device described above can refer to the corresponding process in the foregoing method examples, and will not be elaborated here.
[0095] Figure 8 The following is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. As Figure 8 shown, the electronic device includes: a processor 801, a communication interface 802, a memory 803, and a communication bus 804. Among them, the processor 801, the communication interface 802, and the memory 803 complete communication with each other through the communication bus 804.
[0096] The memory 803 is used to store a computer program;
[0097] In an embodiment of the present application, when the processor 801 is used to execute the program stored in the memory 803, it implements the steps of the intelligent operation and maintenance monitoring method provided in any one of the foregoing method embodiments.
[0098] The electronic device provided by the embodiment of the present application has the same implementation principle and technical effects as the above embodiments, and will not be elaborated here.
[0099] The above-mentioned memory 803 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. The memory 803 has a storage space for program codes for executing any of the method steps in the above-mentioned methods. For example, the storage space for program codes may include respective program codes for implementing each of the steps in the above method. These program codes may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, optical discs (CDs), memory cards, or floppy disks. Such computer program products are typically portable or fixed storage units. The storage unit may have a storage segment or storage space, etc., arranged similarly to the memory 803 in the above-mentioned electronic device. The program codes may be compressed in an appropriate form, for example. Generally, the storage unit includes a program for executing the method steps according to the embodiments of the present application, that is, codes that can be read by a processor such as 801, and when these codes are run by the electronic device, the electronic device is caused to execute each of the steps in the method described above.
[0100] Embodiments of the present application also provide a computer-readable storage medium. A computer program is stored on the above-mentioned computer-readable storage medium, and when the computer program is executed by a processor, the steps of the intelligent operation and maintenance monitoring method described above are implemented.
[0101] The computer-readable storage medium may be included in the device / device described in the above embodiments; or it may exist separately without being assembled into the device / device. The above-mentioned computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0102] According to the embodiments of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in conjunction with an instruction execution system, device, or device.
[0103] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0104] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. An intelligent operation and maintenance monitoring method, characterized in that, Including: Constructing a global resource topology relationship diagram of the IT system; Grouping each resource of the IT system according to a preset type, setting a first weight for each resource, and setting a second weight for each grouping label; Before performing an operation and maintenance change operation on a target resource in the IT system, determining the associated resources of the target resource according to the global resource topology relationship diagram; Determining a global impact value of the operation and maintenance change operation according to the associated resources and the corresponding first weights, the grouping labels to which the associated resources belong, and the corresponding second weights.
2. The method according to claim 1, wherein The constructing of the global resource topology relationship diagram of the IT system includes: Sorting out the basic resources and application resources of the IT system, where the basic resources include at least one of servers and cabinets; Determining the physical location of the cabinet, and associating the physical location of the cabinet with the management IP address and network IP address of the server to form a basic resource topology relationship diagram; Associating the network IP address of the server with the application resources to form an application resource topology relationship diagram; Associating the basic resource topology relationship diagram and the application resource topology relationship diagram through the network IP address of the server to form a global resource topology relationship diagram.
3. The method according to claim 1, characterized in that, The determining of the associated resources of the target resource according to the global resource topology relationship diagram includes: Performing a depth-first search based on the global resource topology relationship diagram to determine the first resources directly affected by the target resource, where the first resources include the target resource itself; Performing a breadth-first search based on the global resource topology relationship diagram to determine the second resources indirectly associated with the target resource, and the first resources and the second resources constitute the associated resources.
4. The method according to any one of claims 1 to 3, characterized in that, The preset type includes at least one of the following: cloud service type, cluster type, service priority type; The grouping labels corresponding to the cloud service type include at least one of the following: Infrastructure as a Service, Platform as a Service, Software as a Service; The grouping labels corresponding to the cluster type include at least one of the following: application cluster, data cluster, message cluster; The grouping labels corresponding to the service priority type include at least one of the following: basic service, critical business service, non-critical business service.
5. The method according to claim 4, characterized in that, The calculation formula of the global impact value is as follows: Among them, f(x) represents the global influence value, n represents the total number of resources of associated resources, m represents the total number of grouping labels of the grouping labels to which the associated resources belong, R i represents the i-th associated resource, Y i represents the first weight corresponding to the i-th associated resource, T j represents the j-th grouping label, W j represents the second weight corresponding to the j-th grouping label, C1 and C2 respectively represent preset constants, and Δx represents the artificial correction offset.
6. The method according to any one of claims 1 to 3, characterized in that, The method further includes at least one of the following: Determining at least one preset quantitative evaluation index value according to the global impact value; Outputting a risk warning message when the global impact value is greater than a preset threshold.
7. An intelligent operation and maintenance monitoring device, characterized in that, Including: A resource topology construction module for constructing a global resource topology relationship diagram of the IT system; A grouping and weighting module for grouping each resource of the IT system according to a preset type, setting a first weight for each resource, and setting a second weight for each grouping label; An associated resource loading module for determining the associated resources of the target resource according to the global resource topology relationship diagram before performing an operation and maintenance change operation on the target resource in the IT system; An operation and maintenance impact calculation module for determining a global impact value of the operation and maintenance change operation according to the associated resources and the corresponding first weights, the grouping labels to which the associated resources belong, and the corresponding second weights.
8. The device according to claim 7, characterized in that, The resource topology construction module is specifically used for: Sort out the basic resources and application resources of the IT system, where the basic resources include at least one of servers and cabinets; Determine the physical location of the cabinet, and associate the physical location of the cabinet with the management IP address and network IP address of the server to form a basic resource topology diagram; Associate the network IP address of the server with the application resources to form an application resource topology diagram; Associate the basic resource topology diagram and the application resource topology diagram through the network IP address of the server to form a global resource topology diagram.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is used to implement the steps of the intelligent operation and maintenance monitoring method described in any one of claims 1-6 when executing the program stored on the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it realizes the steps of the intelligent operation and maintenance monitoring method described in any one of claims 1-6.