Method and system for realizing fault diagnosis of IT equipment and application system based on fault analysis model
By using a fault analysis model-based approach, the problems of decentralized monitoring and lack of asset integration in oilfield information system construction were solved, enabling rapid and accurate fault location and diagnosis, and improving the efficiency of oilfield information operation and maintenance.
Patent Information
- Application Number
- CN202111605152.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2041-12-24
AI Technical Summary
Traditional fault analysis methods in oilfield information construction suffer from problems such as fragmented monitoring, lack of in-depth analysis tools, and lack of integration between assets and monitoring, making fault location difficult and hindering the rapid and accurate identification of problems.
A fault analysis model-based approach is adopted. By acquiring and classifying asset types, defining asset summary information and monitoring information configuration, establishing asset relationship types, using topological relationships to determine the scope of fault impact, and generating a fault analysis topology diagram for diagnosis.
It improves fault location efficiency, enhances the intuitiveness and accuracy of fault analysis, reduces manpower requirements, provides rapid fault diagnosis methods, and improves the efficiency of information operation and maintenance.
Smart Images

Figure CN116389225B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of oil and gas industry information operation and maintenance technology, and relates to a method and system for realizing fault diagnosis of IT equipment and application systems based on a fault analysis model. BACKGROUND
[0002] With the construction and development of oilfield informatization, the equipment and systems of the oilfield data center have undergone tremendous changes in architecture and management, and each IT infrastructure is being pooled, virtualized, or clouded. This results in increasingly complex system application structures. In the traditional application deployment mode, a fault in a set of application systems can determine a small range for finding and solving the problem. However, with the current application deployment cloudification, structure componentization, data platformization, and interface publicization modes, all applications are made into a whole, forming a unified overall structure and a highly abstracted deployment mode. For a single application, fault positioning and analysis also change from individual to overall analysis. Therefore, the number of influencing factors for fault analysis increases, such as CPU, memory, disk, operating system, port, backup, database, data platform, public component, cloud platform, network, web service, application program code, SQL statement, etc., and sometimes even voltage, temperature, dust, etc. The traditional fault analysis in information operation and maintenance cannot solve these problems, mainly in the following aspects:
[0003] 1) Dispersed monitoring, information like an island
[0004] Each type of oilfield information infrastructure has its own professional monitoring and management software, such as communication monitoring systems, virtualization monitoring systems, and dynamic environment monitoring systems. These systems have their own dedicated monitoring platforms and cannot achieve unified monitoring and management.
[0005] 2) Lack of in-depth fault analysis means
[0006] In fault analysis, the monitoring alarm is usually analyzed and checked through the log system. For large-scale and complex applications, the structure is complex, the data volume is large, and the types are many, making it difficult for operation and maintenance personnel to quickly and accurately find the problem.
[0007] 3) Assets and monitoring are not combined
[0008] Information asset management software and monitoring software are two independent individuals without data correlation and interface between each other. The management of various information infrastructures and their configuration information requires a large number of manpower to complete asset entry, attribute entry, and manual asset information change management in the later maintenance process. These changes are not sent to the monitoring software, resulting in information inequality.
[0009] The change of application environment also determines the direction of the change of information operation and maintenance service, and how to quickly find application system faults is an important problem to be solved in the field of information operation and maintenance technology in the oil and gas industry in the future. SUMMARY
[0010] The purpose of the present application is to solve the problems in the prior art, and to provide an IT equipment and application system fault diagnosis method and system based on a fault analysis model, aiming to solve the defect of the difficulty of analyzing and troubleshooting faults in the prior art information operation and maintenance fault analysis.
[0011] To achieve the above purpose, the following technical solutions are adopted:
[0012] The IT equipment and application system fault diagnosis method based on the fault analysis model comprises the following steps:
[0013] The asset type is obtained, and the asset type is classified, and the definition of asset summary information, the definition of asset basic attributes and asset monitoring information configuration are realized according to the classification result;
[0014] The asset monitoring data is obtained, the asset relationship type is determined according to the definition of asset summary information, the definition of asset basic attributes and asset monitoring information configuration, the relationship between monitoring information is matched according to the asset monitoring data and the asset relationship type, and the fault influence range is determined according to the relationship between monitoring objects and the topological relationship;
[0015] The asset monitoring index is established according to the asset monitoring data, the object fault analysis topology graph is obtained according to the fault influence range and the asset monitoring index, and the fault diagnosis is realized.
[0016] Preferably, the asset monitoring information comprises a subclass association identifier; the asset monitoring data is stored in an elastic search cluster, and the subclass association identifier is stored in a my sql database.
[0017] Preferably, when monitoring data in the elastic search cluster, a strategy needs to be established to search data from the elasticsearch cluster at regular intervals.
[0018] Preferably, the object is found according to the content of the specified field in the data to the asset table in the database.
[0019] Preferably, the asset monitoring data is collected through Elastic.
[0020] Preferably, the asset relationship type comprises dependency, backup, belonging, virtualization, data exchange, running on, installation, use and connection.
[0021] Preferably, the topological relationship is an ANTV G6 topological relationship.
[0022] The application provides an IT equipment and application system fault diagnosis method based on a fault analysis model.
[0023] A data acquisition module is configured to acquire asset types, classify the asset types, and define asset summary information, basic attribute definition of the assets and asset monitoring information configuration according to the classification results.
[0024] A data processing module is configured to acquire asset monitoring data, determine asset relationship types according to the definition of the asset summary information, the basic attribute definition of the assets and the asset monitoring information configuration, match the relationship between the monitoring information and the asset monitoring data, and determine the fault influence range according to the relationship between the monitoring objects.
[0025] A fault diagnosis module is configured to establish asset monitoring indexes according to the asset monitoring data, acquire object fault analysis topology according to the fault influence range and the asset monitoring indexes, and realize fault diagnosis.
[0026] A terminal device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor realizes the steps of the IT equipment and application system fault diagnosis method based on the fault analysis model when executing the computer program.
[0027] A computer readable storage medium stores a computer program, and the computer program realizes the steps of the IT equipment and application system fault diagnosis method based on the fault analysis model when executed by a processor.
[0028] Compared with the prior art, the application has the following beneficial effects:
[0029] The IT equipment and application system fault diagnosis method based on the fault analysis model finds out the fault influence range according to the relationship and the topological relationship between the monitoring objects, adopts the topological technology to manage the relationship, improves the fault positioning efficiency, and increases the intuitive analysis effect. The object fault analysis topology is generated according to the object monitoring and asset configuration information, the running states of various objects can be intuitively reflected, and the method provides a decision basis for information production command. The topological technology is comprehensively applied in the information operation fault analysis business, a unique information system fault diagnosis method is formed, and a powerful fault analysis means is provided for the fault problems occurring in the information operation of the oil and gas industry.
[0030] This invention proposes a system for diagnosing IT equipment and application systems based on a fault analysis model. By dividing the system into a data acquisition module, a data processing module, and a fault diagnosis module, the modular approach makes each module independent of the others, facilitating unified management of each module. Attached Figure Description
[0031] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 The flowchart illustrates the fault analysis model implementation method for IT equipment and application system fault diagnosis in this invention.
[0033] Figure 2 This is a diagram of the fault analysis model of the present invention;
[0034] Figure 3 This is a model diagram of the fault location system of the present invention;
[0035] Figure 4 This is a dynamic model diagram for fault location in this invention;
[0036] Figure 5 The flowchart of the fault diagnosis system for IT equipment and application systems is shown in the figure below, which is the implementation of the fault analysis model of the present invention. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0038] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0039] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0040] In the description of the embodiments of the present application, it should be noted that if the terms "upper", "lower", "horizontal", "inner" and the like indicating the orientation or position relationship are based on the orientation or position relationship shown in the drawings, or the orientation or position relationship when the product of the present application is usually placed, which is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first", "second" and the like are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0041] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly inclined. For example, "horizontal" only means that its direction is relatively more horizontal than "vertical", and does not mean that the structure must be completely horizontal, but can be slightly inclined.
[0042] In the description of the embodiments of the present application, it should be noted that unless otherwise explicitly specified and limited, if the terms "arrangement", "installation", "connection", "connection" appear, they should be understood in a broad sense, for example, they can be fixedly connected, or can be detachably connected, or integrally connected; can be mechanically connected, or can be electrically connected; can be directly connected, or can be indirectly connected through an intermediate medium; can be the communication between two elements inside. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0043] The present application will be described in further detail below in conjunction with the accompanying drawings:
[0044] The present application changes the traditional fault analysis means and mode. In the traditional fault analysis process, a professional operation and maintenance personnel is required to analyze the application fault. The operation and maintenance personnel needs to understand the system architecture and be familiar with each component of the system, and then check the problem one by one from the application data or business flow, and the difficulty and time to solve the problem are greatly related to the level of the operation and maintenance personnel and the system scale.
[0045] The IT equipment and application system fault diagnosis method based on the fault analysis model proposed by the present application can be completed in three steps, i.e., firstly, mapping the relationship of the assets and defining the indexes, then collecting the asset monitoring data, and finally, associating the indexes of the assets with the monitoring data. This analysis mode can quickly and accurately locate the fault point without the need of professional operation and maintenance analysis personnel.
[0046] As shown in Figure 1 The asset type is acquired, and the asset type is classified. According to the classification result, the definition of asset summary information, the definition of basic attributes of the asset, and the configuration of asset monitoring information are realized.
[0047] Obtaining asset monitoring data, determining asset relationship type according to definition of asset summary information, basic attribute definition of asset and asset monitoring information configuration, matching relationship between monitoring information according to asset monitoring data and asset relationship type, judging fault influence range according to relationship between monitoring objects and topological relationship;
[0048] Establishing asset monitoring index according to asset monitoring data, obtaining object fault analysis topological graph according to fault influence range and asset monitoring index, and realizing fault diagnosis.
[0049] The application utilizes information topological technology to uniformly associate and manage asset configuration information, operation log data and monitoring index data of various information infrastructure of oil fields, and mainly realizes two fault diagnosis methods of system model and dynamic model. 1) the system model can display all components of an application at one time, and give state information, configuration information, monitoring information of each component and the association relationship between components in the system; 2) the dynamic model can find problems step by step by drilling through a certain object, and this method can extend to other objects outside the application range to find problems of the application caused by the objects, for example, a switch failure causes application access failure.
[0050] The application provides an IT equipment and application system fault diagnosis method based on a fault analysis model, mainly including the following steps: asset and configuration management development; data acquisition method and big data platform building; monitoring index system establishment; topological design and algorithm realization; system model search syntax realization; function interface design and graphical realization.
[0051] Business operation steps:
[0052] Selecting an object or system and opening an analysis model of the object or system;
[0053] Drilling object relationship on the generated topological structure;
[0054] Clicking an object name to view object configuration information and monitoring information cards;
[0055] Clicking a connecting line between objects to view object relationship description;
[0056] After the informationization construction is completed, various information infrastructures and application systems enter a running maintenance stage, and diagnosing faults occurring to the same is an important content for quickly locating fault points and influence scope, which not only can improve application availability, but also can greatly reduce fault recovery time of equipment or system. In the fault diagnosis, object relationship mapping and monitoring index are important for accurately performing fault diagnosis. The function is to use an ANTV G6 topology, collect monitoring data through Elastic, develop an asset configuration dynamic management library, and realize a fault diagnosis analysis function on the topology. Embodiments of the application include five steps of asset configuration management, asset relationship mapping, asset monitoring data collection, monitoring data and asset correlation, and quick fault positioning analysis, and details are as follows:
[0057] 1), asset configuration management
[0058] Asset type division: in the management of information infrastructures, the information infrastructures are first classified, and according to the present situation of oilfield information infrastructures, are divided into six layers, 16 types and more than 70 sub-classes. The asset type division is used for clarifying system relationship and helping users to accurately understand system architecture in fault positioning analysis.
[0059]
[0060] Asset summary information configuration: the summary information of the above various types of hardware and software assets is defined, such as which application the object belongs to, which computer room it is in, and who is the operation and maintenance personnel. The application in the information has an important role in the system model in the fault analysis model, and when all monitoring objects in a system are configured, all objects under the application can be searched according to the application.
[0061] Asset basic attribute definition: the basic attributes of the above various types of hardware and software are defined, for example, the basic attributes of a small computer can include asset number, manufacturer, model, CPU, memory capacity, hard disk capacity, and cabinet attribute, and so on. The basic attributes of all information infrastructures are configured, and this function is used for viewing configuration information of monitoring objects in fault positioning and is used as one of fault analysis judgment conditions.
[0062] Asset monitoring information configuration: the sub-class correlation identifier contained in the monitoring information is used for correlating with the collected monitoring data alarm. For information infrastructures, most of them can be configured as IP addresses.
[0063] 2), asset relationship mapping
[0064] Asset relationship type: in the embodiment of the present application, 9 kinds of asset relationship are defined, including: dependency, backup, belong, virtualization, data exchange, run on, install, use, and connection. These relationships describe the relationship between monitoring objects. In fault location analysis, the relationship type can be used to determine the scope of the fault impact. The specific examples of each type of relationship are as follows:
[0065] Middleware, database, runtime - run on - operating system (physical machine, virtual machine); application system - run on - container, middleware; network device, transmission device - connection - network device, transmission device; virtualization cluster - virtualization - operating system (virtual machine); server device, power device, environment device, storage device - installation - cabinet; database, server - backup - database, server; storage device - data exchange - database; application system - use - database; operating system (physical machine) - run on - server.
[0066] Asset relationship line: in the embodiment of the present application, the relationship line between individual monitoring objects is defined as one. In actual production environment, the relationship between objects may be multiple. The present application defines the relationship as one, and adds a description field to allow operation and maintenance personnel to describe the relationship, as shown in figure 3.
[0067] 3) Asset monitoring data collection Asset data collection According to different types, different technical means, collected data, and established monitoring indicators are different. At the same time, in order to ensure the rapidity and accuracy of fault location analysis of monitoring objects, as much data as possible should be collected. The data collected in the embodiment of the present application includes 7 types of data: state, performance, warning, capacity, energy consumption, configuration, and log. According to the requirements, 23 types of data are collected, including:
[0068]
[0069]
[0070] The indicators created according to the data include:
[0071]
[0072]
[0073] In the embodiment of the present application, the purpose of collecting various types of data and creating various types of indicators is to monitor and warn, and to provide rapid positioning service for faults. The more data collected by the object, the wider the coverage of the indicators.
[0074] 4) Association of monitoring data and assets
[0075] In the embodiment of the present application, the association of monitoring data and assets is completed by relying on the subclass association identifier in asset configuration. SeeFigure 2 For example, if the sub-class association identifier of a certain switch is 192.168.1.19, and a specified field in the data collected by the device is also 192.168.1.19, the data can be regarded as the monitoring data of the device, and the monitoring index can be established according to the data.
[0076] In the embodiment of the application, all the monitoring data are stored in the elastic search cluster, the asset data are stored in the my sql database, and the sub-class association identifier is stored in the my sql database. When the monitoring data in the elastic search are associated, a strategy needs to be established to search the data from the elastic search at regular time, and the object is found in the asset table in the database according to the content of the specified field in the data. Therefore, the configuration of the sub-class association identifier must have the uniqueness of the sub-class in which the asset is located.
[0077] 5), rapid fault location analysis
[0078] In the embodiment of the application, when the system model is used for fault location analysis, the application attribute of the asset during the asset configuration is associated, and the application related object topology can be found, as shown in Figure 3 .
[0079] In the above model, the influence range of the fault object can be quickly found, and the fault can also be located through the dynamic model, that is, the influence domain is completed by manual drilling in the object topology interface, as shown in Figure 4 .
[0080] The system for realizing the IT device and application system fault diagnosis method based on the fault analysis model provided by the application, as shown in Figure 5 , comprises:
[0081] The data acquisition module is used to acquire the asset type, classify the asset type, realize the definition of asset summary information, the definition of basic attributes of the asset and the configuration of asset monitoring information according to the classification result;
[0082] The data processing module is used to acquire the asset monitoring data, determine the asset relationship type according to the definition of asset summary information, the definition of basic attributes of the asset and the configuration of asset monitoring information, match the relationship between the monitoring information according to the asset monitoring data and the asset relationship type, and judge the fault influence range according to the relationship between the monitoring objects;
[0083] The fault diagnosis module is used to establish the asset monitoring index according to the asset monitoring data, acquire the object fault analysis topology graph according to the fault influence range and the asset monitoring index, and realize the fault diagnosis.
[0084] The terminal device provided by the embodiment of the present application comprises a processor, a memory and a computer program stored in the memory and executable on the processor. The processor implements the steps in each of the method embodiments when executing the computer program. Alternatively, the processor implements the functions of each module / unit in each of the device embodiments when executing the computer program.
[0085] The computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present application.
[0086] The terminal device can be a desktop computer, a notebook computer, a palm computer, a cloud server and other computing devices. The terminal device can include, but is not limited to, a processor and a memory.
[0087] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0088] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the terminal device by running or executing the computer program and / or modules stored in the memory, and calling data stored in the memory.
[0089] The modules / units integrated in the terminal device, if realized in the form of software function units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of each method embodiment when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer-readable medium can include any entity or device, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code. It should be noted that the computer-readable medium can include or exclude contents according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0090] The application discloses an IT equipment and application system fault diagnosis method and system based on a fault analysis model, and object fault analysis topology graphs are generated through object relationship mapping and object monitoring and asset allocation information. The function provides two topology modes of a dynamic fault analysis model and a system fault analysis model. Through object topology analysis, the running states, correlation relationships and influence scopes of various objects can be directly reflected, and a decision basis is provided for information production command. The application has rich asset allocation management information and fast and flexible object relationship mapping configuration, meets asset, state and relationship management of various IT infrastructures in oil fields, and provides an important guarantee for good operation of oil field data centers.
[0091] The above only describes the preferred embodiments of the application and is not intended to limit the application. The application can be modified and changed in various ways by those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the application shall be included in the protection scope of the application.
Claims
1. A method for fault diagnosis of IT equipment and application systems based on a fault analysis model, characterized in that, Includes the following steps: The system retrieves and categorizes asset types. Based on the categorization results, it defines asset summary information, basic asset attributes, and configures asset monitoring information. Asset monitoring information includes subclass association identifiers. Asset monitoring data is stored in an Elasticsearch cluster, while subclass association identifiers are stored in a MySQL database. Asset relationship types include dependency, backup, belong, virtualization, data exchange, running on, installed, used, and connected. Acquire asset monitoring data, determine asset relationship types based on the definition of asset summary information, the definition of basic asset attributes, and the configuration of asset monitoring information, match the relationship between monitoring information based on asset monitoring data and asset relationship types, and determine the scope of fault impact based on the relationship and topology relationship between monitored objects; collect asset monitoring data through Elastic. Asset monitoring indicators are established based on asset monitoring data. Fault diagnosis is achieved by obtaining an object fault analysis topology map based on the fault impact range and asset monitoring indicators. The topology relationship is the ANTV G6 topology relationship.
2. The method for fault diagnosis of IT equipment and application systems based on a fault analysis model according to claim 1, characterized in that, When monitoring data in an Elasticsearch cluster, a strategy needs to be established to periodically search for data from the Elasticsearch cluster.
3. The method for fault diagnosis of IT equipment and application systems based on a fault analysis model according to claim 2, characterized in that, The database searches for objects in the asset table based on the specified field content in the data.
4. A system employing the fault analysis model-based fault diagnosis method for IT equipment and application systems as described in any one of claims 1 to 3, characterized in that, include: The data acquisition module is used to acquire asset types, classify asset types, and define asset summary information, basic asset attributes, and configure asset monitoring information based on the classification results. The data processing module is used to acquire asset monitoring data, determine the asset relationship type based on the definition of asset summary information, the definition of basic asset attributes and the configuration of asset monitoring information, match the relationship between monitoring information based on asset monitoring data and asset relationship type, and determine the scope of fault impact based on the relationship between monitoring objects. The fault diagnosis module is used to establish asset monitoring indicators based on asset monitoring data, and to obtain an object fault analysis topology map based on the fault impact range and asset monitoring indicators to achieve fault diagnosis.
5. A terminal device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for diagnosing faults in IT equipment and application systems based on a fault analysis model as described in any one of claims 1 to 3.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for diagnosing faults in IT equipment and application systems based on a fault analysis model as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Integrated analysis platform for integrated network management system of information system
CN107046481A
Complex system-oriented monitoring and fault self-healing system and method thereof
CN111181767A