Fault determination method and device, non-volatile storage medium, and computer terminal
By obtaining target detection indicators in the ICT operation platform and combining the linkage analysis of IT and CT infrastructure services, the problem of inaccurate infrastructure failure location in 5G private network is solved, and efficient isolation and recovery of faults is achieved.
Patent Information
- Application Number
- CN202110235249.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-03
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2041-03-03
AI Technical Summary
In 5G private networks, the existing technology cannot effectively detect the causes of infrastructure failures, resulting in inaccurate fault location and affecting network quality.
By obtaining target detection indicators in the ICT operation platform, combining the linkage analysis of IT and CT infrastructure services, fault location and isolation are achieved, and logs are collected using Syslog and DPI protocols to identify and isolate fault types.
It realizes efficient and accurate positioning and isolation of infrastructure failures, and improves the efficiency and accuracy of network operation and maintenance.
Smart Images

Figure CN115087000B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information communication technology (ICT), and in particular to a fault determination method and device, a non-volatile storage medium, and a computer terminal. Background Art
[0002] In the case of 5G private networks, since private network devices are cloud-based and carried by a unified infrastructure, network quality is not only affected by communication functions but also by the infrastructure itself. As a result, when a 5G private network failure occurs, it is difficult to identify the cause by simply screening communication indicators. Furthermore, in addition to supporting private network communication functions, the infrastructure also needs to support private network industry applications, necessitating monitoring of the dynamic resource availability of the infrastructure.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present application provide a fault determination method and apparatus, a non-volatile storage medium, and a computer terminal to at least solve the technical problem that there is no detection of infrastructure in related technologies.
[0005] According to one aspect of an embodiment of the present application, a fault determination method is provided, including: obtaining a target detection indicator of a target application provided in an ICT operation platform, wherein the target detection indicator is used to reflect the operation status of the target application, and the ICT operation platform includes a variety of infrastructure services, and the multiple infrastructure services include at least: CT infrastructure services and IT infrastructure services; determining whether a target application has failed based on the target detection indicator. It should be noted that, since the infrastructure services of the ICT platform include both IT basic services and CT basic services, in order to accurately determine whether it is the IT basic service or the CT basic service that has failed, it is necessary to troubleshoot both the IT basic service and the CT basic service at the same time, that is, when determining whether a target application has failed based on the target detection indicator, the embodiment of the present application will adopt a method of joint analysis of IT and CT equipment. The embodiment of the present application achieves efficient and accurate fault location through joint troubleshooting of IT equipment and CT equipment, and realizes unified operation and maintenance of IT / CT equipment.
[0006] According to another aspect of an embodiment of the present application, a fault handling method is also provided, including: when at least one fault occurs in an ICT operation platform, collecting current operation status information in the ICT operation platform and generating operation log information; based on the operation log information, determining the fault type of each fault in at least one fault, and locating each fault, wherein the fault types include IT faults and CT faults; and adopting a fault isolation strategy corresponding to the fault type of each fault to isolate the infrastructure service corresponding to each fault.
[0007] According to another aspect of an embodiment of the present application, a fault elimination device is also provided, including: a first acquisition module, used to obtain target detection indicators of a target application provided in an ICT operation platform, wherein the target detection indicators are used to reflect the operation status of the target application, and the ICT operation platform includes multiple infrastructure services, and the multiple infrastructure services include at least: CT infrastructure services and IT infrastructure services; a first determination module, used to determine whether a target application fails based on the target detection indicators; a second acquisition module, used to obtain the operation status of multiple infrastructure services in the ICT operation platform when a target application fails; a second determination module, used to determine the cause of the failure of the target application based on the operation status of the multiple infrastructure services.
[0008] According to another aspect of an embodiment of the present application, a non-volatile storage medium is further provided. The non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above fault determination method.
[0009] According to another aspect of an embodiment of the present application, a computer terminal is also provided, which includes: a processor; and a memory connected to the processor, for providing the processor with instructions for processing the following processing steps: obtaining target detection indicators of a target application provided in an ICT operating platform, wherein the target detection indicators are used to reflect the operating status of the target application, and the ICT operating platform includes multiple infrastructure services, and the multiple infrastructure services include at least: CT infrastructure services and IT infrastructure services; determining whether a target application fails based on the target detection indicators; in the event that the target application fails, obtaining the operating status of the multiple infrastructure services in the ICT operating platform; and determining the cause of the failure of the target application based on the operating status of the multiple infrastructure services.
[0010] In an embodiment of the present application, when a fault is determined to have occurred based on the target detection indicators of the target application in the ICT operation platform, the cause of the fault is determined based on the operating status of multiple infrastructure services in the ICT operation platform, thereby realizing the detection of the infrastructure, and further solving the technical problem that there is no detection of the infrastructure in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0012] Figure 1 This is a traditional 5G To C network operation and maintenance architecture diagram based on relevant technologies;
[0013] Figure 2 This is a schematic diagram of an ICT unified operation and maintenance architecture for industry-specific networks according to an embodiment of the present application;
[0014] Figure 3 This is a schematic diagram of the principle of ICT unified operation and maintenance for industry-specific networks according to an embodiment of the present application;
[0015] Figure 4 is a structural diagram of a computer terminal according to an embodiment of the present application;
[0016] Figure 5 is a flowchart of a fault determination method according to an embodiment of the present application;
[0017] Figure 6 is a flowchart of a fault handling method according to an embodiment of the present application;
[0018] Figure 7 is a structural diagram of a fault determination device according to an embodiment of the present application;
[0019] Figure 8 is a schematic diagram of an interactive interface according to an embodiment of the present application;
[0020] Figure 9 is a flowchart of another fault determination method according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:
[0024] 5G industry-specific networks: Private network communications provide services such as emergency communications, command and dispatch, and daily work communications for government and public security, utilities, industry, and commerce. These networks are built within certain industries, departments, or organizations to meet their organizational management, production safety, and dispatch and command needs.
[0025] ICT: The combination of information technology and communications technology. Information technology refers to the various technologies used to manage and process information. It primarily applies computer science and communications technology to the design, development, installation, and implementation of information systems and application software. Communications technology encompasses transmission access, network switching, mobile communications, wireless communications, optical communications, satellite communications, support management, and private network communications. Currently, popular technologies include 5G, LTE, IPTV, VoIP, NGN, and IMS.
[0026] Resource Orchestrator: Major cloud vendors have also launched their own Resource Orchestration (ROS) services. The concept of ROS is "Infrastructure as Code." On the one hand, it uses code-based version management to record infrastructure changes. On the other hand, it uses code to automate operations and maintenance, simplifying the complexity of writing code. Users use Json / Yaml format templates to describe the configuration and dependencies of multiple cloud computing resources (such as ECS, RDS, and SLB). It automatically deploys and configures all cloud resources in multiple regions and accounts, just like Lego building blocks, making it easy for operations and maintenance personnel to complete the construction.
[0027] Computing-network integration: Cloud-network synergy is the integration of cloud and network services. Cloud and network were once relatively independent, providing the integrated provision of cloud computing and network services, not a unified network architecture. Computing-network integration is cloud-network convergence 2.0, breaking through network bottlenecks in computing and the tidal effect of computing power. Combining technologies such as 5G, MEC, and AI, the network serves computing, while the increase in computing power is also transforming the network, creating a deep fusion of the two.
[0028] Simple Network Management Protocol (SNMP): SNMP is a standard application-layer protocol designed specifically for managing network nodes (servers, workstations, routers, switches, and hubs) on IP networks. SNMP enables network administrators to manage network performance, identify and resolve network problems, and plan network growth. By receiving random messages (and event reports) through SNMP, network management systems are notified of network problems.
[0029] DPI: DPI (Deep Packet Inspection) is a data packet-based deep inspection technology that performs deep inspection on different network application layer payloads (such as HTTP, DNS, etc.) and determines the legitimacy of the message by inspecting the payload.
[0030] The DPI system is mainly responsible for parsing binary network transmission data into visible messages, then performing layer-by-layer feature analysis on massive amounts of messages, and finally presenting them visually to the operator's network management and operation service units in the form of software, to help operators perform more refined network traffic management and other related services.
[0031] like Figure 1 As shown in the figure, the traditional 5G to C (user-oriented 5G technology) network operation and maintenance architecture mainly includes three layers: dedicated hardware infrastructure layer, network equipment layer and professional operation and maintenance platform layer (i.e. Figure 1The unified operation and maintenance center in the network) includes the professional operation and maintenance platform layer, which includes: wireless workstation, core network workstation, transmission workstation, wireless operation and maintenance center (OMC, Operation and Maintenance Center), core network OMC and transmission OMC. Correspondingly, the network equipment layer includes: wireless network, core network and transmission network; the dedicated hardware infrastructure layer includes: wireless network dedicated hardware, core network dedicated hardware and transmission network dedicated hardware. In the traditional 5G to C (user-oriented 5G technology) large-scale network operation and maintenance architecture, since the network equipment involved in the wireless network, transmission network and core network all have dedicated hardware to ensure performance and are deployed completely independently from the upper-layer applications, network operation and maintenance only focuses on communication network indicators. However, in the case of 5G private networks, which are carried by unified infrastructure, the impact of network quality is not only on the communication function, but also on the infrastructure; secondly, in addition to carrying the private network communication function, the infrastructure also needs to carry the private network's industry applications, and the dynamic resource guarantee of the infrastructure also needs to be tested.
[0032] In addition, the following issues will arise when this operation and maintenance system is applied in 5G to B private networks: The IT infrastructure will have resource scheduling issues: Because the private network infrastructure carries both 5G industry applications and 5G communication functions, it requires a unified resource orchestrator that integrates computing and networking, which may cause conflicts or other failures. IT infrastructure issues may affect communication indicators: The original communication indicators tested (such as registration success rate and session establishment success rate) are only related to changes in communication service logic or terminal environment. The infrastructure is dedicated and will not affect the indicators. However, since the private network may use a general-purpose hardware platform for carrying the network, failures of the general-purpose hardware platform will also affect communication indicators.
[0033] To solve the above problems, the embodiments of the present application provide corresponding solutions, namely: when a fault occurs, the cause of the fault is determined based on the target detection indicators of the target application in the ICT operation platform, and the operating status of multiple infrastructure services in the ICT operation platform, thereby realizing the detection of the infrastructure, which is described in detail below.
[0034] Example 1
[0035] Figure 2 This is a schematic diagram of an ICT unified operation and maintenance architecture for industry-specific networks according to an embodiment of the present application. Figure 2 As shown in the figure, the operation and maintenance architecture mainly includes three layers: general infrastructure, service functions and professional operation and maintenance platform (i.e. Figure 2The unified operations and maintenance center (OMC) within the platform includes CT infrastructure services (radio operations OMC, core network OMC, and transmission OMC) and IT infrastructure monitoring. Service functions encompass wireless network functions, core network functions, and industry applications, while general infrastructure includes computing, storage, and networking services.
[0036] Depend on Figure 2 As can be seen, 5G communication functions (including 5G RAN and 5GC) and industry applications are uniformly hosted on a common hardware platform, each managed by a dedicated O&M system. Compared to traditional large-scale networks, common infrastructure is also monitored and managed. This allows the upper-level ICT converged O&M center to centrally monitor metrics and manage O&M. If a communication metric issue or failure occurs, a coordinated investigation across all specialized O&M systems will be conducted, locating any communication metric issues that may be caused by infrastructure failures or insufficient resources.
[0037] The interaction processes of the above layers are as follows: Figure 3 As shown, it can be roughly divided into the following parts: fault discovery, fault location, fault isolation, and fault recovery, among which:
[0038] Fault detection: Through multiple data collection methods and based on specific identifiers, a unified log is generated, and faults are screened according to different rules for IT and CT. For example, IT equipment uses Syslog SNMP, while CT equipment uses the proprietary DPI protocol.
[0039] Syslog, often referred to as system log or system logging, is a standard for transmitting log messages over Internet Protocol (TCP / IP) networks. The term often refers to the actual syslog protocol, or the applications or databases that submit syslog messages. The syslog protocol is a client-server protocol: a syslog sender sends a small text message (less than 1024 bytes) to a syslog receiver. The receiver is typically named "syslogd," "syslogdaemon," or "syslog server." Syslog messages can be sent using the UDP protocol and / or the TCP protocol. This data is sent in clear text. However, since SSL encryption wrappers (such as Stunnel, sslio, or sslwrap) are not part of the syslog protocol itself, they can be used to provide a layer of encryption via SSL / TLS.
[0040] Fault location: Under the unified timestamp log, IT faults and CT faults are analyzed in conjunction.
[0041] Fault isolation: All CT service states are persistently stored and can be isolated at any time.
[0042] Fault recovery: IT water level detection, on-demand capacity expansion, and CT elastic expansion are guaranteed.
[0043] Based on the above principles, an embodiment of the present application provides a method embodiment of a fault determination method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0044] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 4 FIG1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a fault determination method. Figure 4 As shown, the computer terminal 40 (or mobile device 40) may include one or more (402a, 402b, ..., 402n are used to illustrate) processors 402 (the processor 402 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 404 for storing data, and a transmission module 406 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 4 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 4 More or fewer components than shown, or with Figure 4 Different configurations shown.
[0045] It should be noted that the one or more processors 402 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 40 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0046] The memory 404 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method in the embodiment of the present application. The processor 402 executes various functional applications and data processing by running the software programs and modules stored in the memory 404, that is, implementing the vulnerability detection method of the above-mentioned application. The memory 404 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 404 may further include a memory remotely located relative to the processor 402, and these remote memories may be connected to the computer terminal 40 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0047] The transmission module 406 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 40. In one embodiment, the transmission module 406 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 406 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0048] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 40 (or mobile device).
[0049] In the above operating environment, if Figure 5 As shown, the fault determination method provided in the embodiment of the present application includes the following processing steps:
[0050] Step S502: Obtain target detection indicators of target applications provided in the ICT operation platform, wherein the target detection indicators are used to reflect the operation status of the target application. The ICT operation platform includes a variety of infrastructure services, and the multiple infrastructure services include at least: CT infrastructure services and IT infrastructure services; wherein the above-mentioned target applications include but are not limited to various types of industry applications.
[0051] ICT converged devices refer to the integration of IT (Information Technology) equipment and CT (Communication Technology) equipment. IT equipment can be devices such as servers, while CT equipment can be routers and switches. ICT converged devices can use the same hardware architecture as CT devices, which includes an Ethernet switching chip and a CPU. The Ethernet switching chip implements Layer 2 switching functions. The CPU runs software to form a router operating system, which implements Layer 3 switching functions such as routing. Furthermore, ICT converged devices run software on the CPU to create virtual applications based on virtual machines, implementing the functions of IT devices.
[0052] There are various ways to obtain target detection indicators. For example, a request message can be sent to the user side, and the user device can respond to the request message and feedback the target detection indicators of the target application to the ICT operation platform. Alternatively, target detection indicators can be received from the user side on a regular basis. These target detection indicators include, but are not limited to, communication quality parameters such as the application's data transmission rate and bit error rate.
[0053] Step S504, determining whether a target application fails based on the target detection indicator;
[0054] In some embodiments, this step can be implemented by comparing the target detection index with a preset threshold; and determining whether a fault has occurred based on the comparison result, wherein if the comparison result indicates that the target detection index is less than the preset threshold, the target application is determined to have failed; and if the comparison result indicates that the target detection index is greater than the preset threshold, the target application is determined to have failed. For example, if the data transmission rate of the target application is less than the rate threshold, the target application is determined to have failed.
[0055] In some embodiments of the present application, it is also possible to Figure 9 The fault determination method shown is to determine the fault, such as Figure 9 As shown, the fault determination method includes:
[0056] Step S902: Obtain target detection indicators of target applications provided in the ICT operation platform, wherein the target detection indicators are used to reflect the operation status of the target application. The ICT operation platform includes a variety of infrastructure services, and the multiple infrastructure services include at least: CT infrastructure services and IT infrastructure services; wherein the above-mentioned target applications include but are not limited to various types of industry applications.
[0057] Step S904, determining whether a target application fails based on the target detection indicator;
[0058] It should be noted that the above steps S902 and S904 are Figure 5 There is a one-to-one correspondence between steps S502 and S904, so the explanations of steps S502 and S504 also apply to steps S902 and S904, and are not repeated here.
[0059] Step S906 , when a target application fails, obtaining the operating status of multiple infrastructure services in the ICT operating platform;
[0060] In some embodiments, operation logs of multiple infrastructure services are obtained, wherein different infrastructure services use different log collection methods, that is, the log collection methods used by different infrastructure services can be independent of each other, for example, IT equipment uses syslog snmp to collect logs, and CT equipment uses DPI proprietary protocol to collect logs; the operation status of multiple infrastructure services is determined based on the operation logs.
[0061] Among them, the above-mentioned operation log can be determined in the following way: determine the time information when the target application fails and the log identifier of the target application, wherein the log identifier is used to identify the logs generated by the IT infrastructure services and CT infrastructure services associated with the target application; determine the timestamp corresponding to the time information, and determine the log set corresponding to the timestamp; determine the log corresponding to the log identifier from the log set based on the log identifier, and use the log corresponding to the log identifier as the operation log.
[0062] Among them, the log identifier can be jointly determined based on the first identifier of the IT infrastructure service associated with the target application and the second identifier of the CT infrastructure service. There are many specific determination methods. For example, the first identifier and the second identifier can be combined to form a log identifier; or the first identifier and the second identifier can be hashed and the result of the hash operation can be determined as the log identifier.
[0063] Step S908: determining the cause of the target application failure based on the operating status of the various infrastructure services.
[0064] Specifically, this step can be implemented as follows: evaluating the operating status of multiple infrastructure services to obtain evaluation indicators for each infrastructure service, where the evaluation indicators are used to evaluate the operating status of each infrastructure service; determining a target infrastructure service from the multiple infrastructure services based on the evaluation indicators, and determining that the fault is caused by the target infrastructure service. When evaluating the operating status of multiple infrastructure services, it is necessary to determine an evaluation method corresponding to each infrastructure service, i.e., different infrastructure services have different evaluation methods; and evaluating the multiple infrastructure services using the evaluation method corresponding to each infrastructure service. Evaluation indicators include, but are not limited to, communication rate, bit error rate, etc.
[0065] In some embodiments of the present application, users can select appropriate evaluation indicators according to their needs, such as Figure 8 shown. Figure 8 is an interactive interface according to an embodiment of the present application, wherein: Figure 8 The upper left corner of the screen displays the private network topology. During normal operation, the topology can be green or another user-specified color. When a fault occurs, the color of the node corresponding to the faulty device changes after the cause is determined. The color also varies depending on the severity of the fault. Figure 8 The upper right corner shows a variety of preset evaluation criteria. Users can directly select at least one evaluation criterion from the preset evaluation criteria based on their own needs to evaluate the infrastructure services of the private network; the lower right corner shows the evaluation indicators of the infrastructure services of the private network based on the selected evaluation criteria. Figure 8 The lower left corner of the display shows the fault summary, such as Figure 8 As shown, all faults are divided into IT faults and CT faults according to the fault category, and important information such as the fault cause, fault level (severity), and duration of each fault are displayed to facilitate users to operate and maintain the private network.
[0066] In other embodiments of the present application, after determining the cause of the target application failure based on the operating status of multiple infrastructure services, the failure can be isolated, for example:
[0067] If the fault is caused by a CT infrastructure service, operating status indication information of the CT infrastructure service is stored during a first time period, where the operating status indication information indicates that the CT infrastructure service is unavailable. If the fault is caused by an IT infrastructure service, operating status indication information of the IT infrastructure service is stored during a second time period, where the operating status indication information indicates that the IT infrastructure service is unavailable. Since CT infrastructure service failures often have a significant impact and are complex to detect, it is necessary to permanently isolate the failed CT infrastructure service before recovery. However, IT infrastructure failures often involve a specific program segment or script, making them easier to detect and therefore readily available for isolation. Therefore, the duration of the first time period needs to be greater than the duration of the second time period.
[0068] In some embodiments, a target application failure may be caused by insufficient capacity of an infrastructure service. After determining the cause of the target application failure based on the operating status of multiple infrastructure services, a prompt message is generated to prompt capacity expansion if the cause is due to insufficient capacity of multiple infrastructure services. Both IT infrastructure services and CT infrastructure services may experience target application failures due to insufficient capacity. This capacity includes, but is not limited to, the remaining operating resources of the infrastructure corresponding to each infrastructure service, such as remaining memory space.
[0069] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0070] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0071] Example 2
[0072] According to the embodiment of the present application, there is also provided a Figure 6 The troubleshooting method shown includes:
[0073] Step S602: When at least one fault occurs in the ICT operation platform, current operation status information of the ICT operation platform is collected and operation log information is generated;
[0074] In some embodiments, operation logs of multiple infrastructure services are obtained, wherein different infrastructure services use different log collection methods, that is, the log collection methods used by different infrastructure services can be independent of each other, for example, IT equipment uses syslog snmp to collect logs, and CT equipment uses DPI proprietary protocol to collect logs; the operation status of multiple infrastructure services is determined based on the operation logs.
[0075] Among them, the above-mentioned operation log can be determined in the following way: determine the time information when the target application fails and the log identifier of the target application, wherein the log identifier is used to identify the logs generated by the IT infrastructure services and CT infrastructure services associated with the target application; determine the timestamp corresponding to the time information, and determine the log set corresponding to the timestamp; determine the log corresponding to the log identifier from the log set based on the log identifier, and use the log corresponding to the log identifier as the operation log.
[0076] Among them, the log identifier can be jointly determined based on the first identifier of the IT infrastructure service associated with the target application and the second identifier of the CT infrastructure service. There are many specific determination methods. For example, the first identifier and the second identifier can be combined to form a log identifier; or the first identifier and the second identifier can be hashed and the result of the hash operation can be determined as the log identifier.
[0077] Step S604: determining the fault type of each of the at least one fault based on the operation log information, and locating each fault, wherein the fault type includes an IT fault and a CT fault;
[0078] Specifically, this step can be implemented as follows: evaluating the operating status of multiple infrastructure services to obtain evaluation indicators for each infrastructure service, where the evaluation indicators are used to evaluate the operating status of each infrastructure service; determining a target infrastructure service from the multiple infrastructure services based on the evaluation indicators, and determining that the fault is caused by the target infrastructure service. When evaluating the operating status of multiple infrastructure services, it is necessary to determine an evaluation method corresponding to each infrastructure service, i.e., different infrastructure services have different evaluation methods; and evaluating the multiple infrastructure services using the evaluation method corresponding to each infrastructure service. Evaluation indicators include, but are not limited to, communication rate, bit error rate, etc.
[0079] Step S606: Isolate the infrastructure service corresponding to each fault using a fault isolation strategy corresponding to the fault type of each fault.
[0080] In other embodiments of the present application, after locating the fault, the fault may be isolated, for example:
[0081] If the fault is caused by a CT infrastructure service, operating status indication information of the CT infrastructure service is stored during a first time period, where the operating status indication information indicates that the CT infrastructure service is unavailable. If the fault is caused by an IT infrastructure service, operating status indication information of the IT infrastructure service is stored during a second time period, where the operating status indication information indicates that the IT infrastructure service is unavailable. Since CT infrastructure service failures often have a significant impact and are complex to detect, it is necessary to permanently isolate the failed CT infrastructure service before recovery. However, IT infrastructure failures often involve a specific program segment or script, making them easier to detect and therefore readily available for isolation. Therefore, the duration of the first time period needs to be greater than the duration of the second time period.
[0082] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0083] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0084] Example 3
[0085] According to the embodiment of the present application, there is also provided a Figure 7 The fault determination device shown includes: a first acquisition module 70, which is used to obtain target detection indicators of the target application provided in the ICT operation platform, wherein the target detection indicators are used to reflect the operating status of the target application, and the ICT operation platform includes multiple infrastructure services, and the multiple infrastructure services include at least: CT infrastructure services and IT infrastructure services; a first determination module 72, which is used to determine whether the target application has a fault based on the target detection indicators; a second acquisition module 74, which is used to obtain the operating status of multiple infrastructure services in the ICT operation platform when the target application has a fault; a second determination module 76, which is used to determine the cause of the fault of the target application based on the operating status of the multiple infrastructure services.
[0086] In some embodiments of the present application, the first acquisition module 70 may acquire the target detection indicator in various ways. For example, the first acquisition module 70 may send a request message to the user side, and the user device may respond to the request message and provide the target detection indicator of the target application to the ICT operation platform. Alternatively, the first acquisition module 70 may periodically receive the target detection indicator from the user side. The target detection indicator may include, but is not limited to, communication quality parameters such as the application's data transmission rate and the application's bit error rate.
[0087] In some embodiments of the present application, the first determination module 72 may determine whether a target application has failed by comparing a target detection indicator with a preset threshold; and determining whether a failure has occurred based on the comparison result, wherein if the comparison result indicates that the target detection indicator is less than the preset threshold, the target application is determined to have failed; and if the comparison result indicates that the target detection indicator is greater than the preset threshold, the target application is determined to have failed. For example, if the data transmission rate of the target application is less than the rate threshold, the target application is determined to have failed.
[0088] In some embodiments, the second acquisition module 74 can obtain the operation logs of multiple infrastructure services, wherein different infrastructure services use different log collection methods, that is, the log collection methods used by different infrastructure services can be independent of each other, for example, IT equipment uses syslog snmp to collect logs, and CT equipment uses DPI proprietary protocol to collect logs; the operation status of multiple infrastructure services is determined based on the operation logs.
[0089] Among them, the above-mentioned operation log can be determined in the following way: determine the time information when the target application fails and the log identifier of the target application, wherein the log identifier is used to identify the logs generated by the IT infrastructure services and CT infrastructure services associated with the target application; determine the timestamp corresponding to the time information, and determine the log set corresponding to the timestamp; determine the log corresponding to the log identifier from the log set based on the log identifier, and use the log corresponding to the log identifier as the operation log.
[0090] Among them, the log identifier can be jointly determined based on the first identifier of the IT infrastructure service associated with the target application and the second identifier of the CT infrastructure service. There are many specific determination methods. For example, the first identifier and the second identifier can be combined to form a log identifier; or the first identifier and the second identifier can be hashed and the result of the hash operation can be determined as the log identifier.
[0091] In some embodiments of the present application, the second determination module 76 can be implemented in the following manner: evaluating the operating status of multiple infrastructure services to obtain evaluation indicators for each infrastructure service, wherein the evaluation indicators are used to evaluate the operating status of each infrastructure service; determining a target infrastructure service from the multiple infrastructure services based on the evaluation indicators, and determining that the fault cause is caused by the target infrastructure service. When evaluating the operating status of multiple infrastructure services, it is necessary to determine the evaluation method corresponding to each infrastructure service within the multiple infrastructure services, i.e., different infrastructure services have different evaluation methods; and evaluating the multiple infrastructure services using the evaluation method corresponding to each infrastructure service. Evaluation indicators include, but are not limited to, communication rate, bit error rate, etc.
[0092] Example 3
[0093] The embodiment of the present application can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.
[0094] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.
[0095] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the vulnerability detection method of the application: obtaining the target detection indicators of the target application provided in the ICT operation platform, wherein the target detection indicators are used to reflect the operating status of the target application, and the ICT operation platform includes multiple infrastructure services, and the multiple infrastructure services include at least: CT infrastructure services and IT infrastructure services; determining whether the target application fails based on the target detection indicators; in the case that the target application fails, obtaining the operating status of the multiple infrastructure services in the ICT operation platform; and determining the cause of the failure of the target application based on the operating status of the multiple infrastructure services.
[0096] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the fault determination method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned fault determination method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories can be connected to the terminal A via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0097] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: obtaining target detection indicators of the target application provided in the ICT operation platform, wherein the target detection indicators are used to reflect the operating status of the target application, and the ICT operation platform includes multiple infrastructure services, and the multiple infrastructure services include at least: CT infrastructure services and IT infrastructure services; determining whether the target application fails based on the target detection indicators; in the event that the target application fails, obtaining the operating status of the multiple infrastructure services in the ICT operation platform; and determining the cause of the failure of the target application based on the operating status of the multiple infrastructure services.
[0098] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0099] Example 4
[0100] The embodiment of the present application further provides a non-volatile storage medium. Optionally, in this embodiment, the non-volatile storage medium can be used to store program codes executed by the fault determination method provided in the embodiment.
[0101] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0102] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: obtaining target detection indicators of a target application provided in an ICT operating platform, wherein the target detection indicators are used to reflect the operating status of the target application, and the ICT operating platform includes a variety of infrastructure services, and the multiple infrastructure services include at least: CT infrastructure services and IT infrastructure services; determining whether a target application fails based on the target detection indicators; in the event that a target application fails, obtaining the operating status of the multiple infrastructure services in the ICT operating platform; and determining the cause of the failure of the target application based on the operating status of the multiple infrastructure services.
[0103] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0104] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0105] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0106] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0107] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0108] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0109] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A fault determination method, wherein: include: Obtaining a target detection indicator of a target application provided in an ICT operation platform, wherein the target detection indicator is used to reflect an operation status of the target application, wherein the ICT operation platform includes multiple infrastructure services, and the multiple infrastructure services include at least: a CT infrastructure service and an IT infrastructure service; Determining whether a failure occurs in the target application based on the target detection indicator; In the event that the target application fails, obtaining the operating status of the multiple infrastructure services in the ICT operating platform includes: Obtain operation logs for the various infrastructure services, including: Determining time information when the target application fails and a log identifier of the target application, wherein the log identifier is used to identify logs generated by IT infrastructure services and CT infrastructure services associated with the target application; determining the operation log based on the time information and the log identifier of the target application; determining the operating status of the plurality of infrastructure services based on the operating log; The cause of the failure of the target application is determined based on the operating status of the multiple infrastructure services.
2. The method according to claim 1, wherein Different infrastructure services use different log collection methods.
3. The method according to claim 1, wherein The operation logs include logs generated by IT infrastructure services and CT infrastructure services.
4. The method according to claim 1, wherein Determine the operation log, including: Determining a timestamp corresponding to the time information, and determining a log set corresponding to the timestamp; The log corresponding to the log identifier is determined from the log set according to the log identifier, and the log corresponding to the log identifier is used as the running log.
5. The method according to claim 1, wherein Determining the cause of the target application failure based on the operating status of the multiple infrastructure services includes: Evaluating the operating status of the multiple infrastructure services to obtain an evaluation index for each infrastructure service, wherein the evaluation index is used to evaluate the operating status of each infrastructure service; A target infrastructure service among the multiple infrastructure services is determined based on the evaluation index, and the cause of the fault is determined to be a fault caused by the target infrastructure service.
6. The method according to claim 5, wherein: After determining that the fault cause is a fault caused by the target infrastructure service, the method further includes: The classification shows the target object the cause of the fault, fault duration, and fault level. The fault level is used to characterize the severity of the fault, and faults of different fault levels are marked with different labels.
7. The method according to claim 5, wherein: Evaluate the operational status of the various infrastructure services, including: determining an evaluation method corresponding to each of the plurality of infrastructure services; The plurality of infrastructure services are evaluated respectively using evaluation methods corresponding to the respective infrastructure services.
8. The method according to claim 6, wherein: Determining an evaluation method corresponding to each of the multiple infrastructure services, including: Display multiple preset evaluation methods in the interactive interface; In response to a selection instruction received by the target object through the interactive interface, an evaluation method corresponding to each infrastructure service is selected from the multiple evaluation methods.
9. The method according to claim 7, wherein: The method further includes: displaying evaluation results of the multiple infrastructure services in an interactive interface, and categorizing and displaying the fault causes, fault durations, and fault levels of the faults, wherein the fault levels are used to characterize the severity of the faults, and faults of different fault levels are marked with different labels.
10. The method according to claim 1, wherein After determining the cause of the failure of the target application based on the operating status of the multiple infrastructure services, the method further includes: In a case where the fault cause is a fault caused by a CT infrastructure service, storing operation status indication information of the CT infrastructure service within a first time period, wherein the operation status indication information is used to indicate that the CT infrastructure service is in an unavailable state; If the fault is caused by an IT infrastructure service, storing operation status indication information of the IT infrastructure service within a second time period, wherein the operation status indication information is used to indicate that the IT infrastructure service is in an unavailable state; The duration corresponding to the first time period is greater than the duration corresponding to the second time period.
11. The method according to any one of claims 1 to 9, wherein: After determining the cause of the failure of the target application based on the operating status of the multiple infrastructure services, the method further includes: When the fault is caused by insufficient capacity of the multiple infrastructure services, prompt information is generated to prompt capacity expansion.
12. A fault determination method, wherein: include: Obtaining a target detection indicator of a target application provided in an ICT operation platform, wherein the target detection indicator is used to reflect an operation status of the target application, wherein the ICT operation platform includes multiple infrastructure services, and the multiple infrastructure services include at least: a CT infrastructure service and an IT infrastructure service; Determining whether a failure occurs in the target application based on the target detection indicator; In the event that the target application fails, obtaining the operating status of the multiple infrastructure services in the ICT operating platform includes: Obtain operation logs for the various infrastructure services, including: Determining time information when the target application fails and a log identifier of the target application, wherein the log identifier is used to identify logs generated by IT infrastructure services and CT infrastructure services associated with the target application; determining the operation log based on the time information and the log identifier of the target application; determining the operating status of the plurality of infrastructure services based on the operating log; The cause of the failure of the target application is determined based on the operating status of the multiple infrastructure services.
13. A fault handling method, wherein: include: When at least one fault occurs in the ICT operation platform, current operation status information of the ICT operation platform is collected and operation log information is generated, including: Obtain operation logs for various infrastructure services, including: Determining time information when a target application fails and a log identifier of the target application, wherein the log identifier is used to identify logs generated by IT infrastructure services and CT infrastructure services associated with the target application; determining the operation log based on the time information and the log identifier of the target application; Determining current operating status information in the ICT operating platform based on the operating log; Determine a fault type of each of the at least one fault based on the operation log information, and locate each of the faults, wherein the fault type includes an IT fault and a CT fault; A fault isolation strategy corresponding to the fault type of each fault is adopted to isolate the infrastructure service corresponding to each fault.
14. A fault determination device, wherein: include: a first acquisition module, configured to acquire a target detection indicator of a target application provided in an ICT operation platform, wherein the target detection indicator is used to reflect an operation status of the target application, wherein the ICT operation platform includes multiple infrastructure services, and the multiple infrastructure services include at least a CT infrastructure service and an IT infrastructure service; A first determination module is configured to determine whether a fault occurs in the target application based on the target detection indicator; The second acquisition module is configured to acquire the operating status of the multiple infrastructure services in the ICT operating platform when the target application fails, including: Obtain operation logs for the various infrastructure services, including: Determining time information when the target application fails and a log identifier of the target application, wherein the log identifier is used to identify logs generated by IT infrastructure services and CT infrastructure services associated with the target application; determining the operation log based on the time information and the log identifier of the target application; determining the operating status of the plurality of infrastructure services based on the operating log; The second determining module is configured to determine a cause of the failure of the target application according to the operating status of the multiple infrastructure services.
15. A non-volatile storage medium, wherein: The non-volatile storage medium includes a stored program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute the fault determination method according to any one of claims 1 to 7.
16. A computer terminal, wherein: include: processor; as well as A memory, connected to the processor, configured to provide the processor with instructions for processing the following processing steps: Obtaining a target detection indicator of a target application provided in an ICT operation platform, wherein the target detection indicator is used to reflect an operation status of the target application, wherein the ICT operation platform includes multiple infrastructure services, and the multiple infrastructure services include at least: a CT infrastructure service and an IT infrastructure service; Determining whether a failure occurs in the target application based on the target detection indicator; In the event that the target application fails, obtaining the operating status of the multiple infrastructure services in the ICT operating platform includes: Obtain operation logs for the various infrastructure services, including: Determining time information when the target application fails and a log identifier of the target application, wherein the log identifier is used to identify logs generated by IT infrastructure services and CT infrastructure services associated with the target application; determining the operation log based on the time information and the log identifier of the target application; determining the operating status of the plurality of infrastructure services based on the operating log; The cause of the failure of the target application is determined based on the operating status of the multiple infrastructure services.
Citation Information
Patent Citations
Abnormity detection and attribution method and device, equipment and computer readable storage medium
CN111901171A