Communication fault detection method, apparatus and device, and computer readable storage medium
By performing transmission layer and application layer detection on the network communication between power equipment and control equipment in the power system, and using execution parameters to quickly determine the cause of the fault, the problem of low efficiency in power system communication fault detection in the existing technology is solved, and rapid and accurate fault location and cause analysis are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN KANGBIDA CONTROL TECH
- Filing Date
- 2025-12-02
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, the methods for detecting communication faults in power systems are inefficient and make it difficult to determine the cause of the fault in real time. This results in anomalies or attacks in the network not being contained in a timely manner, affecting system stability.
By detecting the transmission and application layers of network communication between power equipment and control equipment, the network level of the fault can be quickly determined using execution parameters, and the cause of the fault can be determined based on the triggering of the control strategy.
It enables rapid and accurate identification of communication failures, improves the efficiency and accuracy of fault detection, reduces the time spent troubleshooting one by one, and meets the power system's needs for real-time sensing and rapid response.
Smart Images

Figure CN121887608A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a communication fault detection method, apparatus, device, and computer-readable storage medium. Background Technology
[0002] As power systems rapidly develop towards intelligence and networking, their operation increasingly relies on stable and reliable dedicated communication networks. All aspects of power production, transmission, distribution, and dispatch rely on networks for data exchange and command control; any delay, interruption, or abnormal loss of data packets in any network link can directly affect the stable operation of the power system.
[0003] Among related technologies, analyzing firewall logs of power systems is a method that has poor timeliness and makes it difficult to determine the cause of communication disruptions. Summary of the Invention
[0004] This application provides a communication fault detection method, apparatus, device, and computer-readable storage medium that can quickly determine the cause of communication faults.
[0005] The technical solution of this application embodiment is implemented as follows: This application provides a communication fault detection method, the method comprising: In response to the disruption of network communication between power equipment and control equipment, the transport layer of the network is detected based on the execution parameters of the network communication to obtain a first detection result; In response to the first detection result showing no abnormality, the application layer of the network is detected based on the execution parameters of the network between the power equipment and the control equipment to obtain a second detection result; Based on the first detection result and the second detection result, the network layer where the fault occurred is determined, wherein each network layer is configured with a corresponding control strategy; Based on the execution parameters, the triggering status of the control policy corresponding to the network layer where the fault occurred is determined, and the cause of the communication fault is determined based on the triggering status.
[0006] This application provides a communication fault detection device, including: The first detection module is used to detect the transport layer of the network based on the execution parameters of the network communication in response to the network communication being blocked, and to obtain a first detection result. The second detection module is used to detect the application layer of the network based on the execution parameters of the network between the power equipment and the control equipment in response to the first detection result being normal, and to obtain a second detection result. The first judgment module is used to determine the network layer where the fault occurred based on the first detection result and the second detection result, wherein each network layer is configured with a corresponding control strategy; The second judgment module is used to determine the triggering status of the control policy corresponding to the network layer where the fault occurred based on the execution parameters, and to determine the cause of the communication fault based on the triggering status.
[0007] This application provides an electronic device, the electronic device comprising: Memory is used to store executable instructions or computer programs. When a processor executes computer-executable instructions or computer programs stored in the memory, it implements the communication fault detection method provided in the embodiments of this application.
[0008] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the communication fault detection method provided in this application.
[0009] This application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, they implement the communication fault detection method provided in this application. The embodiments of this application have the following beneficial effects: By analyzing the execution parameters of network communication, the transport and application layers of the network are sequentially examined to initially determine the network layer (transport, application, or session layer) responsible for the communication failure. Based on the triggering status of the control policies corresponding to the identified network layer, the cause of the communication failure is determined. This approach allows for accurate identification of the cause of communication failures without requiring a step-by-step investigation of different faults. Furthermore, compared to methods that rely on analyzing network logs, analyzing execution parameters allows for determining the cause of the failure based on the real-time status of these parameters, thus accelerating the analysis of communication failure causes. Attached Figure Description
[0010] Figure 1 This is a schematic diagram of the architecture of the communication fault detection system provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the communication fault detection method provided in the embodiments of this application. Figure 1 ; Figure 4 This is a flowchart illustrating the communication fault detection method provided in the embodiments of this application. Figure 2 ; Figure 5 This is a flowchart illustrating the communication fault detection method provided in the embodiments of this application. Figure 3 ; Figure 6 This is a flowchart illustrating the communication fault detection method provided in the embodiments of this application. Figure 4 ; Figure 7 This is a flowchart illustrating the communication fault detection method provided in the embodiments of this application. Figure 5 ; Figure 8 This is a flowchart illustrating the communication fault detection method provided in the embodiments of this application. Figure 6 ; Figure 9 This is a flowchart illustrating the communication fault detection method provided in the embodiments of this application. Figure 7 ; Figure 10 This is a flowchart illustrating a communication fault detection method in an application scenario provided in an embodiment of this application. Figure 11 This is a schematic diagram of the structure for setting up network communication in a power system, provided in an embodiment of this application. Figure 12 This is a logical block diagram of a communication fault detection method in an application scenario provided in the embodiments of this application; Figure 13 This is a logical block diagram of the method for confirming a faulty network layer provided in the embodiments of this application; Figure 14 This is a logic block diagram provided in an embodiment of the present application for fault location after determining that a fault has occurred in the transport layer; Figure 15 This is a logic block diagram provided in an embodiment of the present application for determining fault location after an application layer fault occurs; Figure 16 This is a logic block diagram provided in the embodiments of this application for fault location after determining that a session layer fault has occurred. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0012] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0013] In the embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0014] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for descriptive purposes only and is not intended to limit the scope of this application.
[0015] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0016] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0017] 1) Network communication: In power systems, network communication refers to a dedicated data exchange system built between power plants, substations, distribution networks, and control centers to achieve real-time detection, protection, and control of power production, transmission, distribution, and consumption. It is responsible for transmitting various status information and control commands necessary to ensure the safe, stable, and efficient operation of the power grid. Specifically, its core task is to support critical services such as power grid dispatch automation, relay protection, fault recording, and equipment status monitoring, requiring extremely high real-time performance, reliability, and security. Its key difference from ordinary commercial network communication lies in the fact that its communication protocols, network equipment, and channel designs are typically optimized and enhanced for the specific needs of the power industry (such as strong electromagnetic interference environments and millisecond-level response). This communication system is the foundation for realizing smart grids and distribution automation, enabling wide-area detection and precise control from the master station to end users.
[0018] 2) Network layers refer to the structure of network protocols divided into multiple layers. Each layer is responsible for handling specific functions in the communication process, from the underlying physical transmission to the higher-level application interaction, to achieve efficient and reliable data exchange. Specifically, each layer focuses on one aspect of communication, such as signal transmission, data encapsulation, or routing control, and interacts with adjacent layers through standard interfaces, with lower layers providing services to upper layers. This layered design supports modular development, facilitates independent protocol updates and maintenance, and ensures interoperability between different devices or systems. For example, in the OSI reference model, there are the Physical Layer (handles electrical signals and media), Data Link Layer (manages frame transmission and error detection), Network Layer (responsible for addressing and routing), Transport Layer (ensuring end-to-end connections), Session Layer (controlling session establishment), Presentation Layer (handling data format conversion), and Application Layer (providing user services); while in the TCP / IP model, it is simplified to four main layers: Network Interface Layer, Internet Layer, Transport Layer, and Application Layer. When a network communication failure occurs, the network layers that are likely to fail are the Transport Layer, Application Layer, and Session Layer.
[0019] 3) The transport layer is a network layer responsible for providing end-to-end, reliable or unreliable data transmission services for applications running on different hosts. It acts as a bridge between applications and the underlying network infrastructure, ensuring that data is accurately delivered from a process on the source host to the corresponding process on the destination host. The transport layer uses port numbers to identify and address specific applications on a host, thereby enabling the multiplexing and demultiplexing of data from multiple applications. To achieve reliable communication, protocols at this layer (such as TCP) provide a series of safeguards, including packet ordering, error retransmission, flow control, and congestion control. For applications that do not require reliable guarantees, this layer also provides lightweight, best-effort delivery services (such as UDP) to reduce latency.
[0020] 4) The application layer, the highest layer in the communication network hierarchy model, directly provides network services and interfaces to user applications or processes. It is the final carrier for implementing specific network application functions (such as web browsing and email sending / receiving). Utilizing the data transmission capabilities provided by lower layers (such as the transport layer), it focuses on processing the application's own business logic and data content. The application layer defines various application protocols (such as HTTP, SMTP, and DNS), specifying the rules, message formats, and interaction sequences for communication between applications on different hosts. It translates user operations or software requirements into standard, network-transmittable service requests and converts received network data into user-understandable information. This layer is responsible for handling tasks directly related to the application, such as authentication and the representation and conversion of data semantics.
[0021] 5) The Session Layer, located above the Transport Layer in the communication network hierarchy model, is responsible for establishing, managing, and terminating the "dialogue," or session, between two communicating application processes. It ensures that data exchange between processes on different hosts is orderly and synchronous by coordinating the interaction logic. Specifically, the Session Layer is responsible for dialogue control, managing whether the communication mode is half-duplex (alternating send and receive) or full-duplex (simultaneous send and receive). It provides a mechanism for checking and resuming the dialogue by inserting synchronization points into the data stream. In case of transmission interruption, it can resume from the last mutually agreed synchronization point without having to start from scratch. This layer is also responsible for session authentication and maintenance, ensuring the legitimacy of both communicating parties and maintaining the session state during activity.
[0022] 6) Execution parameters refer to a set of specific values or identifiers pre-configured to complete a particular network operation or task. They define the execution target, scope, source, and behavior of the operation. These parameters, as input conditions for network function execution, directly determine the specific object, path, and pace of the operation. Examples include the target node IP address list, source node IP address, the URL path to be detected, and the detection time interval.
[0023] In related technologies, network fault location using firewall-based log analysis methods relies on collecting and parsing communication logs automatically recorded by network boundary security devices (i.e., firewalls). This method reverse-engineers and locates network faults such as communication interruptions or anomalies by retrospectively analyzing key fields in the logs, such as source / destination IP addresses, port numbers, protocol types, and communication statuses (e.g., "deny" or "allow").
[0024] However, the analysis process cannot be performed in real time, resulting in a delay window of several minutes between the occurrence of a fault and its detection. During this period, anomalies or attacks in the network cannot be contained in time, easily escalating from small-scale faults into systemic risks. Therefore, this method is insufficient to support the proactive safety operation and maintenance requirements of power systems for "real-time perception and rapid response." Furthermore, power systems require the deployment of various security protection software, such as encryption software and intranet management systems. There is a possibility that improper configuration or overly strict policies in these systems may mistakenly block legitimate TCP communication. In such cases, on-site personnel cannot pinpoint the cause of the communication disruption, and developers must investigate multiple potential problems one by one, which is extremely time-consuming.
[0025] Based on the problems existing in related technologies, embodiments of this application provide a communication fault detection method, apparatus, device, and computer-readable storage medium, which can locate communication faults in real time, improving the efficiency and accuracy of communication fault detection. The following describes exemplary applications of the communication fault detection device provided in this application embodiment, which is an electronic device used to implement the communication fault detection method. The electronic device provided in this application embodiment can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smartwatches, and vehicle terminals, or it can be implemented as a server. Exemplary applications when the device is implemented as a terminal or server will be described below.
[0026] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the communication fault detection system provided in this application embodiment. To perform communication fault detection operations, a communication fault detection application can be provided. For example, this application can be a dedicated application for communication fault detection, or it can be a functional management module in other applications (such as a communication fault detection module in a power system management application). The communication fault detection system 100 in this application embodiment includes at least a terminal 400, a network 300, a server 200, multiple power devices 500, and at least one control device 600. The server 200 is a server that manages multiple power devices. The server 200 can constitute the communication fault detection device of this application embodiment, that is, the communication fault detection method of this application embodiment is implemented through the server 200. The terminal 400 is connected to the server 200 through the network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.
[0027] See Figure 1 Users can send fault detection requests through the client of the communication fault detection application via terminal 400. After receiving the fault detection request, terminal 400 sends interactive information to server 200. In response to the interruption of network communication between the power equipment and the control equipment, server 200 performs network transport layer detection based on network communication execution parameters to obtain a first detection result. In response to the absence of anomalies in the first detection result, server 200 performs application layer detection based on the network execution parameters between the power equipment and the control equipment to obtain a second detection result. Based on the first and second detection results, server 200 determines the network layer where the fault occurred, where each network layer has a corresponding control policy. Based on the execution parameters, server 200 determines the triggering condition of the control policy corresponding to the faulty network layer, and determines the cause of the communication fault based on the triggering condition.
[0028] In some embodiments, the server 200 may also execute the communication fault detection method of this application embodiment. When the server detects network blockage, the server 200, in response to the blockage of network communication between the power equipment and the control equipment, detects the transport layer of the network based on the execution parameters of the network communication to obtain a first detection result. In response to the absence of abnormality in the first detection result, the server 200, based on the execution parameters of the network between the power equipment and the control equipment, detects the application layer of the network to obtain a second detection result. The server 200 determines the network layer where the fault occurred based on the first detection result and the second detection result, wherein each network layer is configured with a corresponding control strategy. The server 200, based on the execution parameters, determines the triggering condition of the control strategy corresponding to the network layer where the fault occurred, and determines the cause of the communication fault based on the triggering condition.
[0029] As an example, in a power communication network, when a communication fault detection system deployed on a server receives an alarm message indicating "severe lag in the video detection stream of a substation," it can obtain the execution parameters of the substation channel from the power grid business and dispatch management system based on this information, and construct a first state vector such as "TCP retransmission rate and connection latency status." The communication fault detection method in this embodiment determines the network layer where the communication fault occurred based on the execution parameters, and then determines the cause of the communication fault based on the triggering conditions of the control strategy corresponding to the determined network layer.
[0030] In some embodiments, the electronic device may be Figure 1 Server 200, see [link / reference] Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 2 The illustrated electronic device includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components of the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2 The general labeled all buses as Bus System 440.
[0031] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor.
[0032] User interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0033] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.
[0034] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0035] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0036] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc. The presentation module 453 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.). The input processing module 454 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.
[0037] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2A communication fault detection device 455 stored in memory 450 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: a first detection module 4551, a second detection module 4552, a first judgment module 4553, and a second judgment module 4554. These modules are logically connected and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0038] In other embodiments, the apparatus provided in this application can be implemented in hardware. As an example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the communication fault detection method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0039] See Figure 3 , Figure 3 This is a flowchart illustrating the communication fault detection method provided in the embodiments of this application. Figure 1 , will combine Figure 3 The steps shown are explained as follows: Figure 3 As shown, the method for detecting communication faults is illustrated using a server as the executing entity. The method includes the following steps 101 to 104.
[0040] In step 101, in response to the network communication between the power equipment and the control equipment being blocked, the transport layer of the network is detected based on the execution parameters of the network communication to obtain a first detection result.
[0041] Here, the execution parameters of a communication network refer to the specific values, addresses, or identifiers that must be preset or used during execution to drive a particular network function or task (such as data transmission, connection detection, security policy execution, etc.). They constitute the direct basis and quantitative standard for network devices (such as routers and firewalls) or management systems to perform operations, precisely defining who is operating, where they come from, what they do, and how they do it.
[0042] As an example, in an automated operation and maintenance system for a power system, when the task of "checking the HTTP service health status of the core server group every 30 seconds" needs to be executed, the system will use the following execution parameters to specify the task: Target node IP address list: [192.168.1.10, 192.168.1.11, 192.168.1.12] (i.e., the IP address of the core server) Source node IP address: 192.168.2.100 (i.e., the IP address of the device where the detection probe is located) URL path to be checked: / api / health (a specific interface used to determine whether the service is healthy) Detection time interval: 30s (controlling the detection frequency).
[0043] Because generating network logs involves a lengthy pipeline of generation, caching, collection, transmission, storage, indexing, and analysis, by the time the system retrieves key information from the log analysis platform, the fault may have already been ongoing for several minutes or longer, easily missing the optimal repair opportunity. However, by analyzing the real-time readings of the network status counter during parameter execution, communication fault location can begin as soon as a fault occurs, improving the real-time performance of fault detection.
[0044] In some embodiments, see Figure 4 , Figure 4 It is shown that step 101 can be achieved by performing steps 1011 to 1013.
[0045] In step 1011, the control power equipment sends a synchronization packet to the control equipment.
[0046] Here, a synchronization packet is a special control data packet sent by both communicating parties in a communication network to establish and maintain a synchronized connection. It typically does not transmit higher-level application data itself, but rather paves the way and protects the reliable transmission of application data. In this step, the synchronization packet is a SYN packet sent by the power equipment to the control equipment, containing the power equipment's initial sequence number. The power equipment and control equipment complete their initial handshake through the SYN packet.
[0047] In step 1012, in response to the execution parameters including the control device receiving a first confirmation packet from the power device, it is determined that the first detection result is normal.
[0048] The first confirmation packet is a data packet sent by the power equipment after responding to the second confirmation packet sent by the control equipment, and the second confirmation packet is a data packet sent by the control equipment to the power equipment after responding to the synchronization packet sent by the power equipment.
[0049] Here, the second confirmation packet is the SYN-ACK packet sent by the control device to the power device after receiving the SYN packet. The SYN-ACK packet acknowledges the power device's SYN and also contains the control device's initial sequence number. The power device and the control device complete the second handshake through the SYN-ACK packet.
[0050] Here, the first confirmation packet is an ACK packet sent by the power equipment to the control equipment after receiving the SYN-ACK packet, further confirming the control equipment's confirmation of the server's SYN. The power equipment and the control equipment complete the third handshake through ACK packets.
[0051] In step 1013, in response to the absence of a first confirmation packet received by the control device in the execution parameters, the first detection result is confirmed to be abnormal.
[0052] Here, the absence of the first acknowledgment packet received by the control device in the execution parameters indicates an anomaly in the three-way handshake process. This is because the three-way handshake is a dedicated process for the transport layer (taking TCP as an example) to establish a reliable connection. Failure at any stage of this process means the transport layer's fundamental functions cannot be fulfilled, thus confirming a fault in the transport layer of the communication network.
[0053] As an example, in power system communication, if the SCADA server in the dispatch center can successfully ping the IP address of the intelligent monitoring and control device in a remote substation, but the telnet command consistently returns "connection timeout" when initiating a data acquisition connection, this clearly indicates a transport layer failure. The essence of this phenomenon is a failed TCP three-way handshake: the SYN packet sent by the SCADA server can reach the substation network, but due to the strict access control policies configured in the substation's industrial firewall, the connection request from the dispatch center's IP to the monitoring and control device's service port is blocked, causing the device to fail to receive the SYN packet or its returned SYN-ACK packet to be discarded by the firewall.
[0054] This application embodiment achieves rapid confirmation of whether the network layer causing the communication network failure is the transport layer by detecting whether a first confirmation packet is received, based on the execution parameters. If the network layer causing the communication failure is detected to be the transport layer, the cause of the communication network failure can be determined by subsequently detecting potential problems in the transport layer, thereby improving the speed and accuracy of communication failure detection.
[0055] In step 102, in response to the absence of abnormality in the first detection result, the application layer of the network is detected based on the execution parameters of the network between the power equipment and the control equipment to obtain the second detection result.
[0056] In some embodiments, see Figure 5 , Figure 5It is shown that step 102 can be achieved by performing steps 1021 to 1022.
[0057] In step 1021, the data packets in the execution parameters are detected, and in response to the data packets including an acquisition request packet or a submission request packet, it is determined that the second detection result is not abnormal.
[0058] Here, GET requests are primarily used to retrieve data from the server. Their characteristic is that request parameters are appended directly to the URL, making them visible and having a limited length. Furthermore, executing the same GET request multiple times typically does not change the server's data (idempotency). In contrast, POST requests are mainly used to submit data to the server, such as form information and file uploads. The submitted data is encapsulated in the request body, is invisible to the user, has a larger length limit, and multiple submissions may produce different side effects, such as duplicate orders.
[0059] In step 1022, in response to the fact that the data packet does not include an acquisition request packet or a submission request packet, the second detection result is determined to be abnormal.
[0060] In this embodiment, once the underlying network (such as a TCP connection at the transport layer) has been successfully established (i.e., the three-way handshake is successful and the first detection result is normal), it proves that the communication link itself is working properly. At this point, the interaction between the client and the server should enter the dialogue stage of the application layer protocol (such as HTTP). If, at this stage, the expected GET or POST request is not sent, does not arrive, or the server does not respond, this clearly means that the failure occurs in the "application" itself, rather than in network connectivity.
[0061] As an example, application layer failures in communication networks may occur as follows: The application service module of the field device crashes: Although the network port of the intelligent terminal in the substation is online (TCP connection is reachable), the application process carrying the web service or data interface has crashed. This causes the dispatch master station to be able to "connect" to the equipment, but any GET requests to query real-time data will be ignored, and no measurement information can be obtained.
[0062] This application embodiment detects whether the data packet includes an acquisition request packet or a submission request packet, thereby quickly confirming whether the network layer causing the communication network failure is the application layer based on the execution parameters. If the network layer causing the communication failure is detected to be the application layer, the cause of the communication network failure can be determined by subsequently detecting potential problems at the application layer, thus improving the speed and accuracy of communication failure detection.
[0063] In step 103, the network layer where the fault occurred is determined based on the first detection result and the second detection result, wherein each network layer is configured with a corresponding control strategy.
[0064] In some embodiments, see Figure 6 , Figure 6 It is shown that step 103 can be implemented by the following steps 1031 to 1032.
[0065] In step 1031, in response to the inclusion of an abnormal detection result in the first detection result and the second detection result, the network layer corresponding to the abnormal detection result is determined as the network layer where the fault occurred.
[0066] In step 1032, in response to the detection results that both the first and second detection results are normal, the session layer is determined to be the faulty network layer.
[0067] Here, since the detection is initiated when network communication between the power equipment and the control equipment is blocked, if both the first and second detection results are normal, the possibility of faults in the transport layer and application layer is ruled out. Therefore, the faulty network layer is the session layer.
[0068] For example, when the dispatch center collects large-capacity waveform recording files from the substation fault waveform recorder, the transmission process often interrupts unexpectedly after completing a certain percentage (such as 30%). Although the system can automatically rebuild the TCP connection (proving that the transport layer is working properly) and the file transfer service itself is not abnormal (proving that the application layer is normal), the task must start from scratch after each reconnection and cannot resume from the point of interruption.
[0069] This application embodiment quickly determines whether the network causing the communication failure is at the session layer by performing a first detection and a second detection, based on the detection results of the two detections. This enables rapid location of the cause of the network failure. If the network layer causing the communication failure is detected to be the session layer, subsequent detection only needs to target potential problems at the session layer to determine the cause of the communication network failure, thus improving the speed and accuracy of communication failure detection.
[0070] It should be noted that, after obtaining the first and second detection results, business continuity and transaction success rate can be calculated. If the business continuity is lower than the set continuity threshold, or the transaction success rate is lower than the set success rate threshold, it indicates that the level causing the network communication failure is the session layer.
[0071] In step 104, based on the execution parameters, the triggering condition of the control policy corresponding to the network layer where the fault occurred is determined, and the cause of the communication fault is determined based on the triggering condition.
[0072] In some embodiments, see Figure 7 , Figure 7 It is shown that step 104 can be achieved by performing steps 1041 to 1044.
[0073] In step 1041, in response to the fault, the network layer is the transport layer. The rule action field in the execution parameters is matched with the firewall blocking rules in the control policy to obtain the first triggering situation.
[0074] Here, firewall blocking rules are a series of policy entries pre-defined in network perimeter security devices. Their core function is to perform deep inspection and matching of data packets according to predefined security policies, and then execute drop or rejection operations. These rules are typically based on the packet's five-tuple (source IP address, destination IP address, protocol type, source port number, destination port number) and higher-level application layer content or connection state for matching. When the characteristics of a data packet match the matching conditions of a blocking rule, the firewall will interrupt the forwarding process of that data packet, preventing it from entering or leaving the protected network area, thereby achieving access control and threat defense.
[0075] As an example, you can check the firewall policies deployed on the network path. For instance, if you find "deny access to xxxx / tcp port from source IP 192.168.1.10 to destination IP 10.1.1.100" in the firewall's blocking rule list, and the execution parameters (source IP, destination IP, destination port, protocol) match exactly with a specific blocking rule in the firewall's blocking rule list, then you can confirm that the network blocking was caused by this firewall policy, i.e., firewall blocking.
[0076] In step 1042, in response to the absence of a matching blocking rule in the first triggering condition, the return code of the routing table command included in the execution parameters is detected to obtain the second triggering condition.
[0077] The routing table command return code is generated when the control policy meets the first set condition.
[0078] As an example, analyzing the return codes of route tracing commands (such as tracert) can quickly determine the cause of network failures: if the result shows consecutive asterisks (*) after a gateway node before reaching the target IP, it indicates a blockage or packet loss at that node or its subsequent links; if the return message "Destination host unreachable" indicates a missing trailing route or the target host being offline; if the IP address appears in a loop within hops, it confirms the existence of a routing loop in the network. These characteristics can pinpoint the fault to a specific network node or path, i.e., a routing error.
[0079] In step 1043, in response to the second triggering condition of detecting the return code of the routing table command, the driver signature information included in the execution parameters is detected to obtain the third triggering condition.
[0080] The driver signature information is generated when the control policy meets the set second condition.
[0081] The method of determining the cause of communication failure by detecting driver signature information refers to diagnosing communication anomalies in network devices (such as network cards, switch interface cards, etc.) by verifying the digital signature status of their device drivers. If the detection reveals an invalid, revoked, or incompatible driver signature, it indicates that the driver may be corrupted, have an incorrect version, or have been maliciously tampered with. This will directly cause the hardware to malfunction, leading to communication failures such as packet loss, connection interruption, or a sudden drop in performance. This pinpoints the root cause of the problem to the driver level rather than the upper-layer network protocol or physical link, i.e., driver interception. If it is confirmed that it is not driver interception, then it indicates a network failure.
[0082] In step 1044, the cause of the communication failure when the network layer of the fault is the transport layer is determined based on at least one of the first triggering condition, the second triggering condition, and the third triggering condition.
[0083] This application embodiment determines the specific cause of the communication fault by targeting the triggering of the control strategy of the transport layer, after determining that the network layer where the fault occurred is the transport layer. This reduces the steps in communication fault analysis and improves the speed of communication fault analysis.
[0084] In some embodiments, see Figure 8 , Figure 8 It is shown that step 104 can be achieved by performing steps 1045 to 1048.
[0085] In step 1045, the network layer responding to the fault is the application layer. Based on the proxy settings in the execution parameters, the reachability of the proxy address specified by the control policy is detected to obtain the fourth triggering condition.
[0086] As an example, network tools (such as telnet or curl) can be used to directly test the IP address and port connectivity of the proxy server. If the test fails (such as connection timeout or rejection), it indicates that the fault is caused by the proxy server itself crashing, the network path being blocked, or the security policy blocking it. This eliminates problems with client applications and higher-level protocols, focusing the investigation on the proxy service link, i.e., proxy interception.
[0087] In step 1046, in response to the fourth trigger condition that the proxy address is reachable, the first process list set in the control policy and the second process list in the execution parameters are compared to obtain the fifth trigger condition.
[0088] Here, the second process list is the output of isof / procexp, the first process list is the processes allowed in the original environment, and processes outside the first process list may be, for example, detected API hooks or environment variable injections.
[0089] By examining the system output of tools such as Process Explorer (procexp), if it is found that critical network processes (such as browsers and communication services) have loaded unknown modules or have abnormal API hook call chains, it can be determined that their execution flow has been injected and intercepted by third-party code (such as malware or hijacking plugins). When abnormal proxy configurations (such as HTTP_PROXY being tampered with) or malicious DLL paths are detected in the process environment variables at the same time, it can be confirmed that the network failure is caused by traffic being illegally redirected to a proxy or filtering endpoint controlled by the attacker, resulting in connection interruption, data leakage, or access hijacking, i.e., library injection.
[0090] In step 1047, in response to the fifth trigger condition that the second process list includes processes not in the first process list, the completion status of the execution parameters according to the application layer rules in the control policy is checked to obtain the sixth trigger condition.
[0091] Here, if it is found that application-layer rules (such as WAF protection policies, content filtering rules, and authentication processes) in security devices or agent systems fail to complete processing as expected—for example, the rule engine discards legitimate requests due to misconfiguration, or the authentication service times out and fails to return results—it can be determined that the fault originates from an abnormal execution of the application-layer control policy, rather than an underlying network link problem, i.e., application-layer interception. If it is confirmed that it is not application-layer interception, then the cause of the fault is a kernel anomaly.
[0092] In step 1048, the cause of the communication failure when the network layer of the fault is the application layer is determined based on at least one of the fourth, fifth, and sixth triggering conditions.
[0093] This application embodiment determines the specific cause of the communication fault by targeting the triggering of the control policy at the application layer, after identifying the network layer where the fault occurred as the application layer. This reduces the steps involved in communication fault analysis and improves the speed of communication fault analysis.
[0094] In some embodiments, see Figure 9 , Figure 9 It is shown that step 104 can be achieved by performing steps 1049 to 10413.
[0095] In step 1049, in response to the fault, the network layer is the session layer, the time to live in the control execution data is controlled, and the response of network communication routing under the control policy is detected. The seventh trigger condition is obtained based on the routing response.
[0096] Here, after determining that the network session layer is faulty, by checking whether the session state data is prematurely cleared due to the timeout setting of intermediate devices (such as load balancers or firewalls) being too short, and by combining the response path analysis of the tracing route to determine whether there is asymmetric routing (i.e., the round-trip data flows through different nodes), the specific cause of the fault can be located as the session state being lost due to inconsistency in time or path, resulting in the interruption of the persistent connection, i.e., path interruption.
[0097] In step 10410, in response to the detection that the seventh trigger condition is that all routes in network communication generate a response, the eighth trigger condition is obtained by checking the currently effective access control list rules in the control policy based on the execution parameters.
[0098] Here, by examining the currently effective access control list rules, it can be found that if session persistence checks or dynamic allowance mechanisms based on connection state are not configured correctly (such as not correctly recognizing the SYN / ACK state transition of the TCP session), the firewall will still discard subsequent packets to maintain the session after the three-way handshake is successful, thereby interrupting the established legitimate connection. This confirms that the communication failure is due to the mismatch between the access control policy and the session layer state management, i.e., ACL blocking.
[0099] In step 10411, in response to the absence of a blocking rule in the currently effective access control list rules of the eighth triggering condition, a router service quality detection is performed based on the execution parameters to obtain the ninth triggering condition.
[0100] Different router service quality test results are generated when the control policy satisfies different third conditions.
[0101] As an example, by examining the router's QoS configuration and statistics, it was discovered that its bandwidth allocation policy or traffic shaping rules imposed inappropriate delays or drop priorities on session keep-alive packets (such as TCP Keep-alive), causing the continuous transmission of packets maintaining session state to fail in a timely manner. This confirmed that the root cause of the fault was a conflict between the QoS policy and the session layer keep-alive mechanism, leading to session timeout interruption, i.e., QoS rate limiting. If it is confirmed that it is not QoS rate limiting, then the fault lies with the control device.
[0102] In step 10412, the cause of the communication failure when the network layer of the fault is the session layer is determined based on at least one of the seventh triggering condition, the eighth triggering condition, and the ninth triggering condition.
[0103] This application embodiment determines the specific cause of the communication failure by identifying the network layer where the failure occurs as a session after determining the triggering of the control policy for the session. This reduces the steps involved in communication failure analysis and improves the speed of communication failure analysis.
[0104] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0105] In related technologies, network fault location using firewall-based log analysis methods relies on collecting and parsing communication logs automatically recorded by network boundary security devices (i.e., firewalls). This method reverse-engineers and locates network faults such as communication interruptions or anomalies by retrospectively analyzing key fields in the logs, such as source / destination IP addresses, port numbers, protocol types, and communication statuses (e.g., "deny" or "allow").
[0106] See Figure 10 , Figure 10 This is a flowchart illustrating a communication fault detection method in an application scenario provided in an embodiment of this application, including steps 201 to 207.
[0107] In step 201, the detection parameters are read.
[0108] See Figure 11 , Figure 11This is a schematic diagram of the network communication structure in a power system provided in the embodiments of this application. The system is designed with a four-layer architecture of "display layer, control layer, configuration layer, and support layer", with each layer undertaking different dimensions of system functions: Display layer: responsible for human-computer interaction and information presentation, including HMI subsystem (graphic design, HMI plug-in module), WEB subsystem (WEB display, WEB plug-in module), and APP subsystem (APP process, APP plug-in module), which is the "front-end window" for user interaction with the system. Control Layer: Focusing on data processing, acquisition, communication, and service management, it is the system's "core control hub": Data Processing Subsystem: Processes and interacts with data through modules such as "manual operation, data processing, and publish / subscribe interfaces"; Data Acquisition Subsystem: Supports industrial protocols such as Modbus TCP / RTU, CAN, and Profibus to acquire equipment data; Gating Subsystem: Manages data transmission and network topology through modules such as "logic parsing, network services, and topology services"; Service Management Subsystem: Responsible for service scheduling such as "master / standby machine management, process monitoring, and node information forwarding"; Time Synchronization Subsystem: Ensures system time synchronization through the "time synchronization service" module. Configuration Layer: Responsible for system configuration, maintenance, and simulation; it is the system's "customizability and operation and maintenance assurance layer." System Modeling Subsystem: Through modules such as "system modeling, equipment modeling, and configuration management," it completes the digital modeling of the system and equipment. Maintenance Subsystem: Through modules such as "fault diagnosis, project management, and data management," it realizes system operation and maintenance and data management. Simulation Subsystem: Through modules such as "simulation acquisition, simulation management, and project debugging," it supports system simulation verification. Support Layer: Provides underlying support such as data storage, logs, and message bus; it is the system's "operational foundation." Data Storage Subsystem: Through modules such as "database, time-series database, and file storage," it ensures persistent data storage. Log Recording Subsystem: Through modules such as "log interface and log database configuration," it realizes the recording and management of system logs. Message Bus Subsystem: Through modules such as "message bus interface and message bus configuration," it supports the efficient transmission of messages within the system.
[0109] Each layer and subsystem collaborates closely through data flow and functional interaction. Typical logic includes: Data flow: Real-time data collected by the bottom-level "data acquisition subsystem" is processed by the "data processing subsystem" and then transmitted to the "display layer" (HMI / WEB / APP) to be presented to the user; simultaneously, the data is also stored in the "data storage subsystem" of the "support layer" for subsequent analysis. Service and management: The "service management subsystem" is responsible for scheduling system services (such as primary / standby machine switching and process monitoring), the "time synchronization subsystem" ensures time synchronization, and the "gating subsystem" manages network topology and data transmission, jointly supporting the stable operation of the system. Configuration and operation and maintenance: The "system modeling subsystem" of the "configuration layer" completes the digital modeling of the system and equipment, the "maintenance subsystem" realizes fault diagnosis and engineering management, and the "simulation subsystem" verifies the system design through simulation, providing a basis for system optimization.
[0110] Figure 11 It also includes other key functional modules such as log storage: recording system operation logs for fault tracing and system optimization; alarm push: pushing alarm information to users when the system detects an anomaly; Network diagnostic module: Used to detect and analyze network status to ensure communication stability. This architecture clearly demonstrates the entire process logic of an industrial system from "data acquisition → processing → presentation → operation and maintenance," and serves as a typical reference model for system design in fields such as industrial automation and the Internet of Things.
[0111] See Figure 12 This application embodiment determines the specific cause of the fault by pre-determining the network layer at which the fault occurs. (See attached document.) Figure 12 This figure is a logical block diagram of a communication fault detection method in a specific application scenario according to an embodiment of this application. This embodiment elaborates on the complete closed-loop process from receiving alarms to finally reporting diagnostic results. First, the system starts the fault detection task and reads the preset detection parameters through the configuration layer. These parameters are defined in JSON format, specifying the monitoring target (such as the target node IP address list "targetIPs"), the monitoring source (such as the source node IP address "sourceIPs"), the specific monitoring content (such as the URL path to be accessed "urls"), and the monitoring frequency (such as the detection time interval "interval"), providing accurate input for subsequent detection work. Next, based on the read execution parameters, the system enters the information acquisition stage. The three-layer detection engine starts to work actively, capturing key execution parameters and status data in the network communication process in real time through network packet capture, socket status analysis, etc. Subsequently, in step 203, the system enters the core fault level determination stage. This step utilizes Figure 13The three-layer detection logic shown comprehensively analyzes the acquired information to accurately determine whether the fault causing the communication blockage occurs at the transport layer, application layer, or session layer. After the fault level is determined, the system records the preliminary fault network level information and related raw data in the system log for subsequent auditing and in-depth analysis. Then, based on the determined fault level, the system calls the corresponding in-depth diagnostic logic (such as...). Figure 14-16 As shown in the diagram, further analysis is conducted to determine the specific cause of the fault (e.g., "firewall blocking" or "proxy interception"), and this clear diagnostic cause is notified to the alarm processing module. On the one hand, the original network fault alarm information (such as "TCP connection timeout") is pushed to ensure that basic alarms are not lost; on the other hand, the final alarm after in-depth diagnosis, with a clear fault cause, is sent to the human-machine interface (HMI) or upper-level management system to guide maintenance personnel to perform rapid and accurate fault handling.
[0112] In step 202, the detection information is obtained.
[0113] The execution parameters to be detected can be defined by configuring a JSON format file as follows.
[0114]
[0115] The parameters that need to be detected include the target node IP address list, the source node IP address, the URL path to be detected, and the detection time interval.
[0116] In step 203, the fault network level is determined.
[0117] See Figure 12 A three-layer detection engine can be used to detect network layers where faults occur. For example, transport layer detection monitors the TCP three-way handshake state using raw sockets; application layer detection verifies the normal sending of HTTP requests based on feature matching; and session layer detection analyzes business continuity and calculates transaction success rates. (See appendix for details.) Figure 12This figure is a logical block diagram of the three-layer detection engine used to confirm the fault network layer in this embodiment of the application. This detection engine is the core of achieving rapid fault layer location. The first layer is the transport layer detection. The core objective of this detection is to verify the end-to-end basic network connectivity. In specific implementation, the system simulates a client to initiate a TCP connection request to the target server port by creating a raw socket, i.e., sending a SYN packet. Subsequently, the system closely monitors whether the network interface can receive the SYN-ACK packet returned by the server, and whether the client can successfully send the final ACK packet. If this complete three-way handshake process cannot be completed (for example, the SYN packet is sent but not received, or no SYN-ACK response is received), the fault is determined to occur at the transport layer, and the first detection result is set to "abnormal". Conversely, if the three-way handshake successfully establishes a TCP connection, the first detection result is "no abnormality", and the process proceeds to the next layer of detection. The second layer is the application layer detection. After confirming that the transport layer connection is correct, this layer of detection aims to verify whether the interaction of application data is normal. In practice, the system uses feature matching algorithms (e.g., deep packet inspection) on established TCP connections to identify the presence of data packets for specific application protocols, such as HTTP GET or POST request packets. If no application-layer request packet from the client is captured within a preset time window after the TCP connection is successfully established, the system determines that the fault occurs at the application layer and sets the second detection result to "abnormal." This usually means that the client application itself (rather than the network) has failed to generate or send data correctly. If the request packet is successfully detected, the second detection result is "no abnormality." The third layer is session layer detection. If the results of the first two layers are both "no abnormality," but business communication is still interrupted, the system determines that the fault occurs at the session layer. This layer's judgment is based on elimination and combined with analysis of business continuity. In practice, the system will statistically analyze the transaction success rate or session persistence status of specific businesses. For example, in long-term file transfers or data stream interactions, if the connection is frequently and irregularly interrupted and reset, even if the TCP connection can be rebuilt each time (transport layer is normal) and the application request format is correct (application layer is normal), this disruption of business continuity can be attributed to a session layer fault. At this point, the system determines the final fault level to be the session level.
[0118] When network failures occur at different network layers, typical operating conditions are shown in Table 1.
[0119] Table 1
[0120] In step 204, network fault information is recorded.
[0121] Once the network layer where the fault occurred is determined, the cause of the fault is confirmed through the blocking analysis block, and the cause of the fault is sent to the log.
[0122] When the network layer at which the communication failure is determined to be the transport layer, it can be done by means of... Figure 14 The process involves sequentially checking local firewall rules, verifying routing table configurations, and detecting security software drivers to locate the fault. For example, if the network layer where the fault is confirmed to be the transport layer, the local firewall rules (iptables -L) are checked. If a matching blocking rule is found, it is marked as firewall blocking. If no matching blocking rule is found, the routing table (IP route) is further checked. If the routing result shows that the last node route is unreachable, the network fault is marked as a routing error. If the routing result shows that the last node route is reachable, the security driver (Ismod) is further checked. If the security driver result shows that it is a Microsoft-signed driver only, the fault is marked as a network fault. If the security driver result shows that it is not a Microsoft-signed driver, the fault is marked as driver blocking.
[0123] When the network layer at which the communication failure is determined to be the application layer, it can be done through methods such as... Figure 15 The process involves sequentially analyzing LSP chain integrity, checking proxy settings, and detecting application-layer filtering to locate the fault. For example, if the faulty network layer is confirmed to be the application layer, the proxy settings are checked (env|grep-Iproxy). The proxy address is tested for validity using a ping test. If valid, the faulty cloud is marked as being blocked by the proxy. If invalid, LD_PRELOAD (cat / proc / $PID / environ) is further checked to confirm the presence of illegal injection. If illegal injection is found, the fault is marked as library injection. If no illegal injection is found, iptables is further checked (sudo iptables-Lvn) to further confirm the execution status of application-layer rules. If the execution status of application-layer rules is abnormal, the fault is marked as application-layer blocking. If the execution status of application-layer rules is normal, the fault is marked as kernel abnormality.
[0124] When the network layer at which the communication failure is determined to be the session layer, it can be done by, for example... Figure 16The fault location is determined by sequentially performing traceroute path analysis, checking intermediate device ACLs, and verifying QoS policies. For example, if the network layer where the fault is determined to be the session layer, a traceroute (tcptraceroute) test is performed. If the test result indicates that the last hop is unreachable, the fault is marked as a path interruption. If the last hop is reachable, the ACL (tcpdump) is further checked. If the ACL test result indicates a management denial, the fault is marked as an ACL block. If no management denial is found, QoS is further verified. If a sudden increase in latency exceeding 100ms is found, the fault is marked as QoS rate limiting. If no sudden increase in latency exceeding 100ms is found, the fault is marked as a server-side fault.
[0125] Refer to Table 2 for the detection methods used for fault location.
[0126] Table 2
[0127] Typical scene features obtained through the detection methods shown in Table 2 are shown in Table 3.
[0128] Table 3
[0129] In step 205, the cause of the fault diagnosis is notified.
[0130] In step 206, the original network fault alarm information is pushed.
[0131] In step 207, a network failure alarm is sent.
[0132] In this embodiment, monitoring targets are defined through a configuration management module. A three-layer detection architecture is adopted to achieve comprehensive monitoring of the transport layer, application layer, and session layer. A decision tree algorithm is used to analyze the causes of network blockages, and finally, alarms are generated and pushed to the HMI system. This invention can accurately identify more than 90% of network blockage events in power monitoring systems, with an average location time of less than 3 minutes. It is particularly suitable for detecting communication blockages caused by encryption software and intranet management systems, avoiding additional manpower investment due to network blockage problems. Alarms allow users to be notified of network problems and quickly intervene to handle them, improving the efficiency of network problem handling.
[0133] The following continues to describe exemplary results of the implementation of the communication fault detection device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the communication fault detection device 455 in the memory 450 may include: The first detection module 4551 is used to detect the transport layer of the network based on the execution parameters of the network communication in response to the network communication being blocked, and to obtain a first detection result. The second detection module 4552 is used to detect the application layer of the network based on the execution parameters of the network between the power equipment and the control equipment in response to the absence of abnormality in the first detection result, and to obtain the second detection result. The first judgment module 4553 is used to determine the network layer where the fault occurred based on the first detection result and the second detection result, wherein each network layer is configured with a corresponding control strategy. The second judgment module 4554 is used to determine the triggering status of the control strategy corresponding to the network layer where the fault occurred based on the execution parameters, and to determine the cause of the communication fault based on the triggering status.
[0134] In some embodiments, the first detection module 4551 is further configured to control the power equipment to send a synchronization packet to the control equipment; In response to the execution parameters including the control device receiving a first confirmation packet from the power device, it determines that the first detection result is normal. The first confirmation packet is a data packet sent by the power device after responding to a second confirmation packet sent by the control device, and the second confirmation packet is a data packet sent by the control device to the power device after responding to a synchronization packet sent by the power device. In response to the absence of a first confirmation packet received by the control device in the execution parameters, the first detection result is confirmed to be abnormal.
[0135] In some embodiments, the second detection module 4552 is further configured to detect data packets in the execution parameters, and in response to the data packets including an acquisition request packet or a submission request packet, determine that the second detection result is not abnormal; If the data packet does not include an acquire request packet or a submit request packet, the second detection result is determined to be abnormal.
[0136] In some embodiments, the first judgment module 4553 is further configured to, in response to the first detection result and the second detection result including an abnormal detection result, determine the network layer corresponding to the abnormal detection result as the network layer that has failed. In response to the fact that both the first and second detection results are normal, the session layer is identified as the faulty network layer.
[0137] In some embodiments, the second judgment module 4554 is further configured to, in response to the network layer of the fault being the transport layer, match the rule action field in the execution parameters with the firewall blocking rules in the control policy to obtain the first triggering condition; In response to the absence of a matching blocking rule in the first triggering condition, the return code of the routing table command included in the execution parameters is detected to obtain the second triggering condition, wherein the return code of the routing table command is generated when the control policy meets the set first condition; In response to the second triggering condition of detecting the return code of the routing table command, the driver signature information included in the execution parameters is detected to obtain the third triggering condition, wherein the driver signature information is generated when the control policy meets the set second condition; Based on at least one of the first, second, and third triggering conditions, determine the cause of the communication failure when the network layer of the fault is the transport layer.
[0138] In some embodiments, the second judgment module 4554 is further configured to, in response to the network layer of the fault being the application layer, detect the reachability of the proxy address specified by the control policy based on the proxy settings in the execution parameters, and obtain the fourth triggering condition; In response to the fourth trigger condition that the proxy address is reachable, the fifth trigger condition is obtained by comparing the first process list set in the control policy with the second process list in the execution parameters. In response to the fifth trigger condition, which is that the second process list includes processes that are not in the first process list, the execution parameters are checked to see if the application layer rules in the control policy are completed, and the sixth trigger condition is obtained. Based on at least one of the fourth, fifth, and sixth triggering conditions, determine the cause of the communication failure when the network layer of the fault is the application layer.
[0139] In some embodiments, the second judgment module 4554 is further configured to respond to the fault network layer being the session layer, control the time to live in the execution data, and detect the response of network communication routing under the control policy, and obtain the seventh trigger condition based on the routing response. In response to the detection of the seventh trigger condition, which is that all routes in network communication generate a response, the eighth trigger condition is obtained by checking the currently effective access control list rules in the control policy based on the execution parameters. In response to the absence of a blocking rule in the currently effective access control list rules in the eighth triggering condition, a router service quality detection is performed based on the execution parameters to obtain the ninth triggering condition. Different router service quality detection results are generated when the control policy satisfies different third conditions. Based on at least one of the seventh, eighth, and ninth triggering conditions, determine the cause of the communication failure when the network layer of the fault is the session layer.
[0140] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the communication fault detection method described in this application.
[0141] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the communication fault detection method provided in this application. For example, ... Figure 3 The communication fault detection method is shown.
[0142] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0143] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0144] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0145] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0146] In summary, by analyzing the execution parameters of network communication and sequentially detecting the transport and application layers, the network layer (transport, application, or session layer) causing the communication failure can be preliminarily determined. Based on the triggering of the control policies corresponding to the determined network layer, the cause of the communication failure can be identified. This allows for accurate determination of the cause of communication failures without having to troubleshoot each different problem individually. Furthermore, compared to methods that rely on analyzing network logs, analyzing execution parameters significantly speeds up the analysis of communication failure causes.
[0147] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A communication fault detection method, characterized in that, The method includes: In response to the disruption of network communication between power equipment and control equipment, the transport layer of the network is detected based on the execution parameters of the network communication to obtain a first detection result; In response to the first detection result showing no abnormality, the application layer of the network is detected based on the execution parameters of the network between the power equipment and the control equipment to obtain a second detection result; Based on the first detection result and the second detection result, the network layer where the fault occurred is determined, wherein each network layer is configured with a corresponding control strategy; Based on the execution parameters, the triggering status of the control policy corresponding to the network layer where the fault occurred is determined, and the cause of the communication fault is determined based on the triggering status.
2. The method according to claim 1, characterized in that, The first detection result obtained by detecting the transport layer of the network based on the execution parameters of the network between the power equipment and the control equipment includes: The power equipment is controlled to send a synchronization packet to the control equipment; In response to the execution parameters including the control device receiving a first confirmation packet from the power device, it is determined that the first detection result is normal, wherein the first confirmation packet is a data packet sent by the power device after responding to a second confirmation packet sent by the control device, and the second confirmation packet is a data packet sent by the control device to the power device after responding to the synchronization packet sent by the power device; In response to the absence of a first confirmation packet received by the control device in the execution parameters, the first detection result is confirmed to be abnormal.
3. The method according to claim 1, characterized in that, The second detection result is obtained by detecting the application layer of the network based on the execution parameters of the network between the power equipment and the control equipment, including: Detect the data packets in the execution parameters, and in response to the data packets including an acquisition request packet or a submission request packet, determine that the second detection result is normal; If the data packet does not include the get request packet or the submit request packet, the second detection result is determined to be abnormal.
4. The method according to claim 1, characterized in that, Based on the first detection result and the second detection result, the network layer at which the fault occurred is determined, including: In response to the inclusion of an abnormal detection result in the first detection result and the second detection result, the network layer corresponding to the abnormal detection result is determined as the network layer where the fault occurred. In response to the fact that both the first and second detection results are normal, the session layer is identified as the faulty network layer.
5. The method according to claim 1, characterized in that, The step of determining the triggering status of the control policy corresponding to the network layer where the fault occurred based on the execution parameters, and determining the cause of the communication fault based on the triggering status, includes: In response to the fault, which occurs at the transport layer, the rule action field in the execution parameters is matched with the firewall blocking rules in the control policy to obtain the first trigger condition. In response to the absence of a matching blocking rule in the first triggering condition, the routing table command return code included in the execution parameters is detected to obtain the second triggering condition, wherein the routing table command return code is generated when the control policy satisfies the set first condition; In response to the second triggering condition of detecting the return code of the routing table command, the driver signature information included in the execution parameters is detected to obtain the third triggering condition, wherein the driver signature information is generated when the control policy satisfies the set second condition; Based on at least one of the first triggering condition, the second triggering condition, and the third triggering condition, determine the cause of the communication failure when the network layer of the fault is the transport layer.
6. The method according to claim 1, characterized in that, The step of determining the triggering status of the control policy corresponding to the network layer where the fault occurred based on the execution parameters, and determining the cause of the communication fault based on the triggering status, includes: The network layer responding to the fault is the application layer. Based on the proxy settings in the execution parameters, the reachability of the proxy address specified by the control policy is detected to obtain the fourth triggering condition. In response to the fourth trigger condition that the proxy address is reachable, the fifth trigger condition is obtained by comparing the first process list set in the control policy and the second process list in the execution parameters; In response to the fifth triggering condition, which is that the second process list includes processes that are not in the first process list, the completion status of the execution parameters working according to the application layer rules in the control strategy is checked to obtain the sixth triggering condition. Based on at least one of the fourth, fifth, and sixth triggering conditions, determine the cause of the communication failure when the network layer of the fault is the application layer.
7. The method according to claim 1, characterized in that, The step of determining the triggering status of the control policy corresponding to the network layer where the fault occurred based on the execution parameters, and determining the cause of the communication fault based on the triggering status, includes: The network layer responding to the fault is the session layer, which controls the time to live in the execution data and detects the response of the network communication routing under the control policy, and obtains the seventh trigger condition based on the response of the routing. In response to the detection that the seventh trigger condition is generated by all routes in the network communication, the access control list rules currently in effect in the control policy are checked based on the execution parameters to obtain the eighth trigger condition; In response to the absence of a blocking rule in the currently effective access control list rules of the eighth triggering condition, a router service quality detection is performed based on the execution parameters to obtain a ninth triggering condition, wherein different router service quality detection results are generated when the control policy satisfies different third conditions; Based on at least one of the seventh, eighth, and ninth triggering conditions, determine the cause of the communication failure when the network layer of the fault is the session layer.
8. A communication fault detection device, characterized in that, The device includes: The first detection module is used to detect the transport layer of the network based on the execution parameters of the network communication in response to the network communication being blocked, and to obtain a first detection result. The second detection module is used to detect the application layer of the network based on the execution parameters of the network between the power equipment and the control equipment in response to the first detection result being normal, and to obtain a second detection result. The first judgment module is used to determine the network layer where the fault occurred based on the first detection result and the second detection result, wherein each network layer is configured with a corresponding control strategy; The second judgment module is used to determine the triggering status of the control policy corresponding to the network layer where the fault occurred based on the execution parameters, and to determine the cause of the communication fault based on the triggering status.
9. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions or computer programs. The processor, when executing computer-executable instructions or computer programs stored in the memory, implements the communication fault detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the method described in any one of claims 1 to 7.