Network and equipment fault detection method of production control system

Through the bidirectional polling mechanism combined with the collaborative work of redundancy and basic DCS controllers, the problem of insufficient fault warning in redundant cabinet network in DCS system is solved, efficient and accurate fault detection is achieved, and the stability and production efficiency of the system are improved.

CN120508079APending Publication Date: 2025-08-19CNGR ADVANCED MATERIAL CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510572400.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing DCS system has shortcomings in redundant cabinet network fault warning, and it is impossible to realize real-time scanning and fault warning of full-node equipment, resulting in difficulty in troubleshooting and affecting production efficiency and system stability.

Method used

Using a bidirectional polling mechanism, forward and reverse polling scanning commands are sent to the redundant cabinet respectively through the first DCS controller and the second DCS controller. Combined with the status of the feedback from the redundant cabinet, the network fault node and the equipment fault node are determined, and the redundant DCS controller and the basic DCS controller work together to achieve efficient fault detection.

Benefits of technology

It realizes comprehensive, accurate and efficient fault detection of redundant cabinets, reduces troubleshooting time, improves system operation stability and production efficiency, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508079A_ABST
    Figure CN120508079A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of battery material manufacturing, and discloses a network of a production control system and an equipment fault detection method. The method comprises the following steps: a first DCS controller sends a forward polling scanning command to a plurality of redundant cabinets, and each redundant cabinet feeds back a forward polling state to the first DCS controller according to a first frequency; the second DCS controller sends a reverse polling scanning command to a plurality of redundant cabinets, and each redundant cabinet feeds back a reverse polling state to the second DCS controller according to the first frequency; and the first DCS controller and the second DCS controller determine a network fault node and / or an equipment fault node of each redundant cabinet according to the network state fed back by each redundant cabinet. According to the application, comprehensive, accurate and efficient troubleshooting and diagnosis operation can be realized, so that an equipment maintenance engineer can maintain and replace a hidden danger part in time, and the stability of system operation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of battery material manufacturing, and in particular to a network and equipment fault detection method for a production control system. Background Art

[0002] With the rapid development of automation technology, DCS systems (Distributed Control System) have been widely used in industrial production, significantly reducing workers' labor intensity and improving production efficiency.

[0003] Currently, DCS manufacturers face significant deficiencies in early warning of redundant cabinet network failures. Specifically, existing systems only provide status query commands for individual devices, which are complex to operate and incapable of real-time scanning and fault warning for all nodes. Effective methods are also lacking for fault detection of redundant controllers and redundant cabinets. While redundant cabinets can maintain system operation even with a single node failure, system diagnostic messages only indicate a loss of redundancy and fail to pinpoint the fault point, making troubleshooting extremely difficult. Summary of the Invention

[0004] In view of this, an object of the present invention is to overcome the deficiencies in the prior art and to provide a method for detecting network and equipment faults in a production control system.

[0005] The present invention provides the following technical solutions:

[0006] In a first aspect, an embodiment of the present disclosure provides a method for detecting network and device faults in a production control system, the method comprising:

[0007] The first DCS controller sends a forward polling scan command to a plurality of redundant cabinets, and each of the redundant cabinets feeds back a forward polling status to the first DCS controller according to a first frequency;

[0008] The second DCS controller sends a reverse polling scan command to the plurality of redundant cabinets, and each redundant cabinet feeds back a reverse polling status to the second DCS controller according to the first frequency;

[0009] The first DCS controller and the second DCS controller determine the network fault node and / or device fault node of each redundant cabinet according to the forward polling state and the reverse polling state;

[0010] The model of the redundant cabinet includes Profinet, the redundant DCS controller includes the first DCS controller and the second DCS controller, the first DCS controller, multiple redundant cabinets and the second DCS controller form a cascade topology through a communication network, and the multiple redundant cabinets are located between the first DCS controller and the second DCS controller.

[0011] Control system network failures (IP address conflicts, controller disconnection or abnormal shutdown, server downtime) and equipment failures will cause interruptions in the operation of the production workshop control system and trigger major production anomalies, resulting in reduced production efficiency, abnormal product quality, equipment damage, financial losses (re-dissolution, waste treatment cost per ton) and other problems.

[0012] When a network or equipment failure occurs, the DCS control system will immediately start the redundant cabinet closest to the redundant DCS controller and in normal operation to replace the failed cabinet. The replaced redundant cabinet will take over the control of the mechanical equipment to quickly resume production. However, this also makes it difficult for equipment maintenance engineers to quickly and accurately locate the failed network node and / or failed equipment, nor can they immediately know the specific redundant equipment to replace. They can only check the redundant cabinets and their connected networks one by one, which is inefficient and poses a safety hazard to subsequent operations.

[0013] The redundant cabinets are connected to the network in sequence. If a fault occurs between the redundant cabinets, an automatic switching mechanism will be implemented to ensure normal operation of the mechanical equipment controlled by the DCS control system. Therefore, the redundant DCS controller is required to periodically send a first preset polling command to the redundant cabinets. Based on the network status feedback from the redundant cabinets, the redundant DCS controller can comprehensively, accurately, and efficiently determine the network fault nodes and / or equipment fault nodes of each redundant cabinet. This allows equipment maintenance engineers to accurately identify the fault point and perform maintenance, improving the stability of system operation. This eliminates the need to troubleshoot each fault one by one, saving time and costs.

[0014] Specifically, if the communication network connection is normal, the redundant cabinet connected to the redundant DCS controller receives and inserts the feedback network status into the first preset polling command and forwards it to another connected redundant cabinet, and at the same time sends the feedback network status to the redundant DCS controller until the last redundant cabinet returns the feedback network status to the redundant DCS controller to complete the closed-loop communication;

[0015] If the communication network connection is abnormal, the feedback network status is returned to the redundant DCS controller from the redundant cabinet connected to the fault point and capable of normal communication.

[0016] The plurality of redundant cabinets are connected to the network in sequence.

[0017] During forward polling scanning, the system starts from the first redundant cabinet connected to the first DCS controller and polls backward in sequence to the last redundant cabinet or the last redundant cabinet that feeds back the forward polling status; during reverse polling scanning, the system starts from the last redundant cabinet connected to the second DCS controller and polls forward in sequence to the first redundant cabinet or the first redundant cabinet that feeds back the reverse polling status.

[0018] In an optional embodiment, the network and device fault detection method includes:

[0019] The first DCS controller obtains the expected state of each redundant cabinet, and determines whether the forward polling state of each redundant cabinet is the same as its corresponding expected state, and selects redundant cabinets with different expected states and / or redundant cabinets without forward polling state feedback as a first candidate fault cabinet set;

[0020] The second DCS controller obtains the expected state of each redundant cabinet, and determines whether the reverse polling state of each redundant cabinet is the same as its corresponding expected state, and selects redundant cabinets with different expected states and / or redundant cabinets without reverse polling state feedback as the second candidate fault cabinet set;

[0021] Determine the network fault node and / or device fault node of each redundant cabinet according to the first candidate fault cabinet set and the second candidate fault cabinet set;

[0022] and / or,

[0023] The first frequency is 100ms / time to 2s / time.

[0024] In an optional implementation manner, a method for determining a network fault node and / or a device fault node of each redundant cabinet according to the first candidate fault cabinet set and the second candidate fault cabinet set is:

[0025] The redundant cabinets are numbered: starting from 1 and numbering the first redundant cabinet after the first DCS controller to the last redundant cabinet;

[0026] Determine whether the first candidate faulty cabinet set and the second candidate faulty cabinet set have the same redundant cabinet number;

[0027] If there is no identical redundant cabinet number, and both the first candidate fault cabinet set and the second candidate fault cabinet set are not empty, obtaining the starting fault cabinet number in the first candidate fault cabinet set, obtaining the ending fault cabinet number in the second candidate fault cabinet set, and determining the connection network between the starting fault cabinet in the first candidate fault cabinet set and the ending fault cabinet in the second candidate fault cabinet set as a network fault node of the redundant device;

[0028] If the same redundant cabinet exists and both the first candidate fault cabinet set and the second candidate fault cabinet set are not empty, obtain the starting fault cabinet number in the first candidate fault cabinet set, obtain the ending fault cabinet number in the second candidate fault cabinet set, and determine one network connection before the starting fault cabinet of the first candidate fault cabinet set, the network connection between the starting cabinet of the first candidate fault cabinet set and the ending cabinet of the second candidate fault cabinet set, and one network connection after the ending cabinet of the second candidate fault cabinet set as network fault nodes; and determine the starting fault cabinet, the ending fault cabinet, and the redundant cabinets between them as device fault nodes;

[0029] If the first candidate fault cabinet set is empty, determining the connection network between the second DCS controller and the last redundant cabinet as a network fault node;

[0030] If the second candidate fault cabinet set is empty, determining the connection network between the first DCS controller and the first redundant cabinet as a network fault node;

[0031] If both the first candidate faulty cabinet set and the second candidate faulty cabinet set are empty, then there is no network or equipment failure between the first DCS controller and the second DCS controller.

[0032] This application uses a DCS controller to periodically use two-way polling (forward polling and reverse polling) to scan the redundant cabinet Profinet and the network connections between them. By checking whether the network status reported by the redundant cabinet is consistent with the expected status (power on, standby, specific parameter adjustment, etc.) through two pollings, the network fault node and / or equipment fault node of the cabinet can be quickly found.

[0033] In an optional embodiment, before the first DCS controller sends the forward polling scan command to the multiple redundant cabinets, the method further includes:

[0034] The basic DCS controller sends a second preset polling command to the plurality of basic cabinets at a second frequency, receives a network status fed back by each of the basic cabinets, determines whether the network status fed back by each of the basic cabinets is the same as its corresponding expected status, determines a basic cabinet with a status different from the expected status as a faulty basic cabinet, and determines a network connection connected to the basic cabinet with a status different from the expected status as a network fault node of the basic cabinet;

[0035] The basic DCS controller is connected to a plurality of basic cabinets via a communication network. The basic cabinets are connected in a bus connection or a star connection. The basic cabinets include Profibus DP models.

[0036] and / or,

[0037] The second frequency is 100ms / time to 2s / time.

[0038] In terms of timing, fault diagnosis for the basic cabinet precedes fault diagnosis for the redundant cabinet, as the redundant cabinet replacement is initiated only after a basic cabinet failure occurs. If a fault occurs between basic cabinets, there's no automatic failover mechanism, and each basic cabinet is connected to only one basic DCS controller. Therefore, the basic DCS controller only needs to send a second, preset polling command to the basic cabinet. Based on the network status feedback from the basic cabinet, it can comprehensively, accurately, and efficiently determine the network fault node and / or equipment fault node in each basic cabinet. This allows equipment maintenance engineers to precisely maintain the workshop production system, improving system stability and eliminating the need for individual troubleshooting, thereby increasing maintenance efficiency and saving costs. After a basic cabinet and its connected network fail and are replaced with a redundant cabinet, production process engineers and equipment operators promptly visit the production site to assess the impact of the replacement redundant cabinet on production line operations and product quality. The short query interval facilitates expedited resolution. This short query interval facilitates expedited identification of the faulty basic cabinet and prompts timely resolution.

[0039] In an optional embodiment, the method further includes:

[0040] The first OS server and the second OS server respectively send a first detection signal to the DCS controller to be detected according to a third frequency, and receive a first response signal fed back by the DCS controller to be detected; wherein the first OS server and the second OS server are redundant OS servers to each other, and the first OS server and the second OS server are connected to the DCS controller to be detected via a communication network; the DCS controller to be detected includes the first DCS controller, the second DCS controller, and the basic DCS controller;

[0041] and / or,

[0042] The third frequency is 5s / time to 20s / time.

[0043] The first OS server and the second OS server are redundant OS servers. When one OS server fails, the other OS server can continue to be used without affecting the normal operation of the production control system.

[0044] In an optional implementation manner, the method for determining network and device faults of a DCS controller is as follows:

[0045] If the first OS server and the second OS server do not receive the first response signal, the DCS controller that does not feedback the first response signal is determined as a DCS controller device failure node, and the connection network between the DCS controller that does not feedback the first response signal and the first OS server or the second OS server is determined as a network failure node, and information about the network failure node and / or device failure node of the DCS controller is generated.

[0046] The DCS controller is a key connection node between the OS server and the cabinet. Network and equipment failures can significantly impact the ability of production process engineers and equipment operators to understand the operational status of on-site production equipment and issue operational instructions such as recipes. This application proactively polls the base, first, and second DCS controllers periodically (periodically sending a first detection signal) to quickly identify fault points in the relevant DCS controllers.

[0047] Setting the third frequency can effectively strike a balance between real-time performance and system load, enabling timely acquisition of status information without placing excessive pressure on system resources. Shortening the query interval to 5s / time - 20s / time does not affect the normal operation of the equipment, while facilitating quick troubleshooting, repair, and recovery.

[0048] In an optional embodiment, the method further includes:

[0049] The first OS server and the second OS server respectively send a second detection signal to each OS client according to a fourth frequency, and receive a second response signal fed back by each OS client;

[0050] If the first OS server and the second OS server do not receive the second response signal, the OS client that does not feed back the second response signal is determined as a device failure node of the OS client, and the connection network between the OS client that does not feed back the second response signal and the first OS server or the second OS server is determined as a network failure node, and information about the network failure node and / or device failure node of the OS client is generated;

[0051] and / or,

[0052] The fourth frequency is 5s / time to 20s / time.

[0053] The OS server is a key connection point between the OS client and the DCS controller. Network and device failures can significantly impact the ability of production process engineers and equipment operators to understand the operational status of on-site production equipment and issue operational instructions such as recipes. This application proactively polls the first and second OS servers periodically (periodically sending a second detection signal) to quickly identify the fault point on the relevant OS server.

[0054] When assessing the risk of network and device failure on the OS server, shorten the query interval to 5s / time to 20s / time. This will not affect the normal operation of the device and will facilitate rapid fault detection and repair and recovery.

[0055] In an optional embodiment, the method further includes:

[0056] The first DCS controller, the second DCS controller, and the basic DCS controller respectively upload information about the network fault nodes and / or device fault nodes of each redundant cabinet and information about the device fault nodes and / or network fault nodes of the basic cabinet to the first OS server and the second OS server simultaneously;

[0057] The first OS server or the second OS server respectively uploads information of the network fault node and / or the device fault node of the redundant cabinet to each OS client at the same time, and generates a first alarm message;

[0058] The first OS server or the second OS server uploads information about the network fault node and / or the device fault node of the basic cabinet to each OS client at the same time, and generates a second alarm message;

[0059] The first OS server or the second OS server uploads information of the network fault node and / or the device fault node of the DCS controller to each OS client at the same time, and generates a third alarm message;

[0060] The first OS server or the second OS server respectively uploads information about the network fault node and / or device fault node of the OS client to other OS clients operating normally, and generates a fourth alarm message;

[0061] Each of the OS clients displays the first alarm message, the second alarm message, the third alarm message, or the fourth alarm message.

[0062] This application displays information about network fault nodes and / or equipment fault nodes in the form of an interface, which is operator-friendly and convenient for production process engineers and equipment operators to quickly locate the fault point; if an important equipment network fails, the relevant program will be automatically called and an audible and visual alarm will be issued to remind relevant staff to prepare in advance.

[0063] This application can comprehensively monitor the status of OS servers, OS clients, DCS controllers, Profinet and Profibus DP devices, and can monitor in real time from the field control layer to the system management level without missing any network and device failure points.

[0064] In an optional embodiment, the invention is applied to the field of battery material manufacturing;

[0065] The redundant cabinet is used to control battery material production equipment, which includes at least one of a synthesis reactor, monitoring equipment, washing and drying equipment, crushing equipment, sintering equipment and packaging equipment.

[0066] The production process of battery materials, especially the wet synthesis stage of battery positive electrode material precursors, is particularly sensitive to parameters such as pH value, ammonia concentration, oxygen content of the atmosphere in the reactor, reaction temperature, and stirring speed during the reaction process. If the pH value fluctuates by more than 0.2, it may have a great impact on the physical and chemical performance indicators (specific surface area BET, tap density TD, bulk density TD, average pore size, valence states of various metals in the product) and morphology (surface morphology and cross-sectional morphology of secondary particles in scanning electron microscope photos) of the final product, which can easily lead to failure of the synthesis reaction. Unqualified materials need to be re-dissolved (reworked and dissolved and then separated into different metal solutions). Therefore, it is necessary to accurately control the feed flow of mixed metal salt solution, precipitant solution (sodium hydroxide, sodium carbonate, etc.), and complexing agent solution (ammonia water, etc.) and make real-time adjustments.

[0067] When a failure or abnormality occurs in the production system, this application can quickly identify the production nodes that may affect product quality, conduct timely investigations, and handle them as soon as possible, which is conducive to maintaining stable operation of the production line, reducing risks, and ensuring product quality.

[0068] Beneficial effects of this application:

[0069] The embodiment of the present application provides a method for detecting network and equipment faults in a production control system. This application can achieve comprehensive, accurate and efficient detection of network and equipment faults in a production control system, allowing equipment maintenance engineers to promptly perform maintenance, replacement and upkeep on potential hazards, thereby steadily improving the stability of system operation.

[0070] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be considered as limiting the scope. A person of ordinary skill in the art can also derive other relevant drawings based on these drawings without inventive effort. Similar components are numbered similarly in the various drawings.

[0072] Figure 1 A flowchart of a method for detecting network and equipment faults in a production control system provided by an embodiment of the present application is shown;

[0073] Figure 2 A topological diagram of a network structure provided by an embodiment of the present application is shown;

[0074] Figure 3 A topological diagram of another network structure provided in an embodiment of the present application is shown;

[0075] Figure 4 A topological diagram of another network structure provided in an embodiment of the present application is shown;

[0076] Figure 5 A topology diagram of another network structure provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0077] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0078] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0079] Example 1

[0080] like Figure 1 FIG. 1 is a flow chart of a method for detecting network and equipment faults in a production control system according to an embodiment of the present application. The method for detecting network and equipment faults provided by the embodiment of the present application includes:

[0081] Step S110: The first DCS controller sends a forward polling scan command to a plurality of redundant cabinets, and each of the redundant cabinets feeds back a forward polling status to the first DCS controller according to a first frequency;

[0082] Step S120: The second DCS controller sends a reverse polling scan command to the plurality of redundant cabinets, and each redundant cabinet feeds back a reverse polling status to the second DCS controller according to the first frequency;

[0083] Step S130: The first DCS controller and the second DCS controller determine the network fault node and / or device fault node of each redundant cabinet according to the forward polling state and the reverse polling state;

[0084] The model of the redundant cabinet includes Profinet, the redundant DCS controller includes the first DCS controller and the second DCS controller, the first DCS controller, multiple redundant cabinets and the second DCS controller form a cascade topology through a communication network, and the multiple redundant cabinets are located between the first DCS controller and the second DCS controller.

[0085] Redundant cabinets Profinet has been widely used in the field of industrial automation. Profinet is an Ethernet-based industrial communication standard suitable for high-speed data transmission and real-time control.

[0086] Understandably, if Figure 2 As shown, each redundant cabinet is connected to the network in sequence. Since there is an automatic switching mechanism if a fault occurs between the redundant cabinets, the mechanical equipment controlled by the faulty redundant cabinet is in a normal operating state. Since each redundant cabinet is connected to two DCS controllers, two DCS controllers are required to perform forward polling and reverse polling on each redundant cabinet respectively. Only after two pollings can the device fault node and / or network fault node of the redundant cabinet be determined. A first preset polling command is sent to the redundant cabinet by the redundant DCS controller to obtain the network status of each redundant cabinet, thereby providing a basis for subsequent fault detection. The polling mechanism is a commonly used monitoring method in industrial control systems. By sending polling commands regularly, the controller can timely understand the operating status of on-site equipment.

[0087] Specifically, in this embodiment, the first DCS controller sends a forward polling scan command to each redundant cabinet and receives forward polling status periodically fed back from each redundant cabinet at a first frequency. The second DCS controller sends a reverse polling scan command to each redundant cabinet and receives reverse polling status periodically fed back from each redundant cabinet at a first frequency. The first frequency is 100 ms / time to 2 s / time. The specific time can be determined based on actual conditions and is not limited in this embodiment.

[0088] It can be understood that during forward polling scanning, the first redundant cabinet connected to the first DCS controller is polled backward in sequence to the last redundant cabinet or the last redundant cabinet that feeds back the forward polling status. During reverse polling scanning, the last redundant cabinet connected to the second DCS controller is polled forward in sequence to the first redundant cabinet or the first redundant cabinet that feeds back the reverse polling status.

[0089] Exemplarily, each redundant cabinet has an independent ID number, and the redundant cabinets are numbered according to the ID number. For example, there are 8 redundant cabinets numbered 1-8, and the connection order is: first DCS controller-redundant cabinet 1-redundant cabinet 2-redundant cabinet 3-redundant cabinet 4-redundant cabinet 5-redundant cabinet 6-redundant cabinet 7-redundant cabinet 8-second DCS controller;

[0090] During the forward polling scan, the first redundant cabinet connected to the first DCS controller is redundant cabinet 1; if no signal can be fed back starting from redundant cabinet 5, the last redundant cabinet to feed back the forward polling status is redundant cabinet 4;

[0091] During reverse polling scanning, the last redundant cabinet connected to the second DCS controller is redundant cabinet 8; if redundant cabinet 4 fails to feedback a signal, the first redundant cabinet to feedback the reverse polling status is 5.

[0092] During forward polling scanning, the scanning order is: redundant cabinet 1 - redundant cabinet 2 - redundant cabinet 3 - redundant cabinet 4 - redundant cabinet 5 - redundant cabinet 6 - redundant cabinet 7 - redundant cabinet 8. After the first DCS controller sends a forward polling command, redundant cabinet 1, connected to the first DCS controller, receives the forward polling command and inserts its own forward polling status into the forward polling command before passing it to redundant cabinet 2. It also feeds back its own forward polling status to the first DCS controller. Redundant cabinet 2 receives the forward polling command and cabinet 1's forward polling status, inserts its own forward polling status into the forward polling command, then passes it to redundant cabinet 3. It also feeds back the forward polling status of redundant cabinet 1 and its own forward polling status to the first DCS controller. This process repeats until the forward polling command and the forward polling status feedback from each node can no longer be passed to the next node.

[0093] The reverse polling scan is similar to the forward polling scan, except that the scanning order is changed to: redundant cabinet 8 - redundant cabinet 7 - redundant cabinet 6 - redundant cabinet 5 - redundant cabinet 4 - redundant cabinet 3 - redundant cabinet 2 - redundant cabinet 1.

[0094] The combined use of forward and reverse polling scans provides more comprehensive coverage of all redundant cabinets, improving the accuracy and efficiency of status monitoring. Forward polling scans proceed from the first cabinet backward, while reverse polling scans proceed from the last cabinet forward. This design effectively avoids omissions or delays that can occur with one-way polling. Setting the first frequency effectively strikes a balance between real-time performance and system load, ensuring timely status information acquisition without placing excessive pressure on system resources.

[0095] After obtaining the forward polling status and the reverse polling status of each redundant cabinet through bidirectional polling scanning, in an optional embodiment, a method for determining the network fault node and / or the device fault node of each redundant cabinet based on the first candidate fault cabinet set and the second candidate fault cabinet set is as follows:

[0096] The redundant cabinets are numbered: starting from 1 and numbering the first redundant cabinet after the first DCS controller to the last redundant cabinet;

[0097] Determine whether the first candidate faulty cabinet set and the second candidate faulty cabinet set have the same redundant cabinet number;

[0098] Case 1: If there is no identical redundant cabinet number, and both the first candidate fault cabinet set and the second candidate fault cabinet set are not empty, then obtain the starting fault cabinet number in the first candidate fault cabinet set, obtain the ending fault cabinet number in the second candidate fault cabinet set, and determine the connection network between the starting fault cabinet in the first candidate fault cabinet set and the ending fault cabinet in the second candidate fault cabinet set as the network fault node of the redundant device;

[0099] Second case: if the same redundant cabinet exists and both the first candidate fault cabinet set and the second candidate fault cabinet set are not empty, then obtain the starting fault cabinet number in the first candidate fault cabinet set, obtain the ending fault cabinet number in the second candidate fault cabinet set, and determine the network connection before the starting fault cabinet of the first candidate fault cabinet set, the network connection between the starting cabinet of the first candidate fault cabinet set and the ending cabinet of the second candidate fault cabinet set, and the network connection after the ending cabinet of the second candidate fault cabinet set as network fault nodes; and determine the starting fault cabinet, the ending fault cabinet and the redundant cabinets between them as device fault nodes;

[0100] Case 3: If the first candidate fault cabinet set is empty, the connection network between the second DCS controller and the last redundant cabinet is determined as the network fault node;

[0101] Fourth case: if the second candidate fault cabinet set is empty, the connection network between the first DCS controller and the first redundant cabinet is determined as a network fault node;

[0102] Fifth case: if the first candidate faulty cabinet set and the second candidate faulty cabinet set are both empty, then there is no network or equipment failure between the first DCS controller and the second DCS controller.

[0103] It is understandable that this embodiment can determine the network fault node and / or device fault node of each redundant cabinet by determining whether the first candidate fault cabinet set and the second candidate fault cabinet set have the same redundant cabinet.

[0104] In the first case, illustratively, if the forward polling status fed back after the forward polling scan does not include the forward polling status fed back by redundant cabinet 5, redundant cabinet 6, redundant cabinet 7, and redundant cabinet 8, the obtained first candidate fault cabinet set is numbered (5, 6, 7, 8); the reverse polling status fed back after the reverse polling scan does not include the reverse polling status fed back by redundant cabinet 1, redundant cabinet 2, redundant cabinet 3, and redundant cabinet 4, the obtained second candidate fault cabinet set is numbered (1, 2, 3, 4). At this time, the first candidate fault cabinet set and the second candidate fault cabinet set do not have the same redundant cabinet number, which proves that there is a fault in the connection network between the starting fault cabinet 5 in the first candidate fault cabinet set and the ending fault cabinet 4 in the second candidate fault cabinet set. Therefore, the connection network at this location is regarded as a network fault node.

[0105] In the second case, when the first candidate fault cabinet set and the second candidate fault cabinet set contain two or more identical redundant cabinets, for example, if the first candidate fault cabinet set obtained after the forward polling scan is numbered (4, 5, 6, 7, 8), and the second candidate fault cabinet set obtained after the reverse polling scan is numbered (1, 2, 3, 4, 5), then the first candidate fault cabinet set and the second candidate fault cabinet set contain the same redundant cabinet 4 and redundant cabinet 5, then redundant cabinet 4 and redundant cabinet 5 may have equipment failures and are in a disconnected state, so the redundant cabinets are removed. Redundant cabinet 4 and redundant cabinet 5 are regarded as equipment failure nodes. At the same time, a network connection before redundant cabinet 4 (the network connection between redundant cabinet 3 and redundant cabinet 4), a network connection after redundant cabinet 5 (the network connection between redundant cabinet 5 and redundant cabinet 6), and the connection network between redundant cabinets 4 and 5 may also have failures. The network connection before redundant cabinet 4 (the network connection between redundant cabinet 3 and redundant cabinet 4) and the network connection after redundant cabinet 5 (the network connection between redundant cabinet 5 and redundant cabinet 6) are regarded as network failure nodes.

[0106] In the second case, when the first candidate fault cabinet set and the second candidate fault cabinet set only contain one identical redundant cabinet, for example, if the first candidate fault cabinet set obtained after the forward polling scan is numbered (5, 6, 7, 8), and the second candidate fault cabinet set obtained after the reverse polling scan is numbered (1, 2, 3, 4, 5), then the first candidate fault cabinet set and the second candidate fault cabinet set contain the same redundant cabinet 5, which proves that redundant cabinet 5 has a device failure and is in a disconnected state. Therefore, redundant cabinet 5 is regarded as a device failure node, and the two network connections adjacent to redundant cabinet 5 (the network connection between redundant cabinet 4 and redundant cabinet 6) are regarded as network failure nodes.

[0107] In the third case, if the number of the first candidate fault cabinet set obtained after the forward polling scan is empty; illustratively, the number of the second candidate fault cabinet set obtained after the reverse polling scan includes all redundant cabinet numbers, which are (1, 2, 3, 4, 5, 6, 7, 8), then the connection network between the second DCS controller and the last redundant cabinet is determined to be the network fault node of the redundant device;

[0108] In the fourth case, if the second candidate fault cabinet set obtained after the reverse polling scan is empty, illustratively, if the numbers of the first candidate fault cabinet set obtained after the forward polling scan include all redundant cabinet numbers, namely (1, 2, 3, 4, 5, 6, 7, 8); the numbers of the second candidate fault cabinet set obtained after the reverse polling scan are empty, then the connection network between the first DCS controller and the first redundant cabinet is determined as the network fault node of the redundant device;

[0109] In the fifth case, if the first candidate faulty cabinet set and the second candidate faulty cabinet set are both empty, it means that the network connection and equipment between the first DCS controller and the second DCS controller are normal, and there is no network or equipment failure between the first DCS controller and the second DCS controller.

[0110] The above steps utilize a bidirectional verification mechanism to avoid false alarms caused by single-point misjudgment. In the event of a device failure, adjacent connected networks are simultaneously checked to prevent the spread of the fault. This allows for rapid and accurate localization of network and device failures within redundant cabinets, providing targeted detection solutions for different types of faults and improving overall system reliability.

[0111] Understandably, in this embodiment, before determining the fault of the redundant cabinet, it is also necessary to determine the fault of the basic cabinet. The models of the basic cabinet include Profibus DP, which is a mature and reliable fieldbus standard suitable for connecting distributed I / O devices. Figure 3 and Figure 4 As shown, the basic DCS controller is connected to multiple basic cabinets through a communication network. Figure 3 The connection between the basic cabinets is star connection. Figure 4 The basic cabinets are connected via a bus. Since there's no automatic failover mechanism if a fault occurs between the basic cabinets, and each basic cabinet is connected to only one basic DCS controller, the basic DCS controller only needs to poll each basic cabinet once to determine the device fault node and / or network fault node in that basic cabinet.

[0112] First, the basic DCS controller sends a second preset polling command to each basic cabinet at a second frequency and receives network status feedback from each basic cabinet. Next, the basic DCS controller obtains the expected status of each basic cabinet and determines whether the network status feedback from each basic cabinet is consistent with its corresponding expected status. If it is different from the expected status, it indicates that the corresponding basic cabinet has failed. Therefore, the basic cabinet with a different expected status is determined as a failed basic cabinet, and the network connection connected to the basic cabinet with a different expected status is determined as the network failure node of the basic cabinet. The second frequency is 100ms / time to 2s / time. The specific time can be determined based on actual conditions and is not limited in this embodiment of the application.

[0113] Further, if Figure 4 As shown, if the connection between the basic cabinets is a star connection, each faulty basic cabinet is treated as one of the basic cabinet's multiple equipment fault nodes, and the connection network between each faulty basic cabinet is treated as the basic cabinet's network fault node. For example, if there are five basic cabinets numbered 1-5, and the faulty basic cabinets obtained after polling scanning are numbered 3, 4, and 5, then the faulty basic cabinets 3, 4, and 5 are all treated as the basic cabinet's equipment fault nodes, and the faulty basic cabinets 3, 4, and 5, as well as the connection network between basic cabinet 2 and basic cabinet 3, are also treated as the basic cabinet's network fault nodes.

[0114] like Figure 5As shown, if the connection between the basic cabinets is a bus connection, the first faulty basic cabinet among the faulty basic cabinets is used as the basic cabinet's equipment fault node, and the connection network before the first faulty basic cabinet is used as the basic cabinet's network fault node according to the connection sequence. For example, there are five basic cabinets numbered 1-5. If the faulty basic cabinets obtained after a polling scan from basic cabinets 1 to 5 are numbered 3, 4, and 5, the first faulty basic cabinet 3 is used as the basic cabinet's equipment fault node, and the connection network before the first faulty basic cabinet 3 is used as the basic cabinet's network fault node according to the connection sequence.

[0115] It should be noted that in this embodiment, the basic cabinet and the redundant cabinet correspond to cabinets of different batches and models, respectively, and the configurations of the basic DCS controller and the redundant DCS controller are also different. The specific settings can be determined according to actual conditions, and this embodiment of the application does not limit or elaborate on this.

[0116] In the above steps, fault diagnosis for the base cabinet is independent of that for the redundant cabinets, which prevents mutual interference and improves fault detection accuracy. The normal operation of the base cabinet provides stable support for the entire system.

[0117] In an optional embodiment, as Figure 5 As shown in FIG, it is a general diagram including a basic cabinet and a redundant cabinet. In this embodiment, there is also a first OS (Operator Station) server, a second OS server, and multiple OS clients. The first OS server and the second OS server are redundant OS servers. The first OS server and the second OS server are connected to the first DCS controller and the second DCS controller respectively through a communication network. Figure 5 As shown, the first OS server and the second OS server are also connected to the basic DCS controller via a communication network. Each OS server is connected to multiple OS clients via a communication network, and each OS client is connected to a ring network.

[0118] In an optional embodiment, the first OS server and the second OS server also poll and scan for fault statuses of DCS controllers to be detected. Specifically, the first OS server and the second OS server each send a first detection signal (e.g., a heartbeat detection signal) to the DCS controller to be detected at a third frequency, and receive a first response signal fed back by the DCS controller to be detected; the DCS controllers to be detected include the first DCS controller, the second DCS controller, and the basic DCS controller.

[0119] If the first response signal includes an unfed first signal, the DCS controller corresponding to the unfed first signal is determined as the device failure node of the DCS controller, and the network connecting the failed DCS controller and the first OS server or the second OS server is determined as the network failure node of the DCS controller, and information about the network failure node and / or device failure node of the DCS controller is generated. The third frequency is 5s / time to 20s / time, and the specific time can be determined based on actual conditions and is not limited in this embodiment of the application.

[0120] The above steps ensure the normal operation of each DCS controller by regularly polling its status, promptly detecting controller failures and generating fault information. By monitoring each DCS controller, the reliability of the entire system is enhanced, system paralysis caused by DCS controller failure is avoided, maintenance personnel can quickly respond to controller failures, and system downtime is reduced.

[0121] In an optional embodiment, the first OS server and the second OS server further poll and scan the fault status of themselves and each OS client. Specifically, the first OS server and the second OS server each send a second detection signal (e.g., a heartbeat detection signal) to the OS client to be detected at a fourth frequency, and receive a second response signal fed back by each OS client.

[0122] If the second response signal includes an unfed second signal, the OS client corresponding to the unfed second signal is determined as the device failure node of the OS client, the network connecting the OS client and the first OS server or the second OS server is determined as the network failure node of the OS client, and information about the network failure node and / or device failure node of the OS client is generated. The fourth frequency is 5s / time to 20s / time, and the specific time can be determined based on actual conditions and is not limited in this embodiment of the application.

[0123] In the above steps, the first OS server and the second OS server perform self-monitoring on each OS client, thereby ensuring the reliability of the monitoring system. Through self-monitoring and mutual monitoring, the robustness of the system is improved, ensuring the stable operation of the monitoring system.

[0124] In an optional embodiment,

[0125] The first DCS controller, the second DCS controller, and the basic DCS controller respectively upload information about network fault nodes and / or equipment fault nodes of each redundant cabinet and information about equipment fault nodes and / or network fault nodes of the basic cabinet to the first OS server and the second OS server simultaneously;

[0126] The first OS server or the second OS server uploads information about the network fault node and / or device fault node of the redundant cabinet to each OS client at the same time, and generates a first alarm message;

[0127] The first OS server or the second OS server uploads information about the network fault node and / or the device fault node of the basic cabinet to each OS client at the same time, and generates a second alarm message;

[0128] The first OS server or the second OS server uploads information of the network fault node and / or the device fault node of the DCS controller to each OS client at the same time, and generates a third alarm message;

[0129] The first OS server or the second OS server uploads information about the network fault node and / or device fault node of the OS client to other normally operating OS clients simultaneously, and generates a fourth alarm message;

[0130] Each OS client displays the first alarm message, the second alarm message, the third alarm message, or the fourth alarm message.

[0131] The above steps involve a hierarchical upload of fault data: DCS controller → OS server → OS client. The OS server, acting as a central node, centrally collects and manages all fault information, providing a unified monitoring and management platform. The OS server then forwards this information to the OS client, ensuring timely information sharing and transparency. The OS client, connected via the network, enables remote monitoring and fault checking, facilitating the daily work of maintenance personnel.

[0132] Finally, the equipment maintenance engineer can view the screen and alarm information on the display screen of the OS client. Preferably, the OS client displays the information of the network fault node and / or device fault node of each redundant cabinet and the information of the device fault node and / or network fault node of each basic cabinet through a first mark (e.g., a red mark), and generates a first alarm message; the OS client displays the information of the device fault node and / or network fault node of each basic cabinet through a second mark (e.g., an orange mark), and generates a second alarm message; the OS client displays the fault information of the DCS controller through a third mark (e.g., a blue mark), and generates a third alarm message; the OS client displays the fault information of the first OS server and the fault information of the second OS server through a fourth mark (e.g., a yellow mark), and generates a fourth alarm message. The equipment maintenance engineer can find the corresponding machine according to the computer name, which is usually shut down, frozen, or with network line interference.

[0133] If a network fault node and / or device fault node line in the above-mentioned different devices and lines reported a fault after previous detection but does not report a fault after current detection, it will be marked with another color (such as gray) and possible fault information will be generated. The reset will be confirmed after the equipment maintenance engineer has inspected it.

[0134] The entire fault detection system continuously scans for network and device faults within OS servers, OS clients, DCS controllers, redundant cabinets, and basic cabinets. From the field control layer to the system management level, it monitors every control network failure, refreshing the OS client's monitoring screen in real time and archiving alarm messages. When network or device failures occur, the fault point is identified on the screen, making it user-friendly and enabling maintenance personnel to quickly locate the fault. In the event of a critical device network failure, the system automatically invokes relevant programs, issuing audible and visual alarms to alert personnel to prepare in advance.

[0135] In addition, it should be noted that the network and equipment fault detection method provided in the embodiment of the present application is applied to the field of battery material manufacturing, and the redundant cabinets and basic cabinets are used to control battery material production equipment, and the production equipment includes at least one of synthesis reactors, monitoring equipment, washing and drying equipment, crushing equipment, sintering equipment or packaging equipment.

[0136] In a network and equipment fault detection method provided by an embodiment of the present application, a redundant DCS controller is connected to multiple redundant cabinets via a communication network. The redundant DCS controller sends a first preset polling command to the multiple redundant cabinets and receives network status feedback from each redundant cabinet. Based on the network status feedback from each redundant cabinet, the redundant DCS controller determines the network fault node and / or equipment fault node of each redundant cabinet. This application enables comprehensive, accurate, and efficient troubleshooting and diagnosis, allowing equipment maintenance engineers to promptly perform maintenance and replacement on vulnerable areas, thereby significantly improving system operational stability.

[0137] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.

Claims

1. A method for detecting network and equipment faults in a production control system, characterized in that: The network and equipment fault detection method includes: The first DCS controller sends a forward polling scan command to a plurality of redundant cabinets, and each of the redundant cabinets feeds back a forward polling status to the first DCS controller according to a first frequency; The second DCS controller sends a reverse polling scan command to the plurality of redundant cabinets, and each redundant cabinet feeds back a reverse polling status to the second DCS controller according to the first frequency; The first DCS controller and the second DCS controller determine the network fault node and / or device fault node of each redundant cabinet according to the forward polling state and the reverse polling state; The model of the redundant cabinet includes Profinet, the first DCS controller, the multiple redundant cabinets and the second DCS controller form a cascade topology through a communication network, and the multiple redundant cabinets are located between the first DCS controller and the second DCS controller.

2. The network and device fault detection method according to claim 1, characterized in that: The plurality of redundant cabinets are connected to the network in sequence; During forward polling scanning, the system starts from the first redundant cabinet connected to the first DCS controller and polls backward in sequence to the last redundant cabinet or the last redundant cabinet that feeds back the forward polling status; During reverse polling scanning, the system starts from the last redundant cabinet connected to the second DCS controller and polls forward in sequence to the first redundant cabinet or the first redundant cabinet that feeds back the reverse polling status.

3. The network and device fault detection method according to claim 2, characterized in that: The network and equipment fault detection method includes: The first DCS controller obtains the expected state of each redundant cabinet, and determines whether the forward polling state of each redundant cabinet is the same as its corresponding expected state, and selects redundant cabinets with different expected states and / or redundant cabinets without forward polling state feedback as a first candidate fault cabinet set; The second DCS controller obtains the expected state of each redundant cabinet, and determines whether the reverse polling state of each redundant cabinet is the same as its corresponding expected state, and selects redundant cabinets with different expected states and / or redundant cabinets without reverse polling state feedback as the second candidate fault cabinet set; Determine the network fault node and / or device fault node of each redundant cabinet according to the first candidate fault cabinet set and the second candidate fault cabinet set; and / or, The first frequency is 100ms / time to 2s / time.

4. The network and device fault detection method according to claim 3, characterized in that: The method for determining the network fault node and / or device fault node of each redundant cabinet according to the first candidate fault cabinet set and the second candidate fault cabinet set is: The redundant cabinets are numbered: starting from 1 and numbering the first redundant cabinet after the first DCS controller to the last redundant cabinet; Determine whether the first candidate faulty cabinet set and the second candidate faulty cabinet set have the same redundant cabinet number; If there is no identical redundant cabinet number, and both the first candidate fault cabinet set and the second candidate fault cabinet set are not empty, obtaining the starting fault cabinet number in the first candidate fault cabinet set, obtaining the ending fault cabinet number in the second candidate fault cabinet set, and determining the connection network between the starting fault cabinet in the first candidate fault cabinet set and the ending fault cabinet in the second candidate fault cabinet set as a network fault node of the redundant device; If the same redundant cabinet exists and both the first candidate fault cabinet set and the second candidate fault cabinet set are not empty, obtain the starting fault cabinet number in the first candidate fault cabinet set, obtain the ending fault cabinet number in the second candidate fault cabinet set, and determine one network connection before the starting fault cabinet of the first candidate fault cabinet set, the network connection between the starting cabinet of the first candidate fault cabinet set and the ending cabinet of the second candidate fault cabinet set, and one network connection after the ending cabinet of the second candidate fault cabinet set as network fault nodes; and determine the starting fault cabinet, the ending fault cabinet, and the redundant cabinets between them as device fault nodes; If the first candidate fault cabinet set is empty, determining the connection network between the second DCS controller and the last redundant cabinet as a network fault node; If the second candidate fault cabinet set is empty, determining the connection network between the first DCS controller and the first redundant cabinet as a network fault node; If both the first candidate faulty cabinet set and the second candidate faulty cabinet set are empty, then there is no network or equipment failure between the first DCS controller and the second DCS controller.

5. The network and device fault detection method according to claim 2, characterized in that: Before the first DCS controller sends the forward polling scan command to the multiple redundant cabinets, the method further includes: The basic DCS controller sends a second preset polling command to the plurality of basic cabinets at a second frequency, receives a network status fed back by each of the basic cabinets, determines whether the network status fed back by each of the basic cabinets is the same as its corresponding expected status, determines a basic cabinet with a status different from the expected status or a basic cabinet that has not fed back a network status as a faulty basic cabinet, and determines a network connection connected to the basic cabinet with a status different from the expected status or the basic cabinet that has not fed back a network status as a network fault node of the basic cabinet; The basic DCS controller is connected to a plurality of basic cabinets via a communication network. The basic cabinets are connected in a bus connection or a star connection. The basic cabinets include Profibus DP models. and / or, The second frequency is 100ms / time to 2s / time.

6. The network and device fault detection method according to claim 5, characterized in that: Also includes: The first OS server and the second OS server respectively send a first detection signal to the DCS controller to be detected according to a third frequency, and receive a first response signal fed back by the DCS controller to be detected; wherein the first OS server and the second OS server are redundant OS servers to each other, and the first OS server and the second OS server are connected to the DCS controller to be detected via a communication network; the DCS controller to be detected includes the first DCS controller, the second DCS controller, and the basic DCS controller; and / or, The third frequency is 5s / time to 20s / time.

7. The network and device fault detection method according to claim 6, characterized in that: The method for judging network and equipment faults of DCS controller is as follows: If the first OS server and the second OS server do not receive the first response signal, the DCS controller that does not feedback the first response signal is determined as a DCS controller device failure node, and the connection network between the DCS controller that does not feedback the first response signal and the first OS server or the second OS server is determined as a network failure node, and information about the network failure node and / or device failure node of the DCS controller is generated.

8. The network and device fault detection method according to claim 6, characterized in that: Also includes: The first OS server and the second OS server respectively send a second detection signal to each OS client according to a fourth frequency, and receive a second response signal fed back by each OS client; If the first OS server and the second OS server do not receive the second response signal, the OS client that does not feed back the second response signal is determined as a device failure node of the OS client, and the connection network between the OS client that does not feed back the second response signal and the first OS server or the second OS server is determined as a network failure node, and information about the network failure node and / or device failure node of the OS client is generated; and / or, The fourth frequency is 5s / time to 20s / time.

9. The network and device fault detection method according to claim 8, characterized in that: Also includes: The first DCS controller, the second DCS controller, and the basic DCS controller respectively upload information about the network fault nodes and / or device fault nodes of each redundant cabinet and information about the device fault nodes and / or network fault nodes of the basic cabinet to the first OS server and the second OS server simultaneously; The first OS server or the second OS server respectively uploads information of the network fault node and / or the device fault node of the redundant cabinet to each OS client at the same time, and generates a first alarm message; The first OS server or the second OS server uploads information about the network fault node and / or the device fault node of the basic cabinet to each OS client at the same time, and generates a second alarm message; The first OS server or the second OS server uploads information of the network fault node and / or the device fault node of the DCS controller to each OS client at the same time, and generates a third alarm message; The first OS server or the second OS server respectively uploads information about the network fault node and / or device fault node of the OS client to other OS clients operating normally, and generates a fourth alarm message; Each of the OS clients displays the first alarm message, the second alarm message, the third alarm message, or the fourth alarm message.

10. The network and device fault detection method according to any one of claims 1 to 9, characterized in that: Applied in the field of battery material manufacturing; The redundant cabinet is used to control battery material production equipment, which includes at least one of a synthesis reactor, monitoring equipment, washing and drying equipment, crushing equipment, sintering equipment and packaging equipment.