Fault control system and method in complex network environment
By introducing NIPU and the first judgment module into the subnodes of complex network systems, combining multi-agent learning algorithms, differentiating fault types and optimizing processing flow, the problem of inefficiency of traditional fault processing technology in complex networks is solved, and efficient and economical fault processing and system flexibility are achieved.
Patent Information
- Application Number
- CN202510670706.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-08
AI Technical Summary
Traditional fault processing technology is difficult to adapt to the diverse characteristics of complex networks. Software-oriented scheme prediction accuracy is insufficient and it is easy to cause cascade failures. The hardware-oriented scheme processing process is mechanized and costly, resulting in inefficient system decision-making strategies.
The node intelligent processing unit (NIPU) is introduced into the child nodes of complex network systems, and the time threshold and fault type judgment are combined with the first judgment module, and the fault type is judged, and the fault type is distinguished and different processing processes are adopted. Multi-agent reinforcement learning algorithm is used to update the coordinated strategy and optimize fault processing.
It realizes efficient and economical handling of faults in complex network environments, reduces cascading faults, improves system flexibility and processing efficiency, and reduces hardware costs.
Smart Images

Figure CN120455244A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of fault processing of complex network systems, and more particularly to a fault control system and method for complex networks. Background Art
[0002] With the rapid development of technologies such as the Internet of Things, Industrial Internet, and smart grids, modern network systems are characterized by massive scale, strong node heterogeneity, and complex dynamic interactions. Traditional fault handling technologies struggle to adapt to the diverse characteristics of complex networks.
[0003] From the perspective of general detection and diagnosis technologies, traditional fault resolution solutions can be broadly categorized as software-driven or hardware-driven. Software-driven solutions rely heavily on traffic analysis algorithms (such as deep learning models) and protocol parsing tools, tracing faults through large-scale data mining. Their accuracy for detecting configuration errors and traffic anomalies typically reaches 75%-85%. Hardware-driven solutions, on the other hand, rely heavily on deploying intelligent sensors to collect historical or relevant data signals from the physical layer in real time, integrating them with edge computing nodes for rapid response.
[0004] However, there are currently several problems that need to be addressed with this type of solution. Software-oriented solutions rely heavily on a variety of comprehensive factors, such as algorithm accuracy and data noise processing, for predictions or decision-making. Such judgments usually cannot avoid insufficient prediction accuracy and cascading failures (causing secondary failures) caused by prediction errors. Hardware-oriented solutions, although combining software and hardware brings some advantages, still have problems with the mechanization and inefficiency of the processing process and the sharp increase in overall costs, which seriously affect the system decision-making strategy. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to at least partially consider the above problems. In order to solve these problems, a fault control system and method in a complex network are proposed. For more details, please refer to the following:
[0006] One aspect of the present invention provides a fault control system in a complex network environment, comprising a master node device and multiple sub-node devices, characterized in that:
[0007] The sub-node device is equipped with a node intelligent processing unit NIPU, which has partial or complete master node link processing center functions;
[0008] The sub-node is provided with a first judgment module, which receives the fault data signal and performs time threshold judgment and fault type judgment; and initiates different fault handling processes according to the judgment result of the first judgment module.
[0009] Preferably, the fault data signal received by the first judgment module comes from a sensor module or a transceiver module in the sub-node device.
[0010] Preferably, the first judgment module performs a first time threshold judgment on the fault data. If the first time threshold limit is reached, the NIPU of the current node is directly triggered to initiate rapid fault processing; if the first time threshold limit is not reached, the specific type of the fault event is further evaluated.
[0011] Preferably, the first judgment module initiates a fault type identification judgment, and the fault type is identified as the first fault type or the second fault type.
[0012] Preferably, if the fault type identification result is the first fault type, the fault is reported to the master node link processing center for processing; during the master node processing process, if a time urgency indication is detected, the sub-node device NIPU is triggered to perform fault processing.
[0013] Preferably, if the fault type identification result is the second fault type, the current sub-node device NIPU is triggered to perform fault processing, and at the same time, at least one NIPU with complete master node link processing center function in the control system network is searched for joint processing.
[0014] Preferably, if it is detected during the second type of fault handling process that the fault handling time exceeds the second time threshold limit, the fault data sharing of the faulty sub-node is stopped or a fault handling request signal is sent to other nodes.
[0015] Preferably, if during the second type of fault handling process, it is detected that the system network generates a secondary fault indication based on the current fault, the fault data sharing of the faulty sub-node is stopped or a fault handling request signal is sent to other nodes.
[0016] Preferably, the number of NIPUs in the control system is not greater than the number of all sub-nodes, and the number of NIPUs with complete master node link processing center functions is greater than 1.
[0017] A second aspect of the present invention provides a method for fault control in a complex network environment, comprising the following steps:
[0018] Step 1: After receiving the fault data, the sub-node device judgment module of the network system initiates the judgment of the fault data;
[0019] Step 2: Determine the first time threshold for processing the current fault data. If the time threshold is reached, trigger the NIPU of the current node to initiate rapid fault processing. If not, proceed to the next step.
[0020] Step 3: Determine the specific fault type in the current fault event. If it is a Class I fault, report the fault to the master node link processing center for processing. During the master node processing, if a time urgency indication is detected, the sub-node device NIPU is triggered to perform fault processing.
[0021] Step 4: If the fault is of the second type, the current sub-node device NIPU is triggered to process the fault, and at the same time, at least one NIPU with complete master node link processing center functions in the control network system is searched for to perform joint processing;
[0022] Step 5: If, during the second type of fault handling process, it is detected that the fault handling time exceeds the second time threshold limit, and / or it is detected that the system network generates a secondary fault indication based on the current fault, then the fault data sharing of the faulty sub-node is stopped or a fault handling request signal is sent to other nodes.
[0023] The present invention discloses a fault network control system and method in a complex network environment. By setting an NIPU and a specific judgment module in the sub-nodes of the network system and combining the fault handling logic in a specific scenario, it effectively responds to the traditional single fault handling system and method, properly integrates the advantages of simple software and hardware processing processes, and brings good technical effects. The value of the solution of the present invention has been verified in actual application scenarios and has good technical prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 A simplified architecture diagram of the main nodes and sub-nodes of a control system in a complex network provided by an embodiment of the present invention;
[0025] Figure 2 A general processing flow framework for fault handling of a fault control system provided by an embodiment of the present invention;
[0026] Figure 3 This is an operational flow chart of the first judgment module performing time threshold processing in the fault control system provided by an embodiment of the present invention;
[0027] Figure 4 This is an operational flow chart of the first judgment module in the fault control system provided by an embodiment of the present invention performing fault type judgment processing. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0029] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0030] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0031] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer" and the like indicate positions or locations based on the positions shown in the accompanying drawings, or the positions or locations in which the inventive product is typically placed when in use. These terms are intended solely to facilitate the description of the present invention and to simplify the description, and are not intended to indicate or imply that the devices or components referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third," etc., are used solely to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0032] Furthermore, terms such as "horizontal" and "vertical" do not necessarily mean that a component must be absolutely horizontal or overhanging, but rather that it can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal than "vertical," and does not mean that the structure must be completely horizontal, but rather that it can be slightly tilted.
[0033] In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0034] Traditional network node failure problem handling usually chooses a single priority management-related method, and uses simple priority management to screen the specific actions of node processing. However, such actions usually focus on the single and complete recovery function of the node. For example, in CN109787795A, traditional priority identification technology is used to mechanically find the final node to operate as the master node. That is, at all times, although there can be multiple backup master nodes (the backup master node is only used for monitoring, not for specific operation instructions), there is always only one master node working, which greatly reduces the processing time of complex tasks. Whenever a failure occurs or the processing capacity decreases, the efficiency of the system will be reduced exponentially. Moreover, because the functions of the sub-node itself are exactly the same as those of the master node, in complex network scenarios (usually a complete master device is expensive and complex), if the sub-node is chosen to be exactly the same as the master node, the cost will be greatly increased, limiting the application promotion prospects.
[0035] Refer to the attached Figure 1 In the technical solution of the present invention, a simplified processing module with certain link processing center functions is embedded in the underlying sub-node controlled by the complex network system. The simplified processing module can have complete link processing center functions or partial link processing center functions to deal with specific fault problems in emergency scenarios.
[0036] The network control system of the present invention includes multiple sub-nodes (such as radar system intelligent units, sub-device intelligent units, separate network processing nodes, etc.), and at least one main node working device, wherein the sub-nodes can be used to monitor and receive sensor data or adjacent node fault processing data.
[0037] More specifically, the subnode of the network control system includes a transceiver module 11, an antenna module 12, a node intelligent processing unit (NIPU) 13, a first judgment module 14, a sensor module 15, and a power supply 16. Antenna module 12 is connected to transceiver module 11. The transceiver module transmits data to the node intelligent processing unit for processing and is responsible for data communication with the master node or subnodes. The node intelligent processing unit can be connected to the sensor module, the first judgment module, and the power supply module. The first judgment module can judge the data provided by the sensor module and the relevant data received from the transceiver module, and provide the judgment conclusion to the node intelligent processing unit.
[0038] The main node working equipment module includes a link processing center module, a second judgment module, a parameter set processing module, a parameter set database (static parameters, dynamic parameters, mixed parameters, etc.) and an operation center; the operation center is connected to the parameter set processing module, and the link center processing module is connected to the parameter set processing module, the second judgment module and the main equipment transceiver module.
[0039] NIPU can be preferably a multi-agent system, using the enhanced distributed computing features of MAS to cope with the specific processing of more complex tasks; under acceptable cost conditions, through the multi-agent reinforcement learning algorithm, the present invention can centrally learn the environmental data and task execution strategies of all NIPUs. During the data learning and training period, the environmental data and task execution strategies of each NIPU will be saved, and the NIPU will also obtain the relevant fault data and fault handling strategies of other adjacent NIPUs. That is to say, while each NIPU learns its own strategy, it will take into account the strategies of all related and working NIPUs and then learn its own response strategy. It will update its own strategy through the joint behavior composed of actual changes in the environment and actual task execution strategies. Through collaborative strategies, it can respond to specific tasks in specific scenarios and achieve the most suitable automatic control for processing task instructions.
[0040] Although the addition of a simplified Node Intelligent Processing Unit (NIPU) with the functionality of a link processing center module in a separate sub-node entity has significantly improved the inefficiency and complex technical defects that may occur in the solution using a single link processing center, the specific situations faced in actual network systems are more diverse, the addition and modification of network systems occur frequently, and the types of faults faced are also evolving. This has led to reduced recognition rates, recognition delays, and secondary faults when using NIPU units in some special usage environments.
[0041] The NIPU itself offers several advantages. Compared to a complete link processing unit, it can be a simplified processing module, saving hardware costs while ensuring timely and efficient fault identification and processing. When designing subnodes within the system, they can be placed simultaneously or integrated into a single unit. The simplified module can omit the more complex processing modules (such as specific data filtering or allocation) and instead be configured with only functions related to communication link data processing, status monitoring of adjacent processing units, collaborative preprocessing, switching between normal operating states, and interference suppression. This can be considered a simplified node intelligent processing unit. The NIPU enables real-time dynamic parameter learning of the actual environment of discrete subnodes, while flexibly managing the specific static parameters of fixed tasks. This enables a wider range of collaborative methods to meet actual needs and allows for the updating of these methods. In extreme cases, if the link processing center fails due to an unexpected event, one or more NIPUs can replace and overwrite the original link processing center's functions, significantly maintaining the continuity and flexibility of the entire link system.
[0042] In some specific scenarios, the introduction of NIPU has brought about better problem-solving effects; for example, in vehicle traffic light networks, local power grid equipment networks, etc., the types of faults that occur are relatively stable and uniform, and the fault identification, monitoring and processing strategies are relatively clear. They pursue economic processing effects and rapid response, and can be processed according to general topology / hierarchical priority rules. Therefore, a small number of NIPUs are generally introduced to quickly respond to and replace local networks or system networks, which can solve fault diagnosis and resolution at a relatively low cost but more efficiently, and is more effective than traditional single processing solutions.
[0043] In relatively complex real-world device control networks, such as collaborative control networks for industrial robotic devices (dark factories), radar network control systems, and complex neural network control devices, fault types often exhibit complex characteristics such as non-scaling, nonlinearity, strong coupling, and dynamic evolution. Fault handling in these scenarios using simple NIPUs can lead to inaccurate results and unstable processing, such as reduced recognition rates, secondary faults, and repeated processing of a single fault. This instability is primarily due to the diverse nature of faults in complex network systems and the overfitting of the NIPUs used to handle these diverse faults. While hardware improvements have improved processing efficiency in some scenarios, the increased hardware footprint in more complex scenarios has led to more diverse and complex problems. The significant increase in the number of NIPUs simultaneously triggered not only increases power consumption, but also, due to the close communication and data sharing between NIPUs, leads to significant conflicts between multiple NIPUs. This leads to repetitive and overfitting processing of the same fault or secondary fault.
[0044] Reference Figure 2 In the solution of the present invention, a new solution is proposed for troubleshooting in such practical scenarios.
[0045] By setting up NIPU in the sub-node to monitor, judge and handle node failures, the system performance can be improved through fault diagnosis and resolution.
[0046] Different fault types lead to different system processing results. As mentioned above, in a typical network system scenario, for relatively stable and controllable fault types, the child node will first report an error, triggering the activation of the current node and nearby nodes based on the basic functional characteristics of the NIPU. This activates the NIPUs in multiple child nodes, replacing the remote link center processing module to quickly resolve and control the network system fault. However, the NIPU itself does not determine the fault type. Therefore, in practice, this significantly leads to further problems.
[0047] After the NIPU detects and identifies a fault, it triggers a fault. Typically, NIPUs learn and share data parameters from surrounding nodes. Therefore, even a relatively simple fault (such as a single, brief signal interruption and recovery process) that wouldn't itself result in serious consequences can trigger the activation of multiple NIPUs within the current node and the local system network. This is clearly unnecessary. While it's undeniable that in extreme situations, such as natural disasters, such an activation solution is necessary to resolve network faults as quickly as possible and achieve the fastest results, the investment is worthwhile. However, in typical industrial network scenarios, addressing relatively simple faults with this approach is uneconomical. Furthermore, in industrial scenarios, the addition of an additional parallel processing unit can easily amplify the fault, leading to a run on the overall system's fault handling synchronization.
[0048] Therefore, in the solution of the present invention, faults in specific scenarios are classified and screened, and judgments are made in advance based on the actual classifications, so that the diagnosis and processing of faults by the entire system is more compact and efficient. Faults are called abnormal operating conditions in safety engineering. From the perspective of system operation, faults actually also include inefficient states that deviate from the optimal operating state. In industrial processes, faults are divided into equipment type, control type (elements, actuators, sensors), process abnormality type faults (process variable offset, abnormal changes, over-limit type faults, etc.) according to the location of the process; and are divided into slow-changing faults, sudden faults, etc. according to the time characteristics. It can be seen that in fact, different industrial scenarios have various types of faults, and specific analysis and judgment can be made for specific fault types in different scenarios. In order to illustrate actual problems and facilitate the expression of technical solutions, the analysis of specific types of faults is simplified in the solution of the present invention. The fault types are briefly classified into the first type of faults and the second type of faults in the solution of the present invention.
[0049] The first type of fault is characterized by stability and uniformity. The fault identification, monitoring, and handling strategies are relatively clear. The goal is to achieve economical handling and rapid response, and can be handled according to general topology / hierarchical priority rules.
[0050] The second type of fault characterization: Fault types are characterized by non-scaling, high coupling, and strong correlation. This results in a one-to-one correspondence between fault occurrence and symptoms. A single fault is likely to trigger another more complex fault, and faults are interrelated and affect each other, making it difficult to detect and address them using general solutions.
[0051] For the processing elements of the first type of fault, a threshold monitoring method using static rules can be used in the judgment module, or a topology-based priority rule can be selected, or the hardware sensor recognition results can be used more intuitively. The above methods are not specifically limited, and their purpose and effect is to quickly judge and identify relatively stable and single fault types.
[0052] As an example, for the first type of fault, you can select network fault status inspection. By traversing the system network, you can quickly find the key fault nodes in the system network:
[0053] for node in network_data['nodes']:
[0054] if node['status'] < self.rule_based_rules['node_failure']['threshold']:
[0055] if network_data['topology_map'][node['id']] == 'critical':
[0056] alarms.append({'type': 'NODE_FAILURE', ...})
[0057] The processing factors for the second type of fault are relatively complex. Specific methods based on graph feature extraction can be used, including graph data preprocessing, GNN reasoning and identification, and coupled fault analysis. Similarly, there are no specific restrictions on the methods for processing the second type of fault. Their purpose and effect is to more accurately judge and identify highly coupled and highly correlated fault types, providing accurate processing basis for subsequent judgment modules.
[0058] As an example, for the second fault type, you can choose to build a graph structure based on real-time network topology data. By extracting features from historical fault records, traffic characteristics, topology attributes, and even real-time hardware features, and using GNN analysis and reasoning, you can conduct in-depth fault analysis. The general execution process can be referred to as follows:
[0059] # Initialize the model
[0060] model = CoupledFaultGNN(in_feats=6, hid_feats=64, out_feats=2)
[0061] model.load_state_dict(torch.load('fault_gnn.pth'))
[0062] # Input real-time data
[0063] graph_data = build_graph(real_time_data)
[0064] # Inference prediction
[0065] with torch.no_grad():
[0066] logits = model(graph_data)
[0067] probs = torch.softmax(logits, dim=1)
[0068] # Filter suspicious nodes (probability > 70%)
[0069] suspects = np.where(probs[:,1] > 0.7)[0].tolist()
[0070] # Output analysis results
[0071] alarms = analyze_fault_spread(graph_data, suspects)
[0072] In addition to the aforementioned fault type considerations, the present invention's fault identification and processing also prioritizes time, as urgency can sometimes be a primary consideration in troubleshooting. If a particular fault type has a relatively low urgency and tolerates a relatively high time error, then the time requirement for this type of fault is lower. If a particular fault type has a relatively high urgency and tolerates a low time error, then time is crucial for this type of fault.
[0073] Reference Figure 2 In the NIPU, each child node can be equipped with a NIPU unit. Multiple child nodes can be equipped with NIPU units in the entire system network. That is, a network system with N nodes can have a maximum of N NIPU units. These NIPU units can have either partial or full link processing center functions of the master node. Preferably, the entire system control network has at least one NIPU unit with full link center functions to cope with the event of master node failure.
[0074] Although having multiple NIPUs online at the same time can greatly improve the flexibility of the network system's work processing, in some scenarios, the NIPU in the child node can replace the function of the main node's working equipment, and can better monitor and resolve fault events. However, for common fault situations, the frequent use of multiple NIPUs will lead to overlapping event processing and easily cause secondary faults.
[0075] Although NIPUs currently generally use multiple agents or are multi-agent, and they do offer advantages such as rapid response, this often comes at the expense of overall system efficiency. A simple interruption can trigger rapid responses from multiple sub-nodes simultaneously. Each sub-node quickly triggers the NIPU and continuously monitors parameter changes within itself and surrounding nodes. The triggering event itself changes the operating parameters of the nodes in the scenario. If one of the nodes fails to transmit or interpret the information correctly, or if there is a conflict between the current node's processing tasks, or if there is manual intervention by staff at the current node, short-term fault data is likely to occur. This fault data will be self-learned and shared multiple times by the NIPU itself. After several processing steps, the NIPUs of surrounding nodes learn and transmit the data, occasional but systemic cascading failures may occur. In other words, a minor interruption fault may be multiplied, triggering a secondary serious fault type. At this time, the rapid response feature also helps the fault transmission itself, which is obviously contrary to the original intention of intelligent network system design.
[0076] Among the existing technical solutions, there are some solutions to this problem. One of them is to reduce the number of NIPUs or reduce the intelligence level of NIPU units. For example, the layout of NIPU nodes in the entire system network can be appropriately reduced by 70% or even lower. In this way, the entire system will be limited to a small number of nodes for handling occasional systematic cascading failures. Even if a misjudgment occurs, it will only be limited to a small range of node networks. After that, the link processing center module of the main node device will provide remote computing and assistance to the child nodes to solve such problems. However, this treatment brings about a significant reduction in efficiency, which will cause the original advantages of the NIPU system network to be lost.
[0077] In the solution of the present invention, a first judgment module is set at the sub-node and combined with the NIPU unit. The basic categories of faults and related condition thresholds are judged and screened based on the first judgment module, thereby achieving optimized processing of faults in specific scenarios. While maintaining the advantages of the traditional NIPU system control network, the system's ability to handle faults is further improved.
[0078] When the sensor module or transceiver module in the system network node monitors or receives abnormal data, indicating a possible fault event, the sensor module data or the received data of the transceiver module is preferentially sent to the first judgment module, and a preliminary judgment is made based on the type identification unit in the first judgment module to screen whether the fault event belongs to the first category or the second category.
[0079] Based on specific practices in industrial scenarios, the present invention also considers time factors. The distinction between the first and second fault types primarily considers the complexity of the fault diagnosis process itself. In certain scenarios, some simple faults may require extremely tight time constraints for troubleshooting, while some complex faults may require moderate time constraints for handling. The time threshold determination unit within the first determination module is used to perform time threshold assessment and analyze the time urgency of the current fault handling process.
[0080] Based on the configuration of at least the type identification unit and the threshold judgment unit in the first judgment module, the fault diagnosis and processing of the network system with the NIPU module can be more targeted.
[0081] Reference Figure 3 When the judgment module receives data from the sensor module or data from the transceiver module of an adjacent node (including the master node and subnodes), the time threshold unit in the judgment module is triggered first. This unit determines the time urgency of resolving the current fault type. If the result of the fault resolution judgment function F(t) exceeds the set threshold, the relevant fault handling request is directly sent to the NIPU unit, triggering the NIPU unit of the current node to perform real-time fault handling and resolve the fault, thereby ensuring the fastest possible resolution. If the fault resolution judgment function F(t) does not reach the set threshold, further determination of the specific fault type is required.
[0082] Time threshold determination usually needs to be considered as the primary factor, especially for sudden failures (such as line loss caused by disasters or self-destruction). The specific determination of the time threshold can be uniformly specified within the system control network of each sub-node; or the preset time threshold under the corresponding node can be set according to the actual node location scenario. For example, for the same type of "catastrophic" failure, if the node position in the network system is not a critical node position, then its time threshold may change accordingly, that is, the time threshold itself is determined by the impact of the fault event to be processed on the overall system network. The greater the impact of the fault event on the system network, the lower the threshold, and the easier it is to trigger a rapid processing process. Conversely, the smaller the impact of the fault event on the system network, the higher the threshold, and the less likely it is to trigger a rapid processing process. In addition, the time threshold determination can also request support from the main node device unit, especially when the topology structure in the system network increases or decreases or the function of the local node changes significantly, it is necessary to receive key information from the main node in a timely manner.
[0083] If the calculated value of the fault resolution judgment function F(t) does not reach the set threshold, the fault type identification unit in the first judgment module will be further triggered. According to the simplified fault description above, the type identification unit must now screen and determine whether the current fault type is a Class I or Class II fault. Based on the specific fault type, it will proceed to the next node action and assign instructions to other related nodes.
[0084] Reference Figure 4 If the result of the type identification unit's judgment is a first-class fault, then the current fault type is judged to be a relatively regular and simple fault type, and the time urgency required for processing is not high. It can be considered that the processing level of the current fault is a weak priority. If the current fault does not cause the node to be paralyzed, then the NIPU module of the current node can be not triggered. Instead, the weak priority indication can be directly transmitted to the main node link processing center through the transceiver processing module for processing. The main node processing center can respond appropriately based on the general process of handling this type of fault (because appropriate delayed processing is acceptable at this time), such as the basic response to faults in the traversal dictionary library or the fault handling manual of the inspection system network database. Although the timeliness will be reduced, appropriately reducing the timeliness for this type of fault will not have a significant adverse effect on the normal operation of the network system.
[0085] It should be noted that the general method for handling this type of fault is to be solved by the link processing center, but in practice, if the fault event itself changes sporadically during the processing of the first fault event, or the overlapping occurrence of fault events will also lead to changes in the fault handling requirements of the entire network system; when the cyclic detection system reports an urgency warning or the judgment module receives an indication of time urgency, that is, the current fault has a considerable impact on the operation of the node (for example, it has seriously affected the work efficiency of the current node or the node is paralyzed), then it is necessary to trigger the NIPU unit of the current node for processing, search for solutions through the NIPU local library of the current node, or jointly solve them through the NIPUs connected to the surrounding nodes, and stop subsequent processing operations of the main node processing center, appropriately increase the priority of the current fault processing without affecting the main node equipment too much, while ensuring appropriate processing efficiency and preventing the escalation of the current type of fault from causing cascading system failures.
[0086] If the type identification unit determines that the fault is a Class II fault, then the current fault type is a complex type fault characterized by non-scaling, high correlation, and strong coupling. The current fault processing level can be considered high priority. Although the time threshold judgment unit determines that the fault type is of low urgency, the characteristics of this type of fault are dynamic and unstable, and if not handled in a timely manner, it can easily lead to larger-scale network cascade failures. This type of fault usually accounts for the vast majority of fault types. Therefore, it is difficult to find a satisfactory solution if the first type of master node processing or the NIPU unit processing is directly used. The lack of timeliness of the complete master node processing will cause a large number of fault events to squeeze and conflict, while the complete NIPU processing can easily cause the sub-nodes to overload.
[0087] Therefore, for this type of fault, the judgment module will quickly respond and send instructions to the NIPU module of the current node, triggering and activating the local NIPU. The current NIPU starts the processing flow of the fault event and starts the NIPU module detection of the associated node. The purpose of the detection is to search for a NIPU module with complete functions. By triggering and activating the complete NIPU module, it jointly participates in the fault event processing of the faulty node. At this time, the complete NIPU can replace the function of the link processing center of the main node and enhance the processing of fault events. This can not only balance the processing burden of the sub-nodes, but also improve the fault processing efficiency through the link center replacement function of the complete NIPU and reduce the investment in the main node.
[0088] Preferably, a time limit for handling the second type of fault is necessary. If the local NIPU and the complete NIPU module found do not resolve the fault within a limited time frame, if the fault continues to persist within the network system, it will greatly increase the cascading problems caused by such fault events. For example, key parameter drift caused by system aging and other issues. Since parameter drift data itself is extremely susceptible to environmental influences, the severity of data drift will be accelerated in a faulty environment. The longer the processing time, the more likely it is to cause cascading problems. Therefore, in the solution of the present invention, a time limit for handling the second type of fault is used to prevent the impact of the fault event from being further expanded. This time limit can be achieved through two solutions.
[0089] Solution 1: The processing time limit of the local node NIPU module and the searched complete NIPU module is set. By setting the time limit for this type of fault processing, when the time limit for this type of fault processing is exceeded, the complete NIPU module can replace the main node link processing center to issue a mandatory instruction, instructing the current fault node to stop issuing fault data sharing or fault processing requests, and only perform fault processing within the current node range, to prevent fault data and secondary related data generated based on the fault data from interfering with other surrounding sub-nodes, occupying resources or polluting data until the fault is eliminated.
[0090] Option 2: The NIPU module of the local node detects or searches for a secondary fault indication fed back by the NIPU module in the system network. If it is detected that a secondary fault has occurred after the first second-type fault, for example, the processing of the first fault is more complicated, resulting in a significant increase in the node data processing resource utilization rate and causing node overload, etc., then the current faulty node can be directly instructed to stop issuing data sharing or fault processing requests, and only perform fault processing within the current node until the fault is eliminated.
[0091] If no indication of a processing time exceeding or secondary fault risk is detected during the secondary fault handling process, the fault is being resolved in an orderly manner and will be resolved until it is complete. This process allows relatively complex second-type faults to be resolved by leveraging the inherent system advantages of the NIPU while also eliminating mechanization, overfitting, and inefficiencies that may occur during the resolution process, thereby strictly controlling the complex fault handling process.
[0092] Based on the above, the solution of the present invention first improves the priority processing of emergency faults by setting up NIPU modules in the sub-node units of complex networks, thereby improving the effectiveness of the system network; and further, by setting up classification processing for specific fault units in the system network and providing corresponding solution mechanisms, it effectively solves the possible cascading faults caused by the inherent characteristics of complex networks.
[0093] In order to implement the above embodiment, the present invention also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for processing a train network master node failure as shown in the above embodiment.
[0094] In order to implement the above embodiments, the present invention also proposes a non-temporary computer-readable storage medium on which a computer program is stored, characterized in that the program is executed by a processor to implement the method for processing complex network node failures as shown in the above embodiments.
[0095] In order to implement the above embodiments, the present invention further provides a computer program product. When instructions in the computer program product are executed by a processor, the method for processing complex network node failures as shown in the above embodiments is executed.
[0096] In addition, in order to implement the above-mentioned invention solution, those skilled in the art may improve and supplement some details not specifically mentioned in the above-mentioned solution. For example, in the control system fault diagnosis method under a complex dynamic network environment, those skilled in the art may construct relevant network models of the control system (such as dynamic or static, etc.), or may choose to utilize the complex network topology and the correlation between nodes to monitor system status changes in real time, identify potential fault sources, and then perform the technical processing in this case for specific fault sources. Alternatively, through technologies such as data mining and pattern recognition, the processing efficiency of this case in the fault handling process can be further improved, helping to more efficiently and accurately locate the location of the fault and limit the scope of the fault's impact, thereby improving the reliability and intelligence level of the control system. If these relevant technical details do not deviate from the main technical concept of this case, they should not be considered as breakthrough innovations in the invention of this case.
[0097] In summary, although the present invention has been described in detail through the above preferred embodiments, it should be understood that the above description is not intended to limit the present invention. After reading the above, various modifications and substitutions of the present invention will become apparent to those skilled in the art. Therefore, the scope of protection of the present invention is defined by the appended claims.
Claims
1. A fault control system in a complex network environment, comprising a master node device and multiple sub-node devices, characterized in that: The sub-node device is equipped with a node intelligent processing unit NIPU, which has partial or complete master node link processing center functions; The sub-node is provided with a first judgment module, which receives the fault data signal and performs time threshold judgment and fault type judgment; and initiates different fault handling processes according to the judgment result of the first judgment module.
2. The fault control system according to claim 1, characterized in that: The fault data signal received by the first judgment module comes from a sensor module or a transceiver module in the sub-node device.
3. The fault control system according to claim 1, characterized in that: The first judgment module performs a first time threshold judgment on the fault data. If the first time threshold limit is reached, the NIPU of the current node is directly triggered to initiate rapid fault processing; if the first time threshold limit is not reached, the specific type of the fault event is further evaluated.
4. The fault control system according to claim 3, characterized in that: The first judgment module initiates a fault type identification judgment, and the fault type is identified as the first fault type or the second fault type.
5. The fault control system according to claim 4, characterized in that: If the fault type identification result is the first fault type, the fault is reported to the master node link processing center for processing; During the master node processing, if a time urgency indication is detected, the sub-node device NIPU is triggered to perform fault processing.
6. The fault control system according to claim 5, characterized in that: If the fault type identification result is the second fault type, the current sub-node device NIPU is triggered to perform fault processing, and at the same time, at least one NIPU with complete master node link processing center function in the control system network is searched for joint processing.
7. The fault control system according to claim 6, characterized in that: If, during the second type of fault handling process, it is detected that the fault handling time exceeds the second time threshold limit, the fault data sharing of the faulty sub-node is stopped or a fault handling request signal is sent to other nodes.
8. The fault control system according to claim 6, characterized in that: If, during the second type of fault handling process, it is detected that the system network has generated a secondary fault indication based on the current fault, the fault data sharing of the faulty sub-node is stopped or a fault handling request signal is sent to other nodes.
9. The fault control system according to claim 1, characterized in that: The number of NIPUs in the control system is no greater than the number of all sub-nodes, and the number of NIPUs with complete master node link processing center functions is greater than 1.
10. A fault control method in a complex network environment, characterized in that: The steps include: Step 1: After receiving the fault data, the sub-node device judgment module of the network system initiates the judgment of the fault data; Step 2: Determine the first time threshold for processing the current fault data. If the time threshold is reached, trigger the NIPU of the current node to initiate rapid fault processing. If not, proceed to the next step. Step 3: Determine the specific type of fault in the current fault event. If it is a type 1 fault, report the fault to the master node link processing center for processing; During the master node processing, if a time urgency indication is detected, the sub-node device NIPU is triggered to perform fault processing; Step 4: If the fault is of the second type, the current sub-node device NIPU is triggered to process the fault, and at the same time, at least one NIPU with complete master node link processing center functions in the control network system is searched for to perform joint processing; Step 5: If, during the second type of fault handling process, it is detected that the fault handling time exceeds the second time threshold limit, and / or it is detected that the system network generates a secondary fault indication based on the current fault, then the fault data sharing of the faulty sub-node is stopped or a fault handling request signal is sent to other nodes.
Citation Information
Patent Citations
Train network master node fault processing method, node and electronic equipment
CN109787795A