Fiber channel network failure prediction method supporting asynchronous state awareness and cadence driven refresh

By employing asynchronous state awareness and rhythm-driven refresh methods in the FC network, the processing cycle and reporting cycle are dynamically adjusted, solving the response time problem caused by node state differences in the vehicular network. This enables more efficient fault risk prediction and control, and improves the safety and stability of vehicle operation.

CN120499023BActive Publication Date: 2026-06-26COMP APPL TECH INST OF CHINA NORTH IND GRP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2026-06-26

Smart Images

  • Figure CN120499023B_ABST
    Figure CN120499023B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of FC network fault prediction methods of supporting asynchronous state perception and rhythm-driven refresh, belong to vehicle network communication technical field.The present application method includes: each network node respectively link state information dynamic determination reporting period and report state report frame;Center management node obtains the state information of each node connected link in the current processing period based on the state report frame and historical state information received in the current processing period, dynamically adjusts processing period based on the state information of the current processing period, and carries out network fault risk prediction, and the network optimization instruction frame carrying network optimization instruction is issued to optimize network.The present application method realizes the linkage response between the adaptive regulation and control of state information reporting rhythm, the real-time refresh of asynchronous state completion estimation, risk prediction and optimization control for center management node, improves the timeliness and stability of vehicle optical fiber channel network under the condition of asynchronous state upload.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle network communication technology, and in particular to a fault prediction method for FC networks that supports asynchronous state perception and rhythm-driven refresh. Background Technology

[0002] With the rapid integration of complex electronic systems such as autonomous driving, vehicle-to-everything (V2X) communication, and smart cockpits into vehicles, the requirements for communication bandwidth, real-time performance, and reliability of in-vehicle networks have significantly increased. Fiber Channel (FC), due to its high bandwidth, low latency, and high stability, is increasingly being used in critical communication paths such as in-vehicle central controllers, camera buses, high-speed sensors, and decision-making systems, and is particularly suitable for autonomous driving domain control systems with extremely high latency and reliability requirements.

[0003] However, traditional FC network management systems are mostly based on synchronous communication models of data centers and storage networks. Under complex vehicular network topologies and dynamic link states, these systems suffer from the following main problems: 1. Fixed node status reporting rhythm: Nodes in FC networks typically upload link status at a uniform period. In real-world network environments, node health states vary significantly, and the uniform period mechanism cannot balance the detection frequency of high-risk nodes with the communication overhead of low-risk nodes. 2. In existing FC networks, the central management node (CMN) must wait for all nodes to upload data before assessing the network status, resulting in overall response time being hampered by the "slowest node." 3. Inability to flexibly handle network packet loss and other anomalies, making it difficult to build a robust, continuous network status perception model, easily leading to misjudgments, missed reports, or slow responses, threatening vehicle safety. Summary of the Invention

[0004] Based on the above analysis, this invention aims to disclose a fault prediction method for FC networks that supports asynchronous state perception and rhythm-driven refresh. This method enables the coordinated response between adaptive adjustment of the state information reporting rhythm, central management node estimation of asynchronous state completion, risk prediction, and real-time refresh of optimization control. This improves the network's processing efficiency and overall system stability when facing dynamic loads, local degradation, and sudden failures, thereby ensuring vehicle driving safety.

[0005] This invention provides a fault prediction method for FC networks that supports asynchronous state awareness and rhythm-driven refresh, specifically including the following steps:

[0006] The status information of the links connected to each network node in the current processing cycle is obtained based on the status report frames received in the current processing cycle and the historical status information; wherein, the historical status information is obtained based on the historical status report frames.

[0007] The processing cycle is dynamically adjusted based on the status information of the current processing cycle, and network fault risk prediction is performed.

[0008] Based on the predicted failure risk, a network optimization instruction frame carrying network optimization instructions is issued to optimize the network.

[0009] Furthermore, the central management node receives status report frames from each network node;

[0010] The status report frames received by the central management node during the current processing cycle are reported by each network node based on its reporting cycle; the reporting cycle is dynamically determined by each network node based on the status information of the connected links.

[0011] Furthermore, the reporting period is dynamically determined by each network node based on the status information of the connected links, including:

[0012] Each network node calculates its current link health status value based on the current status information of the linked link.

[0013] Each network node updates its reporting cycle based on the current link health status value and the link status risk threshold.

[0014] Furthermore, the status information includes the link frame loss rate, frame loss rate differential trend value, link round-trip time, weighted moving average of high-priority link round-trip time, trend differential value of high-priority link round-trip time delay, and bit error rate of the connected link.

[0015] The step of obtaining the status information of each network node's connected link corresponding to the current processing cycle based on the status report frame received in the current processing cycle and historical status information includes:

[0016] If a status report frame from a network node is received in the current processing cycle, the status information of the link connected to the corresponding node is obtained from the status report frame.

[0017] If no status report frame from the network node is received in the current processing cycle, then determine whether a status report frame from the same network node was received in the previous processing cycle.

[0018] If so, the estimated value of the state information of the link connected to the network node is calculated based on the state information of the previous two processing cycles of the network node, and used as the state information of the current processing cycle;

[0019] If not, then calculate the estimated value of the state information of the link connected to the node based on the node's historical state information as the state information of the current processing cycle; or, calculate the estimated value of the state information of the link connected to the node based on the node's state information and penalty coefficient in the previous processing cycle as the state information of the current processing cycle; or, set the state information of the node in the current processing cycle to null.

[0020] Furthermore, the step of dynamically adjusting the processing cycle based on the status information of the current processing cycle and predicting network failure risks includes:

[0021] The central management node determines whether the network is in a high-risk state based on the status information of the current processing cycle. If so, it immediately performs network failure risk prediction and ends the current processing cycle.

[0022] If not, then:

[0023] The activity level of the status report frame is calculated based on the number of network nodes corresponding to the status report frame received in the current processing cycle and the total number of network nodes.

[0024] A new processing cycle is calculated based on the activity level of the status report frame and the basic refresh cycle.

[0025] And network failure risk prediction is performed within the current processing cycle.

[0026] Furthermore, the failure risk prediction includes:

[0027] The central management node calculates the corresponding overall network health index value, Bayesian probability outlier value, and Bayesian probability outlier trend value of each network node based on the status information of the current processing cycle.

[0028] When the link frame loss rate is greater than or equal to the frame loss rate threshold, or when the cumulative value of the frame loss rate differential trend value within a consecutive first set number of processing cycles is greater than the frame loss rate differential trend threshold, it is predicted that there is a risk of intermittent single-point link failure.

[0029] When the link round-trip delay exceeds the link round-trip differential threshold within a second consecutive set number of processing cycles, it is predicted that there is a risk of abnormal link transmission performance.

[0030] When the weighted moving average of round-trip delay for high-priority links is greater than or equal to the weighted moving average threshold, or when the cumulative value of the round-trip delay trend difference for high-priority links within the third consecutive set number of processing cycles is greater than the cumulative threshold for high-priority trends, it is predicted that there is a risk of performance degradation of high-priority links.

[0031] When the overall network health index value is greater than or equal to the set health threshold, or its growth trend is greater than the exponential growth rate threshold, network congestion risk is predicted.

[0032] When the outlier value of the Bayesian probability of a network node is greater than or equal to the set threshold for outlier probability, or when the cumulative value of the outlier trend value of the Bayesian probability is greater than the set threshold for outlier trend in the fourth consecutive set number of processing cycles, a hidden or intermittent node failure risk is predicted.

[0033] Furthermore, the central management node optimizes the network by issuing network optimization instruction frames carrying network optimization instructions based on the predicted failure risk, including:

[0034] When a single point of failure is predicted to be intermittent, the central management node issues a network optimization instruction frame carrying link isolation or link activation instructions.

[0035] When a risk of abnormal link transmission performance is predicted, the central management node issues a network optimization instruction frame carrying link activation or topology reconfiguration instructions.

[0036] When a risk of performance degradation of a high-priority link is predicted, the central management node issues a network optimization instruction frame carrying link activation or topology reconfiguration instructions.

[0037] When a network congestion risk is predicted, the central management node issues a network optimization instruction frame carrying link activation or topology reconfiguration instructions.

[0038] When a hidden or intermittent node failure risk is predicted, the central management node issues a network optimization instruction frame carrying link isolation or topology reconfiguration instructions.

[0039] Furthermore, the method for calculating the link health status value is as follows:

[0040] H (i) (k)=w1·FLR (i) (k)+w2·RTT (i) (k)+w3·BER (i) (k);

[0041] Where, the superscript i represents the network node number; k represents the reporting period number of network node i; H (i) (k) represents the link health status value of network node i during reporting period k; FLR (i) (k), RTT (i) (k), BER (i) (k) represents the link frame loss rate, link round-trip time, and bit error rate of network node i in reporting period k; w1, w2, and w3 are the corresponding weighting factors.

[0042] Furthermore, each network node updates its reporting cycle based on the current link health status value and the link status risk threshold, including:

[0043]

[0044] in: T represents the (k+1)th and kth reporting periods, respectively; min Minimum reporting period; T defaultThe default period; ΔT is the shrinking step size during updates; δT is the increasing step size; H th This is the threshold for link state risk.

[0045] Furthermore, the central management node determines whether the network is in a high-risk state based on the status information of the current processing cycle, including:

[0046] The number of high-risk nodes in the network is determined based on the status information of the current processing cycle.

[0047] If the number of high-risk nodes exceeds a set threshold, the network is judged to be in a high-risk state.

[0048] A network node is considered a high-risk node if it meets one of the following conditions:

[0049] The link frame loss rate is greater than or equal to the frame loss rate threshold.

[0050] The cumulative value of the frame drop rate differential trend value is greater than the cumulative threshold of the frame drop rate differential trend within a first set number of consecutive processing cycles;

[0051] The link round-trip delay exceeds the link round-trip differential threshold within a second consecutive set number of processing cycles;

[0052] The round-trip weighted moving average of high-priority links is greater than or equal to the weighted moving average threshold;

[0053] The cumulative value of the high-priority link round-trip delay trend difference is greater than the high-priority trend cumulative threshold within the third consecutive set number of processing cycles;

[0054] The Bayesian probability outlier is greater than or equal to the outlier probability threshold.

[0055] The link health status value is greater than the health status threshold for two consecutive processing cycles.

[0056] The present invention can achieve at least one of the following beneficial effects:

[0057] By asynchronously reporting status report frames from each network node, the central management node obtains status information based on the status report frames received in the current processing cycle, and estimates the missing status information based on the sliding window mechanism, dynamically adjusts the processing cycle, and performs network fault risk prediction. This achieves a coordinated response of status reporting, risk prediction, and optimized control, improving the timeliness and stability of the vehicle-mounted fiber optic channel network under asynchronous status uploading conditions.

[0058] By employing an adaptive adjustment mechanism where each network node locally senses and drives the reporting cycle based on the status information of the connected links, the reporting cycle can be dynamically adjusted. This allows for the rapid acquisition of more information when the link status deteriorates, and reduces the monitoring load when the link is stable. It guides the rhythm of status report frames to exhibit a non-uniform distribution strongly correlated with link risk, which is beneficial for system resource scheduling. This solves the problem in existing technologies where the frequency cannot be dynamically adjusted based on link health status, resulting in no difference in reporting frequency between "healthy nodes" and "abnormal nodes," wasting network resources and failing to provide more frequent attention to high-risk nodes.

[0059] By employing a "sliding window estimation mechanism," the problem of information loss caused by asynchronous status report frame uploads is resolved, enabling the complete construction of the status matrix even in scenarios where the rhythms of multiple nodes are not aligned. This improves the continuity and timeliness of the overall network status assessment, preventing the system from being "dragged down by the slowest node." The risk prediction process is triggered as soon as a high-risk node uploads a status report frame, calculating the overall network health index value. This maintains the consistency of status information across processing cycles and provides a basis for dynamically adjusting processing cycles.

[0060] By dynamically adjusting the processing cycle based on the activity of status report frames in the current processing cycle, calculating the overall network health index, Bayesian probability outliers and Bayesian probability outlier trend values ​​for each network node, and performing fault risk prediction, this technology breaks away from the passive strategy of "fixed-cycle refresh" in existing technologies. It achieves rhythm-driven adaptive adjustment, enabling high-risk nodes that upload frequently to dominate the system's judgment rhythm, significantly improving system response efficiency. The fault risk prediction cycle is decoupled from but linked to the reporting cycle rhythm of each node, improving the system's adaptability. Furthermore, the dynamic adjustment of the processing cycle is coordinated with the sliding "sliding window estimation mechanism," allowing for accurate fault risk prediction even when some nodes fail to report status report frames, thus improving system stability and security.

[0061] By constructing dedicated fault diagnosis frames, status report frames, and network optimization command frames, the structure exhibits good compatibility with existing FC network standard frame structures, is easy to implement, and is suitable for widespread industrial applications. Through proactive preventative measures, network optimization command frames carrying network optimization instructions are issued based on predicted fault risks to optimize the vehicular network. Based on fault prediction, proactive preventative measures can rationally allocate network resources according to priority, prioritizing the transmission of critical data. This improves the overall operational efficiency and reliability of the vehicular network, avoids vehicle malfunctions caused by network communication failures, reduces vehicle operation and maintenance costs, and more accurately identifies early signs of faults. Compared to traditional fault detection methods, it can detect potential problems earlier and reduce the impact of network communication failures on vehicle operation.

[0062] Other features and advantages of the invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained from what is particularly pointed out in the description, claims, and drawings. Attached Figure Description

[0063] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.

[0064] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0065] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0066] An embodiment of the present invention discloses a fault prediction method for FC networks that supports asynchronous state awareness and rhythm-driven refresh, specifically including steps S1-S3, and step S0 is included before steps S1-S3.

[0067] Step S0: The central management node receives status report frames from each network node; the status report frames received by the central management node in the current processing cycle are reported by each network node based on its reporting cycle; the reporting cycle is dynamically determined by each network node based on the status information of the connected links. Step S0 includes steps S01-S03.

[0068] S01. The network node calculates the status information of the connected link based on the multiple fault diagnosis frames (FDF) sent and the corresponding response frames received in each reporting cycle.

[0069] The status information includes the link frame loss rate, link round-trip time, and bit error rate of the connected links. The status information also includes the frame loss rate differential trend value, the weighted moving average of round-trip times for high-priority links, and the round-trip time differential trend value for high-priority links.

[0070] Specifically, network nodes calculate the status information of connected links based on one or more fault diagnosis frames (FDFs) sent to other network nodes and the corresponding response frames received during each reporting cycle. The fault diagnosis frame is defined based on the standard FC frame format, including: defining the TYPE field in the frame header as a dedicated type identifier and marking it as a fault diagnosis frame type; defining a priority identifier using a portion of the Parameter field from the standard FC frame; and adding a high-precision timestamp field placed before the payload. The fault diagnosis frame format is shown in Table 1.

[0071] Table 1. Fault Diagnosis Frame Format

[0072] Field Name Length (bytes) Explanation (Comparison with standard FC frame format) SOF 4 Adhere to the standard R_CTL 1 The standard frame header follows the standard. DID 3 The standard frame header follows the standard. SID 3 The standard frame header follows the standard. TYPE 1 Define a special type (0xFE) as an FDF special frame. F_CTL 3 The standard frame header follows the standard. SEQ_ID 1 Standard fields, following the standard DF_CTL 1 Standard fields, following the standard SEQ_CNT 2 Standard field, can be directly used as a serial number. OX_ID 2 Standard fields, following the standard RX_ID 2 Standard fields, following the standard Parameter 4 Standard field, available for frame extension parameters Priority 1 Add a new field, implemented using part of the Parameter field space. Timestamp 8 Payload Prefix Fields (New) Payload Optional Extended data fields store link information. CRC 4 Standards remain unchanged. EOF 4 Standards remain unchanged.

[0073] Furthermore, FDF frames employ an active interaction mode. Network nodes actively send FDF frames to neighboring nodes. Within a reporting cycle, a node can send one or more FDF frames to other network nodes; that is, a reporting cycle includes one or more dynamic FDF sending cycles. A node receiving an FDF frame must immediately return an FDF response frame (ACK frame). The ACK frame has the same frame structure as the original FDF frame, but the ACK frame is marked as a response frame type (0xEF) in the TYPE field and carries the corresponding receive and send timestamps (merged and compressed in the Payload prep field). The active interaction cycle is differentially set according to different service priorities; higher priority links use shorter sending cycles to improve the real-time diagnostics of critical service links.

[0074] Furthermore, the network node calculates the link frame loss rate of the connected link based on the total number of Fault Diagnosis Frames (FDF) sent in the current reporting period and the total number of corresponding ACK frames received; the formula for calculating the link frame loss rate is:

[0075] FLR(k)=1-[N recv (k) / N send (k)];

[0076] Where FLR(k) represents the link frame loss rate in the k-th reporting period; N send (k) represents the total number of FDF frames sent in this period; N recv (k) The total number of ACK frames received in this period;

[0077] The round-trip time of the connected link is calculated based on the time of sending the fault diagnosis frame and the time of receiving the corresponding response frame in the current reporting period. If multiple fault diagnosis frames are sent in the current reporting period, the round-trip time is the average of the round-trip times of all fault diagnosis frames and the corresponding response frames.

[0078] The bit error rate (BER) of the connected link is determined based on the cumulative bit error rate per unit time or unit bit transmission volume within the current reporting period; the bit error rate BER (i) (k) Statistically obtained based on the built-in error counters in the physical layer or MAC layer of the link.

[0079] Furthermore, the frame loss rate difference trend value of the connected links is calculated based on the link frame loss rate of the current reporting period and the previous reporting period.

[0080] The weighted moving average of the high-priority links in the current reporting period is calculated based on the link round-trip delay in the current period and the weighted moving average of the high-priority links in the previous reporting period.

[0081] The round-trip delay trend difference value of the high-priority link is calculated based on the round-trip weighted moving average of the high-priority link in the current reporting period and the previous reporting period.

[0082] Furthermore, the formula for calculating the frame drop rate differential trend value ΔFLR(k) is as follows:

[0083] ΔFLR(k)=FLR(k)-FLR(k-1).

[0084] Furthermore, the formula for calculating the round-trip weighted moving average of high-priority links is as follows:

[0085] RTT EWMA (k)=β·RTT(k)+(1-β)·RTT EWMA (k-1);

[0086] Among them, RTT EWMA (k) is the round-trip weighted moving average of the high-priority link in the k-th reporting period; β is the weighting factor, with an optimal value of 0.6.

[0087] Furthermore, the high-priority link round-trip delay trend difference ΔRTT EWMA The formula for calculating (k) is:

[0088] ΔRTT EWMA (k)=RTT EWMA (k)-RTT EWMA (k-1).

[0089] S02. Each network node calculates the current link health status value based on the current link frame loss rate, link round-trip time, and bit error rate of the connected link.

[0090] Specifically, the calculation method is as follows:

[0091] H (i) (k)=w1·FLR(i) (k)+w2·RTT (i) (k)+w3·BER (i) (k);

[0092] Where, the superscript i represents the network node number; k represents the reporting period number of network node i; H (i) (k) represents the link health status value of network node i during reporting period k; FLR (i) (k), RTT (i) (k), BER (i) (k) represents the link frame loss rate, link round-trip time, and bit error rate of network node i in reporting period k; w1, w2, and w3 are the corresponding weight factors, and the sum of the three is 1.

[0093] Furthermore, the values ​​of weighting factors w1, w2, and w3 can be set according to the link type and criticality. For example, high-priority links involve critical real-time control, and their requirements for link frame loss rate and link round-trip latency are relatively high. The preferred values ​​for w1, w2, and w3 are (0.48, 0.42, 0.1). For high-bandwidth links of high-definition video or sensor data, the bit error rate has a significant impact on service perception. The preferred values ​​for w1, w2, and w3 are (0.3, 0.33, 0.37). For redundant backup links and low-priority auxiliary links, the weights of link round-trip latency and bit error rate can be reduced, and the focus should be on sensitive detection of link frame loss rate. The preferred values ​​for w1, w2, and w3 are (0.64, 0.18, 0.18). For the core switching network, i.e., the distributed mesh topology with symmetrical link structure, the values ​​for w1, w2, and w3 are uniformly (1 / 3, 1 / 3, 1 / 3). The above preferred values ​​can be preset during the system design phase or dynamically adjusted by the central management node according to network traffic characteristics.

[0094] S03. Each network node updates its reporting cycle based on the current link health status value and the link status risk threshold.

[0095] Specifically, updating the reporting cycle based on the current link health status value and link status risk threshold includes:

[0096]

[0097] in: T represents the (k+1)th and kth reporting periods, respectively; min Minimum reporting period (e.g., 10ms); T default The default reporting period is 100 sm; ΔT is the shrinking step size during updates; δT is the increasing step size; H th The threshold value for link state risk is 0.18.

[0098] It should be noted that the link state risk threshold H th The design principle is to be sensitive and lightweight, serving as an early warning system. Its goal is not to directly trigger network-level optimization, but rather to guide the node on whether to increase monitoring frequency in advance, thereby obtaining more link performance samples before potential failures occur. In actual configurations, the link status risk threshold serves to: amplify the sensitivity to small fluctuations or minor anomalies; rapidly increase the frequency of status report frame transmission, enhancing coverage density; collect multi-cycle statistical data of potentially faulty links in advance; and provide richer input for fault diagnosis at the central management node.

[0099] Furthermore, each network node reports a status report frame based on a reporting cycle. The frame format of the status report frame in this invention is shown in Table 2:

[0100] Table 2. Status Report Frame Format

[0101]

[0102]

[0103] This embodiment implements an adaptive adjustment mechanism for each network node to drive the reporting cycle based on the status information of the connected links. This mechanism enables dynamic adjustment of the reporting cycle, allowing for the rapid capture of more information when the link status deteriorates, and reducing the monitoring load when the link is stable. It also guides the rhythm of status report frames to present a non-uniform distribution that is strongly correlated with link risk, which is beneficial for system resource scheduling.

[0104] Step S1: The central management node of the vehicle-mounted fiber optic channel network obtains the status information of the links connected to each node in the current processing cycle based on the status report frames and historical status information received in the current processing cycle; wherein, the historical status information is obtained based on the historical status report frames.

[0105] Specifically:

[0106] If a status report frame from a network node is received in the current processing cycle, the status information of the link connected to the corresponding node is obtained from the status report frame.

[0107] If no status report frame from the network node is received in the current processing cycle, then determine whether a status report frame from the same network node was received in the previous processing cycle.

[0108] If so, the estimated value of the state information of the link connected to the network node is calculated based on the state information of the previous two processing cycles of the network node, and used as the state information of the current processing cycle;

[0109] If not, then calculate the estimated value of the state information of the link connected to the node based on the node's historical state information as the state information of the current processing cycle; or, calculate the estimated value of the state information of the link connected to the node based on the node's state information and penalty coefficient in the previous processing cycle as the state information of the current processing cycle; or, set the state information of the node in the current processing cycle to null.

[0110] Furthermore, if a status report frame from a network node is received in the current processing cycle, the status information of the link connected to the corresponding node is obtained from the status report frame. This status information includes the link frame loss rate, frame loss rate differential trend value, link round-trip time, weighted moving average of high-priority link round-trip time, high-priority link round-trip time delay trend differential value, and bit error rate. It should be noted that if multiple status report frames from the same network node are received in the current processing cycle, the status information is the average of the results reported by each status report frame.

[0111] Furthermore, if no status report frame from the network node is received in the current processing cycle, it is determined whether a status report frame from that network node was received in the previous processing cycle.

[0112] If so, the formula for calculating the estimated values ​​of each state information is:

[0113]

[0114] in, Let γ1 and γ2 be the m-th state information of network node i in the current processing period k', and let γ1 and γ2 be the weight information corresponding to the processing periods k'-1 and k'-2, respectively. γ1 + γ2 = 1, and γ1 > γ2.

[0115] If not, then for ordinary network nodes in the vehicular fiber channel network, the estimated value of the state information of the links connected to the node is calculated based on the node's historical state information. The calculation method is as follows:

[0116] Where, x m,neutral It can be the historical average of the m-th state information, or it can be the system's set value;

[0117] For critical nodes or high-risk nodes in the vehicular fiber channel network (critical nodes can be initially set by the system, while high-risk nodes are determined by the central management node through calculation, the calculation method of which will be disclosed later in this document), the estimated value of the state information of the links connected to the node is calculated based on the state information and penalty coefficient of the node in the previous processing cycle. The calculation formula is as follows:

[0118]

[0119] Where, α penaltyIt is a penalty factor, and α penalty A value greater than 1 indicates that "continuous missing reports may mean a hidden fault," and therefore the risk indicated by the status information of the corresponding node is amplified.

[0120] For ordinary network nodes in a vehicle-mounted fiber channel network, if the system evaluation mechanism can tolerate the condition of a small number of missing node data, the status information of the node in the current processing cycle can be set to null, thereby ensuring that the status information in the current processing cycle is completely based on the actual reported status report frame and improving the purity of the global evaluation.

[0121] In this embodiment, when no status report frame is received in the current processing cycle, a method is adopted to calculate the estimated value of the status information in the current processing cycle using historical status information. This application refers to this technical means as the "sliding window estimation mechanism" to solve the problem of information loss caused by asynchronous status report frame upload. It can still completely construct the status matrix in the scenario of non-aligned rhythm of multiple nodes; improve the continuity and timeliness of the overall network status assessment, and the system is no longer "dragged down by the slowest node"; and maintain the consistency of status information in each processing cycle.

[0122] Step S2: Dynamically adjust the processing cycle based on the status information of the current processing cycle, and predict network failure risks.

[0123] Specifically, it includes:

[0124] The central management node determines whether the network is in a high-risk state based on the status information of the current processing cycle. If so, it executes step S21, immediately performs network fault risk prediction, and ends the current processing cycle.

[0125] If not, proceed to steps S22-S23:

[0126] S22. The activity level of the status report frame is calculated based on the number of network nodes corresponding to the status report frame received in the current processing cycle and the total number of network nodes.

[0127] S23. Calculate a new processing cycle based on the activity level of the status report frame and the basic refresh cycle; and perform network fault risk prediction within the current processing cycle.

[0128] Specifically, the central management node determines whether the network is in a high-risk state based on the status information of the current processing cycle, including:

[0129] The number of high-risk nodes in the network is determined based on the status information of the current processing cycle.

[0130] If the number of high-risk nodes exceeds a set threshold, the network is judged to be in a high-risk state.

[0131] A network node is considered a high-risk node if it meets one of the following conditions:

[0132] The link frame loss rate FLR(k') is greater than or equal to the frame loss rate threshold FLR. threshold This is expressed as FLR(k')≥FLR threshold ;

[0133] Frame Drop Rate Differential Trend Value ΔFLR threshold If the cumulative value exceeds the cumulative threshold of the frame drop rate differential trend within a consecutive first set number of processing cycles (typically 3-5), it is expressed as: M is the first set number;

[0134] The link round-trip time (RTT)(k') exceeds the link round-trip differential threshold (RTT) within a second consecutive set number of processing cycles (preferably 3). threshold (k'), where RTT threshold The calculation method for (k') is: RTT threshold (k')=μ RTT (k')+α·σ RTT (k'), μ RTT (k') is the historical statistical mean. σ RTT (k') represents the historical standard deviation. W is the length of the historical window, which refers to the number of cycles before the k'th processing cycle, typically 20 processing cycles. α represents the differential sensitivity factor, with a preferred value of 3.

[0135] High-priority link round-trip weighted moving average RTT EWMA (k') is greater than or equal to the weighted moving average threshold RTT EWMA,threshold , represented as RTT EWMA (k)≥RTT EWMA,threshold ;

[0136] High-priority link round-trip delay trend difference ΔRTT EWMA (k') The cumulative value within the third consecutive set number (preferred value is 3) of the processing cycle is greater than the high priority trend cumulative threshold ΔRTT. EWMA,threshold , represented as Q is the third consecutive set number;

[0137] Bayesian probability outliers Greater than or equal to the abnormal probability set threshold P th ,in Let be the Bayesian probability outlier of network node i;

[0138] The link health status value is greater than the health status threshold H for two consecutive processing cycles. th .

[0139] Furthermore, the methods for calculating the Bayesian probability outliers of each network node include:

[0140]

[0141] Where X is a feature vector consisting of the link frame loss rate, link round-trip delay, weighted moving average of high-priority link round-trip delay, overall network health index, bit error rate, node CPU load, and node temperature of the calculated node.

[0142] P(fault) is the prior probability, which represents the estimated probability that the central management node is in a faulty state for the network node or link before observing the current data X. It is assigned a value based on historical statistics, empirical data or failure rate estimates in system design.

[0143] It should be noted that when calculating the probability of the observed current data X, it refers to jointly calculating the probability distribution model of each indicator value in X, that is, calculating the joint probability of the observed feature combination under the fault state. Even if the conditional probability of each feature may be small, their joint probability can effectively capture the overall fault risk through Gaussian distribution. For example, suppose a link connected to a network node has a minor performance problem, but can still send and receive data. In this case, the link round-trip delay and the link frame loss rate may still be within the normal range. However, if combined with other features (such as the overall network health index value), the combination of the entire feature vector X may still effectively capture the trend of performance degradation, thus giving a higher "fault probability" assessment in Bayesian inference.

[0144] P(X|fault) represents the probability of observing the current data X when a network node or link is actually in a fault state; the value is determined based on online learning or adaptive updates of the central management node.

[0145] In the numerator of the above equation, P(fault) and P(X|fault) are multiplied together to obtain the weighted probability of observing X under the fault state. That is, assuming that the current node is really in a fault state, and given the degree of prior belief in the fault state itself, the overall probability of observing X is as follows.

[0146] P(normal) represents the estimated probability that the network node or link will remain normal when the current data X is not observed;

[0147] P(X|normal) represents the probability of observing X when a network node or link is normal. For example, if the value of P(X|normal) is large, it means that "even if it looks like a sign of failure, it may often occur under normal conditions", so the final probability of failure will be reduced. The value is determined based on the online learning or adaptive update of the central management node.

[0148] The denominator of the above formula means that all possible sources of X have been observed, which guarantees that the final result is a normalized probability value in the interval [0,1].

[0149] Furthermore, the calculation method for the overall network health index is as follows:

[0150]

[0151] Where TSI(k') represents the overall network health index value in the k'th processing cycle; N represents the number of network nodes; i and j represent the i-th and j-th network nodes, respectively; l ij (k') represents the link state of link (i,j); p ij This represents the priority weight of link (i,j).

[0152] Furthermore, the central management node determines the link status l based on the status information of each stage of the current processing cycle. ij (k'), including:

[0153] A link is considered abnormal if it meets at least one of the following conditions; otherwise, the link is considered normal. These conditions include:

[0154] The link frame loss rate is greater than or equal to the frame loss rate threshold, such as 5% for example;

[0155] The link round-trip delay exceeds the delay threshold for N processing cycles; N is a positive integer, preferably 3, and the delay threshold is preferably three times the standard deviation of the historical delay mean; it should be noted that within a reporting cycle of a network node, regardless of whether the node sends one or more FDF frames, as long as the link round-trip delay exceeds the delay threshold once, it is recorded as the link round-trip delay exceeding the delay threshold within that reporting cycle.

[0156] If a fault diagnosis frame sent through this link fails to receive a response frame twice within a specified time, specifically, if a fault diagnosis frame sent through this link fails to receive a response frame within a specified time, it is characterized by ACK_Missed_Counts and stored in the link status field of the status reporting frame.

[0157] Within the period of sending two fault diagnosis frames and receiving the corresponding response frames through the link, the bit error rate exceeds a preset multiple of the current link benchmark. For example, the preset multiple can be 100%.

[0158] It should be noted that when the central management node determines that the network is in a high-risk state based on the status information of the current processing cycle, and immediately performs network failure risk prediction, it is necessary to recalculate the Bayesian probability outlier value and the Bayesian probability outlier trend value of each network node.

[0159] Furthermore, in step S21, the fault risk prediction includes:

[0160] The central management node calculates the corresponding overall network health index value, Bayesian probability outlier value, and Bayesian probability outlier trend value of each network node based on the status information of the current processing cycle.

[0161] When the link frame loss rate is greater than or equal to the frame loss rate threshold, or when the cumulative value of the frame loss rate differential trend value within a consecutive first set number of periods is greater than the frame loss rate differential trend threshold, it is predicted that there is a risk of intermittent single-point link failure.

[0162] When the link round-trip delay exceeds the link round-trip differential threshold within a second consecutive set number of periods, it is predicted that there is a risk of abnormal link transmission performance.

[0163] When the weighted moving average of round-trip delays of high-priority links is greater than or equal to the weighted moving average threshold, or when the cumulative value of the round-trip delay trend difference of high-priority links within the third consecutive set number of periods is greater than the cumulative threshold of high-priority trend, it is predicted that there is a risk of performance degradation of high-priority links.

[0164] When the overall network health index value TSI(k') is greater than or equal to the set health threshold TSI threshold When the growth trend exceeds the exponential growth rate threshold γ, a network congestion risk is predicted; the formula for calculating the growth trend is: The preferred value for γ is 0.1;

[0165] When the Bayesian probability of a network node is an outlier The abnormality probability is greater than or equal to the set threshold, or the cumulative value of the Bayesian probability abnormality trend value within a set number of consecutive periods is greater than the set abnormality trend threshold ΔP. threshold , represented as The risk of hidden or intermittent node failures is predicted, where R is the fourth set number.

[0166] Furthermore, the Bayesian probability anomaly trend value of the network node in the k-th period is calculated using the following formula:

[0167] ΔP(fault|X) k =P(fault|X) k -P(fault|X) k-1 ;

[0168] It should be noted that ΔP(fault|X) k The trend monitoring indicates that if the posterior probability continues to increase and approaches the threshold, it means that the risk of node failure is very likely to erupt in the short term.

[0169] Furthermore, in step S22, the formula for calculating the activity level of the status report frame is as follows: Where Λ(k') is the number of network nodes corresponding to the status report frame received in the current processing cycle, and N is the total number of network nodes.

[0170] Furthermore, in step S23, the new processing cycle is calculated based on the activity level of the status report frame and the basic refresh cycle, including:

[0171] T TSI (k')=T min +ρ·(1-α(k'))·T base

[0172] Among them, T min ρ is the minimum refresh interval (e.g., 10ms); ρ is the adjustment factor (e.g., 0.8), used to control the delay adjustment magnitude; T base Based on the refresh cycle (e.g., 100ms);

[0173] The calculation formula satisfies:

[0174] If α(k')≈1 (most nodes have reported status report frames), then T TSI (k)→T min Refresh quickly;

[0175] If α(k') is low, it indicates insufficient state information. Therefore, the refresh cycle should be appropriately delayed to avoid invalid judgments.

[0176] In this embodiment, the central management node dynamically adjusts the processing cycle based on the activity level of the status report frames in the current processing cycle, calculates the overall network health index, the Bayesian probability outlier value and the Bayesian probability outlier trend value of each network node, and performs fault risk prediction. This breaks away from the passive strategy of "fixed-cycle refresh" in the prior art, realizing rhythm-driven adaptive adjustment. This allows high-risk nodes that upload frequently to dominate the system's judgment rhythm, significantly improving system response efficiency. The fault risk prediction cycle is decoupled from the reporting cycle rhythm of each node but is adjusted in conjunction with it, improving the system's adaptability. Furthermore, the dynamic adjustment of the processing cycle is coordinated with the sliding "sliding window estimation mechanism," enabling accurate fault risk prediction even when some nodes do not report status report frames, thus improving the system's stability and security.

[0177] Step S3: Based on the predicted fault risk, issue a network optimization instruction frame carrying network optimization instructions to optimize the network.

[0178] include:

[0179] When a single point of failure is predicted to be intermittent, the central management node issues a network optimization instruction frame carrying link isolation or link activation instructions.

[0180] When a risk of abnormal link transmission performance is predicted, the central management node issues a network optimization instruction frame carrying link activation or topology reconfiguration instructions.

[0181] When a risk of performance degradation of a high-priority link is predicted, the central management node issues a network optimization instruction frame carrying link activation or topology reconfiguration instructions.

[0182] When a network congestion risk is predicted, the central management node issues a network optimization instruction frame carrying link activation or topology reconfiguration instructions.

[0183] When a hidden or intermittent node failure risk is predicted, the central management node issues a network optimization instruction frame carrying link isolation or topology reconfiguration instructions.

[0184] Furthermore, this invention defines a network optimization instruction frame sent by the central management node, which adopts an FC-2 layer frame structure. The specific frame format definition is shown in Table 3:

[0185] Table 3. Network optimization instruction frame format

[0186]

[0187] For example, the detailed correspondence between the predicted fault risks and the issued network optimization command frames is shown in Table 4:

[0188] Table 4. Correspondence between predicted fault risks and issued network optimization command frames

[0189]

[0190]

[0191] For example, the central management node sets extended parameters in the network optimization instruction frame to generate specific instructions. Each network node extracts specific instructions from the network optimization instruction frame to optimize the network. Table 5 shows examples of sub-instructions for optimization:

[0192] Table 5. Examples of Sub-instructions

[0193]

[0194]

[0195] In this embodiment, by constructing dedicated fault diagnosis frames, status report frames, and network optimization instruction frames, the structure is highly compatible with the existing FC network standard frame structure, easy to implement, and suitable for widespread industrial applications. By taking proactive preventive measures, the vehicle network is optimized by issuing network optimization instruction frames carrying network optimization instructions based on predicted fault risks. Based on fault prediction, proactive preventive measures can rationally allocate network resources according to priority, prioritizing the transmission of critical data, thereby improving the overall operational efficiency and reliability of the vehicle network, avoiding vehicle operation failures caused by network communication failures, reducing vehicle operation and maintenance costs, and more accurately identifying early signs of faults. Compared with traditional fault detection methods, it can detect potential problems earlier and reduce the impact of network communication failures on vehicle operation.

[0196] It should be noted that the above embodiments are based on the same inventive concept, and any parts not described repeatedly can be referenced from each other.

[0197] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A fault prediction method for FC networks supporting asynchronous state awareness and rhythm-driven refresh, characterized in that, Includes the following steps: The central management node receives status report frames from each network node; The status information of the links connected to each network node in the current processing cycle is obtained based on the status report frames received in the current processing cycle and historical status information; wherein, the historical status information is obtained based on historical status report frames; the status report frames received in the current processing cycle are reported by each network node according to its reporting cycle; the reporting cycle is dynamically determined by each network node based on the status information of the links connected to it; the network node calculates the status information of the connected links based on one or more fault diagnosis frames sent to other network nodes and the corresponding response frames received in each reporting cycle; The processing cycle is dynamically adjusted based on the status information of the current processing cycle, and network fault risk prediction is performed. This includes: the central management node determines whether the network is in a high-risk state based on the status information of the current processing cycle. If so, network fault risk prediction is performed immediately, and the current processing cycle ends. If not, the following steps are taken: the activity level of the status report frame is calculated based on the number of network nodes corresponding to the status report frame received in the current processing cycle and the total number of network nodes; a new processing cycle is calculated based on the activity level of the status report frame and the basic refresh cycle; and network fault risk prediction is performed within the current processing cycle. Based on the predicted failure risk, a network optimization instruction frame carrying network optimization instructions is issued to optimize the network.

2. The FC network fault prediction method according to claim 1, characterized in that, The reporting period is dynamically determined by each network node based on the status information of the connected links, including: Each network node calculates its current link health status value based on the current status information of the linked link. Each network node updates its reporting cycle based on the current link health status value and the link status risk threshold, including: ;in, , They represent the first and the One reporting cycle; Minimum reporting period; The default reporting cycle; To reduce the step size during updates; To increase the step size; For network nodes During the reporting period The link health status value; This is the threshold for link state risk.

3. The FC network fault prediction method according to claim 2, characterized in that, The status information includes the link frame loss rate, frame loss rate differential trend value, link round-trip time, weighted moving average of high-priority link round-trip time, trend differential value of high-priority link round-trip time, and bit error rate of the connected link. The step of obtaining the status information of each network node's connected link corresponding to the current processing cycle based on the status report frame received in the current processing cycle and historical status information includes: If a status report frame from a network node is received in the current processing cycle, the status information of the link connected to the corresponding node is obtained from the status report frame. If no status report frame from the network node is received in the current processing cycle, then determine whether a status report frame from the same network node was received in the previous processing cycle. If so, the estimated value of the state information of the link connected to the network node is calculated based on the state information of the previous two processing cycles of the network node, and used as the state information of the current processing cycle; If not, then calculate the estimated value of the state information of the link connected to the node based on the node's historical state information as the state information of the current processing cycle; or, calculate the estimated value of the state information of the link connected to the node based on the node's state information and penalty coefficient in the previous processing cycle as the state information of the current processing cycle; or, set the state information of the node in the current processing cycle to null.

4. The FC network fault prediction method according to claim 3, characterized in that, The fault risk prediction includes: The central management node calculates the corresponding overall network health index value, Bayesian probability outlier value, and Bayesian probability outlier trend value of each network node based on the status information of the current processing cycle. When the link frame loss rate is greater than or equal to the frame loss rate threshold, or when the cumulative value of the frame loss rate differential trend value within a consecutive first set number of processing cycles is greater than the cumulative threshold of the frame loss rate differential trend, it is predicted that there is a risk of intermittent single-point link failure. When the link round-trip delay exceeds the link round-trip differential threshold within a second consecutive set number of processing cycles, it is predicted that there is a risk of abnormal link transmission performance. When the weighted moving average of round-trip delay for high-priority links is greater than or equal to the weighted moving average threshold, or when the cumulative value of the round-trip delay trend difference for high-priority links within the third consecutive set number of processing cycles is greater than the cumulative threshold for high-priority trends, it is predicted that there is a risk of performance degradation of high-priority links. When the overall network health index value is greater than or equal to a set health threshold, or its growth trend exceeds the exponential growth rate threshold, a network congestion risk is predicted. The calculation method for the overall network health index value is as follows: ;in, Indicates the first The overall network health index value for each processing cycle; N represents the number of network nodes; , They represent the first one. , Network nodes, ; Indicates link The link status; Indicates link Priority weights; When the outlier value of the Bayesian probability of a network node is greater than or equal to the set threshold for outlier probability, or when the cumulative value of the outlier trend value of the Bayesian probability is greater than the set threshold for outlier trend in the fourth consecutive set number of processing cycles, a hidden or intermittent node failure risk is predicted.

5. The FC network fault prediction method according to claim 4, characterized in that, The central management node optimizes the network by issuing network optimization instruction frames carrying network optimization instructions based on the predicted failure risk, including: When a single point of failure is predicted to be intermittent, the central management node issues a network optimization instruction frame carrying link isolation or link activation instructions. When a risk of abnormal link transmission performance is predicted, the central management node issues a network optimization instruction frame carrying link activation or topology reconfiguration instructions. When a risk of performance degradation of a high-priority link is predicted, the central management node issues a network optimization instruction frame carrying link activation or topology reconfiguration instructions. When a network congestion risk is predicted, the central management node issues a network optimization instruction frame carrying link activation or topology reconfiguration instructions. When a hidden or intermittent node failure risk is predicted, the central management node issues a network optimization instruction frame carrying link isolation or topology reconfiguration instructions.

6. The FC network fault prediction method according to any one of claims 4-5, characterized in that, The central management node determines whether the network is in a high-risk state based on the status information of the current processing cycle, including: The number of high-risk nodes in the network is determined based on the status information of the current processing cycle. If the number of high-risk nodes exceeds a set threshold, the network is judged to be in a high-risk state. A network node is considered a high-risk node if it meets one of the following conditions: The link frame loss rate is greater than or equal to the frame loss rate threshold; The cumulative value of the frame drop rate differential trend value is greater than the cumulative threshold of the frame drop rate differential trend within a first set number of consecutive processing cycles; The link round-trip delay exceeds the link round-trip differential threshold within a second consecutive set number of processing cycles; The round-trip weighted moving average of high-priority links is greater than or equal to the weighted moving average threshold; The cumulative value of the high-priority link round-trip delay trend difference is greater than the high-priority trend cumulative threshold within the third consecutive set number of processing cycles; The Bayesian probability outlier is greater than or equal to the outlier probability threshold. The link health status value is greater than the link status risk threshold for two consecutive processing cycles.

Citation Information

Patent Citations

  • Intelligent operation and maintenance terminal based on data analysis

    CN119892227A

  • Internet of Things central control host data communication method and system based on protocol

    CN119946030A