Hierarchical collaborative fault-tolerant method and system based on ship operation
Patent Information
- Application Number
- CN202611092267.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-22
- Publication Date
- 2026-09-25
AI Technical Summary
对于高度依赖物理力矩平衡的船舶而言,这种瞬态的控制权跳变会诱发执行器之间的逻辑冲突(例如舵机反向抢占或推进器功率突跳),导致船舶在极短时间内产生难以补偿的偏航力矩
[0026]1、本发明将多源异构的运行状态数据统一映射至同一量化尺度下,使得系统能够以“场”的形式直观感知各计算节点的健康态势,从而实现了对节点潜在故障的早期识别与量化评估,为后续的容错决策提供了准确且全面的数据支撑。
Smart Images

Figure CN122816083A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ship automatic control technology, specifically to a hierarchical collaborative fault-tolerant method and system based on ship operation. Background Technology
[0002] In the architectural design of modern intelligent ships, distributed control systems have become the core foundation for achieving high-precision trajectory tracking and dynamic positioning. These systems rely on real-time information exchange between sensors, computing nodes, and high-power actuators (such as azimuth thrusters and side thrusters) deployed across different areas. To mitigate the impact of single-point failures on navigation safety, existing fault-tolerant designs often employ a static backup mode based on hardware redundancy. This means that when the main control unit experiences a logical failure or communication interruption, a heartbeat monitoring mechanism forcibly transfers control to a backup unit.
[0003] However, ship maneuvering is inherently a strongly coupled, highly inertial, and significantly time-delayed nonlinear physical process. Existing fault-tolerant paradigms often completely separate the logical switching in the "information domain" from the dynamic response in the "physical domain," neglecting the dynamic permeation of potential sub-optimal states within the system onto maneuvering stability. In real-world operating environments (such as high sea states or narrow waterways), localized load surges in computing nodes or random disturbances in communication links do not immediately cause system failure, but rather accumulate fault potential energy in the form of "implicit delays" or "data jitter." Due to a lack of macroscopic awareness of the propagation path of this fault potential energy, traditional redundancy switching logic is typically in a "lagging trigger" state.
[0004] The combination of this perception gap and physical lag can lead to severe control imbalances when ships execute critical control maneuvers. Due to subtle differences in command processing cycles among distributed heterogeneous nodes, the system often experiences "discontinuities" or "overlaps" in command output during the transient window of control transfer. For ships that heavily rely on physical torque balance, such transient control jumps can induce logical conflicts between actuators (such as reverse rudder preemption or sudden power spikes in the propellers), causing the ship to generate an uncompensable yaw moment within a very short time. This "secondary disturbance" caused by the fault-tolerance mechanism itself can easily lead to closed-loop instability of the control system, seriously threatening the ship's collision avoidance capabilities and structural safety in complex environments. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a hierarchical collaborative fault-tolerant method and system based on ship operations.
[0006] To achieve the above objectives, the technical solution of the present invention is as follows:
[0007] In a first aspect, the present invention discloses a hierarchical collaborative fault-tolerant method based on ship operations, comprising the following steps:
[0008] The system receives raw control commands from the human-computer interaction layer and generates control tasks, which are then executed by each computing node.
[0009] The distributed perception layer collects the computing load of computing nodes, the communication latency between computing nodes, and the frequency of actuator action feedback in real time, and aggregates them into a global state snapshot.
[0010] Based on the global state snapshot, a fault potential energy value representing the sub-healthy state of a node is generated through nonlinear normalization processing, and a global fault potential energy field is constructed accordingly.
[0011] Based on the global fault potential energy field, computing nodes whose fault potential energy values exceed a preset migration trigger threshold are identified as nodes to be migrated, and candidate nodes associated with the nodes to be migrated are determined.
[0012] Based on the candidate node’s own fault potential energy value, the weighted potential energy value of neighboring nodes and the remaining computing capacity, the risk coupling index is calculated through a preset risk coupling function, and the candidate node with the smallest risk coupling index is selected as the target migration node.
[0013] The control task containing the original control instructions is redirected from the node to be migrated to the target migration node, and the scheduling frequency of the target migration node is set to the frequency value corresponding to the preset emergency priority. The reciprocal of the frequency value is used as the time difference benchmark for characterizing the control delay during the task migration process.
[0014] Based on the time difference reference, the original control command is dynamically compensated to generate a compensated control sequence;
[0015] The compensation control sequence is sent to the actuator for physical conflict verification and interception, and the control action that actually drives the ship's actuator is output.
[0016] Secondly, this invention discloses a hierarchical collaborative fault-tolerant system based on ship operations, which uses the aforementioned hierarchical collaborative fault-tolerant method based on ship operations, including:
[0017] Instruction awareness module: used to receive raw control instructions from the human-computer interaction layer and generate control tasks, which are then executed by each computing node.
[0018] Heterogeneous data acquisition module: used to collect computing load of computing nodes, communication latency between computing nodes and actuator action feedback frequency in real time through distributed perception layer, and collect them into a global state snapshot;
[0019] Potential energy modeling module: Based on the global state snapshot, it generates fault potential energy values that characterize the sub-healthy state of nodes through nonlinear normalization processing, and constructs a global fault potential energy field accordingly.
[0020] Migration decision module: Based on the global fault potential energy field, it identifies computing nodes whose fault potential energy values exceed a preset migration trigger threshold as nodes to be migrated, and determines the candidate nodes associated with the nodes to be migrated.
[0021] The target determination module is used to calculate the risk coupling index based on the candidate node’s own fault potential energy value, the weighted value of the potential energy of neighboring nodes, and the remaining computing capacity, and select the candidate node with the smallest risk coupling index as the target migration node.
[0022] The collaborative scheduling module is used to redirect control tasks containing original control instructions from the node to be migrated to the target migration node, and set the scheduling frequency of the target migration node to the frequency value corresponding to the preset emergency priority, using the reciprocal of the frequency value as the time difference benchmark for characterizing the control delay during the task migration process.
[0023] Dynamic compensation module: used to dynamically compensate the original control command based on the time difference reference, and generate a compensated control sequence;
[0024] Execution module: Used to send the compensation control sequence to the actuator end for physical conflict verification and interception, and output the actual control actions that drive the ship's actuators.
[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0026] 1. This invention maps multi-source heterogeneous operational status data to the same quantization scale, enabling the system to intuitively perceive the health status of each computing node in the form of a "field". This achieves early identification and quantitative assessment of potential node failures, providing accurate and comprehensive data support for subsequent fault-tolerant decisions.
[0027] 2. The decision-making mechanism of this invention not only focuses on the health status of individual nodes, but also incorporates the mutual influence between nodes into the evaluation system, avoiding the risk of secondary failures caused by the overload of neighboring nodes due to task migration. Simultaneously, after redirecting the control task to the target node, by increasing its scheduling frequency and using the reciprocal of the frequency value as the time difference benchmark, the real-time performance and predictability of task processing during the migration transient period are ensured, achieving a smooth switchover of the control task.
[0028] 3. This invention effectively fills the control gap caused by instruction interruption during task migration through a dynamic compensation mechanism, ensuring the continuity of output control; while physical conflict verification identifies and avoids the risk of mutual exclusion of actions that may be caused by compensation instructions at the actuator level. This application not only ensures the continuous flow of control logic at the node level, but also ensures the safety and reliability of the final control action at the physical execution level, forming a complete fault-tolerant closed loop from fault perception and task migration to instruction compensation and action execution. Attached Figure Description
[0029] The disclosure of this invention is illustrated with reference to the accompanying drawings. It should be understood that the drawings are for illustrative purposes only and are not intended to limit the scope of protection of this invention. In the drawings, the same reference numerals are used to refer to the same parts. Wherein:
[0030] Figure 1 This is a flowchart of the steps of the present invention;
[0031] Figure 2 This is a schematic diagram illustrating the working principle of the present invention;
[0032] Figure 3 This is a system module connection diagram of the present invention. Detailed Implementation
[0033] It is readily understood that, based on the technical solution of this invention, those skilled in the art can propose various interchangeable structural methods and implementations without altering the essential spirit of the invention. Therefore, the following detailed embodiments and accompanying drawings are merely illustrative examples of the technical solution of this invention and should not be considered as the entirety of the invention or as limitations or restrictions on the technical solution of this invention.
[0034] In existing technologies, fault tolerance in ship distributed control systems largely relies on node-level hardware redundancy or master-slave switching mechanisms, making it difficult to balance dynamic coordination of computing resources with the real-time nature of task migration. Traditional methods typically switch directly to a backup node after a preset threshold alarm when a node experiences performance degradation or increased communication latency. Under this approach, the system cannot synchronously perceive the coupled impact of multiple factors such as changes in computing load, communication link status, and actuator feedback on the performance of control tasks. Especially when multiple nodes are simultaneously in a sub-optimal state, a single-dimensional fault detection model may exhibit decision bias, easily leading to inappropriate selection of task migration targets, which in turn can cause control command interruptions or actuator action conflicts, making it difficult to meet the high reliability requirements of the control system for continuous ship navigation.
[0035] To address the aforementioned issues, in-depth research into the operational status of edge computing nodes revealed an inherent correlation between computational load, communication latency, and actuator feedback frequency. A nonlinear mapping model was constructed to fuse this multidimensional heterogeneous data into a quantitative indicator characterizing the sub-health level of nodes. Further research showed that the operational risk of a single node depends not only on its own state but also on the state of its neighboring nodes. Therefore, a strategy was proposed to dynamically perceive the overall health status of the system based on the global fault potential field. Further experimental verification introduced a transient time difference benchmark and dynamic instruction compensation mechanism into the control closed loop, forming a hierarchical collaborative fault-tolerant system encompassing fault perception, task migration, instruction correction, and physical conflict verification.
[0036] After introducing the basic concept of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0037] Example 1:
[0038] like Figure 1 As shown, the hierarchical collaborative fault-tolerant method based on ship operations includes the following steps:
[0039] It receives raw control commands from the human-computer interaction layer and generates control tasks, which are then executed by each computing node.
[0040] The distributed perception layer collects the computing load of computing nodes, the communication latency between computing nodes, and the frequency of actuator action feedback in real time, and aggregates them into a global state snapshot of the edge node.
[0041] Based on the global state snapshot, fault potential energy values representing the sub-health state of nodes are generated locally on each edge computing node through nonlinear normalization processing, and a global fault potential energy field is constructed accordingly.
[0042] Based on the global fault potential energy field, computing nodes whose fault potential energy values exceed the preset migration trigger threshold are identified as nodes to be migrated, and candidate nodes associated with the nodes to be migrated are determined.
[0043] Based on the candidate node’s own fault potential energy value, the weighted potential energy value of neighboring nodes and the remaining computing capacity, the risk coupling index is calculated through a preset risk coupling function, and the candidate node with the smallest risk coupling index is selected as the target migration node.
[0044] The control task containing the original control command is redirected from the node to be migrated to the target migration node, and the scheduling frequency of the target migration node is set to the frequency value corresponding to the preset emergency priority. The reciprocal of the frequency value is used as the time difference benchmark for characterizing the control delay during the task migration process.
[0045] Dynamically compensate the original control commands based on the time difference reference to generate a compensated control sequence;
[0046] The compensation control sequence is sent to the actuator for physical conflict verification and interception, and the actual control actions of the ship's actuators are output.
[0047] like Figure 2 The diagram shown illustrates the working principle of this application. The system first receives raw control commands (such as heading settings and thruster speed commands) input by the operator through the Human-Machine Interface (HMI). The command sensing module converts these analog or digital commands into standard control task packages and distributes them to multiple preset heterogeneous or homogeneous computing nodes according to the current ship control system topology. Each computing node executes the corresponding control algorithm in parallel under normal conditions, generating intermediate variables to drive the actuators.
[0048] While the task is being executed, the distributed perception layer uses a communication bus or industrial Ethernet to collect multi-dimensional operational characteristics of each computing node in real time and at high frequency. Specifically, the system acquires the processor load rate, memory usage, message exchange latency between computing nodes, and action feedback pulse frequency at the actuator end of each node. These discrete, distributed low-level data are time-aligned and spatially fused at the edge processing unit, and nanosecond-level clock alignment is performed through NTP or PTP protocols to generate a global state snapshot reflecting the overall operating status of the system at the current moment, providing a basic data source for subsequent quantitative evaluation.
[0049] The system utilizes a nonlinear normalization algorithm (such as...) The mapping process performs weighted fusion processing on heterogeneous data in the snapshot. By mapping physical indicators such as load, latency, and feedback frequency to a unified quantization range, a "fault potential value" is generated, characterizing the degree to which a single node deviates from its health baseline. This potential value not only reflects the current sub-health state of the node but also, combined with the topological connections between nodes, constructs a global fault potential field in the information logic space. This field model can reveal the penetration trend of fault risk among nodes, transforming traditional Boolean-type fault determination into continuous risk gradient assessment.
[0050] Based on the constructed global fault potential energy field, the system continuously monitors the dynamic changes in the potential energy values of each node. When the potential energy value of a node exceeds a preset migration trigger threshold... When a node is determined to be in an unreliable state, the system automatically identifies it as a node to be migrated. Subsequently, the system searches for neighboring nodes that have communication link associations or task coupling relationships with the node to be migrated, filters out computing entities with the potential to receive tasks, and thus locks in the set of candidate nodes.
[0051] The system further invokes a pre-defined risk coupling function to perform a multi-dimensional evaluation of each node in the candidate set. The evaluation process considers not only the fault potential energy of the candidate node itself, but also the weighted impact of its influence from neighboring nodes' risks, calculated using topological weights, and takes into account the node's remaining computing capacity (such as available step size or remaining bandwidth). Through this multi-indicator coupled calculation, the system selects the node with the lowest overall risk and strongest stability from the candidate set, officially confirming it as the target migration node to ensure environmental safety after the task migration.
[0052] Once the target is identified, the system executes a redirection protocol to seamlessly migrate the original control commands and real-time context data carried by the node to be migrated to the target migration node. Simultaneously, the system temporarily increases the scheduling frequency of the target migration node in the kernel to the emergency level through a priority inheritance mechanism. By adjusting the scheduling weight of the target migration node, its scheduling frequency reaches the frequency value corresponding to the preset emergency priority, typically set to 2 to 5 times the normal scheduling frequency. At this point, the system uses the reciprocal of the scheduling frequency corresponding to this preset emergency priority (i.e., the task processing cycle) as a benchmark to accurately characterize the micro-control delays caused by task migration, data unpacking, and environment reconstruction, forming a time difference benchmark for subsequent compensation.
[0053] The dynamic compensation module uses this time difference as the independent variable offset to perform time-domain prediction of the original control commands. By analyzing the rate-of-change characteristics of the control commands, it uses an extrapolation algorithm to compensate for the delay jumps caused by node switching, generating a set of compensation control sequences that change continuously over time. This sequence can fill the control gaps during the transition transients, ensuring the smoothness and consistency of the control output on the time axis.
[0054] Before the generated compensation control sequence is sent to the actuator, it must be intercepted and audited by the physical conflict verification module. The system compares the predicted commands with the pre-stored mutual exclusion logic of the ship's mechanisms to identify and eliminate illegal actions that may cause interference or mechanical damage. The sequence that passes the verification is finally converted into actual electrical signals to drive actuators such as the steering gear and propellers, achieving precise control after fault tolerance.
[0055] It should be noted that, in order to meet the real-time requirements of the system, each node asynchronously updates its fault potential energy view based on local data, and ultimately achieves global logical consistency through a coordination protocol.
[0056] By modeling the fault potential energy field, early detection of node sub-health states is achieved, effectively reducing the operational risks caused by sudden faults. Utilizing emergency scheduling frequencies to anchor time difference benchmarks and perform predictive compensation eliminates control command gaps and delays caused by transient task migrations, ensuring the smooth dynamic transition of large-inertia vessels during fault-tolerant switching. Simultaneously, the end-point physical conflict verification mechanism blocks unauthorized or interfering commands that may be generated by the fault-tolerant algorithm, providing multi-level safety guarantees from the information layer to the physical layer, and improving the reliability and stability of the ship's automated maneuvering system in extreme environments.
[0057] This application further proposes specific steps for generating fault potential energy values characterizing the sub-health state of nodes. Based on a global state snapshot, this application performs deep feature extraction and nonlinear quantization. Specifically, within each sampling period, the system extracts multi-dimensional feature indicators of each computing node in real time through a pre-set monitoring agent process. These indicators not only cover computational load values reflecting the intensity of computational resource consumption, but also... It also includes communication latency values that characterize the efficiency of data interaction. And the actuator feedback frequency deviation value, which reflects the underlying physical execution state. Preferably, the load value is calculated. The communication latency value can be obtained by weighting the node's kernel utilization and task queue depth. It is defined by the sliding window average of the round-trip time (RTT) between nodes, while the actuator feedback frequency deviation value is... The frequency of the encoder or sensor feedback from the actuator is directly collected to monitor whether there is command backlog or response lag on the physical side.
[0058] In one embodiment, to eliminate the influence of different dimensions on the evaluation results, the system performs linear normalization on the extracted raw indicators and then uses a potential energy mapping function for quantification. The calculation model is shown below:
[0059] ;
[0060] in, The output fault potential energy value is used to intuitively characterize the absolute degree to which the computing node deviates from its healthy state at the current moment;
[0061] This is a preset weighting coefficient, typically set between [0.1, 1.0], with the specific value preset based on the sensitivity of the ship under different operating conditions. For example, under precision berthing conditions, the system will increase the weighting coefficient. The weight is adjusted to increase sensitivity to physical feedback fluctuations; while under ocean cruising conditions, the weight is increased. To cope with the jitter that may occur in long-distance communication;
[0062] All input items are positive real numbers (load, latency, and frequency drop are all non-negative). The application of the function ensures that the potential energy value is strictly limited to the normalized interval [0,1], where 0 approaches absolute health and 1 represents complete node failure. Besides the tanh function, nonlinear normalization can also be achieved using the sigmoid function, etc.
[0063] In one embodiment, to eliminate the influence of different dimensions on the evaluation results, the system first performs linear normalization on the extracted original indicators, mapping them to the [0,1] interval:
[0064] ;
[0065] Quantization calculations were then performed using a potential energy mapping function. .
[0066] Furthermore, the determination of the aforementioned weight parameters and mapping relationships can be optimized based on a pre-trained logistic regression model or a shallow neural network. Model training is completed on an offline server simulating the ship's operating environment. Training conditions include: the sample dataset contains 10 typical operating conditions such as normal navigation, sensor interference, and node overload. A typical ship operation log was used, containing fault injection samples such as normal propulsion, sensor interference, and compute node overload. Training parameters were set with a learning rate of 0.01, 500 iterations, and mean squared error (MSE) as the loss function. Through training, the system can learn the optimal coupling relationship between load, delay, and frequency under different fault modes, enabling the generated potential value to accurately identify the "sub-healthy" critical point of a node. For example, when the system detects that the load of a node exceeds... And the delay fluctuation exceeds When the preset threshold range is met, the fault potential energy value It will quickly slide into the high range, thus triggering subsequent task migration decisions ahead of time.
[0067] Through the aforementioned technical means, this application achieves refined perception and unified quantitative assessment of the operational status of computing nodes. This mapping mechanism effectively transforms heterogeneous and discrete operational data into physically meaningful risk measurement indicators, solving the problems of lag and false alarms in traditional hard threshold judgment methods when identifying sub-healthy system states. Due to the introduction of nonlinear mapping and multi-dimensional feature coupling, the system can capture potential hidden dangers where a single indicator is normal but the overall trend is abnormal. This provides accurate triggering basis for subsequent time difference benchmark anchoring and dynamic prediction compensation, ensuring the continuity and robustness of ship control logic under complex operating conditions.
[0068] After identifying the nodes to be migrated, this application achieves precise selection of target nodes through multi-dimensional risk topology modeling. Specifically, the system first retrieves the physical and logical neighborhood node sets of each candidate node in the ship's communication topology, and uses a distributed sensing protocol to obtain the current fault potential energy value of each node in this neighborhood set in real time. The significance of this step is that the stability of the ship's control system depends not only on the state of individual nodes, but also on the topological environment in which they are located; if the nodes around a candidate node are generally in a high potential energy (sub-healthy) state, then the candidate node is highly susceptible to cascading failures after undertaking the migration task.
[0069] In one embodiment, the system performs a fusion calculation on the aforementioned multidimensional risks using a preset risk coupling function. Its core mathematical model is shown below:
[0070] ;
[0071] in, The output risk coupling index is used to characterize the potential risk level of candidate node j receiving and executing the migration task;
[0072] The fault potential energy value of candidate node j itself;
[0073] Let j be the set of neighboring nodes of candidate node j;
[0074] Let be the fault potential energy value of neighboring node k;
[0075] The coupling weight coefficient between candidate node j and neighboring node k is preset based on the inter-node communication bandwidth or task correlation.
[0076] This is the capacity weighting coefficient, used to adjust the degree of influence of remaining computing capacity on risk coupling indicators. It is usually preset in the range of [0.2, 0.8] based on the communication distance between nodes or the degree of task correlation.
[0077] It is a normalized capacity factor that is positively correlated with the remaining computational capacity of candidate node j.
[0078] The numerator of the formula is obtained by accumulating the potential energy of the candidate nodes themselves. Potential energy of neighboring nodes The weighted sum constructs a measure of spatial "risk concentration," while the denominator introduces a capacity weighting coefficient. With normalized capacity factor ,in It is positively correlated with the node's remaining computing resources (such as remaining memory and idle CPU cycles). When the candidate node's remaining resources are exhausted ( When ), the denominator Approaching the minimum value, the risk coupling index The relative increase thus inhibits the migration of tasks to overloaded nodes.
[0079] Decreasing denominator leads to risk coupling indicators It exhibits a non-linear increase, thereby logically inhibiting the migration of tasks to overloaded nodes.
[0080] Specifically, the settings of the above parameters are not fixed, but dynamically adjusted based on the actual control requirements of the ship. For example, in a typical dynamic positioning (DP) system containing 5 edge computing nodes, when the system determines that node A needs to migrate, if candidate node B is healthy ( However, its neighboring nodes are all under high load. ), then the calculated This will be significantly higher than node C, which has a cleaner neighborhood environment. Furthermore, the capacity factor... The calculation can employ a linear normalization method, mapping the remaining computing capacity of a node to the [0,1] interval. If the CPU remaining rate of a node falls below the 20% warning threshold, the system will... The regulatory effect makes Rapidly exceeding the target selection threshold The node is selected as the target migration node. Preferably, Set it between 0.5 and 0.75 to ensure that the target node is in a low-risk zone.
[0081] This application achieves intelligent screening of mission migration targets through a risk assessment mechanism based on topological coupling and resource constraints. This method overcomes the limitations of traditional migration decisions based solely on the load of a single node, deeply integrating "individual health" and "environmental risk." This multi-dimensional risk quantification ensures that control tasks are always smoothly redirected to the most stable and resource-rich nodes within the system, providing a stable computational foundation for subsequent dynamic compensation based on time difference benchmarks. This fundamentally reduces the probability of secondary crashes during fault-tolerant switching, ensuring the overall safety of the ship's maneuvering logic.
[0082] This application, after determining the micro-delay caused by task migration, achieves seamless connection of control commands through a time-domain prediction algorithm. Specifically, the system obtains the scheduling frequency of the target migration node and anchors it to a time difference benchmark. Then, this offset is used as the independent variable offset in the dynamic evolution and substituted into the preset Taylor expansion prediction model. The purpose of this step is to artificially "compensate" for the damaged control quantity by utilizing the historical evolution trend of the instructions during the transient window period before the target node officially takes over control, thereby eliminating the actuator instruction jump caused by the switching of computing nodes.
[0083] In one embodiment, the system utilizes the raw control commands received at the current moment. The first derivative is extracted in real time using a numerical differentiator. (Representing the rate of change) and the second derivative (Representing acceleration characteristics). Based on the Taylor series expansion principle, when the signal-to-noise ratio of the command signal is below a threshold or the rate of change is drastic, it automatically degenerates into zero-order hold or first-order prediction. The compensated predicted command value. The calculation model is shown below:
[0084] ;
[0085] in, This is the nominal control value issued by the human-computer interaction layer at the current moment;
[0086] in, That is, the generated compensation control sequence in The value at time;
[0087] This serves as a time difference benchmark derived from the emergency scheduling cycle of the target node, representing the cumulative delay offset of the scheduling cycle. Where k is the cycle multiple introduced during the migration process, and its value is a set of positive integers {1,2,3,4}, used to adjust the time delay of fault confirmation. The selection of k depends on the task's tolerance for delay. The preset task processing cycle is the reciprocal of the scheduling frequency f, i.e. ;
[0088] Instruction trends used to compensate for linear changes;
[0089] It specifically addresses the nonlinear correction of acceleration characteristics during severe maneuvering or large inertia turns of ships.
[0090] Preferably, the time difference reference The value range is usually set according to the physical link time of task migration. Between. If the migration delay is extremely small (e.g., below...) The system can automatically degenerate into first-order prediction to reduce computational overhead; if in complex operating conditions, a second-order term is introduced to ensure that the predicted command value and the trajectory of the actual operation intention coincide.
[0091] Specifically, to prevent the prediction model from generating high-frequency oscillations when the input signal is noisy, the system needs to set a cutoff frequency before calculating the first and second derivatives. A low-pass filter. In practical ship dynamic positioning (DP) scenarios, when a task is redirected from the node to be migrated to the target node, if... Determined as The system will extrapolate the command gradient based on the past 3 to 5 sampling periods using the above formula. The expected value of the subsequent instructions. Through this feedforward compensation mechanism, the first frame of instructions issued by the target node will be precisely aligned with the current physical state of the executor, avoiding the "instruction step" phenomenon commonly found in traditional fault-tolerant schemes.
[0092] Through this Taylor expansion-based dynamic prediction mechanism, this application achieves high-fidelity smoothing of control commands during task transition transients. This method not only considers the time delay of control variables but also delves deeper into the dynamic changes in operational intentions. By combining first- and second-order gradient extrapolation, the system can generate predictive compensatory control sequences, compensating for control performance losses caused by network congestion or node reconfiguration. This deep coupling between feedforward control logic and underlying fault migration logic ensures that the ship's actuators are unaware of drastic fluctuations in the logic layer throughout the fault-tolerant cycle, guaranteeing the linearity and safety of ship maneuvering from the perspective of data flow evolution.
[0093] After generating the predicted command value, in order to prevent overcompensation, the system further performs compensation amount saturation limiting processing. This application introduces a nonlinear dynamic constraint mechanism after generating the predicted compensation command to ensure the safety boundary of the ship's physical actuators.
[0094] Specifically, after calculating the predicted command value based on Taylor expansion, the system does not directly convert it into a control current or voltage signal. Instead, the dynamic compensation module monitors in real time the absolute deviation between the predicted command value and the original control command issued by the human-machine interface layer. The significance of this monitoring process is that, although Taylor extrapolation can compensate for time delay, if there are drastic fluctuations in the input signal or transient logic disturbances in the calculation node, an excessively large compensation amount (i.e., a surge in higher-order terms) may cause the actuator (such as a servo motor or side thruster) to produce a response exceeding its mechanical limits, thereby triggering systemic nonlinear oscillations.
[0095] In one embodiment, the system presets a safety ratio threshold. This threshold is typically set between 5% and 15% of the original command amplitude, based on the physical response frequency and inertial time constant of the ship's actuators; for example, 8% for container ships and 12% for fishing vessels. The system invokes the saturation operator. The logic control model for real-time truncation of the predicted compensation term is shown below:
[0096] ;
[0097] in, This is the final instruction value output to the executor;
[0098] Original nominal instructions;
[0099] The original compensation increment consists of first-order and second-order derivative terms.
[0100] Saturation operator Its function is: when the absolute value of the compensation increment is... Less than or equal to When the value is in the specified range, maintain the original output value; however, when it exceeds the specified value... At that time, it will be forcibly anchored to Boundary. Preferably, for large inertial targets such as large container ships, this safety ratio threshold... It is recommended to preset the value to 8% to prevent the "reverse overshoot" phenomenon induced by excessive compensation.
[0101] Specifically, to achieve a smoother limiting transition, the system can use a soft-limiting algorithm when calling the saturation operator. In an embodiment involving thruster speed compensation in a dynamic positioning (DP) system, if the original speed command is... The compensation amount calculated by Taylor extrapolation is (If the deviation reaches 20%), then if Set to 10%, then the saturation operator The compensation amount will be forcibly truncated to Through this limiting mechanism, the system ensures that the rate of change of the compensation sequence on the time axis remains within the range that the actuator can physically track. This approach effectively mitigates the risk of numerical divergence that may occur in the prediction algorithm under extreme nonlinear conditions.
[0102] Through this dynamic compensation correction mechanism based on saturation limiting, this application achieves a deep synergy between algorithm robustness and physical safety. This processing step, acting as the "last valve" in the generation of the compensation sequence, solves the potential overcompensation problem in the feedforward compensation logic. By real-time monitoring and proportional truncation of the deviation amplitude, the system can ensure that the output control actions comply with the safety specifications of ship maneuvering under any task migration transient. This limiting logic, which balances the accuracy of time delay compensation with the smoothness of actuator action, further enhances the anti-interference capability of the hierarchical collaborative fault-tolerant system under complex sea conditions, and provides a closed-loop guarantee for the reliable output of the control logic from the physical execution dimension.
[0103] This application further proposes that the step of setting the frequency value corresponding to the preset emergency priority also includes scheduling enhancement processing. After determining the target migration node, this application ensures a high degree of determinism of the time difference benchmark during the task takeover process through kernel-level resource scheduling intervention. Specifically, when the system redirects the control task to the target migration node, it not only needs to complete the logical migration of data, but also needs to open a "high-priority scheduling channel" for the task at the underlying operating system kernel level. The necessity of this process lies in the fact that ship control tasks have extremely high real-time requirements. If the target node is currently processing non-critical business (such as log uploading or non-core monitoring), its original time sharding mechanism may cause unpredictable scheduling jitter in the control commands after migration, thereby destroying the time benchmark required by the Taylor expansion prediction model.
[0104] In one embodiment, the system monitors the current task queue length of the target migration node in real time through kernel monitoring operators. and average processing latency The system has a preset latency warning threshold. This latency warning threshold is typically set to 85% to 95% of the current task processing cycle (i.e., sampling cycle), such as 90% for berthing and 85% for cruising. When a high-risk penetration trend is detected in cross-level assessments, this latency warning threshold is dynamically reduced to achieve early intervention. When task backlog at the target node is detected, leading to an average processing latency... When the processing cycle approaches or exceeds this time, the system immediately triggers the priority inheritance mechanism. Specifically, this mechanism "inherits" the high-priority attributes of the task to be migrated to the relevant processing process of the target node. By calling kernel APIs (such as sched_setscheduler in the Linux kernel or related interfaces of the real-time patch PREEMPT_RT), the scheduling policy of the task is dynamically adjusted from ordinary round-robin sharding to the highest priority real-time scheduling.
[0105] Preferably, the dynamic adjustment of the scheduling weight adopts a proportional increment model. Its logical control process is as follows:
[0106] ;
[0107] in, The adjusted kernel scheduling weights, As the benchmark weight, This is a preset enhancement factor (the value usually ranges from 0.5 to 2.0). The preset task processing cycle, To average processing latency. In an embodiment involving main push motor control, if the task processing cycle... The average processing latency is 20ms. When the time reaches 18ms, the system will automatically increase the kernel weight of the process by 1.5 times. Through this weight tilt, the target migration node will force other non-real-time processes to suspend, ensuring that the control task completes its calculation and output within each scheduling cycle. This makes the fixed latency (time difference baseline) introduced by the task migration deterministic and predictable.
[0108] Specifically, to prevent prolonged occupation of kernel resources from paralyzing other system functions, the priority inheritance mechanism is equipped with automatic cancellation logic. This occurs when task migration transitions smoothly, control commands are continuously output for more than 10 cycles, and the task queue length is [not specified]. When the weights fall below a preset safety threshold (e.g., queue depth less than 2), the system automatically rolls back the scheduling weights to their initial state. This "transient enhancement, dynamic rollback" approach ensures that the target migration node can strictly align with the time difference benchmark derived from the scheduling frequency within the critical time frame for taking over the task, eliminating the deterministic deviation caused by task switching at the underlying scheduling level.
[0109] Through this enhancement mechanism based on priority inheritance and dynamic adjustment of kernel weights, this application achieves "hard real-time" assurance for control tasks during migration. This processing step resolves the scheduling uncertainty caused by resource contention in distributed systems, transforming the time difference benchmark from a logical prediction value into a kernel-protected physical execution standard. This fault-tolerant intervention, reaching the operating system level, ensures that the control sequences generated by the dynamic compensation module are delivered to the actuators on time and at the correct point, providing a solid computational resource guarantee for continuous operation of ships under high inertia and high dynamic conditions, and further solidifying the overall reliability of the hierarchical collaborative fault-tolerant system.
[0110] This application constructs a real-time safety barrier based on discrete logic operations at the actuator end. Specifically, this step serves as the final audit checkpoint before control commands enter physical execution, performing physical conflict verification and interception at the actuator end. This aims to solve the problem of physical interference or maneuvering conflict between actuators (such as servo motors and side thrusters, port and starboard thrusters) that may occur during task migration or dynamic compensation in distributed fault-tolerant systems due to concurrent algorithm output.
[0111] In one embodiment, the system pre-stores the mutual exclusion action logic of each actuator in physical space and quantifies it into a conflict feature matrix composed of binary Boolean values. Specifically, the rows and columns of this matrix correspond to different actuator actions in the system (such as "left steering", "right pushing", "forward", "reverse", etc.). If action i and action j cannot be executed simultaneously in the physical dimension, then the matrix elements... Set to binary logic 1 (indicating a conflict), otherwise set to 0. Preferably, this matrix is stored in the non-volatile memory of the actuator control unit, and the sampling period is synchronized with the control task.
[0112] Subsequently, the execution module will use the compensation control sequence generated by the dynamic compensation module. The calculation maps the commands to the task motion space of each actuator. Specifically, the calculation process uses a pre-defined ship maneuvering mechanics model to transform abstract torque or power commands into transient state values (such as target rudder angle and target speed) of specific actuators. After mapping, the system generates a transient motion vector representing the current set of actions to be executed. Each of these corresponds to the activation state of an actuator.
[0113] Furthermore, the system will use transient action vectors Conflict Feature Matrix Perform a bitwise AND operation. The processing logic is as follows:
[0114] ;
[0115] in, This is the conflict identification vector. If a non-zero bit appears in the calculation result, the system immediately identifies that the current compensation command has a mutual exclusion conflict at the physical level. Specifically, in an embodiment involving the coordinated operation of the side thrusters and the main thruster, if the compensation algorithm simultaneously requests the side thrusters to push left at full load and the main servo to deflect to the extreme right, the logical AND operation will locate the conflict bit and trigger the interception mechanism.
[0116] Once a conflict is identified, the system determines its severity based on the task's criticality. Transient action vectors are masked. Preferably, the task criticality is preset by the human-machine interface layer based on the current navigation conditions; for example, "collision avoidance task" has a higher priority than "energy-saving cruise task." During masking, the system generates a masking factor that automatically masks the action components corresponding to low-priority tasks, retaining only high-criticality actions. Finally, the system outputs the execution command after conflict suppression. This approach ensures that when logical competition occurs in the algorithm, the system can make real-time adjustments based on physical safety and task level, fundamentally eliminating the risk of mechanical damage or ship loss of control caused by fault-tolerant switching errors.
[0117] This application achieves decoupled collaboration between software algorithms and physical safety constraints through a physical interception mechanism based on bitwise operation matrices and mask processing. This method avoids the computational lag caused by traditional complex conditional branch judgments and utilizes low-level logical operations to achieve millisecond-level security auditing. This "hard logic" barrier established at the actuator front end compensates for instruction synchronization errors that may arise from distributed computing, providing a final physical safety line for hierarchical collaborative fault-tolerant systems and ensuring the deterministic maneuverability of ships under any extreme fault-tolerant transient conditions.
[0118] This application further proposes that the steps of real-time acquisition of computing load, inter-node communication latency, and actuator action feedback frequency through the distributed sensing layer also include implementation logic for data validity verification and dynamic compensation. This application constructs a self-healing preprocessing mechanism between the distributed sensing layer and the edge computing layer. Specifically, due to the high electromagnetic interference and complex salt spray fluctuations in the physical environment of ships, the heterogeneous operational data (such as load, latency, and frequency) transmitted back by sensors are prone to transient packet loss or outlier distortion. Directly inputting such abnormal data into the potential energy modeling module will cause false fluctuations in the fault potential energy field, thereby triggering incorrect migration decisions.
[0119] In one embodiment, the system monitors the sampling frequency of each sensor in real time via a sensing agent. and signal integrity index Specifically, signal integrity is evaluated by checking data packet verification and sequence number continuity. Preferably, the system presets a validity threshold range; for example, if the sampling frequency fluctuation exceeds the rated frequency... If the signal loss rate exceeds 2%, the system will automatically determine that the heterogeneous operating data collected at the current moment is "unreliable data".
[0120] When abnormal or missing sampled data is detected, the system triggers a snapshot interpolation completion protocol. Specifically, the system retrieves a global state snapshot from the previous decision cycle stored in the cache and uses a first-order zero-fold hold or linear interpolation algorithm to fill the current sampling gap with historical stable data. Simultaneously, the system adjusts the weighting factors... This reduces the update weight of the fault potential energy field at the current moment. The update logic model is as follows:
[0121] ;
[0122] in, For the global fault potential energy field, This is the transient field value estimated based on interpolated data. In an embodiment involving thruster feedback frequency loss, the system will... The value was dynamically reduced from the default 1.0 to 0.2. This means that during periods of data anomalies, the system chooses to "trust" historical trends more rather than be disturbed by abnormal fluctuations, thereby ensuring the smooth evolution of the potential energy field in the time domain.
[0123] Specifically, to ensure robust system reset, the data validity verification module is equipped with a counter mechanism. If the duration of continuous data anomalies exceeds a preset tolerance limit (determined based on the safety level of the ship's control system; in one embodiment, for course control tasks critical to navigation safety, the tolerance limit is set to 3 consecutive cycles; for non-critical energy efficiency monitoring tasks, the tolerance limit can be set to 10 consecutive cycles), the system will stop interpolation and output a "perception layer failure" alarm to the decision-making layer. This approach effectively prevents model drift risks caused by prolonged use of interpolated data.
[0124] This application achieves "soft fault tolerance" for the underlying data of the fault-tolerant system through this data verification mechanism based on snapshot interpolation and weight degradation. This method ensures that the fault potential field can filter out non-technical noise, maintaining the continuity of decision-making. This processing logic, which uses historical snapshots as "data anchors," enhances the survivability of the hierarchical collaborative fault-tolerant system under harsh sea conditions, guaranteeing the accuracy and reliability of task migration and dynamic compensation commands from the source.
[0125] This application further proposes that the construction of the global fault potential field also includes cross-level impact assessment, introducing a vertical-dimensional risk penetration analysis mechanism. Specifically, the ship control system is a typical layered architecture consisting of underlying physical hardware (such as sensors, actuator interface cards, and computing cores) and upper-level logical instructions (such as navigation algorithms, fault-tolerant protocols, and control laws). Traditional fault-tolerant schemes often only trigger a response when the logic layer detects an output anomaly, exhibiting significant lag. This application achieves predictive fault tolerance for potential systemic collapses by identifying the transmission path from the sub-health state of the underlying hardware to the upper-level logical operations.
[0126] In one embodiment, the system calculates the propagation gain coefficient in real time through a cross-level monitoring agent. Propagation gain coefficient A gain mapping table is pre-calibrated through offline simulation or fault injection experiments, and then retrieved from the table based on operating conditions during online runtime. Specifically, this coefficient is used to quantify the potential energy of underlying hardware faults. Stability of the upper logic instruction layer The degree of contribution. Preferably, the propagation gain coefficient is determined based on a preset association topology mapping table, and its calculation model is as follows:
[0127] ;
[0128] in, The underlying potential energy is calculated from hardware load, local temperature rise, or clock drift. This is a preset hierarchical coupling factor (typically ranging from [0.5, 2.5]). When an edge node is responsible for performing high-frequency dynamic positioning (DP) tasks, the system increases its coupling factor. Propagation gain coefficient. The higher the value, the more nonlinear the amplification effect of even minor hardware fluctuations at the lower level (such as slight jitter in the main frequency) on the control precision at the upper level.
[0129] Based on the calculated propagation gain coefficient, the system dynamically adjusts the preset migration trigger threshold. (Typically 0.6-0.8, dynamically adjusted according to ship operating conditions). Specifically, the system employs a reverse adjustment mechanism: when the propagation gain coefficient... When the potential energy threshold of the node to be migrated increases, the system automatically lowers it. The adjustment logic is as follows:
[0130] ;
[0131] in, This is the global baseline threshold (e.g., 0.75). This is a sensitivity adjustment parameter, with a value range of [0.1, 1.0]. It is preset based on the ship's current navigation conditions (e.g., entering or leaving port, open water) or dynamically adjusted by the upper-level control module to control the frequency and accuracy of fault detection. In an embodiment involving main engine fuel injection control, if the packet loss rate of the underlying network card increases slightly, but because this network card carries extremely high real-time injection commands, the propagation gain coefficient is calculated to be high. The system will then quickly shift the original preset trigger threshold. The value was lowered to 0.45. At this point, although the node has not yet reached a traditional failure state, it is in a "high-sensitivity risk penetration period," and the system will trigger task migration in advance to achieve early intervention against deep-level failure risks.
[0132] Specifically, this cross-level evaluation mechanism effectively addresses the systemic collapse caused by "hidden faults." By introducing a propagation gain coefficient, the system can identify critical states that, while not yet leading to logical errors, have already resulted in performance degradation at the physical layer. This intervention logic, based on a "risk penetration" perspective, transforms fault-tolerant decision-making from a simple "post-event switching" to "pre-event risk avoidance," significantly extending the fault-free operation time of ship automation systems under extremely complex conditions.
[0133] This invention provides a hierarchical collaborative fault-tolerant method and system based on ship operation. Through potential energy field modeling (sensing), risk-coupled decision-making (transfer), kernel scheduling enhancement (assurance), Taylor prediction compensation (smoothing), and cross-layer risk penetration assessment (prediction), a multi-dimensional and deep fault-tolerant protection system is constructed. This solution not only eliminates control command jumps caused by transients during task transfer, ensuring the physical safety of actuator actions, but also improves the robustness of the ship control system to minor disturbances at lower levels through cross-layer early intervention mechanisms. The entire technical solution aims to address the reliability challenges of large-inertia ships in distributed control environments, achieving seamless control logic integration and absolute safety in physical execution, demonstrating practical value and technological advancement.
[0134] Example 2:
[0135] like Figure 3 As shown, the hierarchical collaborative fault-tolerant system based on ship operations, using the aforementioned hierarchical collaborative fault-tolerant method based on ship operations, includes:
[0136] Command awareness module: Used to receive raw control commands from the human-computer interaction layer and generate control tasks, which are then executed by each computing node.
[0137] Heterogeneous data acquisition module: used to collect computing load, inter-node communication latency and actuator action feedback frequency of computing nodes in real time through the distributed perception layer, and collect them into a global state snapshot of the edge node.
[0138] Potential energy modeling module: Based on the global state snapshot, it generates fault potential energy values that characterize the sub-health state of nodes locally on each edge computing node through nonlinear normalization processing, and constructs a global fault potential energy field accordingly.
[0139] Migration decision module: Based on the global fault potential energy field, nodes whose fault potential energy values exceed the preset migration trigger threshold are identified as nodes to be migrated, and candidate nodes associated with the nodes to be migrated are determined.
[0140] The target determination module is used to calculate the risk coupling index based on the candidate node’s own fault potential energy value, the weighted value of the potential energy of neighboring nodes, and the remaining computing capacity, and select the candidate node with the smallest risk coupling index as the target migration node.
[0141] The collaborative scheduling module is used to redirect control tasks containing original control instructions from the node to be migrated to the target migration node, and set the scheduling frequency of the target migration node to the frequency value corresponding to the preset emergency priority. The reciprocal of the frequency value is used as the time difference benchmark for characterizing the control delay during the task migration process.
[0142] Dynamic compensation module: used to dynamically compensate the original control commands based on the time difference reference, and generate a compensated control sequence;
[0143] Execution module: Used to send the compensation control sequence to the actuator for physical conflict verification and interception, and output the actual control actions that drive the ship's actuators.
[0144] Specifically, the system first obtains raw control commands (such as heading angle and thruster speed) through a human-machine interface layer consisting of an industrial-grade control handle, a touch screen, and a host computer. These commands are converted into digital signals by the human-machine interface circuitry, and the central controller generates corresponding control task packages. At the hardware level, a cluster of computing nodes (such as embedded single-board computers or blade servers) receives these task packages via a redundant bus, and under normal operating conditions, each node executes the computational logic independently or in parallel, outputting the nominal control quantity in real time.
[0145] The distributed sensing layer utilizes hardware sensors deployed throughout the ship (such as built-in current / temperature rise monitoring modules in each node, switch flow monitoring modules, and high-precision encoders at the actuator end) to capture heterogeneous data in real time. The acquisition frequency is synchronized with the control cycle, and the specific data streams include: CPU utilization and task stacking depth of computing nodes (computational load), round-trip time of heartbeat packets between nodes (communication latency), and the pulse count per second fed back by the actuators (feedback frequency). Edge computing nodes perform time alignment of the above three types of timing signals through data preprocessing algorithms and aggregate them in local cache to generate a global state snapshot that can map the real-time operating boundary of the entire system.
[0146] The system's internal potential energy modeling module calls a preset mapping function to perform nonlinear normalization on the multidimensional heterogeneous data in the snapshot. This processing logic maps discrete operating parameters into scalars representing the degree of sub-health, i.e., fault potential energy values. If a node parameter deviates from a preset envelope (such as abnormal load pulses or frequency jitter), the fault potential energy value at that point increases. The system then combines the computational network topology and uses a diffusion model to correlate the potential energy values of each node to form a continuous global fault potential energy field, thereby intuitively demonstrating the penetration gradient of risk between the physical and information layers.
[0147] The migration decision module monitors gradient changes in the potential energy field in real time. Once the fault potential energy value of a node exceeds the preset migration trigger threshold, the system automatically marks it as a node to be migrated. At this time, based on the physical topology of the ship's local area network, the system automatically searches for surrounding nodes that have communication links with the node to be migrated and are in a low potential energy state, and delineates them as a candidate node set as a backup resource pool for task reception.
[0148] The target determination module invokes a risk coupling function to perform a secondary screening of the candidate set. This process involves logical judgment of three dimensions of data: first, the current sub-health level of the candidate node (its own potential energy value); second, the interference level of the node's environment in the topology (neighborhood potential energy weighted value); and third, its redundancy space for handling new tasks (remaining computing capacity). The system outputs a risk coupling index through composite weighted calculation and selects the node with the smallest index as the target migration node. This process ensures that the task is migrated to the physical entity with the most stable performance and optimal communication environment.
[0149] The collaborative scheduling module performs a task switching operation, redirecting the task packet containing the original instructions to the receive buffer of the target migration node via a switch. Simultaneously, the system modifies the kernel scheduling algorithm of the target node, elevating its priority to the emergency level. The hardware clock synchronization mechanism uses the period corresponding to the node's highest execution frequency under emergency priority as a benchmark to extract the time difference benchmark, which characterizes the entire delay from task loss, migration, to re-execution, and transforms it into input parameters for the prediction and compensation stage.
[0150] The dynamic compensation module receives the original control commands and the time difference reference, and executes an extrapolation compensation algorithm in the digital signal processor (DSP). The machine side calculates the rate of change of the commands and their derivative with respect to time, and predicts the target state value that the actuator should reach at the current moment by combining this with the time difference reference, thereby generating a set of time-continuous compensation control sequences. This step realizes the logical transformation from "command jump" to "smooth prediction".
[0151] Finally, the compensation control sequence is sent to the logic controller at the actuator end. The actuator end has a pre-stored physical conflict matrix (covering physical constraints such as servo motor limits and thruster mutual exclusion), and determines whether the compensation sequence triggers a physical conflict through logical comparison. After interception and verification, the safe command is converted into a PWM waveform or current signal to drive the thruster motor, servo motor hydraulic pump, and other actuators to complete the physical action.
[0152] This application effectively shortens the response lag of ship control systems during fault migration by synergizing "sub-health perception" and "deterministic compensation." Compared to traditional passive switching modes, the fault potential energy field-based prediction mechanism can identify potential risk points in advance. The time difference benchmark anchored by emergency frequencies improves the accuracy of command compensation and eliminates nonlinear shock fluctuations caused by switching transients. Finally, the conflict interception mechanism ensures the physical safety of ship actuators during fault-tolerant actions, achieving a smooth transition and high consistency of the control system at both the hardware and software levels.
[0153] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.
Claims
1. A hierarchical collaborative fault-tolerant method based on ship operations, characterized in that: Includes the following steps: The system receives raw control commands from the human-computer interaction layer and generates control tasks, which are then executed by each computing node. The distributed perception layer collects the computing load of computing nodes, the communication latency between computing nodes, and the frequency of actuator action feedback in real time, and aggregates them into a global state snapshot. Based on the global state snapshot, a fault potential energy value representing the sub-healthy state of a node is generated through nonlinear normalization processing, and a global fault potential energy field is constructed accordingly. Based on the global fault potential energy field, computing nodes whose fault potential energy values exceed a preset migration trigger threshold are identified as nodes to be migrated, and candidate nodes associated with the nodes to be migrated are determined. Based on the candidate node’s own fault potential energy value, the weighted potential energy value of neighboring nodes and the remaining computing capacity, the risk coupling index is calculated through a preset risk coupling function, and the candidate node with the smallest risk coupling index is selected as the target migration node. The control task containing the original control instructions is redirected from the node to be migrated to the target migration node, and the scheduling frequency of the target migration node is set to the frequency value corresponding to the preset emergency priority. The reciprocal of the frequency value is used as the time difference benchmark for characterizing the control delay during the task migration process. Based on the time difference reference, the original control command is dynamically compensated to generate a compensated control sequence; The compensation control sequence is sent to the actuator for physical conflict verification and interception, and the control action that actually drives the ship's actuator is output.
2. The hierarchical collaborative fault-tolerant method based on ship operation as described in claim 1, characterized in that: The specific steps for generating the fault potential energy value representing the sub-health state of a node are as follows: Extract multidimensional feature indicators of the computing node at the current time; the multidimensional feature indicators include at least the computing load value. Communication delay value and actuator feedback frequency deviation value ; Using potential energy mapping function Perform quantitative calculations and output fault potential energy values normalized to a specified range. The fault potential energy value Used to characterize the degree to which a computing node deviates from its healthy state at the current moment; where α, β, and γ are preset weighting coefficients.
3. The hierarchical collaborative fault-tolerant method based on ship operation according to claim 1, characterized in that: The specific steps for calculating the risk coupling index of each candidate node are as follows: Retrieve the set of neighboring nodes of each candidate node in the communication topology, and obtain the fault potential energy value of each node in the set of neighboring nodes; Calculate risk coupling index The formula is as follows: ; in, Let be the fault potential energy value of candidate node j itself. Let be the set of neighboring nodes of candidate node j. Let be the fault potential energy value of neighboring node k. The coupling weight coefficient between candidate node j and its neighboring node k. This is a capacity weighting coefficient used to adjust the degree of influence of remaining computing capacity on risk coupling indicators. It is a normalized capacity factor that is positively correlated with the remaining computational capacity of candidate node j.
4. The hierarchical collaborative fault-tolerant method based on ship operation according to claim 1, characterized in that: The specific steps for dynamically compensating the original control command based on the time difference reference include: A prediction model based on Taylor expansion is established, using the time difference benchmark as the offset of the independent variable; Extrapolation is performed by combining the first and second derivative terms of the original control command to generate a predicted command value to offset the effect of migration delay, which serves as the basis for the compensation control sequence.
5. The hierarchical collaborative fault-tolerant method based on ship operation according to claim 4, characterized in that: The step of generating the compensation control sequence also includes compensation amount saturation limiting processing, specifically including: Real-time monitoring of the deviation of the predicted command value from the original control command; If the deviation exceeds the preset safety ratio threshold, the saturation operator is invoked to truncate the compensation term.
6. The hierarchical collaborative fault-tolerant method based on ship operation according to claim 1, characterized in that: The step of setting the frequency value corresponding to the preset emergency priority also includes scheduling enhancement processing, specifically including: Monitor the current task queue length and processing latency of the target migration node; If the task processing latency is close to the task processing cycle, the priority inheritance mechanism is triggered to dynamically adjust the scheduling weight of the target migration node in the operating system kernel.
7. The hierarchical collaborative fault-tolerant method based on ship operation according to claim 1, characterized in that: The specific steps for physical conflict verification and interception at the actuator end include: Pre-store the mutual exclusion action logic of each actuator in physical space, and construct a conflict feature matrix composed of binary Boolean values; The generated compensation control sequence is solved and mapped to the task action space of each actuator to generate a transient action vector to be executed; The transient action vector is processed by the conflict feature matrix to identify whether there is a physical mutual exclusion conflict; If a conflict exists, the transient action vector is masked according to the task criticality to mask low-priority conflict components, and the execution instruction after conflict suppression is output.
8. The hierarchical collaborative fault-tolerant method based on ship operation according to claim 1, characterized in that: The step of collecting real-time data on the computing load of computing nodes, communication latency between nodes, and actuator action feedback frequency through the distributed perception layer also includes data validity verification, specifically including: Monitor the sampling frequency and signal integrity indicators of each sensor in the sensing layer; If abnormal or missing sampled data is detected, the data is interpolated and supplemented using the global state snapshot from the previous decision cycle, and the update weight of the fault potential energy field at the current moment is reduced.
9. The hierarchical collaborative fault-tolerant method based on ship operation according to claim 1, characterized in that: The construction of the global fault potential field also includes cross-level impact assessment, specifically including: Calculate the propagation gain coefficient of the potential energy of a hardware failure at the lower level to the upper logic instruction layer; The migration trigger threshold is adjusted based on the propagation gain coefficient to enable early intervention against deep-level fault risks.
10. A hierarchical collaborative fault-tolerant system based on ship operations, characterized in that: The hierarchical collaborative fault-tolerant method based on ship operation as described in any one of claims 1 to 9 includes: Instruction awareness module: used to receive raw control instructions from the human-computer interaction layer and generate control tasks, which are then executed by each computing node. Heterogeneous data acquisition module: used to collect computing load of computing nodes, communication latency between computing nodes and actuator action feedback frequency in real time through distributed perception layer, and collect them into a global state snapshot; Potential energy modeling module: Based on the global state snapshot, it generates fault potential energy values that characterize the sub-healthy state of nodes through nonlinear normalization processing, and constructs a global fault potential energy field accordingly. Migration decision module: Based on the global fault potential energy field, it identifies computing nodes whose fault potential energy values exceed a preset migration trigger threshold as nodes to be migrated, and determines the candidate nodes associated with the nodes to be migrated. The target determination module is used to calculate the risk coupling index based on the candidate node’s own fault potential energy value, the weighted value of the potential energy of neighboring nodes, and the remaining computing capacity, and select the candidate node with the smallest risk coupling index as the target migration node. The collaborative scheduling module is used to redirect control tasks containing original control instructions from the node to be migrated to the target migration node, and set the scheduling frequency of the target migration node to the frequency value corresponding to the preset emergency priority, using the reciprocal of the frequency value as the time difference benchmark for characterizing the control delay during the task migration process. Dynamic compensation module: used to dynamically compensate the original control command based on the time difference reference, and generate a compensated control sequence; Execution module: Used to send the compensation control sequence to the actuator end for physical conflict verification and interception, and output the actual control actions that drive the ship's actuators.