Aerospace distributed multi-flight control node fault recovery method and system

Through the aerospace distributed multi-flight control node failure recovery method, representative nodes are dynamically elected and computing node logs are updated, which solves the problems of single-point failure recovery and decision-making differences in multi-mode flight control systems, and achieves efficient system recovery and robustness improvement.

CN120407255AActive Publication Date: 2025-08-01北京天兵科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510480667.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-01
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

There is a lack of effective distributed control models in multimode flight control systems, especially the single-point failure recovery solution for multimode flight control systems lacks mature methods. The existing technology cannot effectively solve the voting and decision-making problems when decision-making results are diverged, and the calculation volume increases exponentially with the increase in the number of processors, and the synchronization relationship is complex.

Method used

The fault recovery method of aerospace distributed multi-flight control nodes is adopted. The representative nodes are voted on by online flight control nodes, and the local command log of the computing node is updated using the representative nodes to realize the consistency of command logs of all nodes in the system. The sensor information is sampled in each solution cycle to generate control commands, and the representative nodes are dynamically elected to improve system robustness.

Benefits of technology

It realizes the robustness of the multi-mode flight control system, reduces the resource dependence on fault recovery, ensures that the system can be effectively controlled after a single point of failure, improves the system processing efficiency and anti-interference ability, and supports multi-mode system recovery of three modes or above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407255A_ABST
    Figure CN120407255A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a spaceflight distributed multi-flight-control-node fault recovery method and system, and relates to the technical field of spaceflight, the method is adopted by a spaceflight distributed multi-flight-control-node system, and the method comprises the following steps: under the condition that no online representative node exists in the system, performing fault recovery on the representative node; each online flight control node in the system sends voting information to other online flight control nodes in the system; each on-line flight control node in the system sets the on-line flight control node as a representative node or a computing node according to the voting information received by the on-line flight control node under the constraint of a preset representative node election rule; and the representative node uses the local command log of the representative node to update the local command log of the computing node, so that the local command logs of all the computing nodes in the system are the same as the local command log of the representative node. And recovery of a single-point fault of the multi-mode flight control system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of space flight technology, and particularly to a method and system for fault recovery of a space distributed multi-flight control node. Background Art

[0002] There is only one core flight control unit in the flight control computer on a traditional space launch vehicle. At present, with the progress of computer technology, it will soon become the mainstream technology to use multi-mode flight control units for parallel computing on the flight control system of a launch vehicle. In the prior art, the simple pairwise synchronization between three processors is used in the triple modular redundancy software data synchronization method, and simple software data synchronization is achieved by relying on the consistency of clocks. It only supports the redundancy of three processors. When the number of processors increases, the number of synchronization relationships between them will increase exponentially, resulting in an explosive growth in the computing amount for synchronization. In the prior art, there is also a method of using a cache queue to save the triple modular data for voting and distribution. The three modes are completely equal in relationship, and there is no real core decision maker, and the problem of how to vote and make a decision when there are differences in decision results is not really solved. Moreover, it is just a simple queue cache and basically does not have the ability of disaster recovery.

[0003] In the process of implementing the present invention, the applicant found that there are at least the following problems in the prior art:

[0004] There is currently a lack of an effective distributed control model in the multi-mode flight control system, especially a mature method for the recovery scheme of single-point faults in the multi-mode flight control system. Summary of the Invention

[0005] Embodiments of the present invention provide a method and system for fault recovery of a space distributed multi-flight control node to solve the technical problem that there is currently a lack of an effective distributed control model in the multi-mode flight control system, especially a mature method for the recovery scheme of single-point faults in the multi-mode flight control system.

[0006] To achieve the above object, on the one hand, an embodiment of the present invention provides a method for fault recovery of a space distributed multi-flight control node, which is adopted by a space distributed multi-flight control node system. The method includes:

[0007] In the case that there is no online representative node in the system, each online flight control node in the system sends voting information to other online flight control nodes in the system; the representative node is one of the multiple flight control nodes in the system, and the representative node serves as the main node of the system;

[0008] Each flight control node online in the system sets itself as a representative node or a computing node according to the voting information received by each, under the constraint of a preset representative node election rule; wherein, the preset representative node election rule is used to constrain that only one flight control node in the system can set itself as a representative node, and other flight control nodes set themselves as computing nodes;

[0009] The representative node uses its own local command log to update the local command logs of the computing nodes, so that the local command logs of all computing nodes in the system are the same as the local command log of the representative node.

[0010] Further, the method further includes:

[0011] In response to each solution cycle, the representative node samples sensor information, schedules the representative node and all the computing nodes to process the sensor information, generates a control command corresponding to the current solution cycle, and executes the control command to update the local command logs of the representative node and all the computing nodes.

[0012] Further, in response to each solution cycle, the representative node samples sensor information, schedules the representative node and all the computing nodes to process the sensor information, generates a control command corresponding to the current solution cycle, and executes the control command to update the local command logs of the representative node and all the computing nodes, including:

[0013] In response to each solution cycle, the representative node samples sensor information and sends the sampled sensor information to all the computing nodes;

[0014] The representative node and all the computing nodes respectively solve the flight control state according to the sensor information, and respectively obtain corresponding flight control state solution results;

[0015] All the computing nodes send their corresponding flight control state solution results to the representative node;

[0016] The representative node receives the corresponding flight control state solution results from all the computing nodes, merges them with the flight control state solution result corresponding to the representative node, and generates a control command corresponding to the current solution cycle according to the merged flight control state solution result;

[0017] The representative node executes the control command corresponding to the current solution cycle, generates a command log of the control command corresponding to the current solution cycle, and saves it to the local command log of the representative node, and the command log includes an incrementing command sequence number.

[0018] Further, while sending the sampled sensor information to all the computing nodes, the representative node also sends the command log corresponding to the control command executed in the previous resolution cycle to all the computing nodes;

[0019] All the computing nodes save the command log corresponding to the control command executed in the previous resolution cycle received by them into the local command log of the computing nodes.

[0020] Further, each flight control node online in the system sends voting information to other flight control nodes online in the system, including:

[0021] Each flight control node online in the system generates first voting information voting for itself and sends the first voting information to other flight control nodes online in the system;

[0022] In response to the first voting information of all other flight control nodes online in the system received by each flight control node online in the system, based on a preset election priority rule, according to the first voting information of the flight control node and all other flight control nodes, in the case of determining that there is a flight control node superior to itself, the flight control node generates second voting information and sends the second voting information to all other online flight control nodes; the second voting information indicates that the flight control node will vote for the flight control node that is superior to itself among other flight control nodes.

[0023] Further, the preset election priority rule includes: the flight control node with a larger maximum command sequence number in the local command log has priority to obtain votes. When there are multiple flight control nodes with the largest and equal maximum command sequence numbers in the local command log, the flight control node with the largest number is selected to have priority to obtain votes; where each flight control node in the system has a unique number preset.

[0024] Further, the encoding format of the first voting information and the second voting information is: <elected node number, maximum command sequence number of the elected node, its own node number>;

[0025] Each flight control node online in the system generates first voting information voting for itself, specifically generating first voting information with the following content: <the number of the flight control node, the maximum command sequence number in the local command log of the flight control node, the number of the flight control node>.

[0026] Further, the generating of the second voting information by the flight control node in the case of determining that there is a flight control node superior to itself based on the preset election priority rule according to the first voting information of the flight control node and all other flight control nodes includes:

[0027] The flight control node compares the field value of the maximum command sequence number of the elected node in the first voting information of the flight control node and all other flight control nodes;

[0028] If the maximum value of the field value of the maximum command sequence number of the elected node in the first voting information of the flight control node and all other flight control nodes is not the maximum command sequence number in the local command log of the flight control node, the following second voting information is generated: <the number of the flight control node to which the first voting information with the largest field value of the maximum command sequence number of the elected node belongs, the maximum value of the field value of the maximum command sequence number of the elected node in the first voting information of the flight control node and all other flight control nodes, the number of the flight control node>.

[0029] Furthermore, each online flight control node in the system, according to the voting information received by each, under the constraint of the preset representative node election rule, sets itself as a representative node or a computing node, including:

[0030] Each online flight control node in the system counts the field values of the elected node numbers in all the first voting information and the second voting information obtained by itself. When it is determined that the number of times the field value of the elected node number obtained by the flight control node is equal to the number of the flight control node is greater than half of the number of all online flight control nodes in the system, the flight control node sets itself as a representative node, otherwise the flight control node sets itself as a computing node;

[0031] The representative node sends representative information to all computing nodes;

[0032] Each computing node, in response to the received representative information, sends node information to the representative node;

[0033] Among them, the fields of the representative information include: the representative node number and the maximum command sequence number in the local command log of the representative node; the fields of the node information include: the representative node number obtained from the received representative information, the maximum command sequence number in the local command log of the representative node obtained from the received representative information, and the maximum command sequence number of the local command log of the computing node.

[0034] Furthermore, the method further includes:

[0035] After any flight control node fails and restarts, when the restart node reconnects to enter the system as a restart node, the restart node establishes a communication connection with other flight control nodes that are already online in the system and sends the first voting information that votes for the restart node to the other nodes that are already online;

[0036] In the case that there is an online representative node in the system, in response to a communication connection initiated by the restart node or receipt of the first voting information of the restart node, the representative node sends representative information to the restart node;

[0037] In response to the received representative information, the restart node returns the node information of the restart node to the representative node;

[0038] The representative node sends the local command log of the representative node to the restart node;

[0039] The restart node updates the local command log of the restart node using the received local command log of the representative node and sets itself as a computing node;

[0040] The fields of the representative information include: the representative node number and the maximum command sequence number in the local command log of the representative node; the fields of the node information include: the representative node number obtained from the received representative information, the maximum command sequence number in the local command log of the representative node obtained from the received representative information, and the maximum command sequence number of the local command log of the computing node; the encoding format of the first voting information is: <elected node number, maximum command sequence number of the elected node, its own node number>.

[0041] On the other hand, an embodiment of the present invention provides a space distributed multi-flight control node system, including: a plurality of flight control nodes for parallel distributed computing that are interconnected and communicate with each other; the space distributed multi-flight control node system adopts the space distributed multi-flight control node fault recovery method as described in any of the previous ones.

[0042] The above technical solutions have the following beneficial effects: By having multiple flight control nodes vote to elect a representative node, dynamic election of the representative node is achieved, and support for node log recovery improves the robustness of the entire flight control system. Redundancy is achieved through software and algorithms, without relying on other hardware and systems on the rocket, minimizing the resources relied on for recovery from faults and increasing the probability of successful recovery from faults. A multi-mode system with more than three modes can be implemented according to the embodiments of the present invention, that is, it can be applied to a multi-mode redundant system and is not limited to three modes; the collapse of a single module node will not cause the system to collapse, and in any case, an elected representative node can effectively control the entire system. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0044] Figure 1 is a flowchart of a method for fault recovery of a single node in an aerospace distributed flight control system according to one embodiment of the present invention;

[0045] Figure 2 is a schematic structural diagram of a system for fault recovery of a single node in an aerospace distributed flight control system according to one embodiment of the present invention;

[0046] Figure 3 is a schematic diagram of the state transition of distributed nodes in a system for fault recovery of a single node in an aerospace distributed flight control system according to one embodiment of the present invention;

[0047] Figure 4 is another flowchart of a method for fault recovery of a single node in an aerospace distributed flight control system according to one embodiment of the present invention;

[0048] Figure 5 is a schematic diagram of the content exchange of the first round of voting in the stage of electing representatives according to one embodiment of the present invention;

[0049] Figure 6 is a schematic diagram of the content exchange of the second round of voting in the stage of electing representatives according to one embodiment of the present invention. Detailed implementation manners

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0051] On the one hand, as Figure 1 shown, an embodiment of the present invention provides a method for fault recovery of multiple flight control nodes in an aerospace distributed system, which is adopted by an aerospace distributed multiple flight control node system. The method includes:

[0052] Step S10, when there is no online representative node in the system, each online flight control node in the system sends voting information to other online flight control nodes in the system; the representative node is one of the multiple flight control nodes in the system, and the representative node serves as the main node of the system;

[0053] Step S11: Each flight control node online in the system, according to the voting information received by each of them, under the constraint of a preset representative node election rule, sets itself as a representative node or a computing node; wherein, the preset representative node election rule is used to constrain that only one flight control node in the system can set itself as a representative node, and other flight control nodes set themselves as computing nodes.

[0054] Step S12: The representative node updates the local command logs of the computing nodes with its own local command log, so that the local command logs of all computing nodes in the system are the same as the local command log of the representative node.

[0055] Wherein, the aerospace distributed multi-flight control node system includes multiple flight control nodes for parallel distributed computing that are interconnected and communicate with each other, and each flight control node has its own processor and memory.

[0056] In some embodiments, the situations where there is no online representative node in the system include but are not limited to the initial state of the system, and the current representative node in the running system fails, etc.; the failure of the current representative node in the running system includes but is not limited to the representative node dropping off the system connection, or the representative node still maintaining the system connection but not sending new communication information to the computing nodes within a preset waiting time. The preset waiting time can be determined according to the solution cycle of the system. For example, the preset waiting time can be set to the solution cycle or half of the solution cycle, etc. After any flight control node in the system fails and is restarted or otherwise processed and then reconnects to the system, it is considered an online flight control node. When each flight control node sends voting information to other flight control nodes, it can also send the same voting information to itself at the same time, so that each flight control node can hold the voting information of all online flight control nodes in the system after completing the communication related to voting. Each flight control node can also directly save the voting information sent by itself locally, and at the same time receive the voting information from other flight control nodes.

[0057] By having multiple flight control nodes vote to elect a representative node, the dynamic election of the representative node is realized, and the support for node log recovery improves the robustness of the entire flight control system. It adopts redundancy implemented in software and algorithms, does not rely on other hardware and systems on the rocket, minimizes the resources relied on for recovery from failures, and increases the probability of successful recovery from failures. A multi-mode system with more than three modes can be implemented according to the embodiments of the present invention, that is, it can be applied to a multi-mode redundant system, not limited to three modes; the collapse of a single module node will not cause the system to collapse, and in any case, the elected representative node can effectively control the entire system.

[0058] Further, the method further includes:

[0059] In response to each solution cycle, the representative node samples sensor information, schedules the representative node and all the computing nodes to process the sensor information, generates a control command corresponding to the current solution cycle, executes the control command, and updates the local command logs of the representative node and all the computing nodes.

[0060] In some embodiments, after selecting the representative node and the computing nodes and completing the synchronization of the local command logs, at the beginning of each solution cycle, the representative node samples the sensors and sends the sensor information to all the computing nodes. All the computing nodes and the representative node will process the sensor information in parallel and distributively. The representative node will generate a control command corresponding to the current solution cycle based on the sensor processing results of the computing nodes and the representative node itself, execute the control command, and update the local command logs of the representative node and all the computing nodes.

[0061] The embodiments of the present invention have the following technical effects: The representative node, as the main node of the system, solves the problem in the traditional technology that there is no core decision maker in the fully peer-to-peer design of a triple-mode or multi-mode system, resulting in the inability to resolve how to vote on decision results when there are differences in decision results.

[0062] Further, the "In response to each solution cycle, the representative node samples sensor information, schedules the representative node and all the computing nodes to process the sensor information, generates a control command corresponding to the current solution cycle, executes the control command, and updates the local command logs of the representative node and all the computing nodes" includes:

[0063] In response to each solution cycle, the representative node samples sensor information and sends the sampled sensor information to all the computing nodes;

[0064] The representative node and all the computing nodes respectively solve the flight control state according to the sensor information and obtain corresponding flight control state solution results;

[0065] All the computing nodes send their respective corresponding flight control state solution results to the representative node;

[0066] The representative node receives the corresponding flight control state solution results from all the computing nodes, merges them with the flight control state solution result of the representative node, and generates a control command corresponding to the current solution cycle based on the merged flight control state solution result;

[0067] The representative node executes the control command corresponding to the current solution cycle, generates a command log of the control command corresponding to the current solution cycle, and saves it to the local command log of the representative node. The command log includes an incrementing command sequence number.

[0068] The embodiments of the present invention have the following technical effects: The representative node serves as the main node of the system, solving the problem in the traditional technology that there is no core decision maker in the fully peer-to-peer design of a triple-mode or multi-mode system, resulting in the inability to resolve how to vote on the decision when there are differences in the decision results.

[0069] Further, while sending the sensor information obtained by sampling to all the computing nodes, the representative node also sends the command log corresponding to the control command executed in the previous solution cycle to all the computing nodes.

[0070] All the computing nodes save the command log corresponding to the control command executed in the previous solution cycle received into the local command log of the computing nodes.

[0071] The embodiments of the present invention have the following technical effects: By using the repeatedly occurring solution cycle, the command log corresponding to the control command completed in the previous solution cycle is synchronized to each computing node in an asynchronous manner while sending the sensor information, which can significantly reduce the number of communication initiations, improve the processing efficiency of the system, increase the silent period of the system, and improve the anti-interference ability of the system.

[0072] Further, each flight control node online in the system sends voting information to other flight control nodes online in the system, including:

[0073] Each flight control node online in the system generates first voting information voting for itself and sends the first voting information to other flight control nodes online in the system.

[0074] Each flight control node online in the system, in response to the received first voting information of all other flight control nodes online in the system, based on a preset election priority rule, according to the first voting information of the flight control node and all other flight control nodes, in the case of determining that there is a flight control node superior to itself, the flight control node generates second voting information and sends the second voting information to all other online flight control nodes; the second voting information indicates that the flight control node will vote for the flight control node that is the best among other flight control nodes and superior to itself.

[0075] In some embodiments, each flight control node votes for itself first and sends the first voting information to other flight control nodes, so that each flight control node can fully understand the current situation of all flight control nodes in the system, such as the maximum command sequence number, etc. Each flight control node determines, according to the first voting information of all flight control nodes obtained locally and the election priority rules, whether there are other flight control nodes that are superior to itself under the election priority rules among other flight control nodes. If so, it finds the flight control node that best meets the election priority rules among other flight control nodes, and re - sends the second voting information to vote for the flight control node that is superior to itself among other flight control nodes.

[0076] The embodiments of the present invention have the following technical effects: The voting process for electing the representative node can be completed through two rounds of voting, with high voting efficiency, enabling the system to complete system recovery in a shorter time.

[0077] Further, the preset election priority rules include: The flight control node with a larger maximum command sequence number in the local command log has priority to obtain the vote. When there are multiple flight control nodes with the largest and equal maximum command sequence numbers in the local command log, the flight control node with the largest number is selected to have priority to obtain the vote; wherein, each flight control node in the system is preset with a unique number.

[0078] In some embodiments, the flight control node with a larger maximum command sequence number in the local command log records more complete operation information. Therefore, the flight control node with a larger maximum command sequence number in the local command log is the optimal flight control node. If, in the initial state of the system, the maximum command sequence numbers in the local command logs of all flight control nodes are initial values and are equal, or if, when the system fails, multiple flight control nodes have completed the communication of synchronizing the command log from the then representative node, so that the maximum command sequence numbers of the relevant flight control nodes and the then representative node are the largest and equal, at this time, the flight control node with the largest number has priority to obtain the vote.

[0079] The embodiments of the present invention have the following technical effects: By the preset election priority rules, it is ensured that the vote is cast for the flight control node that records the most complete current operation information of the system. On this basis, through the numbers of the flight control nodes, a quick and unique decision can be made.

[0080] Further, the encoding format of the first voting information and the second voting information is: <elected node number, elected node maximum command sequence number, own node number>;

[0081] Each flight control node online in the system generates first voting information that votes for itself, specifically generating first voting information including: <the number of the flight control node, the maximum command sequence number in the local command log of the flight control node, the number of the flight control node>.

[0082] Further, based on the preset election priority rule, according to the first voting information of the flight control node and all other flight control nodes, when it is determined that there is a flight control node superior to itself, the flight control node generates second voting information, including:

[0083] The flight control node compares the field value of the maximum command sequence number of the elected node in the first voting information among the flight control node and all other flight control nodes;

[0084] If the maximum value of the field value of the maximum command sequence number of the elected node in the first voting information among the flight control node and all other flight control nodes is not the maximum command sequence number in the local command log of the flight control node, then generate the following second voting information: <the number of the flight control node to which the first voting information with the largest field value of the maximum command sequence number of the elected node belongs, the maximum value of the field value of the maximum command sequence number of the elected node in the first voting information among the flight control node and all other flight control nodes, the number of the flight control node>.

[0085] Further, each flight control node online in the system, according to the voting information received by each, under the constraint of the preset representative node election rule, sets itself as a representative node or a computing node, including:

[0086] Each flight control node online in the system counts the field value of the elected node number in all the first voting information and second voting information it obtains. When it is determined that the number of times the field value of the elected node number obtained by the flight control node is equal to the number of the flight control node is greater than half of the number of all flight control nodes online in the system, the flight control node sets itself as a representative node, otherwise the flight control node sets itself as a computing node;

[0087] The representative node sends representative information to all computing nodes;

[0088] Each computing node, in response to the received representative information, sends node information to the representative node;

[0089] Among them, the fields representing information include: the representative node number and the maximum command sequence number in the local command log of the representative node; the fields of the node information include: the representative node number obtained from the received representative information, the maximum command sequence number in the local command log of the representative node obtained from the received representative information, and the maximum command sequence number of the local command log of the computing node.

[0090] In some embodiments, since each flight control node independently determines whether it is a representative node or a computing node, through the embodiments of the present invention, since each flight control node can only vote for one flight control node when voting for itself or other flight control nodes, it is impossible to vote for more than one flight control node with the same vote. Therefore, there can only be one flight control node that can be voted for more than half. So if a flight control node finds that its number has been voted for more than half, then the flight control node sets itself as the representative node; if a flight control node finds that its number has not been voted for more than half, then it sets itself as the computing node.

[0091] The embodiments of the present invention have the following technical effects: Each flight control node distributes and independently judges whether it is a representative node or a computing node, and the election efficiency is high.

[0092] Furthermore, the method further includes:

[0093] After any flight control node fails and restarts, when the restart node reconnects to the system as a restart node, the restart node establishes a communication connection with other flight control nodes that are already online in the system, and sends first voting information for voting for the restart node to the other nodes that are already online;

[0094] When there is an online representative node in the system, in response to the communication connection initiated by the restart node or receiving the first voting information of the restart node, the representative node sends representative information to the restart node;

[0095] The restart node, in response to the received representative information, returns the node information of the restart node to the representative node;

[0096] The representative node sends the local command log of the representative node to the restart node;

[0097] The restart node updates the local command log of the restart node using the received local command log of the representative node, and sets the restart node itself as a computing node;

[0098] The fields representing information include: the representative node number and the maximum command sequence number in the local command log of the representative node; the fields of the node information include: the representative node number obtained from the received representative information, the maximum command sequence number in the local command log of the representative node obtained from the received representative information, and the maximum command sequence number of the local command log of the computing node; the encoding format of the first voting information is: <elected node number, maximum command sequence number of the elected node, its own node number>.

[0099] In some embodiments, in the initial state of the system, or when the current representative node fails and goes offline or loses its identity as a representative node after restarting, and there is no representative node in the system, the system will conduct a voting election and log synchronization through steps S10 - S12. During the operation of the system, if a computing node fails and then restarts, as a restart node, it needs to rejoin the system. When rejoining the system, the restart node will initiate connections to all currently online flight control nodes, including the representative node and the computing nodes. After the connection, since the restart node is not yet a computing node, the restart node will send the first voting information voting for itself; after being connected by the restart node or receiving the first voting information from the restart node, the representative node will also send representative information to the restart node; here, there is a competition between the representative node sending representative information and the restart node sending the first voting information, and their order is random. However, the representative node only takes the connection establishment of the restart node and the first voting information as a signal to detect the entry of the restart node into the system, and does not process the first voting information. Other computing nodes, since they are already working as computing nodes, will also ignore the first voting information sent by the restart node. After receiving the representative information, the restart node will return the node information of the restart node to the representative node and set itself as a computing node.

[0100] The embodiments of the present invention have the following technical effects: It provides a method for a computing node to re - enter the system after restarting after a failure, allowing the faulty computing node to rejoin the system and realizing the node failure recovery of the system.

[0101] On the other hand, as Figure 2 shown, the embodiments of the present invention provide a space - borne distributed multi - flight control node system, including: a plurality of flight control nodes for parallel distributed computing that are interconnected and communicate with each other; the space - borne distributed multi - flight control node system adopts the space - borne distributed multi - flight control node failure recovery method as described in any of the previous ones.

[0102] For the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, refer to the partial description of the method embodiments.

[0103] The above technical solutions of the embodiments of the present invention will be described in detail below in combination with specific application examples. For technical details not introduced during the implementation process, reference may be made to the relevant descriptions in the foregoing text.

[0104] The embodiment of the present invention provides a distributed flight control single-node fault recovery method. Based on the flight control system of the aerospace system, multiple flight control nodes are used for multi-flight control parallel computing at the same time. Each flight control node is a flight control computer independently equipped with a CPU and a storage space. The embodiment of the present invention solves the problem in the rocket distributed control system with multiple flight control nodes that when a single node fails, how to recover from the failure, and after the node restarts, rejoin the distributed system to act as a normal flight control computing node.

[0105] The following takes a flight control system with 3 flight control nodes as an example for illustration. The embodiment of the present invention stipulates that the rocket control system uses a fixed flight control solution period (for example, a fixed solution period of 10 ms). Within one period, the flight control node needs to complete basic operation steps such as sensor data acquisition, flight control solution (including navigation, guidance, and attitude control operations), and control output.

[0106] In the calculation process of the flight control algorithm, data filtering processing is required. The filtering process not only uses the data of the current period but also uses the data of multiple previous periods. If an effective historical data cannot be obtained during the fault recovery process of the flight control node, the computing node performing the fault recovery cannot perform the flight control solution. At the same time, there are multiple flight control nodes in the system, and each node has a historical record. How to form a unified and correct historical record based on multiple historical records is also one of the key problems solved by the embodiment of the present invention.

[0107] Each flight control node constituting the distributed flight control system of the embodiment of the present invention may be in one of the four states: election state, representative state, recovery state, and calculation state. The state transitions are as Figure 3 shown. Each state will be described separately below.

[0108] Election state: By default, the flight control node will first enter the election state until the system elects a representative node, or the existing representative node in the system notifies this node through a member discovery message that there is already a representative node in the system (occurring when the node reconnects to the system after a fault restart), then this node can leave the election state;

[0109] Representative state: The elected representative node (flight control node) will leave the election state and enter the representative state, starting to control the entire control system. In the representative state, data synchronization needs to be performed for other nodes first, and data synchronization is carried out according to the log status of each node. The representative node in the representative state not only has to perform the duties of a representative but also runs the ordinary flight control calculation function, performs navigation, guidance, attitude control, etc. calculations based on various sensing states of the collected rocket, then counts the calculation results of each node, issues control commands, and announces the commands executed this time to all nodes, generating log data for storage;

[0110] Recovery state: Before each flight control node enters the calculation state, it needs to perform data synchronization with the newly elected representative node. The representative node compares the highest action number (highest command number or highest command sequence number) of each calculation node, and notifies the calculation node to delete the redundant unsubmitted command logs and supplement the missing command logs. Through the current representative node, the command logs of all calculation nodes are unified, enabling the entire distributed flight control system to enter a consistent state; The situation where there are redundant unsubmitted command logs in the calculation node can be as follows: The previous representative node crashed abnormally before issuing a command. The command has been recorded in the local command log, but it didn't have time to be sent to other (flight control) nodes before the representative node crashed. After the representative node restarts and becomes an ordinary flight control node, and other flight control nodes have obtained a new representative node through a new round of election, when this flight control node joins the network of the new representative node, it needs to delete the commands in the unexecuted logs.

[0111] Calculation state: Each flight control node is in the calculation state after synchronizing the logs with the representative node, sending the calculation results of each cycle to the representative node. The representative node decides the final issued control command and notifies all calculation nodes of the calculation results of this cycle, forming operation logs for storage. In addition, the representative node in the representative state also needs to implement the functions of the calculation state, and comprehensively process its own calculation results together with other calculation nodes.

[0112] The following explains the data types involved.

[0113] Node number: Each flight control node in the system is assigned a unique number in the system. During the introduction of the embodiments of the present invention, the three flight control nodes are respectively assigned node numbers 1, 2, and 3;

[0114] Command sequence number: The flight control system needs to complete a round of flight control calculations in each cycle. The calculation result is the control command instruction sent to each system of the rocket. The method introduced in the embodiments of the present invention assigns a command sequence number to each round of command instructions to uniquely represent a round of command instructions;

[0115] Node data: The data content that each flight control node needs to record includes the node number, the current command sequence number, and the operation log.

[0116] Operation log: An operation log includes the command sequence number, sensor input information, calculation process parameters, and result command output information. Among them, the command sequence number is a globally unique auto-incrementing number representing node management.

[0117] Election vote: The content of the vote between nodes is <elected node number, maximum command sequence number of the elected node, its own node number>.

[0118] Delegate information: <delegate node number, maximum command sequence number of the delegate node's log>.

[0119] Node information: <delegate node number obtained from the received delegate information, maximum command sequence number of the delegate node's log obtained from the received delegate information, maximum command sequence number of this node's log>.

[0120] The following describes the system operation process. In the multi-flight control system fault recovery method according to the embodiments of the present invention, reliable network connections are used between each flight control node to ensure that nodes can determine whether the peer node is online. For example, reliable network connection methods such as TCP are used. If an abnormal situation occurs at the peer end, the network will be disconnected, and nodes can determine whether the other party is operating normally through the network connection status. As Figure 4 shown, the system process includes a representative selection stage, a member discovery stage, a log recovery stage, and an instruction synchronization stage.

[0121] Representative selection stage: In the distributed flight control system with single-node fault recovery ability according to the embodiments of the present invention, all nodes need to elect a representative node during system initialization, or during the system operation, when the original representative node fails (for example, it can be determined whether there is a fault through the network connection status. If the network connection is still normal, but no calculation data or commands are sent at the normal cycle node, it can be considered that a fault has occurred), all flight control nodes need to enter the election state and re-elect a representative node.

[0122] After the representative selection stage starts, all flight control nodes are in the election state, count their own maximum command number (maximum command sequence number), vote for themselves, and distribute the vote content to all other flight control nodes; as Figure 5As shown, for example, node 1 sends voting content <elected node number, maximum command sequence number of the elected node, its own node number> to nodes 1, 2, and 3, corresponding to <1, 3, 1> respectively; node 2 sends voting content <elected node number, maximum command sequence number of the elected node, its own node number> to nodes 2, 1, and 3, corresponding to <2, 2, 2> respectively; node 3 sends voting content <elected node number, maximum command sequence number of the elected node, its own node number> to nodes 3, 1, and 2, corresponding to <3, 2, 3> respectively;

[0123] After the flight control node receives the voting content sent by other nodes, it updates its own voting content according to the election priority judgment method below. If it needs to change the voting decision, it sends a new election vote to other nodes.

[0124] The election priority judgment method is as follows: First, compare the command numbers. The node with the larger command number has priority to become the representative node. If the command numbers are the same, compare the device numbers. The node with the larger device number has priority to be elected as the representative node.

[0125] Figure 5 Among the three computing nodes, each sends the first-round election vote to itself and other devices. After the first-round voting ends, each node updates the vote according to the priority judgment method. If the vote is updated, a new vote is sent to other nodes again; as Figure 6 As shown, node 2 finds that the node with the larger command number is node 1 according to the result of the first-round voting. So node 2 sends updated voting content <elected node number, maximum command sequence number of the elected node, its own node number> to nodes 2, 1, and 3, corresponding to <1, 3, 2> respectively; node 3 finds that the node with the larger command number is node 1 according to the result of the first-round voting. So node 3 sends updated voting content <elected node number, maximum command sequence number of the elected node, its own node number> to nodes 3, 1, and 2, corresponding to <1, 3, 3> respectively.

[0126] In the stage of electing representatives, all nodes with reliable network connections have received the election votes sent by each node. Then, the nodes that receive more than half of the votes are counted. The nodes voted by more than half of the nodes immediately become the representative nodes, and the other nodes become computing nodes, leave the election state, and enter the member discovery stage. As Figure 6 As shown, after the second-round voting, node 1 will become the representative node.

[0127] Member discovery phase: The node that has become a representative sends its own log information to other nodes: representative information <representative node number, maximum command sequence number of the local log>. The computing node that receives the representative information returns its own node information <confirmed representative node number, maximum command sequence number of the representative node, maximum command sequence number of this node> to confirm the relationship between the representative node and the computing node.

[0128] During the system operation, if a computing node fails and restarts, after re - establishing a connection with other nodes in the system through the network connection, the computing node can directly enter the member discovery phase, communicate with the representative node about the command sequence numbers, and complete the member discovery phase.

[0129] Log recovery phase: After completing the member discovery phase, the representative node sends the different logs between the representative node and the computing nodes to each computing node. Through the log recovery process, all nodes can reconstruct all control commands during the flight control operation life cycle. The log recovery supports two cases: differential update and invalid log deletion.

[0130] Instruction synchronization phase: After completing the log recovery, the historical working logs of all nodes are consistent. At this time, the entire flight control system can start working according to a fixed calculation cycle. The representative node obtains the sensor inputs of the entire flight control system and notifies all computing nodes. After the computing nodes complete the calculation of the flight control, they send the calculation process parameters and calculation results to the representative node. The representative node counts the calculation results of all nodes, forms control commands and outputs them to the rocket system, and notifies the control commands of this cycle to all computing nodes and stores them in each node to complete a complete instruction synchronization.

[0131] The embodiments of the present invention have the following technical effects: This distributed control system has the function of dynamic election of representative nodes, supports the log recovery of a single node, allows a failed computing node to re - join the system, and improves the robustness of the entire flight control system. It relies on a triple - modular or multi - modular computer itself to achieve triple - modular redundancy, which is realized by software and algorithms and does not depend on other hardware and systems on the rocket, minimizing the resources relied on for recovery from faults and increasing the probability of successful recovery from faults. A multi - modular system with more than triple - modularity can be implemented according to the embodiments of the present invention, that is, it can be applied to a multi - modular redundant system and is not limited to triple - modularity. The embodiments of the present invention are not only aimed at the data processing process, but also design a method for the multi - modular redundant control scheme to recover from disasters after a single module fails. The collapse of a single module node will not cause the system to collapse, and in any case, an effective control of the entire system can be achieved by electing a representative node.

[0132] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of the present disclosure. The appended method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy recited.

[0133] In the foregoing detailed description, various features are combined in a single embodiment to simplify the present disclosure. This method of disclosure should not be interpreted as reflecting an intention that the embodiments of the claimed subject matter require more features than are expressly recited in each claim. Rather, as reflected by the appended claims, the invention lies in less than all of the features of a single disclosed embodiment. Accordingly, the appended claims are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate preferred embodiment of the invention.

[0134] The foregoing embodiments have been described so as to enable any person skilled in the art to make or use the present invention. For those skilled in the art, various modifications to these embodiments will be readily apparent, and the general principles defined herein may be applied to other embodiments without departing from the spirit and scope of the present disclosure. Thus, the present disclosure is not limited to the embodiments given herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0135] The foregoing description includes examples of one or more embodiments. Of course, it is not possible to describe all possible combinations of components or methods for the purpose of describing the above embodiments, but those of ordinary skill in the art should recognize that each embodiment can be further combined and arranged. Accordingly, the embodiments described herein are intended to cover all such changes, modifications, and variations that fall within the scope of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, this term is inclusive in a manner similar to the term "including". Further, any use of the term "or" in the claims or specification is to mean "non-exclusive or".

[0136] Those skilled in the art can also understand that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention can be implemented by electronic hardware, computer software, or a combination of both. To clearly show the interchangeability of hardware and software, the above-mentioned various illustrative components, units, and steps have been generally described in terms of their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art can use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present invention.

[0137] In the embodiments of the present invention, the various illustrative logical blocks or units can be implemented or operate the described functions through a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of the above designs. The general-purpose processor can be a microprocessor. Optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented through a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.

[0138] The steps of the methods or algorithms described in the embodiments of the present invention can be directly embedded in hardware, software modules executed by a processor, or a combination of both. The software modules can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be disposed in an ASIC, and the ASIC can be disposed in a user terminal. Optionally, the processor and the storage medium can also be disposed in different components of the user terminal.

[0139] In one or more exemplary designs, the functions described in embodiments of the present invention may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, these functions may be stored on a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. A computer-readable medium includes both computer storage media and communication media that facilitate transfer of a computer program from one place to another. The storage media may be any available media that can be accessed by a general-purpose or special-purpose computer. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store program code in the form of instructions or data structures and that can be accessed by a general-purpose or special-purpose computer, or a general-purpose or special-purpose processor. Additionally, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless means such as infrared, radio, and microwave, it is included in the definition of computer-readable medium. Disk and disc include compact disc, laser disc, optical disc, DVD, floppy disk, and Blu-ray disc, where disks usually reproduce data magnetically and discs usually reproduce data optically with lasers. The above combinations may also be included within computer-readable media.

[0140] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for fault recovery of aerospace distributed multi-flight control nodes, characterized in that Adopted by the aerospace distributed multi-flight control node system, the method includes: In the case where there is no online representative node in the system, each online flight control node in the system sends voting information to other online flight control nodes in the system; the representative node is one of the multiple flight control nodes in the system, and the representative node serves as the master node of the system; Each online flight control node in the system, according to the voting information received by each, under the constraint of a preset representative node election rule, sets itself as a representative node or a computing node; wherein, the preset representative node election rule is used to constrain that only one flight control node in the system can set itself as a representative node, and other flight control nodes set themselves as computing nodes; The representative node uses its own local command log to update the local command logs of the computing nodes, so that the local command logs of all computing nodes in the system are the same as the local command log of the representative node.

2. The aerospace distributed multi-flight control node fault recovery method according to claim 1, characterized in that, The method further includes: In response to each solution cycle, the representative node samples sensor information, and schedules the representative node and all the computing nodes to process the sensor information, generates a control command corresponding to the current solution cycle, and executes the control command to update the local command logs of the representative node and all the computing nodes.

3. The aerospace distributed multi-flight control node fault recovery method according to claim 2, characterized in that, The "In response to each solution cycle, the representative node samples sensor information, and schedules the representative node and all the computing nodes to process the sensor information, generates a control command corresponding to the current solution cycle, and executes the control command to update the local command logs of the representative node and all the computing nodes" includes: In response to each solution cycle, the representative node samples sensor information and sends the sampled sensor information to all the computing nodes; The representative node and all the computing nodes respectively perform a solution of the flight control state according to the sensor information, and respectively obtain corresponding flight control state solution results; All the computing nodes send their respective corresponding flight control state solution results to the representative node; The representative node receives the corresponding flight control state solution results from all the computing nodes, merges them with the flight control state solution result corresponding to the representative node, and generates a control command corresponding to the current solution cycle according to the merged flight control state solution result; The representative node executes the control command corresponding to the current solution cycle, generates a command log of the control command corresponding to the current solution cycle, and saves it to the local command log of the representative node, and the command log includes an incrementing command sequence number.

4. The aerospace distributed multi-flight control node fault recovery method according to claim 3, characterized in that, While sending the sampled sensor information to all the computing nodes, the representative node also sends the command log corresponding to the control command executed in the previous solution cycle to all the computing nodes; All the computing nodes save the command log corresponding to the control command executed in the previous solution cycle received to the local command log of the computing node.

5. The aerospace distributed multi-flight control node fault recovery method according to claim 1, characterized in that, Each online flight control node in the system sends voting information to other online flight control nodes in the system, including: Each flight control node online in the system generates first voting information voting for itself and sends the first voting information to other flight control nodes online in the system; Each flight control node online in the system, in response to the first voting information received from all other flight control nodes online in the system, based on a preset election priority rule, according to the first voting information of the flight control node and all other flight control nodes, when it is determined that there is a flight control node superior to itself, the flight control node generates second voting information and sends the second voting information to all other flight control nodes online; the second voting information indicates that the flight control node will vote for the flight control node that is superior to itself among other flight control nodes.

6. The aerospace distributed multi-flight control node fault recovery method according to claim 5, wherein, The preset election priority rule includes: the flight control node with a larger maximum command sequence number in the local command log has priority to obtain votes. When there are multiple flight control nodes with the largest and equal maximum command sequence numbers in the local command log, the flight control node with the largest number is selected to have priority to obtain votes; among them, each flight control node in the system is preset with a unique number.

7. The aerospace distributed multi-flight control node fault recovery method according to claim 6, wherein The encoding format of the first voting information and the second voting information is: <elected node number, maximum command sequence number of the elected node, its own node number>; Each flight control node online in the system generates first voting information voting for itself, specifically generating first voting information with the following content: <the number of the flight control node, the maximum command sequence number in the local command log of the flight control node, the number of the flight control node>.

8. The aerospace distributed multi-flight control node fault recovery method according to claim 7, characterized in that Based on the preset election priority rule, according to the first voting information of the flight control node and all other flight control nodes, when it is determined that there is a flight control node superior to itself, the flight control node generates second voting information, including: The flight control node compares the field value of the maximum command sequence number of the elected node in the first voting information of the flight control node and all other flight control nodes; If the maximum value of the field value of the maximum command sequence number of the elected node in the first voting information of the flight control node and all other flight control nodes is not the maximum command sequence number in the local command log of the flight control node, then the following second voting information is generated: <the number of the flight control node to which the first voting information with the largest field value of the maximum command sequence number of the elected node belongs, the maximum value of the field value of the maximum command sequence number of the elected node in the first voting information of the flight control node and all other flight control nodes, the number of the flight control node>.

9. The aerospace distributed multi-flight control node fault recovery method according to claim 1, wherein The method further includes: After any flight control node fails and restarts, when the restart node reconnects to enter the system as a restart node, the restart node establishes a communication connection with other flight control nodes that are already online in the system and sends the first voting information voting for the restart node to the other nodes that are already online; In the case that there is an online representative node in the system, in response to a communication connection initiated by the restart node or upon receiving the first voting information of the restart node, the representative node sends representative information to the restart node; In response to the received representative information, the restart node returns the node information of the restart node to the representative node; The representative node sends the local command log of the representative node to the restart node; The restart node updates its local command log with the received local command log of the representative node and sets itself as a computing node; The fields of the representative information include: the representative node number and the maximum command sequence number in the local command log of the representative node; the fields of the node information include: the representative node number obtained from the received representative information, the maximum command sequence number in the local command log of the representative node obtained from the received representative information, and the maximum command sequence number of the local command log of the computing node; the encoding format of the first voting information is: <elected node number, maximum command sequence number of the elected node, its own node number>.

10. A spaceborne distributed multi-flight control node system, characterized in that, Including: Multiple interconnected and communicating parallel distributed computing flight control nodes; The aerospace distributed multi-flight control node system adopts the aerospace distributed multi-flight control node fault recovery method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Power network controller control method and device based on Zookeeper

    CN118983795A

  • Large-scale data consistency method and device based on multi-stage tree structure cluster

    CN119512692A

  • Systems and methods for service replication, validation, and recovery in cloud-based systems

    US20190146884A1