Satellite-borne computing cluster with automatic operating system repairing function
By designing automatic repair solutions for PXE servers, computing nodes and control nodes in the satellite-based computing cluster, automatic identification and repair of operating system failures is achieved, the problem of lack of automatic repair solutions in the existing technology is solved, and the reliability of the system and the continuity of task execution are improved.
Patent Information
- Application Number
- CN202510163541.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The lack of automatic repair solutions for existing satellite computing clusters when handling operating system failures, resulting in limited system reliability and task continuity.
A satellite-based computing cluster is designed, equipped with PXE server, computing node and control node. The computing node transmits heartbeat signals and operating system status information to the control node in real time. The control node determines the faulty node through monitoring and diagnosis, and configures it into PXE network boot mode through the remote management interface, and sends system reinstallation instructions to automatically repair the operating system.
Automatic identification and repair of operating system failures of the satellite-based computing cluster is realized, which improves the reliability of the system and the continuity of task execution, and reduces the dependence of manual intervention.
Smart Images

Figure CN120104420A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of satellite-borne computing technology, in particular to the field of system repair of satellite-borne computing clusters, and more particularly to a satellite-borne computing cluster with an automatic operating system repair function. Background Art
[0002] It should be noted that this background technology is only used to introduce the relevant information of the present invention to help understand the technical solution of the present invention, but it does not mean that the relevant information is necessarily the prior art. If there is no evidence that the relevant information has been disclosed before the application date of the present invention, the relevant information shall not be regarded as the prior art.
[0003] As onboard missions become increasingly complex and diverse, onboard computing clusters play an increasingly critical role in the aerospace field. In the field of onboard computing clusters, system stability and fault recovery capabilities are crucial. Current onboard computing nodes mostly use hardware redundancy, automatic restart and other means to improve system stability, and use hardware health monitoring (such as heartbeat signals) and basic fault detection mechanisms. However, there is no mature automatic repair solution for operating system-level failures in the existing technology, which to a certain extent restricts the reliability of onboard computing clusters and the continuity of mission execution, and further technological breakthroughs and innovations are urgently needed.
[0004] There are two main ways to handle faults in existing onboard computing clusters: one is through hardware redundancy switching, that is, when a computing node fails, it continues to run by switching to a redundant node. This method mainly relies on the stability of the hardware and ignores the repair and maintenance at the operating system level. The other is to perform remote repairs through the remote monitoring and maintenance interface (such as IPMI) equipped by the onboard computing cluster, such as manually logging in remotely and executing specific repair commands to repair the fault, but this method is mainly used to restart the system and cannot automatically identify and solve the faults in the operating system startup phase. Therefore, the two existing ways of handling onboard computing cluster faults have obvious shortcomings. Although redundant switching can temporarily solve the problem of the faulty node, it cannot fundamentally repair the faulty node. The faulty node may still have problems again in subsequent tasks, and manual intervention depends on the remote operation channel. For onboard computing clusters that are unattended for a long time, this is not realistic, and the timeliness and accuracy of manual operation are also difficult to guarantee. The overall performance of the system and the smooth execution of tasks may be affected due to delays or misoperations.
[0005] In summary, the existing fault handling methods of onboard computing clusters have obvious defects. Although hardware redundancy switching can temporarily deal with faults, it cannot cure the problem of faulty nodes. The manual remote operation method cannot guarantee the timeliness and accuracy of the operation, which can easily affect the system performance and task execution, and cannot automatically identify and solve faults in the operating system startup phase. Therefore, there is an urgent need for a solution that can automatically identify and handle faults in the operating system of onboard computing clusters. Summary of the invention
[0006] Therefore, the purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a satellite-based computing cluster with an automatic repair function of an operating system.
[0007] The objective of the present invention is achieved through the following technical solutions:
[0008] According to a first aspect of the present invention, a satellite-borne computer cluster is proposed, the satellite-borne computer cluster comprising a plurality of PXE servers, a plurality of computing nodes and a control node, wherein: each PXE server is used to store boot files and image files of an operating system; each computing node is configured with a heartbeat service, a serial port output mechanism and a remote management interface, and is configured to transmit a heartbeat signal to the control node in real time during the process of starting the operating system, and to send operating system status information to the control node in real time based on the configured serial port output mechanism; the control node is used to monitor the heartbeat signals of all computing nodes, and when a heartbeat abnormality is detected, the control node is configured to perform a heartbeat control according to the operating system status of the computing node. The invention discloses a method for diagnosing whether a computing node has an operating system failure based on status information, configuring the computing node with the operating system failure to a PXE network boot mode through a remote management interface, and sending a system reinstallation instruction to the computing node with the operating system failure; wherein each computing node is configured with a predefined installation script and is configured to: communicate with all PXE servers, send a heartbeat signal and operating system status information to the control node in real time, and upon receiving the system reinstallation instruction fed back by the control node, obtain the boot file and image file of the operating system from an available PXE server and run the predefined installation script to reinstall and start the operating system.
[0009] Preferably, the serial port output mechanism is configured with operating system status information corresponding to each startup step of the operating system, wherein the startup steps of the operating system include BIOS / UEFI startup step, MBR / GPT loading step, GRUB boot step, kernel loading step, Init process startup step, system service startup step and user login step in sequence.
[0010] Preferably, each computing node is configured to send operating system status information corresponding to a startup step of an operating system to the control node after the computing node completes the execution of the startup step of the operating system.
[0011] Preferably, the control node is configured to perform fault diagnosis on each computing node in the following manner: the control node monitors in real time whether the computing node periodically sends heartbeat signals at preset time intervals; if so, it determines that the computing node has no faults; otherwise, it reads the latest operating system status information sent by the computing node; if the startup step corresponding to the latest operating system status information sent by the computing node is the MBR / GPT loading step or any step after the MBR / GPT loading step, it is determined that the computing node has an operating system fault.
[0012] Preferably, when a computing node needs to obtain the boot file and image file of the operating system from a PXE server, the computing node interacts with the PXE server as follows: the computing node sends a boot file request to each PXE server one by one in a preset order until a response to the boot file request is received from a PXE server; the computing node pulls the boot file from the responding PXE server, and executes the boot file so that the computing node downloads the image file from the PXE server.
[0013] Preferably, the computing node communicates with each PXE server using a PXE protocol.
[0014] Preferably, the computing node is configured to: download the image file from the PXE server using a simple file transfer protocol.
[0015] Preferably, the remote management interface is any one of IPMI, RDP, RFB, Telnet, and SSH.
[0016] Preferably, the predefined installation script is configured with operation instructions for sequentially executing image loading, automated partitioning and formatting, operating system installation, operating system configuration, and service startup.
[0017] Preferably, the parameters of each operation instruction in the predefined installation script can be pre-configured according to the user's needs.
[0018] Compared with the prior art, the advantages of the present invention are:
[0019] The present invention proposes a solution for automatically identifying and processing operating system failures of a satellite computing cluster. The solution accurately determines operating system failures by monitoring serial port output information and heartbeat service status of computing nodes, and reinstalls and repairs the operating system to repair failed nodes based on an automated network boot mechanism without human intervention. In addition, the satellite computing cluster is configured with multiple PXE servers to ensure that computing nodes accurately obtain boot files and image files, thereby ensuring the reliability of the repair operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The embodiments of the present invention are further described below with reference to the accompanying drawings, in which:
[0021] Figure 1 A schematic diagram of connecting a computing node and a control node according to an embodiment of the present invention;
[0022] Figure 2 A schematic diagram of connecting a computing node and a PXE server according to an embodiment of the present invention;
[0023] Figure 3 A schematic diagram of control node fault diagnosis according to an embodiment of the present invention;
[0024] Figure 4 A schematic diagram of automatic repair of computing nodes according to an embodiment of the present invention;
[0025] Figure 5 Schematic diagram of a boot file request according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail by specific embodiments below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0027] As mentioned in the background technology section, the fault handling methods of existing satellite computing clusters have obvious defects. Although hardware redundancy switching can temporarily deal with faults, it cannot fundamentally solve the problem of faulty nodes. The manual remote operation method cannot guarantee the timeliness and accuracy of the operation, which can easily affect system performance and task execution, and cannot automatically identify and solve faults in the operating system startup phase.
[0028] To solve the above problems, the present invention proposes a solution that can automatically identify and handle operating system failures of onboard computing clusters. The solution accurately determines operating system failures by monitoring the serial port output information and heartbeat service status of computing nodes, and reinstalls and repairs the operating system based on an automated network boot mechanism without human intervention to repair failed nodes. In addition, by configuring multiple PXE server redundancy mechanisms, it is ensured that computing nodes accurately obtain boot files and image files, thereby ensuring the reliability of repair operations.
[0029] In order to better understand the present invention, the present solution is described in detail below in conjunction with specific embodiments.
[0030] Under existing technologies, satellite-based computing clusters lack fault diagnosis capabilities. Once a fault occurs, they can only be repaired through manual intervention or hardware switching. However, for satellite-based computing clusters that are unattended for a long time, the timeliness of manual intervention is difficult to guarantee. In addition, although hardware switching can temporarily bypass the faulty node, it cannot fundamentally repair the faulty node. To solve the above problems, in the solution of the present invention, the computing nodes and control nodes in the satellite computing cluster are improved, and the computing nodes are equipped with a heartbeat service, a serial port output mechanism, and a remote management interface, so that the computing nodes send heartbeat signals and operating system status information to the control nodes periodically during the startup of the operating system, so that the control nodes can grasp the operating status of each computing node in real time, and the control nodes perform fault diagnosis on each computing node based on the data reported by the computing nodes. Once it is detected that a computing node has an operating system failure, a system instruction is sent to it to trigger the computing node to reinstall the operating system, so as to realize the fault self-repair of the satellite computing cluster. In addition, the present invention enhances the fault tolerance and robustness of the satellite computing cluster by deploying multiple PXE servers storing operating system boot files and image files in the satellite computing cluster. Even if some PXE servers fail suddenly, the computing nodes can still stably and reliably obtain the required boot files and image files when reinstalling the operating system.
[0031] According to an embodiment of the present invention, in this solution, each computing node is connected to a control node, such as Figure 1 As shown, each computing node sends a heartbeat signal to the control node in real time, and whenever the operating system completes a startup step, the operating system status corresponding to the startup step is sent to the serial port of the control node based on the serial port output mechanism, so that the control node can obtain the running status of the computing node in real time. It should be noted that the serial port output mechanism is configured with operating system status information corresponding to each startup step of the operating system, wherein the startup steps of the operating system include BIOS / UEFI startup steps, MBR / GPT loading steps, GRUB boot steps, kernel loading steps, Init process startup steps, system service startup steps and user login steps in sequence. It should be understood that the startup steps of the operating system are well known to those skilled in the art, and the startup steps of the operating system are not described in detail here.
[0032] According to one embodiment of the present invention, in this scheme, when the heartbeat signal of a computing node fails to reach the control node on time, the control node will deem that the computing node has a fault. In order to more accurately determine the type of fault, the control node further parses the operating system status information in the serial port to determine in which startup step the computing node fails. For example, if the latest operating system status is the BIOS / UEFI startup step, it is determined that the computing node fails during BIOS / UEFI startup; if the latest operating system status is the MBR / GPT loading step, it is determined that the computing node fails during MBR / GPT loading.
[0033] According to an embodiment of the present invention, in this solution, multiple PXE servers are configured in the onboard computing cluster, and the same computing node is connected to multiple PXE servers. Figure 2 As shown in the figure, each computing node has established a connection with multiple PXE servers. On this basis, the computing node uses the PXE communication protocol to exchange data with the PXE server. This solution adopts a multi-PXE server redundancy mechanism with strong fault tolerance. Even if a PXE server fails due to hardware failure, network interruption, etc., the computing node can still quickly obtain the boot file and image file of the operating system from other PXE servers connected to it. It is crucial to ensure that the computing nodes can stably and reliably obtain the boot file and image file for the fault repair of the onboard computing cluster. This solution configures multiple PXE servers to effectively ensure that the onboard computing cluster can continue to operate stably in the complex and changeable space environment, greatly improving the overall reliability and anti-interference ability of the onboard computing cluster, and providing solid technical support for the smooth execution of space missions.
[0034] According to an embodiment of the present invention, in this scheme, the control node is used to perform fault diagnosis on each computing node based on the heartbeat information reported by each computing node and the operating system status transmitted by the computing node through the serial port. Specifically, when the control node detects that the heartbeat signal of a computing node is abnormal (that is, the heartbeat signal is not sent periodically), it is determined that the computing node has a fault. In order to further determine the cause of the fault of the computing node, the control node will read the latest operating system status transmitted by the computing node in the serial port to determine at which startup step the computing node has a problem. If the latest operating system status transmitted by the computing node is the MBR / GPT startup step or any step after the MBR / GPT startup step, it indicates that the computing node has an operating problem. For system-level failures, the control node controls the computing node to enter the BIOS setup interface through the computing node's remote management interface (such as IPMI, RDP, RFB, Telnet, SSH or other remote management interfaces), and selects the PXE network boot option so that the computing node's boot mode will be switched to network boot. Then, a system reinstallation instruction is sent to the computing node so that the computing node automatically reinstalls the operating system. On the contrary, if the computing node transmits the latest operating system status BIOS / UEFI startup steps, the failure of the computing node does not belong to an operating system failure and cannot be repaired by reinstalling the system. For nodes with such failures, this solution will include them in the list to be checked and wait for manual intervention for debugging. This solution automatically diagnoses computing node failures based on the computing node's heartbeat information and operating system status, does not rely on physical contact and manual intervention, greatly improves the stability and reliability of the cluster, and reduces the risk of system failures.
[0035] According to an embodiment of the present invention, when a computing node with an operating system failure receives a reinstallation instruction sent by a control node, it enters an automatic repair process to reinstall the operating system. Figure 2 As shown, the automatic repair process shown in the figure is as follows: step S1, requesting a PXE boot file, that is, the computing node uses the PXE protocol to communicate with the PXE server and obtains the boot file from an available PXE server; step S2, the computing node loads the PXE boot program and uses the TFTP protocol (Trivial File Transfer Protocol) to download the operating system image from the PXE server that obtains the boot file; step S3, automatically installing the operating system, the computing node automatically executes a predefined installation script based on the downloaded operating system image to install the operating system, wherein the installation script can be personalized according to user needs, such as selecting the operating system version, installed system components, etc.; step S4, the computing node automatically restarts the operating system and sends a heartbeat signal and the operating system status to the control node.
[0036] According to an embodiment of the present invention, in step S2 of the automatic repair process, the communication process between the computing node and the PXE server is as follows: Figure 3 As shown in the figure, the computing node obtains the boot file from the PXE server. Specifically, when a computing node requests a boot file, the computing node will use the PXE protocol and send requests to multiple PXE servers in a pre-set order. If a PXE server fails to respond in time due to failure, network anomaly or other reasons, the computing node will automatically trigger the switching mechanism and seamlessly connect to the next available PXE server until an available PXE server responds and obtains the boot file from the responding PXE server. If all PXE servers do not respond, the onboard computing cluster will not be able to repair the computing node. At this time, the onboard computing cluster reports a PXE server error and waits for manual processing. This solution uses a multi-PXE server redundancy mechanism to significantly enhance the fault tolerance and robustness of the system. Even if some PXE servers suddenly fail, the computing nodes can stably and reliably obtain the required files, thereby laying a solid foundation for the operating system repair process, and effectively maintaining the stability and continuous operation of the entire computing environment.
[0037] According to one embodiment of the present invention, the predefined installation script in this scheme is used to automatically install the operating system based on the acquired image file, and is configured with multiple operation instructions to perform the following steps: Step S31, boot image loading, the operating system image downloaded from the PXE server will be loaded into the node memory, and the boot installation program will be started; Step S32, automatic partitioning and formatting, the installation script will automatically perform disk partitioning and formatting operations to ensure that the operating system can be correctly installed on the faulty node; Step S33, system installation, the file of the operating system image is copied to the hard disk of the node, and the installation operation is performed, wherein all configurations in the installation operation are completed by the predefined script; Step S34, system configuration and service startup, after the operating system is installed, the operating system is automatically started, and network configuration, user settings and other operations are performed to ensure that the node can normally join the onboard computing cluster. It should be understood that the implementation of the operating system installation script is a well-known technology for those skilled in the art and is not described in detail here. In addition, this scheme aims to realize the function of automatically installing the operating system on the computing node through the predefined installation script. The steps contained in the above installation script are only illustrative, and the user can add or delete the steps in the installation script according to the needs, or adjust the order of each installation step in the installation script.
[0038] In order to ensure the accuracy of system recovery, according to an embodiment of the present invention, after the reinstalled operating system of a computing node is started, the control node continues to monitor the computing node to ensure that its various services can operate normally. Specifically, after the computing node starts the reinstalled operating system, the heartbeat service will also work again to report the heartbeat information to the control node. If the control node receives a normal heartbeat signal, it means that the operating system of the computing node has been successfully restored. If the control node does not receive a normal heartbeat signal, it means that the computing node has failed to be repaired, and the control node sends a system reinstallation instruction to the computing node again to repair the computing node.
[0039] This solution improves the existing onboard computing cluster. By real-time monitoring of the serial port output information and heartbeat service status of the computing node, it can accurately determine the operating system failure and automatically trigger the system reinstallation and repair process to repair the faulty computing node, solving the problem of poor timeliness of manual intervention. In addition, the onboard computing cluster is configured with multiple PXE servers that store operating system boot and image files to ensure that even if some servers are abnormal, the files can be obtained stably and reliably with the redundant configuration, thereby ensuring the reliability of repair and enhancing the fault tolerance and stability of the onboard computing cluster.
[0040] It should be noted that although the above describes the various steps in a specific order, it does not mean that the various steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.
[0041] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A satellite-borne computer cluster, comprising a plurality of PXE servers, a plurality of computing nodes and a control node, wherein: Each PXE server is used to store the boot files and image files of the operating system; Each computing node is configured with a heartbeat service, a serial port output mechanism and a remote management interface, and is configured to transmit a heartbeat signal to the control node in real time during the process of starting the operating system, and to send operating system status information to the control node in real time based on the configured serial port output mechanism; The control node is used to monitor the heartbeat signals of all computing nodes, and when an abnormal heartbeat is detected, diagnose whether the computing node has an operating system failure based on the operating system status information of the computing node, configure the computing node with the operating system failure to the PXE network boot mode through the remote management interface, and send a system reinstallation instruction to the computing node with the operating system failure; Each computing node is configured with a predefined installation script and is configured as follows: Communicate with all PXE servers and send heartbeat signals and operating system status information to the control node in real time. When receiving the system reinstallation instruction fed back by the control node, obtain the boot file and image file of the operating system from an available PXE server and run the predefined installation script to reinstall and start the operating system.
2. The onboard computer cluster according to claim 1, characterized in that: The serial port output mechanism is configured with operating system status information corresponding to each startup step of the operating system, wherein the startup steps of the operating system include BIOS / UEFI startup step, MBR / GPT loading step, GRUB boot step, kernel loading step, Init process startup step, system service startup step and user login step in sequence.
3. The onboard computer cluster according to claim 2, characterized in that: Each compute node is configured as: After each operating system startup step is completed, the computing node sends the operating system status information corresponding to the startup step to the control node.
4. The onboard computer cluster according to claim 3, characterized in that: The control node is configured to perform fault diagnosis on each computing node in the following manner: The control node monitors in real time whether the computing node sends heartbeat signals periodically at preset time intervals. If so, it is determined that the computing node has no faults. Otherwise, the latest operating system status information sent by the computing node is read. If the startup step corresponding to the latest operating system status information sent by the computing node is the MBR / GPT loading step or any step after the MBR / GPT loading step, it is determined that the computing node has an operating system failure.
5. The onboard computer cluster according to claim 1, characterized in that: When a computing node needs to obtain the boot file and image file of the operating system from the PXE server, the computing node interacts with the PXE server as follows: The computing node sends a boot file request to each PXE server one by one in a preset order until a response to the boot file request from a PXE server is received; The computing node pulls the boot file from the responding PXE server and executes the boot file to cause the computing node to download the image file from the PXE server.
6. The onboard computer cluster according to claim 5, characterized in that: The computing node communicates with each PXE server using the PXE protocol.
7. The onboard computer cluster according to claim 5, characterized in that: The computing nodes are configured as follows: Use Simple File Transfer Protocol to download the image file from the PXE server.
8. The onboard computer cluster according to claim 1, characterized in that: The remote management interface is any one of IPMI, RDP, RFB, Telnet, and SSH.
9. The onboard computer cluster according to claim 1, characterized in that: The predefined installation script is configured with operation instructions for sequentially executing image loading, automated partitioning and formatting, operating system installation, operating system configuration, and service startup.
10. The onboard computer cluster according to claim 9, characterized in that: The parameters of each operation instruction in the predefined installation script can be pre-configured according to the needs of the user.
Citation Information
Patent Citations
Cluster system and deployment method thereof
CN106502797A
OpenStack large-scale deployment method and system based on bare computer server
CN111198696A
Automatic deployment method and system supporting multiple domestic operating systems
CN112230942A
Method and system for switching operating system to execute test task and medium
CN112749095A
Cabinet test method and device, electronic equipment and storage medium
CN115756978A
Cited By
Multistage fault-tolerant starting method for spaceborne computer
CN120631634A
Satellite computer multi-stage fault-tolerant starting method
CN120631634B