System automatic repair and fault environment saving method
Patent Information
- Application Number
- CN202611135548.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-08-28
AI Technical Summary
1)现有技术中尽管可在系统无法正常启动的情况下利用组合按键输入快速启动指令fastboot进入恢复出厂设置模式,如图1所示,但是需要依赖手动操作,在无人值守的部署场景下无法实现自动恢复
[0017]相较于现有技术,本发明能够实现的技术效果包括:
Smart Images

Figure CN122653907A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of system faults and repair, and specifically to a method for automatic system repair and fault environment preservation. Background Technology
[0002] Existing systems typically contain two systems: a normal operating system and a recovery system. The recovery system allows users to perform operations such as factory resets and upgrades.
[0003] However, the aforementioned existing technologies still have the following defects and shortcomings in practical use: 1) In existing technologies, although it is possible to enter factory reset mode by inputting the fastboot command using a combination of keystrokes when the system fails to boot normally, such as... Figure 1 As shown, however, it requires manual operation and cannot achieve automatic recovery in unattended deployment scenarios.
[0004] 2) After system recovery, it is impossible to know why the system could not start normally before the recovery, so it is difficult to fix the anomaly to make the system more perfect and robust.
[0005] Therefore, there is an urgent need to provide a new solution to address the defects and shortcomings of the existing technologies. Summary of the Invention
[0006] In order to address the defects and shortcomings of existing technologies, this invention provides a method for automatic system repair and fault environment preservation.
[0007] The specific solution provided by this invention is as follows: A method for automatic system repair and fault environment preservation, characterized in that the method includes the following steps: S100: The system powers on and starts the boot program; S200: Check if there is an abnormal marker in the information partition; if there is an abnormal marker, proceed to step S300; if there is no abnormal marker, proceed to step S400. S300: Automatically loads the system image in the repair partition, determines the cause of the fault, performs corresponding recovery based on the cause of the fault, and saves the fault scene; S400: Set an exception flag to automatically load the boot image from the regular partition and start the system; S500: Clear exception markers.
[0008] As a further preferred embodiment of the present invention, the disk partitions of the system include at least an information partition, a repair partition, and a regular partition; wherein... The anomaly marker is set inside the information partition, and the information partition can adjust and clear the anomaly marker according to the actual situation; The repair partition contains a pre-set startup script, which can execute a pre-set diagnostic program to restore the system and save the fault scene. The regular partition stores a boot image, enabling data storage, system startup, and system installation.
[0009] As a further preferred embodiment of the present invention, in step S300, the system image can mount a root file system through a repair partition. The root file system has a preset startup script, which can execute a preset diagnostic program to perform system recovery and save the fault scene.
[0010] As a further preferred embodiment of the present invention, in step S300, the cause of the failure includes at least the following: 1) The system image file is corrupted; 2) The file system is corrupted; 3) The file was deleted or its content was changed.
[0011] As a further preferred embodiment of the present invention, in step S300, the following steps are used to determine whether the cause of the fault is a corrupted system image file; S301: Determine if the system image file is corrupted by checking if the system image checksum is correct. If the checksum is correct, the system image is considered to be intact. If the checksum is incorrect, the system image is considered to be corrupted. The system image is compressed and saved to the repair partition for subsequent analysis. The system image with the incorrect checksum is replaced with a system image with the correct checksum that was previously saved in the repair partition. Then proceed to the next step.
[0012] As a further preferred embodiment of the present invention, in step S301, the initial verification value corresponding to each system image is calculated and saved to the repair partition; When checking if the system image checksum is correct, calculate the current checksum for each system image and compare the current checksum with the initial checksum: When the current checksum is equal to the initial checksum, the system image checksum is considered correct. If the current checksum is not equal to the initial checksum, the system image checksum is considered to be incorrect.
[0013] As a further preferred embodiment of the present invention, in step S300, the cause of the failure is determined to be file system corruption through the following steps; S302: Determine if the file system is corrupted by detecting changes in file system metadata. If it is determined that the file system is corrupted and cannot be mounted, resulting in boot failure, the original file system data is read, compressed, and saved to the repair partition for subsequent analysis. Then, the file system partition is reformatted, the original file system data is extracted from the repair partition, and copied to the formatted file system partition. If it is determined that the file system is not corrupted, the original data of the file system is read, compressed and saved to the repair partition. At the same time, the file system mounted in the normal boot process is compressed and saved to the repair partition for subsequent analysis. The original data of the file system in the repair partition is used to overwrite the file system mounted in the normal boot process. Then proceed to the next step.
[0014] As a further preferred embodiment of the present invention, in step S300, the cause of the failure is determined by the following steps: S303: The cause of the failure is determined to be a file deletion or content change if and only if the system image file is not corrupted and the file system is not corrupted; At this point, the original file system data in the repair partition is used to overwrite the deleted or altered files.
[0015] As a further preferred embodiment of the present invention, in step S300, the fault scene is saved to the repair partition.
[0016] As a further preferred embodiment of the present invention, the fault scene stored in the repair partition can be exported as needed, and the exported fault scene can be analyzed in conjunction with the fault cause to determine the fault source: 1) When the cause of the failure is determined to be a corrupted system image file, the location of the corruption in the system image file is confirmed by comparing the checksum of the image file; 2) When the cause of the failure is determined to be file system corruption, check the corresponding applications deployed in the system to determine the location of the application that caused the file system corruption; 3) When the cause of the failure is determined to be that the file has been deleted or its content has been changed, check the disk hardware to determine whether the data error was caused by the disk hardware or by the application directly accessing the disk's raw interface and rewriting the data on the disk.
[0017] Compared with existing technologies, the technical effects that this invention can achieve include: 1) This invention provides a method for automatic system repair and fault environment preservation. By detecting whether there are abnormal markers in the information partition, it can automatically determine whether the system has failed and automatically recover without relying on manual operation. It can also achieve automatic recovery in unattended deployment scenarios.
[0018] 2) This invention provides a method for automatic system repair and fault environment preservation, which can preserve the abnormal scene while restoring the abnormality, so as to provide analytical data for subsequent fault investigation. Attached Figure Description
[0019] Figure 1 The diagram shows the working steps of a repair system in the prior art.
[0020] Figure 2 The diagram shown is a flowchart of the steps of the method provided by the present invention.
[0021] Figure 3 The diagram shown is a schematic diagram of the disk partition structure of the system provided by the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," "outer," "front end," "rear end," "both ends," "one end," and "the other end," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0024] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installed," "equipped with," "connected," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0025] [First Embodiment] like Figure 2 The figure shown is a system automatic repair and fault environment saving method provided by the first embodiment of the present invention, which includes the following steps: S100: The system powers on and starts the boot program. In this embodiment, the boot program is uboot. In this embodiment, as Figure 3 As shown, the system's disk partitions include at least an information partition (INFO), a recovery partition (Recovery), and a regular partition; among which, Anomaly markers are set within the information partition, and the information partition can adjust and clear the anomaly markers according to the actual situation; The repair partition contains a pre-set startup script, which can execute a pre-set diagnostic program to restore the system and save the fault scene. The regular partition contains a boot image, which is used for data storage, system booting, and system installation.
[0026] S200: Check if there is an abnormal marker in the information partition; if there is an abnormal marker, proceed to step S300; if there is no abnormal marker, proceed to step S400. S300: Automatically loads the system image in the repair partition, determines the cause of the fault, performs corresponding recovery based on the cause of the fault, and saves the fault scene; In this step, the system image is different from the image that boots normally. It can mount the root file system through the repair partition. In this embodiment, the FAT32 file system is selected. The root file system has a pre-set startup script. The startup script can execute a pre-set diagnostic program to perform system recovery and save the fault scene.
[0027] In step S300, the causes of the fault include at least the following: 1) The system image file is corrupted; 2) The file system is corrupted; 3) The file was deleted or its content was changed.
[0028] In step S300, the following steps are used to determine whether the cause of the fault is a corrupted system image file; S301: Determine if the system image file is corrupted by checking if the system image checksum is correct. If the checksum is correct, the system image is considered to be intact. If the checksum is incorrect, the system image is considered to be corrupted. The system image is then compressed and saved to the repair partition for subsequent analysis. The system image with the incorrect checksum is replaced with a pre-saved system image with the correct checksum in the repair partition. The system image with the correct checksum was placed in the repair partition during the initial flashing for subsequent system repair.
[0029] Then proceed to the next step.
[0030] In this embodiment, an initial verification value is calculated for each system image and saved to the repair partition; When checking if the system image checksum is correct, calculate the current checksum for each system image and compare the current checksum with the initial checksum: When the current checksum is equal to the initial checksum, the system image checksum is considered correct. If the current checksum is not equal to the initial checksum, the system image checksum is considered to be incorrect.
[0031] As one preferred implementation, the SHA256 algorithm can be used to calculate the initial and current checksums for each system image based on the actual content of the image. As long as the image content changes, the SHA256 checksum calculated by the algorithm will also change.
[0032] For any length of the actual content in the image, a 256-bit hash value, called a message digest, can be generated using the SHA256 algorithm, which is also the checksum in this embodiment. This checksum is essentially a 32-byte array, usually represented by a 64-byte hexadecimal string.
[0033] For example, for the mirrored content BlockChain, its hash checksum calculated using the SHA256 algorithm is: 3a6fed5fc11392b3ee9f81caf017b48640d7458766a8eb0382899a605b41f2b9. Once the actual content of the mirrored content changes, this hash checksum calculated using the SHA256 algorithm will also change accordingly.
[0034] In step S300, the following steps are used to determine whether the cause of the failure is file system corruption; S302: Determine if the file system is corrupted by detecting changes in the file system metadata: the file system metadata can be at least one or a combination of inodes, block allocation tables, and directory structures; If the file system is corrupted and cannot be mounted, causing the boot process to fail, the original file system data is read, compressed, and saved to the repair partition for subsequent analysis. Then, the file system partition is reformatted, the original file system data is extracted from the repair partition, and copied to the formatted file system partition. If the file system is not corrupted, read the original file system data, compress it, and save it to the repair partition. At the same time, compress the file system mounted during the normal boot process and save it to the repair partition for subsequent analysis. Use the original file system data in the repair partition to overwrite the file system mounted during the normal boot process. Then proceed to the next step.
[0035] This is because although the file system itself is not corrupted, the presence of anomaly markers indicates that some files within the file system have been damaged. For example, if critical files on the C drive are deleted, the system will not be able to boot normally. However, the file system itself is not damaged; it's just that some critical files are missing. Therefore, this process repairs the damaged files through an overwrite procedure.
[0036] In step S300, the following steps are used to determine whether the cause of the failure is that the file has been deleted or its content has been changed: S303: The cause of the failure is determined to be a file deletion or content change if and only if the system image file is not corrupted and the file system is not corrupted; At this point, the original file system data in the repair partition is used to overwrite the deleted or altered files.
[0037] Based on this, in this embodiment, the fault scene is saved to the repair partition. After the repair is completed and the system can start normally, the saved fault scene can be decompressed from the repair partition, and the cause of the fault can be analyzed (e.g., missing files or incorrect file content).
[0038] It is worth noting that the fault scene saved in the repair partition can be exported as needed, and the exported fault scene can be analyzed in conjunction with the fault cause to determine the source of the fault: 1) When the cause of the failure is determined to be a corrupted system image file, the location of the corruption in the system image file is confirmed by comparing the checksum of the image file; 2) When the cause of the failure is determined to be file system corruption, check the corresponding applications deployed in the system to determine the location of the application that caused the file system corruption; 3) When the cause of the failure is determined to be that the file has been deleted or its content has been changed, check the disk hardware to determine whether the data error was caused by the disk hardware or by the application directly accessing the disk's raw interface and rewriting the data on the disk.
[0039] S400: Set an exception flag to automatically load the boot image from the regular partition and start the system; S500: Clear exception markers.
[0040] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A method for automatic system repair and fault environment preservation, characterized in that: The method includes the following steps: S100: The system powers on and starts the boot program; S200: Check if there is an abnormal marker in the information partition; if there is an abnormal marker, proceed to step S300; if there is no abnormal marker, proceed to step S400. S300: Automatically loads the system image in the repair partition, determines the cause of the fault, performs corresponding recovery based on the cause of the fault, and saves the fault scene; S400: Set an exception flag to automatically load the boot image from the regular partition and start the system; S500: Clear exception markers.
2. The method for automatic system repair and fault environment preservation according to claim 1, characterized in that: The system's disk partitions include at least an information partition, a repair partition, and a regular partition; among which, The anomaly marker is set inside the information partition, and the information partition can adjust and clear the anomaly marker according to the actual situation; The repair partition contains a pre-set startup script, which can execute a pre-set diagnostic program to restore the system and save the fault scene. The regular partition stores a boot image, enabling data storage, system startup, and system installation.
3. The method for automatic system repair and fault environment preservation according to claim 1, characterized in that: In step S300, the system image can mount the root file system through the repair partition. The root file system has a pre-set startup script, which can execute a pre-set diagnostic program to restore the system and save the fault scene.
4. The method for automatic system repair and fault environment preservation according to claim 1, characterized in that: In step S300, the causes of the fault include at least the following: 1) The system image file is corrupted; 2) The file system is corrupted; 3) The file was deleted or its content was changed.
5. The method for automatic system repair and fault environment preservation according to claim 4, characterized in that: In step S300, the following steps are used to determine whether the cause of the fault is a corrupted system image file; S301: Determine if the system image file is corrupted by checking if the system image checksum is correct. If the checksum is correct, the system image is considered to be intact. If the checksum is incorrect, the system image is considered to be corrupted. The system image is compressed and saved to the repair partition for subsequent analysis. The system image with the incorrect checksum is replaced with a system image with the correct checksum that was previously saved in the repair partition. Then proceed to the next step.
6. The method for automatic system repair and fault environment preservation according to claim 5, characterized in that: In step S301, the initial verification value corresponding to each system image is calculated and saved to the repair partition. When checking if the system image checksum is correct, calculate the current checksum for each system image and compare the current checksum with the initial checksum: When the current checksum is equal to the initial checksum, the system image checksum is considered correct. If the current checksum is not equal to the initial checksum, the system image checksum is considered to be incorrect.
7. The method for automatic system repair and fault environment preservation according to claim 6, characterized in that: In step S300, the following steps are used to determine whether the cause of the fault is file system corruption; S302: Determine if the file system is corrupted by detecting changes in file system metadata. If it is determined that the file system is corrupted and cannot be mounted, resulting in boot failure, the original file system data is read, compressed, and saved to the repair partition for subsequent analysis. Then, the file system partition is reformatted, the original file system data is extracted from the repair partition, and copied to the formatted file system partition. If it is determined that the file system is not corrupted, the original data of the file system is read, compressed and saved to the repair partition. At the same time, the file system mounted in the normal boot process is compressed and saved to the repair partition for subsequent analysis. The original data of the file system in the repair partition is used to overwrite the file system mounted in the normal boot process. Then proceed to the next step.
8. The method for automatic system repair and fault environment preservation according to claim 7, characterized in that: In step S300, the following steps are used to determine whether the cause of the fault is that the file has been deleted or its content has been changed: S303: The cause of the failure is determined to be a file deletion or content change if and only if the system image file is not corrupted and the file system is not corrupted; At this point, the original file system data in the repair partition is used to overwrite the deleted or altered files.
9. The method for automatic system repair and fault environment preservation according to claim 8, characterized in that: In step S300, the fault scene is saved to the repair partition.
10. The method for automatic system repair and fault environment preservation according to claim 9, characterized in that: The fault scene stored in the repair partition can be exported as needed, and the exported fault scene can be analyzed in conjunction with the fault cause to determine the source of the fault: 1) When the cause of the failure is determined to be a corrupted system image file, the location of the corruption in the system image file is confirmed by comparing the checksum of the image file; 2) When the cause of the failure is determined to be file system corruption, check the corresponding applications deployed in the system to determine the location of the application that caused the file system corruption; 3) When the cause of the failure is determined to be that the file has been deleted or its content has been changed, check the disk hardware to determine whether the data error was caused by the disk hardware or by the application directly accessing the disk's raw interface and rewriting the data on the disk.