Disaster recovery method in industrial equipment software OTA process

By monitoring the update status and network quality during the OTA update process of industrial equipment, and re-executing the OTA update or recovery backup program, the problem of the failure of OTA updates caused the device to not work properly, improving the reliability and user experience of the update.

CN120196481AActive Publication Date: 2025-06-24KINGWAY FOSHAN ELECTRONICS TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510645109.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-24
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

During the OTA update process of industrial equipment software, the update may fail due to network interruption, insufficient storage space, hardware compatibility issues or software conflicts, resulting in the equipment not working properly, resulting in user experience problems and economic losses.

Method used

By monitoring the OTA update status and device startup status, if an update failure or device startup fails, the network quality level of the current industrial equipment is detected. If the network quality level is above the set threshold, re-execute the OTA update; if it is below or equal to the threshold, select the backup program to restore the device to run.

Benefits of technology

Provides a more comprehensive disaster recovery mechanism, ensuring reliability when OTA updates, reducing the risk of device failure, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196481A_ABST
    Figure CN120196481A_ABST
Patent Text Reader

Abstract

The invention discloses a disaster recovery method in an industrial equipment software OTA process, and belongs to the technical field of data processing, and the method comprises the following steps: S1, obtaining a current running program, generating a backup program, and executing OTA updating; s2, detecting an OTA updating state and an equipment starting state, and if OTA updating failure or equipment starting failure is detected, detecting the network quality grade of the current industrial equipment; if the network quality grade of the current industrial equipment is higher than a set threshold value, executing the step S3; if the network quality grade of the current industrial equipment is lower than or equal to a set threshold value, executing a step S4; s3, OTA updating is executed again; and S4, selecting a backup program to recover the operation of the equipment. According to the disaster recovery method in the industrial equipment software OTA process, the problem of low OTA updating reliability caused by the risk of equipment failure due to updating failure in the existing OTA updating process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a disaster recovery method during the OTA process of industrial equipment software. Background Art

[0002] OTA remote update is a key link in industrial software maintenance. However, in practical applications, it faces a complex technical challenge: any failure during the update process may cause the device to malfunction, resulting in serious user experience problems and potential economic losses. The reasons for update failures are diverse and may be caused by factors such as network interruption, insufficient storage space, hardware compatibility issues, or software conflicts. More seriously, once the update fails, the device may enter a state where it cannot be started, and remote intervention becomes extremely difficult.

[0003] This forms a technical contradiction: on the one hand, we need to perform OTA updates frequently to fix vulnerabilities and add functions; on the other hand, each update potentially increases the risk of device failure. In addition, different types of devices and diverse usage environments further increase the complexity of the problem, making it difficult to ensure the reliability during OTA updates. Summary of the Invention

[0004] In order to overcome the defects existing in the prior art, the present invention provides a disaster recovery method during the OTA process of industrial equipment software to solve the above problems.

[0005] The technical solution adopted by the present invention to solve its technical problems is: a disaster recovery method during the OTA process of industrial equipment software, including the following steps: S1: Obtain the currently running program and generate a backup program, and execute the OTA update to replace the currently running program; S2: Detect the OTA update status and the device startup status. If it is detected that the OTA update fails or the device startup fails, then detect the network quality level of the current industrial device; if the network quality level of the current industrial device is higher than the set threshold, then execute step S3; if the network quality level of the current industrial device is lower than or equal to the set threshold, then execute step S4; S3: Re-execute the OTA update; S4: Select the backup program to restore the device operation.

[0006] It should be noted that in step S2, the preset status file is detected through the recovery process or the exit status code of the OTA update process is detected through the recovery process. If the status file indicates that the update is not completed or the exit status code is the value corresponding to the update failure, then it is determined that the OTA update fails.

[0007] It should be noted that in the step S2, it is determined whether the device starts successfully by detecting the startup flag file through the recovery process or by detecting the running status of the device's key services through the recovery process.

[0008] It should be noted that in the step S2, the steps of detecting the preset status file through the recovery process include: At the beginning of the OTA update, a preset status file is set; after the update is successfully completed, the status file is deleted and a first success flag file is created; When the recovery process detects the status file, if the status file exists and the first success flag file does not exist, it is determined that the OTA update fails; if the first success flag file exists, it is determined that the OTA update is successful.

[0009] Specifically, in the step S2, the steps of detecting the exit status code of the OTA update process through the recovery process include: in the OTA update program, values corresponding to update success and update failure are assigned to the exit status code; the recovery process obtains the exit status code of the OTA update process; if the exit status code is the value corresponding to update success, it is determined that the OTA update is successful; if the exit status code is the value corresponding to update failure, it is determined that the OTA update fails.

[0010] Preferably, in the step S2, the steps of detecting the startup flag file through the recovery process include: When the system starts, a startup flag file is created; after the system starts completely, the startup flag file is deleted and a second success flag file is created; When the recovery process detects the startup flag file, if the startup flag file exists and the second success flag file does not exist, it is determined that the device startup fails; if the second success flag file exists, it is determined that the device startup is successful.

[0011] It should be noted that in the step S2, when detecting the running status of the device's key services through the recovery process, if the device's key services are not running, it is determined that the device startup fails; if the device's key services are running, it is determined that the device startup is successful.

[0012] Optionally, in the step S2, it is detected whether the device starts through the hardware watchdog timer. If the hardware watchdog timer is not reset regularly, it is determined that the device startup fails.

[0013] It should be noted that in the step S1, before performing the OTA update, the program file and related configuration files are obtained from the path where the current running program is located, a backup program is generated and stored in the specified backup path, and the storage location of the backup program is recorded after the storage is completed.

[0014] It should be noted that in the step S2, the steps for re - executing the OTA update are as follows: Call the OTA update tool through the recovery process to re - obtain the update package and execute the OTA update; During the execution of step S3, detect the number of times of executing the OTA update. When the number of times of executing the OTA update exceeds the update threshold, if the OTA update fails or the device fails to start, select the backup program to restore the device operation.

[0015] The beneficial effects of the present invention are as follows: In the disaster recovery method during the OTA process of industrial equipment software, the update result is judged by monitoring the OTA update status and the device start status. If it is detected that the OTA update fails or the device fails to start, then judge whether the network environment is stable by detecting the network quality level of the current industrial device. If the network quality level of the current industrial device is higher than the set threshold, it indicates that the network environment is stable and the OTA update will be re - triggered. If the network quality level of the current industrial device is lower than or equal to the set threshold, it indicates that the network environment is unstable and the device operation will be restored through the backup program. This solution combines re - executing the OTA update and restoring through the backup program to provide a more comprehensive disaster recovery mechanism and ensure the reliability during OTA updates. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of the disaster recovery method during the OTA process of industrial equipment software in an embodiment of the present invention; Figure 2 It is a flowchart of detecting the exit status code of the OTA update process through the recovery process in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following further describes the specific embodiments of the present invention with reference to the drawings. It should be noted here that the description of these embodiments is for helping to understand the present invention, but does not limit the present invention. In addition, the technical features involved in the following various embodiments of the present invention can be combined with each other as long as they do not conflict with each other.

[0018] As Figure 1 and 2 shown, a disaster recovery method during the OTA process of industrial equipment software includes the following steps: S1: Obtain the currently running program and generate a backup program, and execute the OTA update to replace the currently running program; S2: Detect the OTA update status and the device start status. If it is detected that the OTA update fails or the device fails to start, then detect the network quality level of the current industrial device; If the network quality level of the current industrial device is higher than the set threshold, then execute step S3; If the network quality level of the current industrial device is lower than or equal to the set threshold, then execute step S4; S3: Re - execute the OTA update; S4: Select the backup program to restore the device operation.

[0019] In the disaster recovery method during the OTA process of the industrial device software, the update result is judged by monitoring the OTA update status and the device startup status. If it is detected that the OTA update fails or the device startup fails, then the network environment stability is judged by detecting the network quality level of the current industrial device. If the network quality level of the current industrial device is higher than the set threshold, it indicates that the network environment is stable, and the OTA update will be triggered again. If the network quality level of the current industrial device is lower than or equal to the set threshold, it indicates that the network environment is unstable, and the device operation will be restored through the backup program. This solution combines re - executing the OTA update and restoring through the backup program to provide a more comprehensive disaster recovery mechanism and ensure the reliability during OTA updates.

[0020] In this embodiment, detailed logs are recorded for each OTA update and backup program restoration operation, which is convenient for subsequent analysis and troubleshooting.

[0021] In this embodiment, the generation process of the network quality level includes: Obtain network data streams from the target network device. The network data streams include packet loss rate (packet loss rate in the past 5 minutes), latency fluctuation (average latency standard deviation), number of connection interruptions (number of disconnections in the past 1 hour), download rate (average rate of the most recent OTA), and upload rate (average rate reported by the device status); Calculate the network stability score based on the weighted sum of the packet loss rate, latency fluctuation, and number of connection interruptions; where the weight of the packet loss rate is 0.4, the weight of the latency fluctuation is 0.3, and the weight of the number of connection interruptions is 0.3; Calculate the network speed score based on the weighted sum of the download rate and upload rate; where the weight of the download rate is 0.6 and the weight of the upload rate is 0.4; Calculate the comprehensive score based on the weighted sum of the network stability score and the network speed score, where the weight of the network stability score is 0.7 and the weight of the network speed score is 0.3; Set a corresponding network quality level for each preset quality level interval, and determine the current network quality level according to the mapping relationship between the comprehensive score and the preset quality level interval.

[0022] It should be noted that in step S2, the preset status file is detected through the recovery process or the exit status code of the OTA update process is detected through the recovery process. If the status file indicates that the update is not completed or the exit status code is the value corresponding to the update failure, then it is judged that the OTA update fails.

[0023] In this embodiment, in the case of OTA upgrading multiple industrial devices, multiple servers are set to locate OTA update failures, and the steps include: Group all industrial devices by binary encoding to assign a unique binary encoding to each industrial device; In the binary encoding, the two states (0 or 1) of each binary bit are bound to different servers; for example, the state 0 of the k-th bit in the binary encoding is bound to one server, the state 1 of the k-th bit in the binary encoding is bound to one server, and the state 1 of the m-th bit in the binary encoding is also bound to one server; The industrial device communicates and connects with the server corresponding to its own binary encoding; for example, if the binary encoding of the industrial device is 101, this industrial device will communicate and connect with the server bound to the state 1 of the first binary bit, the server bound to the state 0 of the second binary bit, and the server bound to the state 1 of the third binary bit respectively; when the OTA update is successful, the industrial device will send a success signal to the server with which it communicates, such as returning 1, and when the OTA update fails, the industrial device will send a failure signal to the server with which it communicates, such as returning 0; After receiving the failure signal, the server will form a feedback signal; the host computer sorts the servers with feedback signals according to the binary bits bound to the servers themselves to obtain the binary encoding of the industrial device with OTA update failure, so as to locate the industrial device with OTA update failure. For example, if the industrial device with binary encoding 101 fails in OTA update, it will send failure signals to the server bound to the state 1 of the first binary bit, the server bound to the state 0 of the second binary bit, and the server bound to the state 1 of the third binary bit. After receiving the failure signals, these three servers will form feedback signals to the host computer. The host computer traces back according to the feedback signals to obtain the server that sends the feedback signal, and obtains the binary bit bound to the server and the state (0 or 1) of the binary bit, and then recombines the binary encoding according to the binary bit and the state of the binary bit to obtain the binary encoding 101, so as to locate the corresponding industrial device.

[0024] Preferably, the device cannot start usually due to reasons such as system crash, kernel error, or file system corruption. In step S2, the startup flag file is detected through the recovery process or the running state of the device's key services is detected through the recovery process to determine whether the device starts successfully.

[0025] Before OTA update, the system starts an independent recovery process and keeps it running. The specific process is as follows: Write a program to restore the process (such as recovery_daemon), implemented in C, C++, or other languages. This program needs to include monitoring logic and OTA re-triggering logic to create an executable file for the recovery process; Create a systemd service file in the / etc / systemd / system / directory, and the code is as follows: [Unit] Description=Recovery Daemon for OTA Updates After=network.target [Service]Type=simple ExecStart= / usr / bin / recovery_daemonRestart=always RestartSec=5 [Install] WantedBy=multi-user.target In this embodiment, the recovery process is an independent daemon process responsible for monitoring the OTA update status and the device startup status. The recovery process has a high priority to ensure normal operation in case of device anomalies. Through an independent recovery process, automated repair of OTA update failures is achieved, reducing manual intervention.

[0026] Specifically, by combining the detection mechanisms for update failures and device unbootability to form a comprehensive detection mechanism, the recovery process can operate according to the following logic: Detection at startup: Check the startup flag file or key services to determine whether the device starts up normally; Detection during update: Check the OTA update status file or exit status code to determine whether the update is successful; Recovery logic: If an update failure or device unbootability is detected, re-trigger the OTA update.

[0027] If the OTA update is successful, the device starts with the new version program.

[0028] Optionally, in step S2, the step of detecting a preset status file by the recovery process includes: At the start of the OTA update, preset a status file (such as / var / ota_in_progress) to indicate that the update is in progress; after the update is successfully completed, delete this status file and create a first success flag file (such as / var / ota_success); if the update fails, retain the status file; After the OTA update ends, when the recovery process detects the status file, if the status file exists and the first success flag file does not exist, it is determined that the OTA update fails; if the first success flag file exists, it is determined that the OTA update is successful.

[0029] The code is as follows: #include<stdio.h> #include<unistd.h> #include<sys / stat.h> #define OTA_IN_PROGRESS_FILE " / var / ota_in_progress" #define OTA_SUCCESS_FILE " / var / ota_success" int is_ota_failed() { if(access(OTA_IN_PROGRESS_FILE,F_OK) == 0&&access(OTA_SUCCESS_FILE,F_OK) != 0) { return 1; / / OTA failed } return 0; / / OTA succeeded or not started } Specifically, in the step S2, the steps of detecting the exit status code of the OTA update process through the recovery process include: in the OTA update program, assign the values corresponding to successful update and failed update to the exit status code; the recovery process obtains the exit status code of the OTA update process through waitpid or a similar mechanism; if the exit status code is the value corresponding to successful update, it is determined that the OTA update is successful, and if the exit status code is the value corresponding to failed update, it is determined that the OTA update fails.

[0030] The code is as follows: #include<stdio.h> #include<stdlib.h> #include<unistd.h> #include<sys / wait.h> void run_ota_update() { pid_t pid = fork(); if (pid == 0) { / / Child process: run OTA update execl(" / usr / bin / ota_update", "ota_update", NULL); exit(1); / / If execl fails } else if (pid>0) { / / Parent process: wait for OTA update to complete int status; waitpid(pid,&status, 0); if (WIFEXITED(status)&&WEXITSTATUS(status) != 0) { printf("OTA update failed with status %d\n", WEXITSTATUS(status)); } } } It should be noted that in the step S2, the steps of detecting the startup flag file by the recovery process include: When the system starts, create a startup flag file (such as / var / boot_in_progress); after the system starts completely, delete the startup flag file and create a second success flag file (such as / var / boot_success); After the system starts and after a preset time, when the recovery process detects the startup flag file, if the startup flag file exists and the second success flag file does not exist, it is determined that the device startup fails; if the second success flag file exists, it is determined that the device startup is successful.

[0031] The code is as follows: #include<stdio.h> #include<unistd.h> #include<sys / stat.h> #define BOOT_IN_PROGRESS_FILE " / var / boot_in_progress" #define BOOT_SUCCESS_FILE " / var / boot_success" int is_boot_failed() { if (access(BOOT_IN_PROGRESS_FILE,F_OK)==0&&access(BOOT_SUCCESS_FILE,F_OK) != 0) { return 1; / / Boot failed } return 0; / / Boot succeeded or not started } Preferably, in the step S2, when detecting the running status of the device key services (such as network service or file system service, etc.) through the recovery process, if the device key service is not running, it is determined that the device startup fails; if the device key service is running, it is determined that the device startup is successful.

[0032] The code is as follows: #include<stdio.h> #include<stdlib.h> int is_service_running(const char *service_name) { char command

[256] ; snprintf(command, sizeof(command), "systemctl is-active --quiet %s",service_name); return system(command) == 0; } int is_boot_failed() { If(!is_service_running("network.service") ||!is_service_running("filesystem.service")) { return 1; / / Boot failed } return 0; / / Boot succeeded } Optionally, in the step S2, it is detected whether the device starts up through the hardware watchdog timer. If the hardware watchdog timer is not reset regularly, it is determined that the device startup fails.

[0033] Use a hardware watchdog timer to detect whether the system is running properly. If the system does not reset the watchdog regularly, the hardware watchdog timer will trigger a restart to detect whether the device fails to start. The code is as follows: #include<stdio.h> #include<unistd.h> #include<fcntl.h> #define WATCHDOG_DEVICE " / dev / watchdog" void feed_watchdog() { int fd = open(WATCHDOG_DEVICE, O_WRONLY); if (fd>= 0) { write(fd, "1", 1); / / Feed the watchdog close(fd); } } Specifically, in step S1, before performing OTA update, obtain the program file and related configuration files from the path where the currently running program is located, generate a backup program and store it in the specified backup path. After the storage is completed, record the storage location of the backup program.

[0034] Before OTA update, the system makes a complete backup of the currently running program, configuration files, and dependent libraries to the specified path. The backup path can be local storage (such as a hard disk, flash memory) or remote storage (such as a server). When the device starts, the system will detect whether the main program can start normally. If the main program fails to start, the system will automatically switch to the program under the backup path.

[0035] Specifically, in step S3, the steps to re-perform OTA update are as follows: call the OTA update tool through the recovery process to re-obtain the update package and perform OTA update; during the execution of step S3, detect the number of times of performing OTA update. When the number of times of performing OTA update exceeds the update threshold, if OTA update fails or the device fails to start, select the backup program to restore the device operation. If the newly performed OTA update still fails, the recovery process can try multiple times (the upper limit of the number of times is the update threshold). If multiple attempts fail, the recovery process can switch to the backup program and notify the operation and maintenance personnel. This solution backs up the current program before performing the update and realizes automatic recovery of the device by switching to the backup file when multiple updates fail.

[0036] In this embodiment, the recovery process will re-download the update package and verify its integrity. If the update package is complete, the recovery process will re-execute the OTA update process.

[0037] In step S4, detect the OTA update process status. If the OTA update fails after being triggered repeatedly for multiple times, record that the number of failures reaches the threshold and generate a status flag. According to the status flag, switch the system to the backup path and load the program code in the backup file. Start the device through the backup path, obtain the device running log, and determine whether the startup is successful. If the device starts successfully, extract the key parameters in the running log to determine the current running state of the system. Generate an alarm message using the running state data, including the number of failures and the device identifier. Transmit the alarm message to the operation and maintenance system through the system communication module to obtain a transmission confirmation signal. If the transmission confirmation signal is not received, retry sending the alarm message and record the number of retries. Adjust the sending frequency according to the number of retries and determine whether the alarm message is successfully delivered. Obtain the feedback data from the operation and maintenance system, update the local log, and determine that the alarm handling is completed.

[0038] For example, when detecting the OTA update process status, the system uses a polling mechanism to check the update progress every 5 seconds. If the update fails or the device fails to start for 3 consecutive times, it is determined that the number of failures reaches the threshold and a "OTA_FAILED" status flag is generated. According to this flag, the system automatically switches to a predefined backup path, loads the backup program and verifies its integrity. After extracting the key parameters in the log such as memory occupancy rate and CPU load, the system generates an alarm message. The communication module calls the MQTT protocol to push the alarm to the operation and maintenance system address. If no ACK response is received within 5 seconds, retry according to the exponential backoff algorithm (such as intervals of 1s, 2s, 4s). When the number of retries exceeds 3 times, automatically switch to the TCP long connection channel to resend. After the operation and maintenance system returns the processing result, the local log file appends the record timestamp and the processing status.

[0039] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments still fall within the protection scope of the present invention.

Claims

1. A disaster recovery method in the OTA process of industrial equipment software, characterized in that: The following steps are involved: S1: Get the current running program and generate a backup program, and perform an OTA update to replace the current running program; S2: Detect the OTA update status and the device startup status. If the OTA update failure or the device startup failure is detected, the network quality level of the current industrial device is detected; if the network quality level of the current industrial device is higher than the set threshold, step S3 is executed; if the network quality level of the current industrial device is lower than or equal to the set threshold, step S4 is executed; S3: Re-execute OTA update; S4: Select the backup program to restore the device.

2. The method for disaster recovery in the OTA process of industrial equipment software according to claim 1, characterized in that: In step S2, a preset status file is detected through the recovery process or an exit status code of the OTA update process is detected through the recovery process. If the status file indicates that the update is not completed or the exit status code is a value corresponding to an update failure, it is determined that the OTA update has failed.

3. The method for disaster recovery in the OTA process of industrial equipment software according to claim 1, characterized in that: In step S2, a startup flag file is detected by the recovery process or the running status of a key service of the device is detected by the recovery process to determine whether the device is successfully started.

4. The method for disaster recovery in the OTA process of industrial equipment software according to claim 2, characterized in that: In step S2, the step of detecting a preset status file through a recovery process includes: When the OTA update starts, a status file is preset; after the update is successfully completed, the status file is deleted and a first success flag file is created; When the recovery process detects the status file, if the status file exists and the first success flag file does not exist, it is determined that the OTA update has failed; if the first success flag file exists, it is determined that the OTA update has succeeded.

5. The method for disaster recovery in the OTA process of industrial equipment software according to claim 2, characterized in that: In the step S2, the step of detecting the exit status code of the OTA update process through the recovery process includes: in the OTA update program, assigning values ​​corresponding to update success and update failure to the exit status code; the recovery process obtains the exit status code of the OTA update process; if the exit status code is a value corresponding to a successful update, it is judged that the OTA update is successful, and if the exit status code is a value corresponding to an update failure, it is judged that the OTA update has failed.

6. The method for disaster recovery in the OTA process of industrial equipment software according to claim 3, characterized in that: In step S2, the step of detecting the startup flag file through the recovery process includes: When the system starts, a startup flag file is created; after the system is completely started, the startup flag file is deleted and a second success flag file is created; When the recovery process detects the startup flag file, if the startup flag file exists and the second success flag file does not exist, it is determined that the device startup fails; if the second success flag file exists, it is determined that the device startup succeeds.

7. The method for disaster recovery in the OTA process of industrial equipment software according to claim 3, characterized in that: In the step S2, when the running status of the key service of the device is detected through the recovery process, if the key service of the device is not running, it is determined that the device startup fails; if the key service of the device is running, it is determined that the device startup is successful.

8. The method for disaster recovery in the OTA process of industrial equipment software according to claim 1, characterized in that: In step S2, whether the device is started is detected by a hardware watchdog timer. If the hardware watchdog timer is not reset regularly, it is determined that the device fails to start.

9. The method for disaster recovery in the OTA process of industrial equipment software according to claim 1, characterized in that: In the step S1, before executing the OTA update, the program file and related configuration files are obtained from the path where the currently running program is located, a backup program is generated and stored in a specified backup path, and the storage location of the backup program is recorded after the storage is completed.

10. The method for disaster recovery in the OTA process of industrial equipment software according to claim 1, characterized in that: In step S3, the step of re-executing the OTA update is: calling the OTA update tool through the recovery process to re-acquire the update package and execute the OTA update; During the execution of step S3, the number of times the OTA update is executed is detected. When the number of times the OTA update is executed exceeds the update threshold, it is detected that the OTA update has failed or the device has failed to start, and a backup program is selected to restore the device operation.

Citation Information

Patent Citations

  • Online upgrading method of microcontroller, microcontroller and storage medium

    CN112181455A

  • OTA upgrading method and device, equipment and storage medium

    CN117369844A

  • Engine system and method for vehicle end OTA upgrading

    CN117749841A

  • Automatic OTA rescue upgrading method and device, electronic equipment and storage medium

    CN118474720A

  • Logistics vehicle updating method and device

    CN119002955A