Remote flashing and backup system and method for artificial intelligence machines

By setting up an independent read-only recovery subsystem in the AI ​​machine, combined with Rescuezilla and Tauri, remote flashing and backup of the AI ​​machine were realized, solving the problems of inconsistent system upgrades, difficult on-site recovery, and complex processes, and improving the stability and management efficiency of the equipment.

CN121501301BActive Publication Date: 2026-05-26ZHONGKE TIMES (SHENZHEN) COMPUTER SYST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHONGKE TIMES (SHENZHEN) COMPUTER SYST CO LTD
Filing Date
2026-01-12
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The existing system upgrade solutions for AI machines are not standardized, resulting in high management costs, consistency risks, insufficient on-site recovery capabilities, long downtime and high costs due to factory repairs, complex and error-prone upgrade processes, and restrictions on backup methods due to the inability to use USB drives. These factors make it difficult to meet the needs for rapid deployment and unified maintenance.

Method used

A remote flashing and backup system based on Rescuezilla as the backend service and Tauri as the client is adopted. By setting up an independent read-only recovery subsystem in the AI ​​machine, a remote flashing and backup operating environment is provided, the management process is unified, manual operation is reduced, automated verification and image writing are achieved, and adaptation to different system architectures is supported.

Benefits of technology

It enables unified, efficient, and convenient system maintenance of AI machines, shortens on-site recovery time, reduces the rate of misoperation, improves the stability and reliability of equipment, and meets the needs of factories for rapid deployment and traceability management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121501301B_ABST
    Figure CN121501301B_ABST
Patent Text Reader

Abstract

This application provides a remote flashing and backup system and method for AI devices. It includes: a client unit, used to obtain template information corresponding to the image resource based on the image resource identifier, and initiate a verification request to the server unit on the target AI device side; a recovery subsystem, used to provide a runtime environment for remote flashing and backup; a server unit, used to match and verify device information, system status information, and the image resource identifier according to the template information, generating verification information and task summary information, the task summary information indicating the target storage area identifier; a confirmation control unit, used to generate execution instructions and send them to the server unit; and a flashing execution unit, used to perform an image write operation on the target storage area according to the target storage area identifier when receiving the execution instruction and the execution instruction indicates flashing. This application can unify the flashing and backup process, shorten on-site recovery time, and reduce the error rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial automation technology, and in particular to a remote writing and backup system and method for industrial machines. Background Technology

[0002] Industrial robots (ICDs) in factories typically handle tasks such as production control, data acquisition, and human-machine interaction. Their system maintenance mainly involves system upgrades, fault recovery, and system and data backup. Because factories have high requirements for downtime, maintenance efficiency, and on-site compliance, ICDs need to have unified and controllable maintenance methods to ensure continuous production.

[0003] In existing technologies, system upgrades and flashing of AI machines mostly adopt two methods: one is to perform offline upgrades or reinstalls on-site via mobile storage media (such as USB media); the other is to complete the flashing by entering the installation or recovery environment through network booting (such as PXE). At the same time, when the system is damaged or the upgrade fails, it is often necessary to rely on on-site personnel to perform multiple steps according to the documentation, or to return the equipment to the factory for repair to complete system recovery and reinstallation.

[0004] The aforementioned existing technologies have at least the following problems: First, different devices use multiple upgrade solutions such as USB or network booting, resulting in inconsistent maintenance processes and tools, increasing management costs and introducing consistency risks; second, the on-site recovery capability after system damage is insufficient, and returning to the factory for repair leads to long downtime and high costs; third, the upgrade and flashing process is complex, heavily reliant on manual operation, and prone to errors such as incorrect image selection and incorrect target disk writing; fourth, factory safety management restricts the entry of removable media, limiting offline upgrade and backup methods and making it difficult to meet the needs of rapid on-site deployment, unified maintenance, and traceable management. Summary of the Invention

[0005] In view of this, embodiments of this application provide a remote flashing and backup system and method for artificial intelligence machines to solve the problems of inconsistent upgrade and flashing methods, difficulty in on-site recovery of damaged systems, and complex and error-prone flashing and backup processes in the prior art.

[0006] The first aspect of this application provides a remote flashing and backup system for an AI device, comprising: a client unit, configured to receive flashing or backup task parameters input by a user, the flashing or backup task parameters including a mirror resource identifier, the network address of the target AI device, and authentication information; obtain template information corresponding to the mirror resource based on the mirror resource identifier, and initiate a verification request to a server unit on the target AI device side; a recovery subsystem, set in the target AI device, the recovery subsystem being an independent read-only partition, configured to provide a runtime environment for remote flashing and backup within the recovery subsystem; and a server unit, running in the recovery subsystem, configured to obtain device information and system status information of the target AI device after receiving the verification request. The system collects and verifies device information, system status information, and image resource identifiers based on template information, generating verification information and task summary information. The task summary information is used to indicate the target storage area identifier. The confirmation control unit receives user confirmation instructions after the client unit receives the verification information and task summary information, and generates execution instructions to send to the server unit. The flashing execution unit performs image writing operations on the target storage area based on the target storage area identifier when it receives the execution instruction and the execution instruction indicates flashing. The backup execution unit reads system data from the target storage area based on the target storage area identifier and generates backup data when it receives the execution instruction and the execution instruction indicates backup.

[0007] The second aspect of this application provides a method for remote flashing and backup of an AI device based on the system of the first aspect, comprising: receiving flashing or backup task parameters input by a user, the flashing or backup task parameters including a mirror resource identifier, the network address of the target AI device, and authentication information; obtaining template information corresponding to the mirror resource based on the mirror resource identifier, and initiating a verification request to a server unit on the target AI device side based on the network address of the target AI device; in the server unit, obtaining device information and system status information of the target AI device, and performing matching verification between the device information, system status information, and mirror resource identifier according to the template information, generating verification information and task summary information, the task summary information being used to indicate a target storage area identifier; after receiving the verification information and task summary information, receiving a user confirmation instruction, generating an execution instruction and sending it to the server unit, the execution instruction being used to indicate the task type and the target storage area identifier; in the server unit, performing remote flashing or remote backup according to the execution instruction, wherein remote flashing includes performing a mirror write operation on the target storage area according to the target storage area identifier, and remote backup includes reading system data from the target storage area according to the target storage area identifier and generating backup data.

[0008] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects:

[0009] The client unit receives user-input flashing or backup task parameters, including the image resource identifier, the target AI device's network address, and authentication information. Based on the image resource identifier, it obtains the template information corresponding to the image resource and initiates a verification request to the server unit on the target AI device side. The recovery subsystem, located within the target AI device, is an independent read-only partition that provides the operating environment for remote flashing and backup. The server unit, running within the recovery subsystem, obtains the target AI device's device information and system status information upon receiving the verification request and, based on the template information, processes the device information... The system status information is matched and verified with the image resource identifier to generate verification information and task summary information. The task summary information is used to indicate the target storage area identifier. A confirmation control unit receives the user confirmation instruction after the client unit receives the verification information and task summary information, and generates an execution instruction to send to the server unit. A flashing execution unit performs an image write operation on the target storage area according to the target storage area identifier when it receives an execution instruction indicating flashing. A backup execution unit reads system data from the target storage area and generates backup data according to the target storage area identifier when it receives an execution instruction indicating backup. This application can unify the flashing and backup process, shorten on-site recovery time, and reduce the error rate. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of the structural composition of the AI ​​remote flashing and backup system provided in the embodiments of this application;

[0012] Figure 2 This is a flowchart illustrating the remote flashing and backup method for artificial intelligence provided in this application embodiment. Detailed Implementation

[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0014] In factory settings, AI-powered industrial robots currently face several challenges. First, the lack of standardized system upgrade methods significantly complicates equipment management. Some AI-powered robots use USB for system upgrades, while others use PXE. This inconsistency necessitates the use of different tools and methods for equipment maintenance and upgrades, increasing complexity and difficulty. Furthermore, different upgrade methods may lead to varying upgrade outcomes and stability, posing potential risks to factory production.

[0015] Secondly, when the AI ​​machine's system malfunctions, it cannot be restored on-site and must be returned to the factory for repair. This not only delays production time but also increases repair costs. In factory production, time is money, and equipment failure can cause production line shutdowns, affecting the entire factory's production schedule. Moreover, returning the equipment for repair requires transporting it back to the factory, which increases transportation costs and the risk of further equipment damage.

[0016] Furthermore, existing AI machine system upgrade solutions are overly complex. Even following detailed documentation, errors can still occur. This poses a significant challenge for factory technicians. They need to invest considerable time and effort in learning and understanding the upgrade plan, and operate with extreme caution to avoid mistakes. This complex upgrade process not only reduces work efficiency but also increases the likelihood of human error.

[0017] Finally, due to the factory's unique environment and safety requirements, USB drives cannot be brought into the customer's site. This makes it impossible to use USB drives for system upgrades or data backups in certain situations. This further limits the maintenance and management of the AI ​​machines, causing inconvenience to the factory's equipment management.

[0018] To address the pain points of the existing technologies, this application proposes a remote flashing and backup system and method for AI machines based on Rescuezilla as the backend service and Tauri as the client. Through remote flashing and backup, unified, efficient, and convenient system maintenance can be achieved, improving the stability and reliability of AI machines and thus ensuring smooth factory production.

[0019] The specific modules and functions of the AI ​​remote flashing and backup system provided in this application will be described in detail below with reference to the accompanying drawings and specific embodiments. Figure 1 This is a schematic diagram of the structural composition of the AI ​​remote flashing and backup system provided in the embodiments of this application, as shown below. Figure 1 As shown, the system may specifically include the following components:

[0020] The client unit 101 is used to receive the flashing or backup task parameters input by the user. The flashing or backup task parameters include the image resource identifier, the network address of the target AI machine, and authentication information; obtain the template information corresponding to the image resource based on the image resource identifier, and send a verification request to the server unit on the target AI machine side.

[0021] Recovery subsystem 102 is set in the target AI machine. The recovery subsystem is an independent read-only partition, which is used to provide a running environment for remote flashing and backup.

[0022] The server unit 103 runs in the recovery subsystem and is used to obtain the device information and system status information of the target AI machine after receiving the verification request. It also matches and verifies the device information, system status information and image resource identifier according to the template information, and generates verification information and task summary information. The task summary information is used to indicate the target storage area identifier.

[0023] The confirmation control unit 104 is used to receive the user confirmation instruction after the client unit receives the verification information and task summary information, and to generate an execution instruction to send to the server unit.

[0024] The write execution unit 105 is used to perform a mirror write operation on the target storage area according to the target storage area identifier when an execution instruction is received and the execution instruction indicates write.

[0025] The backup execution unit 106 is used to read system data from the target storage area and generate backup data according to the target storage area identifier when it receives an execution instruction that indicates backup.

[0026] In some embodiments, the template information includes flashing configuration parameters that match the target AI device model identifier and storage area mapping rules. The storage area mapping rules are used to map the image content corresponding to the image resource identifier to the target storage area identifier.

[0027] Specifically, after receiving the image resource identifier input by the user, the client unit first locates the corresponding template information based on the image resource identifier. The template information can be packaged and stored together with the image, or stored in the client's template library and indexed by the image resource identifier. The template information is described in a structured manner, and its core fields include at least the compatible device model field, the flashing configuration parameter field, and the storage area mapping rule field. The compatible device model field is used to list the set of compatible device identifiers or device identifier patterns, such as limiting it to a certain product code range or a certain hardware version segment, to prevent the image from being applied to incompatible devices.

[0028] The flash configuration parameters field specifies the key strategies to be used during flashing, such as defining flashing strategies for real-time and non-real-time domains, whether to allow overwriting the boot partition, whether to perform consistency checks after writing, and whether to allow file-level or partition-level writing to virtual machine image files in non-real-time domains. The storage region mapping rules field describes the correspondence between image content and target storage regions. It can be expressed using a mapping relationship of "image content identifier → target storage region identifier," thus decomposing the abstract image content into locatable write objects.

[0029] Furthermore, to adapt to different system architectures, the storage area mapping rules in this embodiment can cover at least two types of storage formats: the first type is the real-time domain corresponding area, which is usually the partition or logical volume where the underlying Linux system resides. Its mapping rules are used to indicate the target storage area identifier to which the real-time domain system image should be written. The second type is the non-real-time domain corresponding area, which is usually the storage space where the Windows virtual machine system resides. In some implementations, the non-real-time domain exists in the form of a virtual machine image file, and the mapping rules are used to indicate the target file area identifier to which this virtual machine image file should be written. In other implementations, the virtual machine image file of the non-real-time domain is split into an independent partition, and the mapping rules are used to indicate the partition area identifier corresponding to this independent partition. Through the unified expression of the above mapping rules, the server unit does not need to rely on manual selection of the target disk or target partition when performing a write operation. Instead, it automatically locates the write object based on the target storage area identifier, thereby reducing the probability of erroneous operations.

[0030] In one example, a target AI machine adopts a dual-domain architecture of "real-time domain + non-real-time domain," where the real-time domain is an underlying Linux system and the non-real-time domain is a Windows virtual machine system. The server unit runs in the recovery subsystem of the target AI machine. The recovery subsystem is an independent read-only partition and can access the main system partition directory. When the client unit initiates a verification request to the server unit, the server unit obtains device information to parse the machine model identifier and obtains system status information to determine the system architecture type.

[0031] Subsequently, the server unit selects the flashing configuration parameters and storage area mapping rules that match the device model identifier from the template information, and performs validity verification on the target storage area identifier indicated in the mapping rules, such as checking whether the target storage area exists, whether the capacity is sufficient, and whether the partition type conforms to the flashing policy.

[0032] After successful verification, the server unit generates task summary information, which includes a set of target storage area identifiers determined according to the storage area mapping rules, and sends it back to the client unit to display the execution scope information. After user confirmation, during the flashing execution phase, the server unit performs real-time domain mirror writing on the corresponding area of ​​the real-time domain and non-real-time domain mirror writing on the corresponding area of ​​the non-real-time domain, based on the target storage area identifier set, thereby completing the unified remote flashing of this model.

[0033] In another example, the target AI machine only contains a real-time domain system and its partition structure differs from the aforementioned dual-domain machine model. The adaptable machine model field in the template information is limited to the identifier of this single-domain machine model, and the storage area mapping rules only configure the mapping relationship from the real-time domain system image to the target partition area identifier. After the server unit selects this template by matching the machine model identifier, the task summary information only contains the target storage area identifier corresponding to the real-time domain, thereby achieving unified flashing for different machine models without changing the client operation process.

[0034] As can be seen from the above embodiments, the template information in this embodiment uses the model identifier as the matching entry point and the flashing configuration parameters and storage area mapping rules as the core constraints, so that the image content corresponding to the image resource identifier can be automatically mapped to the target storage area identifier, realizing the automatic location and policy-based execution of the flashing object, thereby providing a unified and scalable implementation foundation for subsequent remote flashing and backup.

[0035] In some embodiments, the recovery subsystem is associated with a boot control component, which provides boot options for entering the recovery subsystem during the boot phase of the target AI machine, so that the server unit can start in the recovery subsystem and provide remote flashing and backup service interfaces to the outside world.

[0036] Specifically, the recovery subsystem refers to a maintenance subsystem independent of the target AI machine's main system. It is stored in an independent partition and can be booted separately at startup. It provides maintenance capabilities such as system flashing, backup, factory reset, system updates, and rollback without relying on the integrity of the main system. The boot control component refers to the control and configuration entity running in the boot chain. It can be composed of a bootloader, boot menu configuration, and corresponding startup scripts. It provides multiple boot entries when the target AI machine is powered on and loads and controls the kernel, initial file system, and partition mounting strategy corresponding to the selected entry. The remote flashing and backup service interface refers to the set of network call interfaces exposed by the server unit. It is used to receive verification requests, execution instructions, and log acquisition requests sent by the client unit and to return verification information, task summary information, and status logs.

[0037] In this embodiment, the recovery subsystem is created and deployed as an independent read-only partition on the target AI machine's disk. Setting this partition to read-only ensures the stability and consistency of the recovery subsystem, enabling it to boot reliably even after multiple write operations and abnormal power outages. The recovery subsystem includes a pre-installed basic operating environment and maintenance tools, as well as a pre-installed server unit for remote interaction. This unit automatically starts and listens on a preset network port after entering the recovery subsystem, providing an interface for remote write and backup services. To achieve this, this embodiment can configure an auto-start item in the recovery subsystem's startup process, allowing the server unit to start with the system; or it can be configured to start the server unit only after detecting network readiness, thereby improving interface availability.

[0038] The association between the boot control component and the recovery subsystem can be configured via the boot menu. In one optional implementation, the boot control component maintains boot menu entries based on the target AI machine's bootloader, configuring at least the main system entry and the recovery subsystem entry in the boot menu. The recovery subsystem entry points to the kernel and initial file system corresponding to the recovery subsystem and carries location information for mounting the recovery subsystem partition in its boot parameters, ensuring that booting can be completed from the recovery subsystem partition after booting. The boot control component can also carry network boot policy parameters in the boot parameters of this entry, enabling the recovery subsystem to configure the network according to a preset strategy after booting, such as using a static or dynamic address acquisition method, so that client units can connect stably.

[0039] In some examples, this embodiment adds a separate read-only partition subsystem to Ubuntu, adds boot menu options, and supports functions such as factory reset, system updates, and rollback. For instance, in a specific implementation, the target AI device integrates the recovery subsystem partition with the main system image during the factory image creation stage, allowing the device to access the recovery subsystem via the boot menu upon delivery. This way, when the main system fails, on-site personnel can select to enter the recovery subsystem during the boot phase without disassembling the device or using external media, ensuring accessibility to maintenance.

[0040] In one example, a target AI machine in a factory experienced a power outage, resulting in corruption of its main system file system. After powering on, the machine was unable to access the business system. During the boot process, on-site personnel selected the "Recovery Subsystem" entry through the boot menu displayed by the boot control component. The device then started and entered the recovery subsystem. Once the recovery subsystem booted, it automatically configured the network address. In a typical network flashing scenario, a preset static address (e.g., a fixed address within a predefined network segment) was used to allow direct access from the operator's Windows client. In a batch flashing scenario in the factory, the recovery subsystem could also obtain a dynamic address or a specified network segment address through network boot configuration, allowing clients to scan and discover devices based on address ranges.

[0041] Subsequently, the server unit automatically starts in the recovery subsystem and provides remote flashing and backup service interfaces. The installer enters the target AI machine's network address on the Windows client, and the client sends a verification request to the server unit. The server unit returns verification information and a task summary. After confirmation by the installer, an execution command is issued. The server unit performs a mirror write to the target storage area in the recovery subsystem, thus completing the recovery of the main system. The entire process does not rely on removable media such as USB drives, nor does it require factory repair.

[0042] As can be seen from the above embodiments, this embodiment sets up an independent read-only recovery subsystem in the target AI machine, and provides a boot option to enter the recovery subsystem during the boot phase through the boot control component. This enables the server unit to start in the recovery subsystem and provide remote flashing and backup service interfaces to the outside world, thereby providing an independent and reliable operating environment for remote flashing and backup, and supporting rapid on-site recovery and unified maintenance processes in the event of main system anomalies.

[0043] In some embodiments, the device information, system status information, and image resource identifier are matched and verified based on the template information to generate verification information and task summary information, including:

[0044] Based on the device information parsing, the model identifier and storage structure information of the target AI machine are obtained. The storage structure information includes the storage device identifier and partition structure information.

[0045] Based on the system status information, determine the system architecture type corresponding to the target AI machine, and select a mapping rule that matches the machine type and system architecture type from the template information;

[0046] The consistency between the image resource identifier and the storage structure information is verified according to the mapping rules, the verification result is obtained, and verification information is generated to characterize the verification result.

[0047] When the verification result meets the preset verification conditions, the target storage area identifier is determined based on the mapping rules, and task summary information containing the target storage area identifier is generated.

[0048] Specifically, the system architecture type is type information used to characterize the composition of the target AI system, and is used to distinguish between a single-domain architecture that only contains a real-domain system and a dual-domain architecture that contains both a real-domain system and a non-real-time domain system; among them, the real-domain system is usually an underlying Linux system, and the non-real-time domain system is usually a Windows system running on virtualization.

[0049] Storage structure information is a set of structured information used to characterize the storage layout of the target AI machine. It includes at least storage device identifiers and partition structure information. The storage device identifier is used to locate the physical or logical storage devices on the target AI machine, and the partition structure information is used to describe the partition list, partition type, capacity range, and mount point of the storage device.

[0050] Mapping rules are the core rules in the template information, used to define the correspondence between the image content corresponding to the image resource identifier and the storage area to be written or read on the target AI machine; the target storage area identifier is the output of the mapping rules, used to uniquely indicate the target storage area to be applied to in subsequent flashing or backup, and can be a partition area identifier and / or a file area identifier.

[0051] In this embodiment, the matching and verification process occurs after the user inputs the image location and the target AI device's network address, but before the user confirms the flashing process, within the remote flashing sequence. Upon receiving the verification request from the client, the server unit retrieves device information and system status information from the recovery subsystem environment and performs the verification. Because the recovery subsystem is an independent read-only partition that can access the main system partition directory, the server unit can read the necessary hardware and storage information without relying on the main system's normal startup, thus meeting the requirement for verification and recovery even in the event of system damage.

[0052] In this embodiment, the implementation method for obtaining the machine model identifier and storage structure information based on device information parsing is as follows: The server unit determines the machine model identifier by reading a preset machine model configuration file, hardware identification field, or model field exported from the firmware; simultaneously, the server unit enumerates the storage devices of the target AI machine to obtain the storage device identifier and reads the partition table to form partition structure information. For example, the partition structure information may include the partition identifier where the real-time domain partition is located, the boot partition identifier, and the identifier of the partition or independent partition where the non-real-time domain virtual machine image file is located; if the non-real-time domain virtual machine image file is split into an independent partition, the partition structure information may also include the capacity and file system type of the independent partition. The above information can be obtained by the recovery subsystem directly reading the target disk metadata, avoiding reliance on the business processes in the main system.

[0053] Furthermore, the implementation method for determining the system architecture type based on system status information and selecting mapping rules from template information is as follows: The server unit determines whether the target AI machine is configured with a non-real-time domain virtualization environment from the system status information. For example, it checks whether there is a virtual machine image file used to host the non-real-time domain, whether there is an independent partition used to store the virtual machine image file, or whether there is a mount path identifier related to the non-real-time domain, thereby determining whether the system architecture type is single-domain or dual-domain. Subsequently, the server unit matches the machine type identifier with the system architecture type in the template information and selects the corresponding set of mapping rules. This set of mapping rules at least distinguishes between real-time domain mapping rules and non-real-time domain mapping rules: real-time domain mapping rules are used to indicate the target partition area identifier corresponding to the real-time domain image; non-real-time domain mapping rules are used to indicate the file area identifier or partition area identifier corresponding to the non-real-time domain image, in order to adapt to the implementation method of "virtual machine image files of non-real-time domain systems can be split into independent partitions" in the specification.

[0054] Furthermore, the implementation method for consistency verification of image resource identifiers and storage structure information according to mapping rules is as follows: The server unit obtains the metadata of the image resource based on the image resource identifier, including the image version, the range of compatible machine models, the image content list, and necessary verification summaries, and performs corresponding verification with the storage structure information. The corresponding verification includes at least: whether the machine model identifier of the target AI machine falls within the range of the template information; whether the target storage area indicated by the mapping rules can be located in the partition structure information; whether the capacity and type of the target storage area meet the write conditions required by the mapping rules; and whether, in a dual-domain architecture, the target storage areas of the real-time domain and the non-real-time domain exist simultaneously without conflict. If any verification item is not met, verification information indicating verification failure is generated, and the failure reason identifier is recorded in the verification information so that the client can display it and prevent it from entering the confirmation stage.

[0055] Furthermore, when the verification result meets the preset verification conditions, this embodiment determines the target storage area identifier based on the mapping rules and generates task summary information in the following way: The server unit summarizes one or more target storage area identifiers determined in the mapping rules to form task summary information. The task summary information may include a set of target storage area identifiers, image content identifiers corresponding to each target storage area identifier, and execution scope description information for client display. After receiving the task summary information, the client can present the execution scope description information to the user, for example, prompting that a write operation will be performed on the real-time domain partition and the area where the non-real-time domain image is located, thereby realizing a unified process of "verification first, confirmation later".

[0056] In one example, after the construction worker selects a specific version of the image resource identifier and enters the network address of the target AI machine on the Windows client, the client initiates a verification request to the server unit on the target AI machine side. The server unit reads the machine model identifier as a certain model in the recovery subsystem and detects the existence of a non-real-time domain virtual machine image file located on an independent partition. This determines the system architecture type to be dual-domain, and the server selects a dual-domain mapping rule from the template information. Subsequently, the server unit verifies that both the real-time domain target partition and the non-real-time domain target partition indicated by the mapping rule exist and their capacity meets the requirements. It also verifies that the compatible machine model range of the image resource identifier includes this machine model identifier. Therefore, it generates verification information indicating successful verification and sends a task summary containing the identifiers of the two target storage areas back to the client. After user confirmation, the flashing execution unit can then perform image writing on the real-time domain and non-real-time domain respectively, avoiding accidental flashing caused by manual selection of the target disk.

[0057] As can be seen from the above embodiments, this embodiment obtains device information and system status information in the recovery subsystem, combines the mapping rules in the template information to complete the consistency verification of the image resource identifier and the target AI machine storage structure, and outputs the target storage area identifier to form task summary information when the verification is successful. This provides a definite write object and a visualized execution scope basis for the subsequent flashing / backup stage, thereby supporting reliable flashing and recovery of different models under a unified remote process.

[0058] In some embodiments, after the client unit receives the verification information and task summary information, it receives a user confirmation instruction and generates an execution instruction to send to the server unit, including:

[0059] The verification results are output based on the verification information, and the execution scope information corresponding to the image resource identifier is output based on the task summary information.

[0060] When the verification result meets the preset execution conditions, the system receives the user's confirmation instruction for the execution scope information and determines the task type identifier that is consistent with the parameters of the flashing or backup task.

[0061] Execution instructions are generated based on the task type identifier and task summary information. The execution instructions include the task type identifier, the image resource identifier, and the target storage area identifier.

[0062] The execution command is sent to the server unit to trigger the flush execution unit or backup execution unit to perform the corresponding operation.

[0063] Specifically, the verification information is structured information used to characterize the matching verification result, which includes at least a verification result identifier and an anomaly reason identifier; wherein the verification result identifier is used to indicate whether the verification passed or failed, and the anomaly reason identifier is used to indicate the main reason category when it fails.

[0064] Task summary information is structured information used to describe the scope of the task to be executed. It at least indicates the target storage area identifier and may further include the image content identifier corresponding to each target storage area identifier and the execution scope description information.

[0065] The execution scope information is a visual description of the scope generated by the client unit based on the task summary information, used to show the user the range of the target storage area to be written to or read from. The preset execution conditions are a set of conditions that the client unit must meet before entering the confirmation phase, which at least include a successful verification result and that the task parameters meet the integrity requirements.

[0066] In this embodiment, the client unit outputs the verification result based on the verification information and the execution scope information based on the task summary information as follows: After the server unit completes the matching verification, it sends the verification information and task summary information back to the client unit. Upon receiving this information, the client unit parses the verification result identifier and the exception reason identifier in the verification information and outputs the verification result on the interface. For example, if the verification result identifier is "failed," the client unit outputs a verification failure message on the interface and, combined with the exception reason identifier, outputs corresponding prompts such as "model mismatch," "target storage area does not exist," or "insufficient capacity," to help on-site personnel quickly locate the problem. If the verification result identifier is "passed," the client unit further parses the target storage area identifier set in the task summary information and generates execution scope information. The execution scope information can be output in a hierarchical description manner, for example, outputting the target storage area ranges corresponding to the real-time domain and the non-real-time domain respectively, or outputting prompts such as "will be written to the underlying system partition" or "will be written to the non-real-time domain mirror area," so that the user understands the scope of impact before confirmation.

[0067] The execution scope information in this embodiment can also be combined with batch operations in factory flashing scenarios. In some implementations, the client unit supports scanning target devices by address range and compiling a list of candidate devices. On-site personnel can select a target AI device from the list for verification and confirmation. In this case, the client unit can simultaneously output the network address and model identifier summary of the target AI device when outputting the execution scope information, thereby reducing the probability of selecting the wrong device in a multi-device environment.

[0068] In some examples, the implementation of receiving a user confirmation instruction and determining the task type identifier when the verification result meets preset execution conditions is as follows: The client unit judges the current status based on the preset execution conditions. The preset execution conditions include at least: the verification result identifier is "passed"; the flash or backup task parameters have been fully entered and are consistent with the image resource identifier that passed verification; and a valid target storage area identifier exists in the task summary information. Only when the above conditions are met does the client unit unlock the confirmation interaction entry, for example, by enabling the "Confirm Flash" or "Confirm Backup" button. Subsequently, the client unit receives a user confirmation instruction regarding the execution scope information. This confirmation instruction can be a click on the confirmation button or a combination of checking the confirmation box and submitting. The client unit determines the task type identifier based on the task type selected by the user or the task type field in the flash or backup task parameters to ensure that the task type identifier in subsequent execution instructions matches the user's intention. Figure 1 To.

[0069] In some examples, the implementation of generating execution instructions based on task type identifiers and task summary information is as follows: After receiving the user's confirmation instruction, the client unit constructs an execution instruction data packet. The execution instruction includes at least a task type identifier, a mirror resource identifier, and a target storage region identifier. The mirror resource identifier indicates the mirror resource version that the server unit should use, and the target storage region identifier indicates the target storage region range that the server unit should operate on. If the task summary information indicates multiple target storage region identifiers, the execution instruction can carry a set of target storage region identifiers to support separate execution of the real-time and non-real-time domains under a dual-domain architecture. To ensure consistency, the client unit can also carry an association identifier consistent with the verification phase in the execution instruction, such as a verification session identifier or a task instance identifier, enabling the server unit to associate the execution instruction with the completed verification results and avoid accidental triggering across tasks.

[0070] In some examples, the execution command is sent to the server unit to trigger the corresponding operation as follows: The client unit sends the execution command to the server unit running in the recovery subsystem through a connection channel established with the target AI machine's network address. After receiving the execution command, the server unit routes the task to either the flash execution unit or the backup execution unit based on the task type identifier. For example, when the task type identifier indicates flashing, the server unit calls the flash execution unit to perform image writing according to the target storage area identifier; when the task type identifier indicates backup, the server unit calls the backup execution unit to read system data according to the target storage area identifier and generate backup data. This routing mechanism unifies flashing and backup under the same remote interface and confirmation process.

[0071] For example, in one scenario, after on-site personnel select the image resource identifier and enter the target AI machine's network address on the Windows client, the client initiates a verification request to the server and receives the returned verification information and task summary information. The client outputs "Verification passed" on the interface and displays the execution scope information: "Write to the real-time domain system partition and the non-real-time domain image area." After the on-site personnel confirm that the scope matches expectations, they click "Confirm Flash Write." Based on this, the client determines the task type identifier as flash write and generates an execution command containing the task type identifier, image resource identifier, and two target storage area identifiers, which is then sent to the server. Upon receiving the command, the server triggers the flash write execution unit to begin image writing and subsequently displays the progress to the client through status log feedback, thus implementing the unified remote flash write process specified in the manual.

[0072] As can be seen from the above embodiments, this embodiment uses "verification result output + execution range output + confirmation after meeting the conditions + generation of execution instructions with key identifiers" as the core means on the client unit side to standardize the pre-execution verification and manual confirmation process of remote flashing and backup, ensuring that the execution instructions correspond one-to-one with the verified image resources and target storage areas, thereby reducing misoperation and improving the controllability of the unified maintenance process.

[0073] In some embodiments, upon receiving an execution instruction that indicates a flush operation, a mirror write operation is performed on the target storage region based on the target storage region identifier, including:

[0074] Parse the execution instructions, obtain the image resource identifier, and determine the image content corresponding to the image resource identifier;

[0075] The write object information of the target storage area is determined based on the target storage area identifier. The write object information includes the target storage device identifier and the partition area identifier and / or file area identifier corresponding to the target storage device identifier.

[0076] Based on the information of the object being written, the image content is written to the target storage area;

[0077] In cases where the target storage area includes both the real-time domain corresponding area and the non-real-time domain corresponding area, real-time domain mirror writing is performed on the real-time domain corresponding area and non-real-time domain mirror writing is performed on the non-real-time domain corresponding area.

[0078] Specifically, the image writing operation refers to the process of writing the image content pointed to by the image resource identifier into the target storage area of ​​the target AI machine. This process may include resource location before writing, determination of the writing object, writing execution, and generation of writing results.

[0079] Write object information is structured information used to guide write execution. It includes at least the target storage device identifier and the partition region identifier and / or file region identifier corresponding to the target storage device identifier, which is used to translate the abstract target storage region identifier into an executable write object.

[0080] The real-time domain corresponds to the storage area that hosts the underlying Linux system of the target AI machine, while the non-real-time domain corresponds to the storage area that hosts the non-real-time virtual machine system of the target AI machine. The non-real-time virtual machine system can exist as a virtual machine image file, or the virtual machine image file can be split into an independent partition to reduce the image size and improve upgrade flexibility. Real-time domain image writing and non-real-time domain image writing refer to the image writing sub-processes executed for the above two types of areas respectively.

[0081] In some examples, when the server unit receives an execution command indicating a write operation, it first parses the command to obtain the image resource identifier. The execution command is generated by the client after successful verification and user confirmation; therefore, the image resource identifier carried by the execution command is consistent with that in the preceding verification stage. The server unit then determines the corresponding image content based on the image resource identifier. In some implementations, the image content can be stored on the client side and retrieved by the server unit on demand, or it can be pre-stored in an internal image repository and downloaded by the server unit over the network. In other implementations, the image content can be directly transmitted from the client to the server unit after establishing a connection. Regardless of the transmission method, the server unit uses the image resource identifier as an index to locate the set of image content to be written.

[0082] Subsequently, the server unit determines the write object information of the target storage area based on the target storage area identifier. This target storage area identifier originates from the task summary information and is generated by mapping rules based on template information during the preceding matching and verification phase, thus avoiding manual selection of the target disk or partition. In this embodiment, the server unit parses the target storage area identifier to determine the target storage device identifier and enumerates storage devices in the recovery subsystem to confirm the reachability of the target storage device identifier; it then further determines the partition area identifier and / or file area identifier corresponding to the target storage device identifier based on the target storage area identifier. For example, when the target storage area is a partition area, the write object information includes the partition number, start and end sectors, or logical volume identifier of the target partition; when the target storage area is a file area, the write object information includes the partition identifier carrying the file, the file path identifier, and the write location identifier of the target file. Through the above write object information, the server unit converts the abstract "target storage area" into an executable write object.

[0083] Furthermore, after determining the write object information, the server unit writes the image content to the target storage area according to the write object information. In some implementations, protective checks can be performed on the write object before writing, such as confirming that the target storage area is not occupied by the current recovery subsystem, confirming that the target storage area does not conflict with the read-only partition where the recovery subsystem is located, and confirming that the target partition is in a writable state, to reduce the risk of accidentally writing to the recovery partition. During the write execution, the server unit can choose partition-level writing or file-level writing according to the type of the write object: for write objects corresponding to partition region identifiers, partition-level writing is used to write the corresponding image segment to the specified partition; for write objects corresponding to file region identifiers, file-level writing is used to write the virtual machine image file or its fragments to the specified file path or file region.

[0084] In some examples, where the target storage area includes both a real-time domain corresponding area and a non-real-time domain corresponding area, this embodiment further splits the write process into two parts: real-time domain image writing and non-real-time domain image writing, and executes them separately. For example, the server unit first locates the real-time domain corresponding area based on the write object information and writes the underlying Linux system image to that area; then, it locates the non-real-time domain corresponding area based on the write object information and writes the non-real-time domain virtual machine system image to that area.

[0085] If the non-real-time domain exists as a virtual machine image file, then writing the non-real-time domain image involves writing the virtual machine image file to the corresponding file area. If the non-real-time domain virtual machine image file is split into an independent partition, then writing the non-real-time domain image involves writing the corresponding image segment to that independent partition. This split execution method can adapt to the specific implementation method in the manual that "MetaOS is divided into real-time domain and non-real-time domain, and the virtual machine image file of the non-real-time domain can be split into an independent partition," thus ensuring compatibility with the differences of various models under a unified flashing process.

[0086] For example, in one scenario, a network flash is performed on a target AI machine with a dual-domain architecture at the factory site. Field personnel access the recovery subsystem via a boot menu. After the recovery subsystem starts, it configures the network address and automatically starts the server unit. Once the Windows client completes verification and receives user confirmation, it sends an execution command containing a set of image resource identifiers and target storage region identifiers to the server unit.

[0087] The server unit parses and executes the instructions, retrieves the image content from the image repository based on the image resource identifier, and then generates write object information based on the target storage area identifier set, including the partition area identifier of the real-time domain system partition and the file area identifier of the area where the non-real-time domain image is located.

[0088] The server unit first performs image writing on the real-time domain system partition, then performs writing on the non-real-time domain image file area. After completing the flashing, it generates flashing result information and subsequently sends it back to the client via status logs so that the client can display the flashing progress and results. This example corresponds to the implementation method in the specification that "all flashing can be performed over the network, and flashing is performed after entering recovery mode, with the client displaying a progress bar."

[0089] As can be seen from the above embodiments, this embodiment uses the execution command as the trigger in the recovery subsystem, uses the target storage area identifier determined in the pre-verification stage to generate write object information, and performs image writing separately according to the real-time domain and non-real-time domain, thereby realizing a unified and controllable remote flashing execution process under different models and storage forms, and providing a repeatable implementation path for rapid on-site recovery.

[0090] In some embodiments, the system further includes a status feedback unit, which is set in the server unit and is used to generate status logs during the process of the brush execution unit performing image writing operations or the backup execution unit generating backup data, and to send the status logs back to the client unit; the client unit is used to output task status information based on the status logs.

[0091] Specifically, this embodiment preferably employs at least one of polling-based and push-based return methods. Polling-based return corresponds to the cyclical request method of the write timing in the specification: after the write or backup task enters the execution state, the client unit sends a log retrieval request to the server unit at preset time intervals, and the server unit returns the latest status log fragment to the client unit. Push-based return, on the other hand, is initiated by the server unit after the status log is generated, used to reduce request overhead and lower status display latency when network conditions permit.

[0092] In some examples, the status log generation by the status backhaul unit is implemented as follows: When the write execution unit starts mirror writing, the status backhaul unit segments the writing process, including at least the stages of "mirror positioning and preparation", "target storage area positioning", "real-time domain writing", "non-real-time domain writing (if it exists)" and "write completion processing", and generates stage change logs at the beginning and end of each stage; within each stage, the status backhaul unit generates a progress log based on the ratio of the amount of data written to the total amount of data, or based on the ratio of the number of written objects to the total number of written objects.

[0093] Similarly, when the backup execution unit generates backup data, the status return unit records at least the following stages: "backup scope determination," "data reading," "backup data generation," "backup data output / transmission," and "backup completion processing," and continuously records progress information during the data reading and generation process. To facilitate subsequent diagnostics, the status return unit can also associate image resource identifiers, target storage area identifiers, or task instance identifiers in the status log, enabling the client to distinguish and aggregate the status logs of different tasks.

[0094] In some examples, to enhance error handling capabilities, the status feedback unit can also generate exception logs. When an exception occurs during the execution of the write or backup execution unit, such as unreachable image resources, inability to write to the target storage area, insufficient capacity, network interruption, or verification failure, the status feedback unit generates an exception log containing an exception identifier and the exception cause category, and may include contextual information related to the exception cause, such as the stage identifier of the exception and the corresponding target storage area identifier. After receiving the exception log, the client unit can output prompt information corresponding to the exception cause category on the interface and switch the task status to failed or pending status so that on-site personnel can take appropriate actions.

[0095] In some examples, the client unit outputs task status information based on the status log as follows: After the task starts, the client unit enters monitoring mode, receives and parses the status log returned by the server, maps stage identifiers to stage prompts on the interface, maps progress identifiers to progress bars or percentage displays, and maps exception identifiers to error message windows or error lists. For batch brushing scenarios in factories, the client unit can also summarize and statistically analyze the status logs of multiple target AI machines, such as separately counting the number of completed tasks, the number of failures, and the categories of failure reasons, so that construction personnel can quickly grasp the overall brushing situation. This is consistent with the implementation method in the design specification of "displaying brushing progress with a progress bar and summarizing and statistically analyzing the number of failures".

[0096] For example, in one scenario, on-site personnel perform network flashing on the target AI machine via a Windows client. After user confirmation, the server-side unit begins image writing, and the status feedback unit generates the first stage log, "Image Location and Preparation," followed by multiple progress logs to characterize the image transmission and preparation progress. After entering the "Real-time Domain Writing" stage, the status feedback unit periodically generates progress logs based on the amount of data written. If the target AI machine has a dual-domain architecture, it enters the "Non-Real-time Domain Writing" stage after the real-time domain writing is completed and continues to generate progress logs.

[0097] The client unit continuously retrieves status logs via polling and refreshes the interface display. When a brief network fluctuation occurs, the status feedback unit generates an exception log indicating a "network link error," which the client unit uses to prompt on-site personnel to check the network connection. After the log is finally written, the status feedback unit generates a completion log, and the client unit outputs the task completion status. This example demonstrates that the "real-time monitoring, error handling, and reduced troubleshooting difficulty" achieved in the specification stem from the specific status log generation and feedback mechanism.

[0098] As can be seen from the above embodiments, this embodiment sets up a status feedback unit in the server unit to generate and feedback the stage, progress and exception of the writing or backup task in the form of status logs to the client unit, so that the client unit can output continuous task status information, thereby supporting visual monitoring and fault location under the unified remote maintenance process, and reducing the operational difficulty and troubleshooting cost of on-site maintenance.

[0099] In some embodiments, when generating backup data, the backup execution unit establishes baseline information corresponding to historical backup points, and determines backup difference data based on the difference between the baseline information and the currently read system data, and uses the backup difference data as backup data or as part of the backup data.

[0100] Specifically, a historical backup point refers to a traceable status identifier formed when a target AI machine completes a backup at a certain historical time. It can be formed by a backup timestamp, backup version number, task instance identifier, or verification summary, and is used to uniquely identify a backup result.

[0101] Baseline information is a set of reference information used for difference calculations corresponding to historical backup points. It is used to characterize the structure and summary features of the data in the target storage area at the time of the historical backup point.

[0102] Backup difference data is a collection of incremental data that has changed relative to baseline information. It may include new data, changed data, and optional deletion markers.

[0103] In this embodiment, the backup execution unit runs within the recovery subsystem of the target AI machine. The recovery subsystem is an independent read-only partition that can access the main system storage area. When a backup is needed, the client unit sends a backup execution command to the server unit. The server unit then calls the backup execution unit to read system data according to the target storage area identifier and generate backup data. The target storage area identifier can be determined by mapping rules based on template information during the preceding matching and verification stage, thus enabling the backup range and the write range to use the same mapping system, achieving unified processing of the real-time and non-real-time domains.

[0104] In some examples, the baseline information corresponding to historical backup points is established as follows: When a historical backup point already exists for the target AI machine, the backup execution unit first locates the corresponding baseline information based on the historical backup point identifier. The baseline information can be generated and stored synchronously by the backup execution unit when the last backup is completed, or it can be generated and returned to the backup execution unit by the backend storage device during storage. In one optional implementation, the baseline information includes a directory structure summary, file-level fingerprint information, and / or block-level summary information for the target storage area: for non-real-time domain virtual machine image files existing as files, the baseline information can include the file's fragment index and fragment summary; for real-time domain systems existing as partitions, the baseline information can include the index and summary of data blocks within the partition. Using the above baseline information, the backup execution unit can determine differences without retransmitting all data during each backup.

[0105] In some examples, the implementation of determining backup difference data based on the difference between baseline information and currently read system data is as follows: The backup execution unit reads the current system data according to the target storage area identifier, generates current summary information at the same granularity as the baseline information, and then compares the current summary information with the baseline information to obtain a difference set. The difference set includes at least newly added items and changed items, and in optional implementations, it may also include marker information for deleted items.

[0106] For file region backups, the backup execution unit can compare file size, update timestamp, and file digest to locate changed files or file fragments, and extract the changed file fragments as backup difference data. For partition region backups, the backup execution unit can compare partition data block digests to locate changed data blocks, and extract the changed data blocks as backup difference data. To ensure reconstructability during recovery, the backup difference data can also carry difference location information, such as file path identifiers, fragment sequence numbers, and block index numbers, enabling subsequent recovery to accurately apply the difference data to the corresponding locations.

[0107] In some examples, backup difference data is implemented as backup data or as part of backup data in the following way: When the difference set is not empty, the backup execution unit encapsulates the backup difference data into the backup data of this backup task, and may attach metadata of this backup, such as the associated historical backup point identifier, the current backup point identifier, difference location information, and verification summary.

[0108] For example, in some implementations, to accommodate scenarios involving initial backups or missing baselines, the backup execution unit switches to full backup mode when it detects the absence of historical backup points or unavailable baseline information, outputting the read system data as backup data. In scenarios where baseline information exists and the difference set is small, it outputs incremental backup data containing only the backup difference data. This allows the system to adaptively select the backup method under different field conditions, meeting the "scheduled backup + incremental backup" functional requirements in the manual.

[0109] In one example, a target AI machine employs a dual-domain architecture: the real-time domain is an underlying Linux system, while the non-real-time domain is a Windows virtual machine system with its virtual machine image files stored separately on independent partitions. During the initial backup, the backup execution unit reads data from both the real-time domain partition and the non-real-time domain independent partition according to the target storage area identifier, generating full backup data and establishing historical backup points. Simultaneously, corresponding baseline information is generated and stored in the database along with the backup data. For the following week, the factory triggers daily backups according to a scheduled task.

[0110] During the second backup, the backup execution unit reads the current system data, generates current summary information, and compares it with the baseline information. It finds that only a small number of blocks within the non-real-time domain independent partition have changed. Therefore, it extracts the changed blocks to form backup difference data, which is then output as the backup data for this backup. The client or backend storage device receives this data and completes the storage. In this way, daily backups only transmit a small amount of difference data, significantly reducing network traffic and storage usage, while still being able to reconstruct the target state of the target storage area based on historical backup points and difference data when recovery is needed.

[0111] As can be seen from the above embodiments, this embodiment establishes baseline information corresponding to historical backup points and calculates backup difference data based on the difference between the baseline information and the current system data, realizing a specific implementation method for incremental backup. This enables remote backup to maintain high availability in factory environments with large image sizes and limited network conditions, and, in conjunction with the independent operating environment of the recovery subsystem, provides basic data support for rapid on-site recovery.

[0112] In some embodiments, the system further includes a backup storage unit for encrypting backup data and storing it in a back-end storage device; and a backup execution unit for performing a recovery operation based on the backup data stored in the back-end storage device to restore the target storage area to a state corresponding to the backup data.

[0113] Specifically, the backup generation and recovery process in this embodiment also follows a unified remote maintenance link: the construction personnel initiate a backup or recovery task on a Windows client, and the client sends the corresponding execution instructions to the server unit running in the recovery subsystem via the network. The server unit calls the backup execution unit to read system data according to the target storage area identifier to generate backup data; or in the recovery task, it calls the backup execution unit to read backup data from the backend storage device and write it back to the target storage area. Since the recovery subsystem is an independent read-only partition and can access the main system partition directory, even if the main system is damaged, the recovery subsystem can still be entered to perform recovery operations, thereby avoiding the need for factory repair.

[0114] In some examples, the backup storage unit encrypts backup data and stores it on the backend storage device as follows: After the backup execution unit generates backup data, it delivers the backup data to the backup storage unit. The backup storage unit first encrypts the backup data. In one optional implementation, the encryption process includes symmetric encryption of the backup data content and generating a verification digest for the metadata used to identify the backup point to prevent tampering during transmission or storage. In another optional implementation, the encryption process may further include segmented encryption of the backup data, so that the backend storage device stores the backup data as segmented objects, facilitating subsequent on-demand retrieval.

[0115] After encryption, the backup storage unit writes the encrypted backup data to the backend storage device and establishes an associated index with the backup point identifier to support subsequent retrieval and reading by backup point identifier. To ensure the traceability of backup data, the backup storage unit can also associate the backup data with metadata such as image resource identifier, machine type identifier, target storage area identifier, and backup timestamp, thereby enabling verification of the compatibility between the backup data and the target AI machine during recovery.

[0116] In some examples, the backup execution unit performs recovery operations based on backup data stored in the backend storage device as follows: When a client initiates a recovery task, the server unit receives the execution instruction and routes the task type to the recovery process of the backup execution unit. The backup execution unit retrieves the corresponding backup data from the backend storage device based on the backup point identifier or associated retrieval information carried in the execution instruction, and performs decryption processing to restore the backup data content.

[0117] Subsequently, the backup execution unit applies the backup data to the target storage area indicated by the target storage area identifier, based on the difference location information or full data structure carried in the backup data. For full backup data, the recovery operation can be manifested as overwriting the target partition or target file area; for incremental backup data, the recovery operation can be manifested as replaying the backup difference data according to the difference location information on the baseline of the historical backup point, thereby reconstructing the state of the target storage area at the backup point. After the recovery is completed, the backup execution unit can generate recovery result information and can send back the recovery stage and progress log through the status feedback unit so that the client can output the recovery progress and results.

[0118] In some examples, a target AI machine is backed up daily on a scheduled basis at the factory site. After the backup execution unit generates backup difference data, it is handed over to the backup storage unit for processing. The backup storage unit encrypts the backup difference data and stores it on the back-end storage device of the factory intranet, and records the association index between the backup point identifier and the target AI machine model identifier and the target storage area identifier. On a certain day, the main system of the target AI machine fails and cannot start. On-site personnel enter the recovery subsystem through the boot menu, and the Windows client initiates a recovery task and selects the most recent backup point identifier.

[0119] The server-side unit calls the backup execution unit to read and decrypt the corresponding backup data from the backend storage device. If the backup is an incremental backup, the backup execution unit replays the difference data based on the corresponding historical baseline, restoring the real-time domain partition and the non-real-time domain mirror area to the backup point state. After the recovery is complete, the device can restart and enter the main system, completing a rapid on-site recovery without needing to return to the factory for repair. This example corresponds to the implementation method described in the manual that "the system can be quickly restored from the backup after damage, shortening the recovery time."

[0120] As can be seen from the above embodiments, this embodiment encrypts the backup data through the backup storage unit and stores it to the back-end storage device, thereby achieving secure storage and traceable management of the backup data; and performs recovery operations based on the backup data of the back-end storage device in the recovery subsystem environment through the backup execution unit, so that the target storage area can be restored to the state corresponding to the backup point, thereby providing a usable remote backup and fast recovery implementation path in a factory controlled environment.

[0121] The following section details the design process of the AI ​​remote flashing and backup system in this application, using a real-world project system design scenario as an example.

[0122] I. Project Overview

[0123] (I) Main functions of the project: The following are the main functions of the "Intelligent Machine Remote Flashing and Backup System":

[0124] 1) Remote flashing function

[0125] 1. Unified flashing method: Regardless of the previous upgrade method (USB or PXE) used by the AI ​​machine, this project achieves a unified remote flashing process, eliminating the need to rely on different hardware tools and greatly reducing management complexity.

[0126] 2. Intelligent Adaptation: It can automatically detect the model and current system status of the AI ​​device and provide the most suitable flashing solution for different AI devices to ensure the accuracy and stability of the flashing process.

[0127] 3. Real-time monitoring: During the flashing process, the client can monitor the flashing progress and status in real time, so that technicians can understand the situation in a timely manner and deal with any problems that may arise.

[0128] 4. Error Handling: If an error occurs during the flashing process, the system will automatically diagnose it and provide a detailed error report and solution, reducing the time and difficulty of manually troubleshooting errors.

[0129] 2) Remote backup function

[0130] 1. Scheduled Backup: Scheduled tasks can be set to automatically back up the AI ​​machine's system and important data to prevent data loss.

[0131] 2. Incremental backup: Only backs up data that has changed since the last backup, improving backup efficiency and reducing storage space usage.

[0132] 3. Secure storage: Backup data is stored on a secure backend server, and encryption technology is used to ensure the security and confidentiality of the data.

[0133] 4. Rapid recovery: When the AI ​​system malfunctions or data is lost, the system and data can be quickly restored from backup through the client, without having to return to the factory for repair, greatly shortening the recovery time.

[0134] (II) Project Platform Description:

[0135] Operating system: All models of AI-powered machines (Ubuntu xx);

[0136] Network architecture: TCP / IP protocol;

[0137] Development tools or technology system: Development languages: HTML, JS, TS, Rust, Python;

[0138] Development tool: Visual Studio Code.

[0139] (III) Software operating environment description:

[0140] Backend GRPC service operation: In the AI ​​machine environment, it exclusively occupies the partition and exists as a flashing subsystem.

[0141] Client software: Windows series, Win10 and above are available.

[0142] II. Software Development

[0143] (I) Software Requirements Analysis:

[0144] The current system suffers from pain points such as inconsistent upgrade schemes, inability to perform on-site recovery after damage, and complex operations. To address these issues, a recovery subsystem solution is proposed, including adding an independent read-only partition subsystem to Ubuntu, adding GRUB menu options, and supporting functions such as factory reset, system updates, and rollback. To address the issue of different system architectures across various machine models, the system integration and partition recovery schemes are unified, with separate recovery schemes defined for real-time and non-real-time domains. The methods for factory flashing and network flashing, as well as the factory reset process, are also clearly defined.

[0145] (II) Software Structure Description:

[0146] 1) Overall Architecture

[0147] The software system mainly consists of MetaOS and the recover subsystem. MetaOS is divided into a real-time domain (underlying Linux) and a non-real-time domain (Windows virtual machine system). The recover subsystem exists as an independent read-only partition in the Ubuntu system.

[0148] 2) MetaOS

[0149] Real-time domain: 1. System images can be backed up to disk; 2. The underlying Linux system provides the basic operating environment for the entire software system.

[0150] Non-real-time domain: Windows virtual machine system. The qcow2 file of the non-real-time domain system can be split into an independent partition (if it exists). Users can choose to back up the non-real-time domain system.

[0151] 3) Recovery subsystem

[0152] Partition structure: As an independent read-only partition, it ensures system stability and security.

[0153] Functional modules: Factory reset: restores the system to its initial state; System update and rollback: supports system upgrades and rollbacks for both real-time and non-real-time domains of MetaOS; GRUB menu integration: adds GRUB menu options, allowing users to select to enter the recovery subsystem.

[0154] File Access: You can access the main system file directory of MetaOS and perform system updates and factory resets by accessing the directory according to specific functions.

[0155] III. Flashing Method

[0156] All flashing can be done via the network. After entering recovery mode through GRUB, the recovery subsystem defaults to a static IP (192.168.1.200). When flashing a project, enter recovery mode via PXE. This IP can be configured through PXE, such as using DHCP or a fixed network segment IP.

[0157] (I) Basic Design

[0158] Installation instructions:

[0159] 1. A read-only system needs to be installed in a separate partition on the AI ​​machine;

[0160] 2. The system flashing client needs to be installed on the construction workers' computers.

[0161] 3. Ensure smooth network connectivity during direct connection.

[0162] (II) Software Module Design

[0163] Overall process:

[0164] 1. Install the recovery client using the MSI installer on the user's Windows system.

[0165] 2. Start the recovery subsystem in the AI ​​machine; it will usually start automatically.

[0166] 3. Windows clients can use existing images to flash the target machine or remotely back up the target machine's system to their local machine.

[0167] The current system faces numerous pain points, such as inconsistent upgrade solutions (some using USB, others PXE); the inability to recover systems after damage, requiring factory return and lacking on-site recovery capabilities; complex and error-prone existing upgrade methods; and varying architectures across different machine models, resulting in large system images, slow downloads, and lengthy and inconsistent backup and restore times. To address these issues, this application designs a series of modules and strategies. The Recover subsystem is a key component, adding an independent read-only partition on Ubuntu to ensure system stability.

[0168] By adding a GRUB menu option, users can easily access this subsystem. The subsystem can access the main system's file directory, enabling important functions such as factory reset, system updates, and rollback. Simultaneously, the virtual machine is used as an independent partition for system upgrades, improving upgrade flexibility and reliability. For various machine architectures, MetaOS is divided into a real-time domain (underlying Linux) and a non-real-time domain (Windows virtual machine system). The real-time domain system image is backed up to disk, while the non-real-time domain system can be backed up by the user. Furthermore, the qcow2 file of the non-real-time domain system is split into an independent partition (if it exists), achieving a unified solution for integrated distribution and partition recovery across all systems.

[0169] In terms of flashing methods, factory flashing and network flashing are combined. During factory flashing, the machine boots the recovery system via PXE and obtains an IP address via DHCP. The operator sets the IP range, scans for devices, and performs the flashing through a Windows interface program. A progress bar displays the flashing progress and summarizes the number of failures, making it easy for the operator to understand the flashing status.

[0170] During network flashing, after entering recovery mode via GRUB, the recovery subsystem's default static IP is 192.168.1.200. For engineering flashing, the IP can be configured via PXE, such as using DHCP or a fixed network segment IP. Furthermore, when restoring factory settings, MetaOS and the recovery system are distributed independently. The recovery partition is integrated when creating the MetaOS image, ensuring convenient and stable system recovery. These module designs and strategies are expected to effectively solve the current system's problems and improve system availability and user experience.

[0171] The above embodiments have described in detail the specific modules and functions of the AI ​​remote flashing and backup system of this application. The implementation process of the AI ​​remote flashing and backup method of this application will be described in detail below with reference to specific embodiments. Figure 2 This is a flowchart illustrating the remote flashing and backup method for artificial intelligence machines provided in this application embodiment, as shown below. Figure 2 As shown, the method may specifically include the following steps:

[0172] S201, Receive user input of flash or backup task parameters, including image resource identifier, target AI machine network address and authentication information;

[0173] S202, obtain the template information corresponding to the mirror resource based on the mirror resource identifier, and initiate a verification request to the server unit on the target AI machine side based on the network address of the target AI machine;

[0174] S203, in the server unit, obtain the device information and system status information of the target AI machine, and match and verify the device information, system status information and image resource identifier according to the template information, and generate verification information and task summary information. The task summary information is used to indicate the target storage area identifier.

[0175] S204 After receiving the verification information and task summary information, receive the user confirmation instruction, generate an execution instruction and send it to the server unit. The execution instruction is used to indicate the task type and the target storage area identifier.

[0176] S205, in the server unit, performs remote flashing or remote backup according to the execution instruction. The remote flashing includes performing a mirror write operation on the target storage area according to the target storage area identifier, and the remote backup includes reading system data from the target storage area according to the target storage area identifier and generating backup data.

[0177] It should be understood that the sequence number of each step in the above method embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0178] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although the technical solutions of this application have been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A remote flashing and backup system for artificial intelligence machines, characterized in that, include: The client unit is used to receive flashing or backup task parameters input by the user. The flashing or backup task parameters include the image resource identifier, the network address of the target AI device, and authentication information. Based on the image resource identifier, obtain the template information corresponding to the image resource, and initiate a verification request to the server unit on the target AI machine side; A recovery subsystem is set in the target AI machine. The recovery subsystem is an independent read-only partition used to provide a running environment for remote flashing and backup. The server unit, running in the recovery subsystem, is used to obtain the device information and system status information of the target AI machine after receiving the verification request, and to match and verify the device information, the system status information and the image resource identifier according to the template information, and generate verification information and task summary information. The task summary information is used to indicate the target storage area identifier. The confirmation control unit is used to receive the user confirmation instruction after the client unit receives the verification information and task summary information, and to generate an execution instruction to send to the server unit. The write execution unit is used to perform a mirror write operation on the target storage area according to the target storage area identifier when the execution instruction is received and the execution instruction indicates write. A backup execution unit is configured to, upon receiving the execution instruction and the execution instruction indicating backup, read system data from the target storage area according to the target storage area identifier and generate backup data; The step of matching and verifying the device information, system status information, and image resource identifier based on the template information to generate verification information and task summary information includes: Based on the device information parsing, the model identifier and storage structure information of the target AI machine are obtained, and the storage structure information includes storage device identifier and partition structure information; Based on the system status information, determine the system architecture type corresponding to the target AI machine, and select a mapping rule from the template information that matches the machine type identifier and the system architecture type; The consistency between the image resource identifier and the storage structure information is verified according to the mapping rules to obtain the verification result, and verification information is generated to characterize the verification result. When the verification result meets the preset verification conditions, the target storage area identifier is determined based on the mapping rules, and task summary information containing the target storage area identifier is generated.

2. The system according to claim 1, characterized in that, The template information includes flashing configuration parameters that match the target AI device model identifier and storage area mapping rules. The storage area mapping rules are used to map the image content corresponding to the image resource identifier to the target storage area identifier.

3. The system according to claim 1, characterized in that, The recovery subsystem is associated with a boot control component, which is used to provide boot options for entering the recovery subsystem during the boot phase of the target AI machine, so that the server unit can start in the recovery subsystem and provide remote flashing and backup service interfaces to the outside world.

4. The system according to claim 1, characterized in that, The step of receiving a user confirmation instruction after the client unit receives the verification information and task summary information, and generating an execution instruction to send to the server unit, includes: Based on the verification information, the verification result is output, and based on the task summary information, the execution scope information corresponding to the image resource identifier is output; When the verification result meets the preset execution conditions, the system receives a confirmation instruction from the user regarding the execution range information and determines a task type identifier that matches the parameters of the flashing or backup task. The execution instruction is generated based on the task type identifier and the task summary information. The execution instruction includes the task type identifier, the image resource identifier, and the target storage area identifier. The execution instruction is sent to the server unit to trigger the write execution unit or the backup execution unit to perform the corresponding operation.

5. The system according to claim 1, characterized in that, Upon receiving the execution instruction and the execution instruction indicating a write operation, the step of performing a mirror write operation on the target storage region based on the target storage region identifier includes: Parse the execution instruction to obtain the image resource identifier, and determine the image content corresponding to the image resource identifier; Based on the target storage region identifier, write object information of the target storage region is determined. The write object information includes the target storage device identifier and the partition region identifier and / or file region identifier corresponding to the target storage device identifier. Based on the written object information, the image content is written to the target storage area; In the case where the target storage area includes a real-time domain corresponding area and a non-real-time domain corresponding area, real-time domain mirror writing is performed on the real-time domain corresponding area and non-real-time domain mirror writing is performed on the non-real-time domain corresponding area.

6. The system according to claim 1, characterized in that, The system also includes a status feedback unit, which is located in the server unit and is used to generate status logs during the process of the brush execution unit performing image writing operations or the backup execution unit generating backup data, and to send the status logs back to the client unit; the client unit is used to output task status information based on the status logs.

7. The system according to claim 1, characterized in that, When generating backup data, the backup execution unit establishes baseline information corresponding to historical backup points, and determines backup difference data based on the difference between the baseline information and the currently read system data, and uses the backup difference data as the backup data or as part of the backup data.

8. The system according to claim 1, characterized in that, The system also includes a backup storage unit, which is used to encrypt the backup data and store it in a back-end storage device; the backup execution unit is used to perform a recovery operation based on the backup data stored in the back-end storage device to restore the target storage area to the state corresponding to the backup data.

9. A method for remote flashing and backup of an AI computer based on the system described in any one of claims 1 to 8, characterized in that, include: Receive user input for flashing or backup task parameters, which include image resource identifier, target AI machine network address, and authentication information; Based on the image resource identifier, obtain the template information corresponding to the image resource, and initiate a verification request to the server unit on the target AI machine side based on the network address of the target AI machine; In the server unit, the device information and system status information of the target AI machine are obtained, and the device information, system status information and image resource identifier are matched and verified according to the template information to generate verification information and task summary information. The task summary information is used to indicate the target storage area identifier. After receiving the verification information and task summary information, the system receives a user confirmation instruction, generates an execution instruction, and sends it to the server unit. The execution instruction is used to indicate the task type and the target storage area identifier. In the server unit, remote flashing or remote backup is performed according to the execution instruction. The remote flashing includes performing a mirror write operation on the target storage area according to the target storage area identifier, and the remote backup includes reading system data from the target storage area according to the target storage area identifier and generating backup data.