A method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage

CN122569847APending Publication Date: 2026-08-14HUNAN TIANSHUO INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

主控制器与备用控制器缺少实时协同机制,主控制器故障时无法快速完成权限切换,掉电保护流程容易中断

Benefits of technology

本发明搭建多重冗余硬件架构,配置独立供电模块、多通道缓存控制器与多分区非易失性存储介质,搭配硬件故障隔离机制,提升系统硬件层面的容错能力,弱化单点故障对整体数据安全的影响。复合供电监测方式结合多级电压阈值与电流变化率检测,融入温度补偿与电网波动识别逻辑,提升掉电检测的准确性,减少误触发与漏触发的情况。多通道缓存冗余设计实现数据多副本存储,缓存内部采用分区存储模式,提升缓存数据的安全性与管理规范性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569847A_ABST
    Figure CN122569847A_ABST
Patent Text Reader

Abstract

This invention relates to the field of electronic digital data processing technology, specifically to a power-loss data protection method for industrial-grade multi-redundant fault-tolerant solid-state storage. The method includes constructing a multi-redundant hardware architecture, real-time monitoring of power supply status and classifying response levels, activating multi-channel cached redundant storage data, performing transactional processing on write operations, dynamically allocating remaining power energy, writing cached data to multi-redundant storage partitions, completing data verification and automatic recovery, recording fault logs, and ensuring data security and continuous stability of the protection process during power outages through collaborative fault tolerance via primary and backup controllers. This invention improves power-loss detection accuracy and hardware fault tolerance capabilities, optimizes energy utilization efficiency and data storage reliability, simplifies the data recovery process, improves the fault tracing mechanism, and ensures data integrity and system operational stability of industrial solid-state storage during power outages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, and in particular to a method for protecting data from power loss in industrial-grade multi-redundant fault-tolerant solid-state storage. Background Technology

[0002] Industrial-grade solid-state storage (SSD) is a core data storage medium for demanding industrial scenarios such as industrial control, intelligent manufacturing, and rail transportation, and its operation places stringent requirements on data integrity and storage continuity. Current power-loss data protection solutions for industrial SSDs mostly employ a basic architecture of a single backup power supply unit paired with a single cache channel. In the event of a power outage, only a small portion of the cached data can be written to the non-volatile storage medium, failing to effectively migrate the entire amount of data to be written. This type of solution lacks hardware-level redundancy; a failure in any module—cache, storage, or power supply—will directly lead to data loss, and its hardware fault tolerance is ill-suited to the complex operating environment of industrial sites.

[0003] Existing power outage protection systems rely on a limited set of methods for monitoring power supply status, depending solely on a fixed voltage threshold. They lack current change rate detection and ambient temperature compensation logic, making them susceptible to interference from instantaneous power grid fluctuations and industrial ambient temperature drift. This can lead to false triggering or missed detection of actual power outages. Furthermore, the power outage response process lacks tiered classification; regardless of the speed of the power outage or the amount of remaining energy, a uniform protection procedure is executed. Energy allocation uses a fixed ratio and cannot be dynamically adjusted based on data importance or the amount of data to be processed. The limited energy of backup power supplies cannot prioritize the writing of core business data.

[0004] Existing data verification and recovery mechanisms rely on a single verification method, which can only determine whether data is corrupted but struggles to accurately pinpoint the error location, resulting in low data recovery efficiency. Write operations lack transactional management mechanisms, making partial writes during power outages prone to occur, leading to data inconsistencies and file system structure anomalies. The lack of real-time coordination between the primary and backup controllers means that access control cannot be quickly switched in the event of a primary controller failure, and the power-down protection process is easily interrupted. Fault logs are incomplete, data recovery does not support resuming interrupted data transfers, and a full data scan is required after system power-on, making the overall recovery process cumbersome and failing to meet the high reliability and fault tolerance requirements of solid-state storage in industrial scenarios. Summary of the Invention

[0005] This invention proposes an industrial-grade multi-redundant fault-tolerant solid-state storage power-loss data protection method to solve the problems mentioned in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: an industrial-grade multi-redundancy fault-tolerant solid-state storage power-loss data protection method, comprising the following steps: Construct a multi-redundant hardware architecture for industrial-grade solid-state storage, integrating multiple independent power supply modules, multi-channel cache controllers, and multi-partition non-volatile storage media. Establish a real-time status interaction mechanism and hardware-level fault isolation mechanism between modules. Each module has independent status monitoring and fault reporting capabilities, limiting the impact of single-point failures on system operation. The system monitors the power supply status in real time and uses a composite detection method that combines multi-level voltage threshold detection and current change rate detection to identify different types of power outage events. Different response levels are defined based on the power outage occurrence speed, remaining energy, and system operating status, and the corresponding data protection process is automatically triggered. A multi-channel data caching redundancy mechanism is activated, which simultaneously writes the data to be written to multiple independent volatile cache channels, and synchronously records the data writing progress, data length and verification information of each channel to achieve real-time redundant storage of multiple copies of cached data. Write operations are processed in a transactional manner, each complete write operation is broken down into indivisible atomic transaction units, the complete process log of transaction execution is recorded, the atomicity and consistency of write operations are maintained, and the structural state of the file system and stored data is stabilized. The system dynamically allocates remaining power energy based on the current amount of data to be written in the cache, the number of write operations to be completed, the write speed of the non-volatile storage medium, and the power consumption of each module. It calculates the energy allocation ratio of each functional module and matches the energy supply requirements for key data writing and storage. The valid data in the cache is written to the pre-divided multi-redundant non-volatile storage partitions. Each data block is written to multiple independent physical storage partitions at the same time, and the data storage location, write time and verification information of each partition are recorded synchronously. After all data writing operations are completed, a full data integrity check is performed. The corresponding data content of each redundant partition is compared with the pre-stored check information to identify partitions with damaged, lost, or inconsistent data. The complete data of the normal partition is then called to complete the automatic recovery of the damaged data. Record the complete process information of the power failure event and the system operating status, generate standardized fault logs and store them in an independent non-volatile log partition to support subsequent fault location, cause analysis and system debugging. A multi-controller collaborative fault-tolerant mechanism is established, in which the main controller and the backup controller synchronize the system status and cached data in real time. When the main controller fails, the backup controller automatically takes over the control of the system, calls the redundantly stored logs and data to complete the system status recovery, and maintains the continuity and reliability of the power failure protection process.

[0007] Furthermore, it also includes a dynamic power allocation optimization step, which optimizes the allocation of remaining energy based on module priority and the amount of data to be processed. The calculation formula is as follows: ; in The energy allocated to the i-th functional module, in joules. Let be the priority coefficient of the i-th functional module, which is dimensionless. The system detects the total remaining energy during a power outage, measured in joules. The effective data volume to be processed by the i-th functional module is expressed in bytes. The total effective data volume to be processed by all modules of the system is expressed in bytes. The priority coefficient is dynamically adjusted according to the importance of the data business. The priority coefficient of the module corresponding to industrial real-time control data is higher than that of ordinary log and configuration data modules. During the power failure warning stage, the energy demand of each module is pre-calculated. When non-essential modules are shut down, the remaining energy is recovered and redistributed to the core data writing module to improve energy utilization efficiency and match the writing needs of key data under limited energy.

[0008] Furthermore, it also includes data integrity verification and single-byte error location steps, using a weighted verification algorithm to improve the accuracy of error detection and location precision. The calculation formula is as follows: ; in This is the final checksum of the data block, dimensionless. Let j be the decimal value of the j-th byte in the data block, dimensionless. Let be the weight coefficient corresponding to the j-th byte in the data block, which is dimensionless. The total number of bytes contained in the data block, dimensionless. The preset check modulus is dimensionless, and the weighting coefficients are adaptively adjusted according to the error distribution characteristics of the storage medium. The bytes corresponding to the high-order storage units of the flash memory are assigned higher weights. The check value of each data block is stored in two independent check partitions at the same time, maintaining the storage stability of the check value itself. Single-byte errors can be directly calculated based on the weighting coefficients and automatically corrected without relying on other redundant partitions, providing a reliable basis for subsequent data recovery.

[0009] Furthermore, the multi-level voltage threshold detection system sets three continuous voltage thresholds, corresponding to normal power supply status, early warning power outage status, and emergency power outage status, respectively. The current change rate detection identifies instantaneous power outage events by sampling the supply current every microsecond and calculating the differential change value. A temperature compensation mechanism is introduced to correct the voltage detection threshold under different industrial ambient temperatures, reducing detection errors caused by temperature drift. A power grid fluctuation identification algorithm is used to distinguish between normal voltage fluctuations and real power outage events, filtering voltage spike interference with a duration of less than ten microseconds. When the supply voltage is lower than the early warning threshold or the current change rate exceeds the preset limit, the pre-protection process is triggered in advance, and non-essential peripheral interfaces and computing modules of the system are shut down in stages. In the early warning state, low-priority debugging interfaces and status indicator lights are shut down first, and in the emergency state, all non-core computing units are shut down, gradually reducing the overall power consumption of the system and reserving sufficient energy for core data protection operations.

[0010] Furthermore, the multi-channel data cache adopts a three-channel fully independent redundant architecture. Each channel is equipped with an independent power management unit, data bus, and cache chip. The channels are physically isolated from each other. When a hardware failure or data corruption occurs in a single channel, it will not affect the normal data read and write of other channels. During normal operation, a dynamic load balancing algorithm is used to allocate write tasks according to the idle status of each channel, thereby improving the overall write throughput of the cache. After each channel writes data, it immediately performs local cyclic redundancy check. If an error is detected, it automatically triggers data rewriting and prevents erroneous data from entering the subsequent storage process. The cache is internally divided into three independent sub-partitions: real-time control data area, log data area, and configuration data area. Data partitions of different business types are stored without interference. When power is lost, data is processed in order according to partition priority, with priority given to completing the write operation of real-time control data.

[0011] Furthermore, the non-volatile storage medium is divided into three independent physical partitions: a data storage area, a transaction log area, and a verification information area. The data storage area adopts a three-copy redundant storage method, with each logical data block stored on three different physical flash memory chips. A dynamic wear leveling mechanism is introduced to dynamically adjust the data storage location based on the number of erases and writes of each flash memory chip, balancing the wear degree of each chip and extending the service life of the medium. Bad blocks in flash memory are detected in real time and automatically marked, and the data is migrated to adjacent healthy physical blocks to prevent the risk of data loss caused by bad blocks. Each data block retains backups of the three most recent valid versions. When the latest version of the data is corrupted, it is automatically rolled back to the previous complete version. The transaction log area records all transaction operations using a circular overwrite writing method. The verification information area stores the verification value, physical address mapping table, and version information of each data block, realizing physical isolation storage of data, logs, and verification information.

[0012] Furthermore, the write operation transaction processing adopts an improved two-phase commit protocol. In the first phase, the transaction data is written to a multi-channel cache and the transaction start log is recorded. In the second phase, the cached data is written in parallel to multiple non-volatile storage partitions. After all redundant partitions have completed data writing and verification, the transaction is committed. It supports the function of nested splitting of large transactions, splitting write operations exceeding the preset size into multiple independent sub-transactions and executing them sequentially. Each sub-transaction records its own log and commits, improving the efficiency of writing large files. A transaction pre-commit mechanism is added. After the data is written to the cache, a pre-commit is completed and a pre-commit log is recorded. In the event of a power failure, the pre-committed data can be directly written to non-volatile storage, shortening the data writing time during emergency power failures. If a power failure or hardware failure occurs at any stage, the system will perform an incremental rollback according to the transaction log after power-on, restoring only the modified data blocks without resetting the entire file system.

[0013] Furthermore, the data integrity verification adopts a two-layer verification mechanism that combines block-by-block content comparison with weighted check value verification. First, the content of the corresponding data blocks in each redundant partition is compared byte by byte to identify data blocks with inconsistent content. Then, the weighted check value of the normal data block is calculated and compared with the pre-check value stored in the verification information area to confirm the correctness of the data. Parallel processing technology is used to perform verification and recovery operations on multiple data blocks simultaneously to improve the efficiency of large-scale data verification. A data corruption degree assessment model is established, and recovery priorities are divided according to the number, location and business importance of the damaged data blocks. The recovery operation of the core industrial control data is completed first. The recovery process supports breakpoint resume. If a power failure occurs again during the recovery process, the recovery process will continue from the last interrupted position after the system is powered on, without having to rescan all data blocks.

[0014] Furthermore, the fault log records include the precise time of power failure, power supply voltage change curve, system remaining energy value, cached data to be written, number of completed write operations, detailed information on incomplete write operations, data recovery process records, and final system status. A hierarchical storage strategy is adopted to store emergency fault logs, ordinary fault logs, and operation logs in different log sub-partitions. Emergency fault logs are permanently retained, while ordinary logs automatically overwrite content that has exceeded the retention period. Log data is encrypted using an irreversible encryption algorithm to maintain the originality and integrity of the log data. After system recovery, the fault logs are automatically uploaded to the cloud-based industrial equipment management platform to support remote fault diagnosis and analysis. After each power-on, the system automatically reads the three most recent fault logs and performs automatic fault analysis to generate a standardized report containing the cause of the fault, the scope of impact, and maintenance recommendations.

[0015] Furthermore, the multi-controller collaborative fault tolerance adopts a primary-backup hot standby real-time synchronization architecture. The primary controller and the backup controller synchronize the system operating status, cached data, and transaction logs in real time through an independent high-speed communication bus. A three-channel heartbeat detection mechanism is used to monitor the status of the primary controller, reducing the probability of false switching caused by single-channel heartbeat interruption. When the primary controller is running normally, the backup controller is in a low-power standby state, only performing status synchronization and heartbeat monitoring functions. When the primary controller experiences hardware failure, communication interruption, or program crash, the backup controller automatically takes over system control within a preset switching time. During the switching process, all cached data and incomplete transaction states are retained without re-initializing the system. After taking over, the system status verification and data integrity check are automatically performed. After confirming that there are no errors, the unfinished power-down protection process is continued. At the same time, a primary-backup switching log is generated to record the entire switching process.

[0016] Compared with existing technologies, the beneficial effects of this invention are: This invention establishes a multi-redundant hardware architecture, configuring independent power supply modules, multi-channel cache controllers, and multi-partition non-volatile storage media, coupled with a hardware fault isolation mechanism, to enhance the fault tolerance of the system hardware and mitigate the impact of single-point failures on overall data security. The composite power supply monitoring method combines multi-level voltage thresholds and current change rate detection, incorporating temperature compensation and grid fluctuation identification logic to improve the accuracy of power failure detection and reduce false triggers and missed triggers. The multi-channel cache redundancy design enables multiple copies of data for storage, and the cache itself employs a partitioned storage mode, improving the security and management standardization of cached data.

[0017] This invention maintains the atomicity and consistency of write operations through transactional processing and an improved two-phase commit protocol, reducing the possibility of data inconsistency and file system corruption. Large transaction splitting and transaction pre-commit mechanisms shorten data write time in emergency power outage scenarios. Dynamic energy allocation determines the remaining energy based on module priority and the amount of data to be processed, improving energy utilization efficiency and matching the write requirements of core data. Multi-redundant partitioned storage, combined with dynamic wear leveling and real-time bad block management, optimizes the usage status of the storage medium and extends the stable operating cycle of the storage units.

[0018] This invention employs a dual-layer data verification mechanism combined with weighted verification logic to improve the accuracy of data error detection and location. The breakpoint resume recovery method simplifies the data recovery process and improves the smoothness of data recovery. Complete fault log recording and hierarchical storage modes provide a solid basis for fault tracing and system debugging, supporting subsequent equipment maintenance. A primary / backup controller hot standby architecture combined with multi-channel heartbeat detection improves the accuracy of controller switching, maintains the continuity of the power-down protection process, and enables the solid-state storage system to maintain stable data protection capabilities in harsh industrial environments. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the overall business process of power-loss protection for industrial-grade solid-state storage in this invention. Figure 2 This is the logic diagram of multi-level composite detection and dynamic energy allocation in this invention; Figure 3 This is a flowchart of the atomic transaction processing and multi-copy storage process in this invention; Figure 4 This is a flowchart illustrating the dual-layer verification mechanism and damaged data recovery process in this invention. Figure 5 This is a flowchart of the multi-controller collaborative fault tolerance and fault log recording process in this invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Reference Figures 1 to 5 A method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage includes the following steps: Construct a multi-redundant hardware architecture for industrial-grade solid-state storage, integrating multiple independent power supply modules, multi-channel cache controllers, and multi-partition non-volatile storage media. Establish a real-time status interaction mechanism and hardware-level fault isolation mechanism between modules. Each module has independent status monitoring and fault reporting capabilities, limiting the impact of single-point failures on system operation. The system monitors the power supply status in real time and uses a composite detection method that combines multi-level voltage threshold detection and current change rate detection to identify different types of power outage events. Different response levels are defined based on the power outage occurrence speed, remaining energy, and system operating status, and the corresponding data protection process is automatically triggered. A multi-channel data caching redundancy mechanism is activated, which simultaneously writes the data to be written to multiple independent volatile cache channels, and synchronously records the data writing progress, data length and verification information of each channel to achieve real-time redundant storage of multiple copies of cached data. Write operations are processed in a transactional manner, each complete write operation is broken down into indivisible atomic transaction units, the complete process log of transaction execution is recorded, the atomicity and consistency of write operations are maintained, and the structural state of the file system and stored data is stabilized. The system dynamically allocates remaining power energy based on the current amount of data to be written in the cache, the number of write operations to be completed, the write speed of the non-volatile storage medium, and the power consumption of each module. It calculates the energy allocation ratio of each functional module and matches the energy supply requirements for key data writing and storage. The valid data in the cache is written to the pre-divided multi-redundant non-volatile storage partitions. Each data block is written to multiple independent physical storage partitions at the same time, and the data storage location, write time and verification information of each partition are recorded synchronously. After all data writing operations are completed, a full data integrity check is performed. The corresponding data content of each redundant partition is compared with the pre-stored check information to identify partitions with damaged, lost, or inconsistent data. The complete data of the normal partition is then called to complete the automatic recovery of the damaged data. Record the complete process information of the power failure event and the system operating status, generate standardized fault logs and store them in an independent non-volatile log partition to support subsequent fault location, cause analysis and system debugging. A multi-controller collaborative fault-tolerant mechanism is established, in which the main controller and the backup controller synchronize the system status and cached data in real time. When the main controller fails, the backup controller automatically takes over the control of the system, calls the redundantly stored logs and data to complete the system status recovery, and maintains the continuity and reliability of the power failure protection process.

[0022] This invention also includes a dynamic power allocation optimization step, which optimizes the allocation of remaining energy based on module priority and the amount of data to be processed. The calculation formula is as follows: ; in The energy allocated to the i-th functional module, in joules. Let be the priority coefficient of the i-th functional module, which is dimensionless. The system detects the total remaining energy during a power outage, measured in joules. The effective data volume to be processed by the i-th functional module is expressed in bytes. The total effective data volume to be processed by all modules of the system is expressed in bytes. The priority coefficient is dynamically adjusted according to the importance of the data business. The priority coefficient of the module corresponding to industrial real-time control data is higher than that of ordinary log and configuration data modules. During the power failure warning stage, the energy demand of each module is pre-calculated. When non-essential modules are shut down, the remaining energy is recovered and redistributed to the core data writing module to improve energy utilization efficiency and match the writing needs of key data under limited energy.

[0023] This invention also includes data integrity verification and single-byte error location steps. A weighted verification algorithm is used to improve the accuracy of error detection and location precision. The calculation formula is as follows: ; in This is the final checksum of the data block, dimensionless. Let j be the decimal value of the j-th byte in the data block, dimensionless. Let be the weight coefficient corresponding to the j-th byte in the data block, which is dimensionless. The total number of bytes contained in the data block, dimensionless. The preset check modulus is dimensionless, and the weighting coefficients are adaptively adjusted according to the error distribution characteristics of the storage medium. The bytes corresponding to the high-order storage units of the flash memory are assigned higher weights. The check value of each data block is stored in two independent check partitions at the same time, maintaining the storage stability of the check value itself. Single-byte errors can be directly calculated based on the weighting coefficients and automatically corrected without relying on other redundant partitions, providing a reliable basis for subsequent data recovery.

[0024] In this invention, a multi-level voltage threshold detection system sets three continuous voltage thresholds, corresponding to normal power supply status, warning power outage status, and emergency power outage status, respectively. The current change rate detection identifies instantaneous power outage events by sampling the supply current every microsecond and calculating the differential change value. A temperature compensation mechanism is introduced to correct the voltage detection threshold under different industrial ambient temperatures, reducing detection errors caused by temperature drift. A power grid fluctuation identification algorithm is used to distinguish between normal voltage fluctuations and actual power outage events, filtering voltage spike interference with a duration of less than ten microseconds. When the supply voltage is lower than the warning threshold or the current change rate exceeds the preset limit, a pre-protection process is triggered in advance, and non-essential peripheral interfaces and computing modules of the system are shut down in stages. In the warning state, low-priority debugging interfaces and status indicator lights are shut down first, and in the emergency state, all non-core computing units are shut down, gradually reducing the overall power consumption of the system and reserving sufficient energy for core data protection operations.

[0025] In this invention, the multi-channel data cache adopts a three-channel fully independent redundant architecture. Each channel is equipped with an independent power management unit, data bus, and cache chip. The channels are physically isolated from each other. When a hardware failure or data corruption occurs in a single channel, it will not affect the normal data reading and writing of other channels. During normal operation, a dynamic load balancing algorithm is used to allocate write tasks according to the idle status of each channel, thereby improving the overall write throughput of the cache. After each channel writes data, it immediately performs local cyclic redundancy check. If an error is found, it automatically triggers data rewriting and prevents erroneous data from entering the subsequent storage process. The cache is internally divided into three independent sub-partitions: real-time control data area, log data area, and configuration data area. Data partitions of different business types are stored without interference. When power is lost, data is processed in order according to partition priority, with priority given to completing the write operation of real-time control data.

[0026] In this invention, the non-volatile storage medium is divided into three independent physical partitions: a data storage area, a transaction log area, and a verification information area. The data storage area adopts a three-copy redundant storage method, with each logical data block stored on three different physical flash memory chips. A dynamic wear leveling mechanism is introduced to dynamically adjust the data storage location based on the number of erases and writes of each flash memory chip, balancing the wear degree of each chip and extending the service life of the medium. Bad blocks in flash memory are detected in real time and automatically marked, and the data is migrated to adjacent healthy physical blocks to prevent the risk of data loss caused by bad blocks. Each data block retains backups of the three most recent valid versions. When the latest version of the data is corrupted, it is automatically rolled back to the previous complete version. The transaction log area records all transaction operations using a cyclic overwrite writing method. The verification information area stores the verification value of each data block, the physical address mapping table, and the version information, realizing the physical isolation storage of data, logs, and verification information.

[0027] In this invention, the transactional processing of write operations adopts an improved two-phase commit protocol. In the first phase, the transaction data is written to a multi-channel cache and the transaction start log is recorded. In the second phase, the cached data is written in parallel to multiple non-volatile storage partitions. After all redundant partitions have completed data writing and verification, the transaction is committed. It supports the nested splitting function of large transactions, splitting write operations exceeding the preset size into multiple independent sub-transactions and executing them sequentially. Each sub-transaction records its own log and commits, improving the efficiency of writing large files. A transaction pre-commit mechanism is added. After the data is written to the cache, a pre-commit is completed and a pre-commit log is recorded. In the event of a power failure, the pre-committed data can be directly written to non-volatile storage, shortening the data writing time during emergency power failures. If a power failure or hardware failure occurs at any stage, the system performs an incremental rollback according to the transaction log after power-on, restoring only the modified data blocks without resetting the entire file system.

[0028] In this invention, data integrity verification adopts a two-layer verification mechanism combining block-by-block content comparison and weighted check value verification. First, the contents of the corresponding data blocks in each redundant partition are compared byte by byte to identify data blocks with inconsistent contents. Then, the weighted check value of the normal data block is calculated and compared with the pre-check value stored in the verification information area to confirm the correctness of the data. Parallel processing technology is used to perform verification and recovery operations on multiple data blocks simultaneously, improving the efficiency of large-scale data verification. A data corruption degree assessment model is established, and recovery priorities are divided according to the number, location, and business importance of the damaged data blocks. The recovery operation of core industrial control data is completed first. The recovery process supports breakpoint resumption. If a power failure occurs again during the recovery process, the recovery process will continue from the last interrupted position after the system is powered on, without having to rescan all data blocks.

[0029] In this invention, the fault log records include the precise time of power failure, the power supply voltage change curve, the remaining energy value of the system, the amount of data to be written in the cache, the number of completed write operations, detailed information on incomplete write operations, data recovery process records, and the final system status. A hierarchical storage strategy is adopted to store emergency fault logs, ordinary fault logs, and operation logs in different log sub-partitions. Emergency fault logs are permanently retained, while ordinary logs automatically overwrite content that has exceeded the retention period. The log data is encrypted using an irreversible encryption algorithm to maintain the originality and integrity of the log data. After the system recovers, the fault logs are automatically uploaded to the cloud-based industrial equipment management platform to support remote fault diagnosis and analysis. After each power-on, the system automatically reads the three most recent fault logs and performs automatic fault analysis to generate a standardized report containing the cause of the fault, the scope of impact, and maintenance recommendations.

[0030] In this invention, the multi-controller collaborative fault tolerance adopts a primary-backup hot standby real-time synchronization architecture. The primary controller and the backup controller synchronize the system operating status, cached data, and transaction logs in real time through an independent high-speed communication bus. A three-channel heartbeat detection mechanism is used to monitor the status of the primary controller, reducing the probability of false switching caused by single-channel heartbeat interruption. When the primary controller is running normally, the backup controller is in a low-power standby state, only performing status synchronization and heartbeat monitoring functions. When the primary controller experiences hardware failure, communication interruption, or program crash, the backup controller automatically takes over system control within a preset switching time. During the switching process, all cached data and incomplete transaction states are retained without re-initializing the system. After taking over, the system status verification and data integrity check are automatically performed. After confirming that there are no errors, the unfinished power-down protection process is continued. At the same time, a primary-backup switching log is generated to record the entire switching process.

[0031] refer to Figure 1 This diagram illustrates the overall logic of solid-state storage power-loss protection. The process begins with the construction of an industrial-grade, multi-redundant hardware architecture, which integrates independent power supply modules, multi-channel controllers, and partitioned media to lay the physical foundation for fault isolation. During operation, the system performs real-time composite power-loss detection, and immediately triggers a tiered response upon detecting an anomaly. The core protection process encompasses multi-channel cache redundancy storage, atomic transaction processing for write operations, and dynamic optimization allocation of remaining energy. Once data is securely written to the non-volatile storage partition, the system automatically performs full integrity verification and data recovery. Finally, by recording standardized fault logs and relying on a multi-controller collaborative mechanism, a closed-loop management system is achieved from power outage to system recovery. This process aims to limit the impact of single-point failures and maintain strong data consistency even in extreme industrial environments.

[0032] refer to Figure 2This diagram details the system's keen energy sensing and precise scheduling logic. The detection end employs a composite algorithm combining voltage threshold and current change rate, along with temperature compensation and grid fluctuation identification, to accurately distinguish between noise interference and actual power outages. The energy allocation logic reflects the prioritization awareness of industrial-grade design: the system calculates the priority coefficients and pending data volume of each functional module in real time. Upon detecting a power outage, the system gradually shuts down non-core computing units such as debugging interfaces and status indicator lights according to three levels: warning, pre-protection, and emergency. The recovered electrical energy is re-injected into the core data writing module, ensuring that real-time industrial control data receives the most sufficient power support. This dynamic allocation mechanism maximizes energy utilization and provides valuable microseconds of time for storing critical data.

[0033] refer to Figure 3 This diagram focuses on the highly reliable data flow logic in the storage chain. Write operations are broken down into indivisible atomic transaction units, following an improved two-phase commit protocol. In the first phase, data is synchronously pushed to three completely physically isolated redundant cache channels, each independently performing cyclic redundancy checks. In the second phase, through a transaction pre-commit mechanism, cached data is distributed in parallel to the data storage area, transaction log area, and verification information area. The data storage area adopts a three-replica architecture and introduces dynamic wear leveling and automatic bad block marking technology. This design ensures that even if a single physical flash memory chip fails or a transaction is interrupted at any stage, incremental rollback can be performed through the transaction log after the system restarts. This multi-partition, multi-replica physically isolated storage scheme fundamentally blocks the risk of data corruption and achieves highly reliable data preservation.

[0034] refer to Figure 4 This diagram illustrates how the system performs a "check-up" and "repair" of the stored data after power-on. The verification system employs a two-tier architecture: content comparison and weighted verification. First, the system scans multiple redundant partitions in parallel, identifying inconsistent data blocks through bit-by-bit content comparison. Then, a weighted checksum is calculated for suspicious blocks. This algorithm specifically assigns higher weights to high-order flash memory storage units to adapt to the physical characteristics of the medium. A unique feature of this algorithm is its support for direct location and correction of single-byte errors, enabling repair without relying on other copies. If data corruption is severe, a multi-partition data recovery process is initiated, prioritizing critical business data based on a damage severity model. The entire recovery process supports breakpoint resumption, ensuring that data repair can continue even in environments with unstable power.

[0035] refer to Figure 5This diagram illustrates the system's self-healing and diagnostic capabilities in controller failure scenarios. High-speed heartbeat detection and real-time status synchronization are maintained between the primary and backup controllers. The primary controller mirrors cached data and transaction progress to the backup controller in real time, keeping it in a hot standby state. If the heartbeat is interrupted or the primary controller program malfunctions, the backup controller quickly takes over control through pre-defined switching logic. The takeover process preserves all unfinished transactions, avoiding the time loss associated with system initialization. Simultaneously, the system generates detailed, standardized fault logs, recording precise data such as voltage curves, pending writes, and recovery processes, and performs irreversible encrypted storage. These logs are used not only for local automatic fault analysis report generation but also uploaded to the cloud management platform. This collaborative fault-tolerance and deep evidence storage design not only maintains the continuity of power-down protection but also provides scientific evidence for subsequent remote maintenance and cause tracing.

[0036] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage, characterized in that, Includes the following steps: Construct an industrial-grade solid-state storage multi-redundancy hardware architecture, integrating multiple independent power supply modules, multi-channel cache controllers and multi-partition non-volatile storage media, and establish a real-time status interaction and hardware-level fault isolation mechanism between modules to limit the impact of single-point failures; The system monitors the power supply status in real time, uses multi-level voltage threshold and current change rate composite detection to identify power outage events, classifies response levels based on power outage rate, remaining energy and operating status, and triggers the corresponding level of data protection process. A multi-channel data caching redundancy mechanism is activated, which synchronously writes the data to be written to multiple independent volatile cache channels and records the writing progress, data length and verification information of each channel. Perform transactional processing on write operations, break down complete write operations into atomic transaction units, record the entire transaction execution process log, maintain the atomicity and consistency of write operations, and stabilize the state of the file system and storage data structure; The system dynamically allocates remaining power based on the amount of data to be written in the cache, the number of write operations to be completed, the write speed of non-volatile storage, and the power consumption of the modules to calculate the energy allocation ratio and match the energy supply requirements for writing and storing critical data. Valid data in the cache is written to pre-divided multi-redundant non-volatile storage partitions. Each data block is synchronously written to multiple independent physical storage partitions, and the data storage location, write time and verification information of each partition are recorded. After all data is written, a full data integrity check is performed. The data corresponding to each redundant partition is compared with the pre-stored check information to identify abnormal partitions. The complete data of the normal partition is then called to automatically recover the damaged data. Record the complete process of the power outage event and the system operating status, generate standardized fault logs and store them in an independent non-volatile log partition to support subsequent fault location, cause analysis and system debugging; Establish a multi-controller collaborative fault-tolerant mechanism, where the main controller and the backup controller synchronize system status and cached data in real time. When the main controller fails, the backup controller automatically takes over control and restores the system status by calling redundantly stored logs and data.

2. The method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage according to claim 1, characterized in that, It also includes a dynamic power allocation optimization step, which optimizes the allocation of remaining energy based on module priority and the amount of data to be processed. The calculation formula is as follows: ; in The energy allocated to the i-th functional module, in joules. Let be the priority coefficient of the i-th functional module, which is dimensionless. The system detects the total remaining energy during a power outage, measured in joules. The effective data volume to be processed by the i-th functional module is expressed in bytes. This represents the total amount of valid data to be processed across all modules of the system, expressed in bytes. The priority coefficient is dynamically adjusted based on the importance of the data service.

3. The method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage according to claim 1, characterized in that, It also includes data integrity verification and single-byte error location steps, using a weighted verification algorithm to improve the accuracy of error detection and location precision. The calculation formula is as follows: ; in This is the final checksum of the data block, dimensionless. Let j be the decimal value of the j-th byte in the data block, dimensionless. Let be the weight coefficient corresponding to the j-th byte in the data block, which is dimensionless. The total number of bytes contained in the data block, dimensionless. The preset parity modulus is dimensionless, and the weighting coefficients are adaptively adjusted according to the error distribution characteristics of the storage medium. The bytes corresponding to the high-order storage units of flash memory are assigned higher weights. The parity value of each data block is stored in two independent parity partitions at the same time to maintain the storage stability of the parity value itself. Single-byte errors can be directly calculated based on the weighting coefficients to calculate the error location and value and be automatically corrected without relying on other redundant partitions.

4. The method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage according to claim 1, characterized in that, The multi-level voltage threshold detection system sets three continuous voltage thresholds, corresponding to normal power supply status, early warning power outage status, and emergency power outage status, respectively. The current change rate detection identifies instantaneous power outage events by sampling the power supply current once per microsecond and calculating the differential change value. A temperature compensation mechanism is introduced to correct the voltage detection threshold under different industrial ambient temperatures, reducing detection errors caused by temperature drift. The system employs a power grid fluctuation identification algorithm to distinguish between normal voltage fluctuations and actual power outages, filtering out voltage spikes lasting less than ten microseconds. When the supply voltage is lower than the warning threshold or the current change rate exceeds the preset limit, the pre-protection process is triggered in advance, and non-essential peripheral interfaces and computing modules of the system are shut down in stages. In the warning state, low-priority debugging interfaces and status indicator lights are shut down first, and in the emergency state, all non-core computing units are shut down.

5. The method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage according to claim 1, characterized in that, The multi-channel data cache adopts a three-channel fully independent redundant architecture. Each channel is equipped with an independent power management unit, data bus and cache chip. The channels are physically isolated. When a single channel has a hardware failure or data corruption, it will not affect the normal data reading and writing of other channels. During normal operation, a dynamic load balancing algorithm is used to allocate write tasks according to the idle status of each channel. After each channel writes data, a local cyclic redundancy check is immediately performed. If an error is detected, the data is automatically rewritten, preventing erroneous data from entering the subsequent storage process. The cache is internally divided into three independent sub-partitions: real-time control data area, log data area, and configuration data area. Data partitions of different business types are stored without interference. In the event of a power failure, data is processed in order of partition priority, with real-time control data write operations completed first.

6. The method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage according to claim 1, characterized in that, The non-volatile storage medium is divided into three independent physical partitions: a data storage area, a transaction log area, and a verification information area. The data storage area adopts a three-copy redundant storage method, with each logical data block stored on three different physical flash memory chips. A dynamic wear leveling mechanism is introduced to dynamically adjust the data storage location based on the number of erases and writes of each flash memory chip, thereby balancing the wear degree of each chip and extending the service life of the medium. Bad blocks in flash memory are detected in real time and automatically marked, and the data is migrated to adjacent healthy physical blocks to prevent the risk of data loss caused by bad blocks.

7. The method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage according to claim 1, characterized in that, The transactional processing of write operations adopts an improved two-phase commit protocol. In the first phase, the transaction data is written to a multi-channel cache and the transaction start log is recorded. In the second phase, the cached data is written to multiple non-volatile storage partitions in parallel. After all redundant partitions have completed data writing and verification, the transaction is committed. It supports the function of nested splitting of large transactions, which splits write operations exceeding the preset size into multiple independent sub-transactions and executes them in sequence. Each sub-transaction records its own log and commits separately.

8. The method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage according to claim 1, characterized in that, Data integrity verification adopts a two-layer verification mechanism that combines block-by-block content comparison with weighted check value verification. First, the contents of the corresponding data blocks of each redundant partition are compared byte by byte to identify data blocks with inconsistent contents. Then, the weighted check value of the normal data block is calculated and compared with the pre-check value stored in the verification information area to confirm the correctness of the data. Parallel processing technology is used to perform verification and recovery operations on multiple data blocks simultaneously. Establish a data corruption severity assessment model, prioritize recovery based on the number, location, and business importance of corrupted data blocks, prioritize the recovery of core industrial control data, support breakpoint resume during the recovery process, and if a power outage occurs during the recovery process, the system will resume the recovery process from the last interrupted position after power-on, without having to rescan all data blocks.

9. The method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage according to claim 1, characterized in that, The fault log records include the precise time of the power outage, the power supply voltage change curve, the remaining energy value of the system, the amount of data to be written in the cache, the number of completed write operations, detailed information on incomplete write operations, data recovery process records, and the final system status. A hierarchical storage strategy is adopted to store emergency fault logs, ordinary fault logs, and operation logs in different log sub-partitions. Emergency fault logs are permanently retained, while ordinary logs automatically overwrite content that has exceeded the retention period. Log data is encrypted and stored using an irreversible encryption algorithm to maintain the originality and integrity of the log data.

10. The method for power-loss data protection of industrial-grade multi-redundancy fault-tolerant solid-state storage according to claim 1, characterized in that, The multi-controller collaborative fault tolerance adopts a primary-backup hot standby real-time synchronization architecture. The primary controller and the backup controller synchronize the system operating status, cached data, and transaction logs in real time through an independent high-speed communication bus. A three-channel heartbeat detection mechanism is used to monitor the status of the primary controller, reducing the probability of false switching caused by single-channel heartbeat interruption. When the primary controller is running normally, the backup controller is in a low-power standby state, only performing status synchronization and heartbeat monitoring functions. When the primary controller experiences hardware failure, communication interruption, or program crash, the backup controller automatically takes over system control within a preset switching time. During the switching process, all cached data and incomplete transaction states are retained without re-initializing the system. After taking over, the system status verification and data integrity check are automatically performed. After confirming that there are no errors, the unfinished power-down protection process is continued. At the same time, a primary-backup switching log is generated to record the entire switching process.