A disk-based big data storage method and system

By obtaining storage capacity in the storage system to divide data into blocks, using consensus nodes to ensure data consistency, and combining real-time monitoring and thermal protection mechanisms, the problems of limited capacity expansion and data reliability in traditional storage methods are solved, and efficient and secure data storage is achieved.

CN120144584BActive Publication Date: 2025-09-19邓志敏
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510156415.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-09-19
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

Traditional storage methods are unable to flexibly cope with the rapid growth of data volume, resulting in limited storage system capacity expansion capabilities and a lack of effective loss assessment and thermal management mechanisms, affecting the reliability and stability of data storage.

Method used

By obtaining the storage capacity of the slave database, data is divided into blocks and nodes for the blocks to be stored are generated. Consensus nodes are used to ensure data consistency and encrypted storage is performed in the slave database. Simultaneously, monitoring tools are used to obtain real-time disk status data, perform storage array failure analysis, generate failure data, analyze flash memory particle loss, and set thermal protection mechanisms to ensure the security and stability of data storage.

Benefits of technology

It improves the reliability and stability of data storage, enhances data security, avoids storage system failures caused by limited capacity expansion, performance degradation and overheating, and provides a flexible, efficient and secure storage solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144584B_ABST
    Figure CN120144584B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of big data storage technology, and in particular to a disk-based big data storage method and system. The method comprises the following steps: obtaining the storage capacity of a slave database; performing data segmentation on the storage capacity of the slave database to generate a first block node to be stored; obtaining the big data to be stored, converting the big data to be stored into a request transaction data packet and sending it to the first block node to be stored; broadcasting the request transaction data packet to other block nodes to be stored during the sending process; using a preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the node consistent data is encrypted and stored in the slave database; and using a monitoring tool to obtain the slave database disk status data in real time. The present invention optimizes the efficiency of big data storage based on big data storage technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data storage, and in particular to a disk-based big data storage method and system. Background Art

[0002] Traditional storage methods typically employ fixed storage structures and capacities, making them inflexible in responding to rapid data growth and limiting the storage system's capacity expansion capabilities. When data volumes increase dramatically, traditional systems often struggle to provide sufficient storage space, leading to insufficient storage system resources and impacting data processing efficiency and stability. In systems using flash storage media, traditional methods often overlook key factors such as flash memory chip wear and erase cycles. Flash memory chips have a limited lifespan, and traditional methods lack effective wear assessment and maintenance mechanisms. This can lead to performance degradation in data storage systems after extended periods of operation, and even damage to the storage media, compromising data storage reliability. Traditional storage systems often lack effective solutions for thermal management of storage devices. Especially in high-load, large-scale data storage scenarios, the lack of dynamic thermal protection mechanisms can easily lead to storage device overheating, potentially damaging the hard drive or storage chips, and compromising the efficiency and stability of the entire data storage system. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a disk-based big data storage method and system to solve at least one of the above technical problems.

[0004] To achieve the above object, a disk-based big data storage method includes the following steps:

[0005] Step S1: Obtain the storage capacity of the slave database; divide the storage capacity of the slave database into blocks to generate a first block node to be stored; obtain the big data to be stored, convert the big data to be stored into a request transaction data packet and send it to the first block node to be stored; during the sending process, broadcast the request transaction data packet to other block nodes to be stored; use the preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the node consistent data is encrypted and stored in the slave database;

[0006] Step S2: using a monitoring tool to obtain slave database disk status data in real time; performing storage array failure analysis on the slave database disk status during the encryption storage process to generate slave database storage array failure data;

[0007] Step S3: Scan the slave database hard disk using a disk imaging tool to obtain a slave database virtual hard disk; perform flash memory particle loss analysis on the slave database virtual hard disk based on the storage array failure data to obtain flash memory particle loss data;

[0008] Step S4: Setting a thermal protection mechanism for the slave database virtual hard disk based on the flash memory particle loss data to generate thermal protection mechanism data; uploading the thermal protection mechanism data to the slave database, and storing the slave database data in the master database to perform the big data storage task.

[0009] By acquiring the storage capacity of a slave database and performing data segmentation, the present invention converts large data to be stored into request transaction data packets. Data consistency is ensured through a multi-node broadcast and consensus mechanism, effectively improving the reliability and stability of data storage. Pre-set consensus nodes handle data inconsistencies between different nodes, ensuring that the final stored data is highly consistent across multiple nodes, effectively reducing the risk of data loss or inconsistency and enhancing the fault tolerance of the storage system. The encrypted storage process not only enhances data security but also ensures the privacy of sensitive data, preventing the risk of data leakage. By monitoring the status data of slave database disks in real time and combining it with fault analysis, potential hardware failures can be identified and predicted in advance, allowing timely preventive measures to be taken, thereby reducing the risk of data loss or service interruption caused by storage array failures. Flash memory chip loss analysis of virtual hard disks helps understand the health of the storage media and predict the hard disk lifespan. This analysis provides an effective basis for subsequent hardware maintenance, preventing performance degradation caused by excessive chip loss. By setting a thermal protection mechanism, not only can the storage system's temperature management be optimized to prevent overheating and damage to the hard disk or storage chip, but the operating conditions of the storage device can also be dynamically adjusted to extend the storage device's lifespan. By uploading thermal protection mechanism data to a slave database and synchronizing it with the master database, the system's monitoring data is effectively stored and available for subsequent analysis, enabling more efficient storage and processing of big data. This approach fully guarantees the stability, reliability, and efficiency of the storage system during big data storage and processing tasks, avoiding the issues of limited capacity expansion, performance degradation, and overheating faced by traditional storage methods, providing a flexible, efficient, and secure storage solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments thereof made with reference to the following drawings:

[0011] Figure 1 A schematic diagram of the steps of the disk-based big data storage method of the present invention;

[0012] Figure 2 Detailed step flow diagram of step S2 in the present invention;

[0013] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0014] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative work are within the scope of protection of the present invention.

[0015] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.

[0016] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.

[0017] To achieve this, please refer to Figures 1 to 2 The present invention provides a disk-based big data storage method, the method comprising the following steps:

[0018] Step S1: Obtain the storage capacity of the slave database; divide the storage capacity of the slave database into blocks to generate a first block node to be stored; obtain the big data to be stored, convert the big data to be stored into a request transaction data packet and send it to the first block node to be stored; during the sending process, broadcast the request transaction data packet to other block nodes to be stored; use the preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the node consistent data is encrypted and stored in the slave database;

[0019] In this embodiment, storage capacity data is obtained from a database to determine the total storage space size. This storage space is expressed in GB. Next, the storage capacity is divided according to certain rules (for example, each data block is set to 1GB). Specifically, the entire storage space is divided into multiple fixed-size storage blocks, referred to as "block nodes." Each block node represents a storage unit that will be used to store the data to be stored. When large data to be stored arrives, data segmentation techniques are used to divide the data into multiple smaller data packets, ensuring that each packet does not exceed the size of a single data block (e.g., 1GB). The key to this step is to appropriately segment the data packets based on their size, preventing any single packet from exceeding the predetermined storage block size, thereby ensuring efficient data storage. After segmentation, each packet is labeled according to its assigned block node, ensuring that each packet can accurately identify its target storage location. All packets are simultaneously sent to multiple block nodes to be stored using broadcast technology. These block nodes form a multi-node network, improving storage reliability and fault tolerance through distributed storage. Broadcast technology ensures that all data packets are transmitted rapidly and in parallel to multiple nodes, thus enabling data distribution in a multi-node system. To ensure data consistency and accuracy, consensus nodes are configured to coordinate data packet consistency verification across all block nodes. When a node receives a data packet, it compares it with the data stored by other nodes to ensure consistency across all nodes. If any inconsistencies are detected during the comparison process (such as data loss or transmission errors), consensus is reached through an algorithm based on a pre-defined consensus mechanism. Common consensus algorithms include the Paxos protocol and the Raft protocol. These algorithms ensure consistency through voting and confirmation among multiple nodes, thus preventing data divergence. Once consensus is reached, all block nodes encrypt and store the data. To ensure data security, the strong encryption algorithm AES-256 (Advanced Encryption Standard with a 256-bit key) is used for data encryption. AES-256 is currently a relatively secure and efficient encryption standard, effectively protecting stored data from unauthorized access. After encryption, the data is stored in the slave database's storage area, ensuring secure storage and a certain degree of tamper resistance.

[0020] Step S2: using a monitoring tool to obtain slave database disk status data in real time; performing storage array failure analysis on the slave database disk status during the encryption storage process to generate slave database storage array failure data;

[0021] In this embodiment, a real-time monitoring tool is used to monitor the disk status data from the database. The monitoring tool obtains data such as the hard disk health status, temperature, read and write speed, and power consumption from the database in real time by connecting to the disk controller interface. Specifically, the monitoring tool reads the SMART (Self-Monitoring, Analysis, and Reporting Technology) information of each hard disk to obtain indicators such as the hard disk's operating temperature, read and write error rate, and remaining life. Based on these indicators, a storage array fault analysis is performed to determine whether there are signs of impending hard disk failure. If the monitoring tool detects an abnormality in a hard disk (such as a temperature exceeding 65°C or multiple read and write errors), a fault data record is generated, including information such as the ID of the failed hard disk, the time the failure occurred, and the type of failure. These fault data records will be used to further analyze potential problems in the storage array and provide a data basis for subsequent flash memory particle loss analysis.

[0022] Step S3: Scan the slave database hard disk using a disk imaging tool to obtain a slave database virtual hard disk; perform flash memory particle loss analysis on the slave database virtual hard disk based on the storage array failure data to obtain flash memory particle loss data;

[0023] In this embodiment, a disk imaging tool (such as Clonezilla) is used to scan the slave database's hard drive and generate a virtual image file of the hard drive. The disk imaging tool performs a comprehensive scan of all partitions on the slave database's hard drive, recording the hard drive's logical partition table and the data contents of each partition, and then generates a virtual hard drive copy of the slave database. During operation, the disk imaging tool obtains the hard drive's physical address and data contents from the hard drive's controller and reads data block by block based on the hard drive's block size (e.g., 512 bytes or 4KB), generating a virtual image file containing all the hard drive's contents. Flash memory chip wear is analyzed based on storage array failure data from the slave database. A specialized hard drive analysis tool (such as SSD Life) is used to read information such as the hard drive's write amplification factor (WAF), number of erase cycles, and number of bad blocks to assess flash memory chip wear. By comparing the usage of different hard drives, the flash memory chip wear data for each drive is calculated and a wear analysis report is generated. This wear data provides a basis for subsequent thermal protection mechanism configuration.

[0024] Step S4: Setting a thermal protection mechanism for the slave database virtual hard disk based on the flash memory particle loss data to generate thermal protection mechanism data; uploading the thermal protection mechanism data to the slave database, and storing the slave database data in the master database to perform the big data storage task.

[0025] In this embodiment, a hard drive temperature threshold is set based on flash memory chip wear data. For example, a thermal protection mechanism is set to activate immediately when the hard drive temperature exceeds 70°C. This temperature threshold, based on the technical specifications provided by the hard drive manufacturer, ensures effective hard drive protection when the temperature exceeds a certain range, preventing excessive temperatures from accelerating flash memory chip wear. Thermal protection mechanisms include dynamically adjusting the hard drive's workload, reducing write frequency, or activating a cooling system (such as an external fan or coolant) to control the hard drive's temperature. A monitoring tool acquires real-time hard drive temperature data and compares it with a set threshold. When the temperature reaches the preset threshold, protective measures are automatically initiated. The generated thermal protection mechanism data is uploaded to a slave database, storing detailed parameters of the mechanism (such as temperature threshold and protection mode). The data from the slave database is then stored in the master database using an encryption protocol to perform big data storage. Specifically, dual encryption is used for storage: data is encrypted using AES-256 and hashed using a hash algorithm (such as SHA-256) for data integrity verification, ensuring data security and reliability during storage.

[0026] It is particularly important that step S4 includes the following steps:

[0027] Step S41: Obtain historical flash memory particle loss data;

[0028] In this embodiment, the health data of the storage device is obtained through a dedicated tool (such as a SMART monitoring tool or an API provided by the hard disk manufacturer). These data include but are not limited to the number of erase and write cycles, the erase and write cycles of each particle, temperature records, error counts, read and write times, etc. The erase and write count data is recorded by the hard disk controller, and the erase and write count of each flash memory particle is usually read through the interface between the log record of the flash memory management unit and the monitoring software. In actual operation, the smartctl command is used to extract historical data from the SMART log of the hard disk, and the accuracy of the data is ensured by comparing the erase and write count of each particle with the current status. The key parameters of this step include the "erase and write count threshold" (for example, an erase and write count greater than 10,000 times indicates that the life limit is approaching) and the "temperature threshold" (for example, when the temperature exceeds 70°C, it starts to enter the attention state).

[0029] Step S42: constructing a loss model based on historical flash memory particle loss data and flash memory particle loss data, thereby obtaining a loss model;

[0030] In this embodiment, the historical data is first sorted out to extract the number of erase and write times, temperature data, and particle status of each flash memory particle. Then, the historical data is fitted using linear regression or multiple regression analysis methods to predict the loss trend of the flash memory particles. The construction of the loss model requires the selection of key factors related to loss, such as the number of erase and write times, operating temperature, power consumption, etc. Assume that for every thousand times the number of erase and write times increases, the particle life decreases by 10%. In this way, the relationship between the loss factor and the particle life is set. The loss model also includes a standard formula for calculating the residual life, such as L = (Lmax - E) *(1 - T / Tmax), where Lmax is the maximum life, E is the consumed life, T is the current temperature, and Tmax is the maximum operating temperature. The output of the loss model will provide a prediction of the remaining life of each particle.

[0031] Step S43: performing life prediction on the flash memory particle loss data according to the loss model to obtain life data;

[0032] In this embodiment, parameters such as the number of writes and erases per chip, operating temperature, and power consumption are input and calculated using a model formula. For each chip, its consumed lifespan (e.g., losses due to writes and erases and operating temperature) is first compared with the preset parameters in the model. The remaining lifespan is then calculated based on the model. For example, if a chip has been written and erased 8,000 times and the current temperature is 65°C, assuming the model parameters indicate a 1% loss of lifespan for every 1,000 writes and erases, and a 0.1% reduction in lifespan for every 1°C increase in temperature, the lifespan is calculated to be 8% + 0.5% (temperature loss). Finally, the remaining lifespan of each chip is calculated. This lifespan data will serve as the basis for determining the hard drive's temperature control strategy and write control.

[0033] Step S44: monitoring the temperature of the slave database virtual hard disk to obtain temperature data;

[0034] In this embodiment, temperature data needs to be obtained through a hard disk controller or a dedicated temperature sensor. The temperature sensor data provided by the hard disk controller (usually part of the SMART data) can be obtained through the smartctl command to read the operating temperature of the hard disk in real time. In addition, if a virtualization platform such as VMware or Hyper-V is used, relevant temperature data can also be obtained through the virtual hard disk control interface. In order to ensure the accuracy of the data, it is necessary to set the temperature sampling frequency, for example, collect temperature data once a minute, and record the maximum, minimum and average temperatures. The key parameters at this time are the upper temperature threshold (for example, 80°C is the high temperature warning threshold), and the standard error of the temperature of each particle (for example, ±2°C). The monitoring system should be set with an automatic alarm function to send an alarm signal when the temperature exceeds the set threshold.

[0035] Step S45: performing statistics on the temperature data and the lifespan data to obtain high temperature data and low lifespan data, and performing time overlap calculation on the high temperature data and the low lifespan data to obtain high-risk overlap time data;

[0036] In this embodiment, based on the temperature data obtained in step S44 and the lifespan data from step S43, statistics are collected to identify high-temperature data exceeding a set threshold and data with a remaining lifespan below a set threshold. For temperature data, if a particle's temperature exceeds 80°C (e.g., a set high-temperature threshold), it is recorded as high-temperature data. For lifespan data, if a particle's remaining lifespan is less than 20% (a set low-lifespan threshold), it is recorded as low-lifespan data. Then, based on the overlap between these two types of data within a time period, a temporal overlap calculation is performed to identify the overlap time for particles with excessively high temperatures and nearing the end of their lifespans. For example, if a particle's temperature exceeded 80°C and its remaining lifespan was less than 20% within the past 30 minutes, these 30 minutes are recorded as high-risk overlap time data. The overlap calculation method is: high-risk overlap time = time spent exceeding the temperature limit + time spent below the lifespan threshold (unit: minutes).

[0037] Step S46: starting a cooling mechanism for the slave database virtual hard disk based on the high-risk overlapping time data to obtain cooling mechanism data;

[0038] In this embodiment, a cooling activation threshold is set. For example, if the high-risk overlap time exceeds 10 minutes, the cooling mechanism is immediately activated. The cooling mechanism is implemented by reducing the hard drive temperature through temperature control systems such as fan regulation, air conditioning, and liquid cooling systems. Monitoring tools control the activation and deactivation of the cooling device. The cooling device should adjust its operating state in real time based on temperature data and record the time and duration of each cooling activation and deactivation, which serves as the basis for cooling mechanism data storage and analysis.

[0039] Step S47: starting a low data write frequency mechanism for the slave database virtual hard disk based on the high-risk overlapping time data, and obtaining low data write frequency mechanism data;

[0040] In this embodiment, for high-risk overlapping time data, if the high-risk overlapping time exceeds a set threshold (for example, more than 10 minutes), the low data write frequency mechanism is activated. In specific implementation, the write rate of the virtual hard disk is adjusted through the hard disk controller or the operating system, such as limiting the number of write requests per second, or reducing the concurrency of data writing. This method can reduce the burden on the flash memory particles, thereby reducing the temperature rise and loss caused by frequent writing. In implementation, a low write frequency threshold needs to be set, which is usually limited to 50% of the normal write rate or lower. This strategy is implemented through the file system of the operating system or the firmware of the hard disk controller, and the time period and write rate when the low data write frequency is enabled are recorded.

[0041] Step S48: Integrate the low data writing frequency mechanism data and the cooling mechanism data to obtain thermal protection mechanism data;

[0042] In this embodiment, data aggregation tools (such as databases or log analysis systems) are used to combine the cooling mechanism's activation time, duration, and cooling effectiveness with data on the activation time and rate of low write frequencies. Analysis of this data can form a unified thermal protection mechanism strategy that incorporates both cooling and write control measures. This thermal protection mechanism data provides system administrators with real-time protection status and can be used for subsequent optimization and adjustment.

[0043] Step S49: Upload the thermal protection mechanism data to the slave database, and store the slave database data to the master database to perform the big data storage task.

[0044] In this embodiment, it is necessary to ensure that the thermal protection mechanism data has been generated and is available for upload in step S48. During implementation, the thermal protection mechanism data is transmitted to the slave database via a network interface (e.g., TCP / IP). The data transmission protocol used should ensure data integrity and transmission stability. Typically, a secure encrypted transmission protocol (e.g., HTTPS or SSL) can be used to prevent tampering or loss during data transmission. The transmitted data should include information such as the cooling mechanism activation time, duration, write frequency adjustment time, and changes in write rate. Appropriate verification mechanisms, such as packet checksums (e.g., CRC or MD5), should be implemented during data transmission to confirm that no data loss or corruption occurred during transmission. Once the thermal protection data is successfully uploaded to the slave database, the next step is to synchronize the data from the slave database with the master database. This step requires setting an appropriate synchronization strategy based on the specific database management system (e.g., MySQL, PostgreSQL, etc.). Generally speaking, database synchronization is divided into two methods: real-time synchronization and periodic synchronization. For real-time synchronization, the database's streaming replication function can be used to maintain real-time consistency between the slave and master databases. Periodic synchronization typically uses scheduled tasks (e.g., cron jobs) to synchronize data in batches. The synchronization process uses incremental updates to reduce data transmission and improve efficiency. The incremental synchronization mechanism compares the timestamps or version numbers of the data in the source and target databases to detect which data has changed and synchronizes only the changed data, thereby reducing unnecessary data transmission. When data in the slave database changes, the master database's synchronization mechanism is automatically triggered to upload the data to the master database. The master database typically has a data storage policy to ensure sufficient storage space and high availability. The use of a distributed storage architecture for data storage can better adapt to the management of large data volumes and ensure reliable data storage. The data stored in the master database can also support subsequent big data analysis, real-time query, report generation, data mining, and other tasks. Appropriate backup and recovery strategies should be configured during data storage to prevent data loss or corruption and ensure data integrity and security.

[0045] Preferably, step S1 is specifically as follows:

[0046] Step S11: Obtain storage capacity from the database;

[0047] In this embodiment, the storage capacity of the slave database is obtained. In order to obtain the accurate storage capacity of the database, the system table of the database management system (such as the INFORMATION_SCHEMA table in MySQL and sys.master_files in SQL Server) can be used for query. The query process executes the SQL command SELECT SUM(data_length) FROMinformation_schema.tables, which will return the total storage space size of all data tables in the current database. Assume that the query result shows that the storage capacity of the slave database is 20TB. At this time, the storage capacity data is recorded as the basis for subsequent database and table sharding and data partitioning. After the storage capacity data is obtained in this step, it will be passed to the next database and table sharding process.

[0048] Step S12: vertically partition the storage capacity of the secondary database, wherein the size of the sub-database in the vertical partition is set to not exceed 5TB, and generate the vertical partition data to be stored;

[0049] In this embodiment, based on the storage capacity from the previous step (e.g., 20TB), the database is partitioned into multiple sub-databases. The maximum capacity of each sub-database is 5TB to avoid performance degradation or increased management complexity caused by excessive capacity of each database. Using database partitioning technology (such as MySQL's PARTITION function, or using the storage engine functions provided by the database management system), the database is divided into multiple logical sub-databases based on storage capacity. For example, a 20TB storage capacity can be divided into four sub-databases, each with a capacity of 5TB. The generation of each sub-database can be configured in the database management system to ensure that each sub-database has an independent physical storage path and configure read and write load balancing for each sub-database based on business needs. At this point, the four generated sub-databases will provide the foundation for subsequent data storage and table sharding operations.

[0050] Step S13: vertically partitioning the table based on the vertical partitioned database data, wherein the vertical partitioned table size is set to no more than 2TB, and the vertical partitioned table data to be stored is obtained;

[0051] In this embodiment, vertical table sharding refers to splitting database tables by fields, with each table containing a subset of data fields. The goal of this step is to ensure that each table does not exceed 2TB in size, to avoid performance bottlenecks or decreased query efficiency caused by excessively large individual tables. During implementation, it is necessary to analyze the data table structure of each sub-database and determine the fields and data division for each table. For example, if a table contains 100 fields and the table data reaches 10TB, these fields can be divided into several sub-tables based on relevance. The data volume of each sub-table does not exceed 2TB. Therefore, if 10TB of data can be divided into five tables, each with a capacity of 2TB. Using the table sharding functionality provided by the database management system (such as MySQL's CREATE TABLE command with the PARTITION BY clause, or sharding technology), perform vertical table sharding in each sub-database, ensuring that the data size and table fields of each table meet the requirements. At this point, each generated sub-table will be partitioned by field and the data volume will not exceed 2TB, providing a foundation for the next step of block node division.

[0052] Step S14: Integrate the vertical sub-library data to be stored and the vertical sub-table data to be stored to obtain the first block node to be stored;

[0053] In this embodiment, the vertical sub-library data and the vertical sub-table data to be stored are integrated. The purpose of the integration is to generate a complete data set to be stored for use in the subsequent storage block node division. The sub-library generated in step S12 and the sub-table data generated in step S13 are integrated through a database merge operation. Use the SQL INSERT INTO SELECT operation to insert the sub-table data into the sub-library. For example, for multiple 2TB sub-table data contained in a sub-library, these data are merged into block data to be stored through the database transaction management. At this point, the merged data set meets the requirements for storage node division, and the first block node to be stored generated contains the integrated data, providing a basis for subsequent partitioning operations and consensus node data verification.

[0054] Step S15: Using the first block node to be stored, the storage capacity of the slave database is divided into storage block nodes. The storage capacity of the slave database is divided according to a fixed ratio of 1:1. The capacity of each node is controlled within 2TB-4TB, and other block nodes to be stored are obtained.

[0055] In this embodiment, the storage capacity of the slave database (such as 20TB) is divided according to the fixed ratio of 1:1. The capacity of each storage block node is set between 2TB and 4TB. Therefore, the total storage capacity of 20TB needs to be divided into multiple block nodes, and the capacity of each node does not exceed 4TB. In order to ensure that the capacity of each block node is reasonable, in actual operation, the storage capacity is divided into 5 nodes, each with a capacity of 4TB, and the remaining 5TB capacity is allocated to other nodes in proportion. At this point, the capacity of all nodes is between 2TB and 4TB, which meets the storage requirements. This operation uses the partition table function of the database or a custom storage partition management tool for division to ensure that each storage node has an independent storage path and can process data read and write requests in parallel.

[0056] Step S16: Obtain the big data to be stored, convert it into a transaction request data packet, and send it to the first block node to be stored. The size of each data packet is fixed to within 500MB. During the sending process, a dedicated transmission protocol (TCP / IP) is used to broadcast the transaction request data packet at a fixed broadcast rate of 500MB / s to other block nodes to be stored.

[0057] In this embodiment, the large data to be stored is split according to the specified packet size (for example, 500MB) to ensure that the size of each request data packet does not exceed 500MB. In order to efficiently transmit these data packets, a dedicated transmission protocol (for example, the TCP / IP protocol) is adopted, and a fixed broadcast rate of 500MB / s is set to ensure that the data packets are not lost during transmission and can be quickly propagated to other storage nodes. During the transmission process, the sending window size and retransmission mechanism of the TCP / IP protocol are set to ensure that the transmission can be effectively restored in the event of network instability. In addition, the data packets will be encrypted during the transmission process to prevent data leakage during the transmission process. During the entire transmission process, it is ensured that each node can receive the corresponding data packet within the predetermined time and store and process the received data packet.

[0058] Step S17: Use the preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node, and the repair error threshold is set to 0.01% until the data of each node is consistent. The node consistent data is encrypted and stored in the slave database.

[0059] In this embodiment, when a data packet arrives at each block node to be stored, it is first compared with the consensus node to ensure data consistency across all nodes. The consensus node compares the data received from each block node. If any inconsistencies are found, a consensus algorithm (such as the Raft protocol or the Paxos protocol) is initiated to repair the data. During this process, the repair error threshold is set at 0.01%. That is, the repair mechanism is only initiated when the data error between nodes exceeds 0.01%. The consensus processing process includes steps such as data verification, data synchronization, and conflict resolution. After the consensus node completes processing, the data of all nodes is encrypted and stored in the slave database. The encryption operation uses the AES-256 encryption algorithm to ensure the security of data storage. At this point, data consistency is guaranteed across all storage nodes, and the data is stored encrypted, ensuring that subsequent operations can be performed in an efficient and secure environment.

[0060] Preferably, step S17 is specifically as follows:

[0061] Step S171: Calculate the data digests of the other block nodes to be stored and the first block node to be stored, wherein the calculation time is set to be controlled within 2 seconds, and obtain the data digests of the other block nodes to be stored and the data digest of the first block node to be stored;

[0062] In this embodiment, hash algorithms such as SHA-256 or SHA-512 are used to calculate the data of each 2TB block node. To ensure that the calculation time is controlled within 2 seconds, the data of the block node will be divided into multiple small blocks, each of which is 64MB in size. The data summary of each small block is processed using parallel computing. After all small blocks are calculated, the summaries of these small blocks are merged to generate a complete data summary. All calculation steps are accelerated by multi-core processors or GPUs to ensure that the calculation time of each data summary does not exceed 2 seconds. During calculation, the summary of each data block is completed in parallel by multiple processing units, thereby improving the overall computing efficiency and ensuring efficient time control.

[0063] Step S172: Integrate the data digests of the other block nodes to be stored and the data digest of the first block node to be stored to obtain a data digest, and broadcast the data digest to the preset consensus node with a timeout of 5 seconds to obtain the consensus node data digest;

[0064] In this embodiment, the data summaries from each block node are merged in sequence to form a large data summary. After the data summary is merged, it is broadcast to the preset consensus nodes through a dedicated network using the TCP / IP protocol. At this time, the broadcast operation needs to ensure stable data transmission. All consensus nodes should receive the data summary within the set timeout period of 5 seconds. If a response is not received in time, the retry mechanism will be triggered. After receiving the data summary, each consensus node will verify it. The key to the entire broadcast process is to ensure the integrity and reliability of data transmission and ensure that the timeout period is controlled within 5 seconds.

[0065] Step S173: Use the consensus node data digest to determine the consistency of the data digests of other to-be-stored block nodes and the first to-be-stored block node. If the data digest difference exceeds 5%, an inconsistent data digest is obtained. The first to-be-stored block node is used to initiate a consensus request to the consensus node. The consensus node responds with the voting result within 5 seconds. When at least 3 / 4 of the nodes that meet the protocol requirements agree to the data update, consensus is confirmed and the node consistent data is obtained.

[0066] In this embodiment, the consensus node uses a hash algorithm to compare the data summaries of each node to determine whether the data summaries are consistent. If the difference in the data summaries exceeds 5%, it is determined to be inconsistent. In this case, the first block node to be stored initiates a consensus request to all consensus nodes. According to distributed protocols, such as the Paxos protocol or the Raft protocol, at least 3 / 4 of the consensus nodes need to vote on data consistency within 5 seconds. If the protocol requirements are met and 3 / 4 of the nodes agree to the data update, the data consistency is determined and the update of the data summary is confirmed. The core of the judgment process is to quickly confirm the correctness of the data through the comparison of hash values ​​combined with the consistency protocol to ensure that there will be no data deviation in the execution of the update operation.

[0067] Step S174: using the node consistent data to repair the inconsistent data digest to obtain a repaired data digest;

[0068] In this embodiment, for digests with discrepancies, the inconsistencies are repaired after confirming consistency with the first block node to be stored and other consensus nodes. This repair operation relies on the consensus node's vote. The key to repairing the data is to obtain the support of at least three-quarters of the consensus nodes during the repair process, ensuring that the data repair process is approved by a majority of nodes. The repaired data digest is rehashed to ensure its consistency and integrity. The repair operation requires coordinating the correct data from each node to ensure that the repaired data no longer differs from the digests of other nodes, thus completing the consistency repair.

[0069] It is particularly important that step S174 includes the following steps:

[0070] Using the node consistent data to identify the inconsistent data summary, thereby obtaining the inconsistent data summary;

[0071] In this embodiment, it is necessary to obtain confirmed consistent datasets from multiple nodes. This data is typically extracted from each node through a database or API. Using a programming language such as the pandas library in Python, this data is loaded into a DataFrame for subsequent operations. The extracted data typically includes multiple fields, such as timestamps, device status, and sensor values. Next, comparison rules are set to determine which data is "consistent." For numeric data, a tolerance range (for example, ±0.01) can be set; data outside this range is considered inconsistent. For string data, field contents are directly compared for exact matches. During data comparison, the merge or join function in pandas is used to align data from different nodes based on common fields (such as timestamps and device IDs) to ensure one-to-one correspondence. These fields are then compared one by one. If the difference in a numeric field exceeds the set tolerance range, the data is considered inconsistent. For string fields, the data is directly compared for consistency. After identifying inconsistent data, it is marked as "inconsistent data." Ultimately, the inconsistent data will be output as an inconsistent data summary, including inconsistent data items, source nodes, field names, and inconsistent values ​​or identifiers, providing a basis for subsequent data repair and analysis.

[0072] Missing item analysis was performed on inconsistent data summaries to obtain missing item targets for the data summaries;

[0073] In this embodiment, missing items typically appear as null values ​​(NaN) or incomplete field content, caused by data transmission errors, missing records, or node synchronization issues. By examining each field in the inconsistent data summary, incompletely recorded data items are marked. Specifically, the inconsistent data summary is first cleaned using the pandas library to remove outliers and empty data other than valid data. Each field is then checked to determine which data items have missing values. If a field value is empty or invalid (such as NaN, None, or an empty string), the field is marked as missing. To further analyze missing items, missing value detection methods can be used, such as the SimpleImputer class in scikit-learn. This tool provides a variety of missing value imputation strategies, such as mean imputation, median imputation, and most frequent value imputation, suitable for different data types. During missing item analysis, the data type of the missing field (e.g., numeric, character, or time) is first determined, and then an imputation strategy appropriate for that data type is used for analysis. For example, for numeric fields, you can choose to calculate the mean or median of the field as the filling value; for character fields, you can use the most frequent value or the specified default value for filling. In addition to the automatic filling strategy, you can also analyze the context of each missing field by comparing it with the consistent data of the node to determine whether there is a systematic missing. In this process, you can also supplement through the relationship between fields. For example, the missing of a field will affect the data judgment of other related fields. In this step, the missing items are finally classified, all missing data items that need to be supplemented are listed, and a target is provided for the subsequent filling process. The data summary missing item target will include the field name, data type, number of missing items and the corresponding filling strategy to provide a reference for the subsequent filling steps.

[0074] Fill in the missing items in the data summary according to the node consistent data to obtain the filled data summary;

[0075] In this embodiment, field data with similar characteristics is extracted from the stored node consistency data, or missing items are filled from the same data source. For example, if the missing item is the operating temperature data of a certain device, it can be filled based on known data of other devices or the same device type. The filling method can be based on regression analysis, mean filling or median filling. For example, for numerical data, if there are many missing items in the field, you can choose to use the mean of adjacent nodes or data blocks for filling; for categorical data, the most common category is selected as the filling value. In order to maintain the accuracy of the data, the sklearn.impute.SimpleImputer toolkit can be used in the filling process, and the filling strategy can be adjusted according to the data distribution. After filling, the missing items are effectively supplemented, and the filled data summary is finally obtained.

[0076] Perform least squares fitting and repair on the inconsistent data summary to obtain a repaired summary;

[0077] In this embodiment, numerical fields such as temperature, pressure, and humidity are extracted from the inconsistent data summary obtained in step S12. These fields contain some outliers, which typically manifest as extreme values ​​far from the mean and are caused by factors such as sensor errors, data transmission issues, or recording errors. After extracting this numerical data, the outliers must first be identified. Statistical methods can be used, such as calculating the mean and standard deviation of each field and setting a reasonable threshold to determine which data is anomalous. For example, if a data point deviates from the mean by more than three times the standard deviation, it is considered an outlier. Visualization methods such as box plots can also be used to assist in identifying outliers. Using these methods, all anomalous data points are marked, preparing for the fitting and repair phase. The anomalous data is repaired using the least squares method. The least squares method is an optimization method whose goal is to minimize the sum of squared errors between the fitted data and the actual data to find the best fitting model. For implementation, libraries such as numpy.polyfit or scipy.optimize.curve_fit can be used for least squares fitting. To perform this operation, first select an appropriate fitting model type, such as linear, quadratic, or a more complex polynomial fit. The choice of model depends on the data's changing trend. If the data changes relatively smoothly, a linear fit is appropriate; if the data exhibits nonlinear characteristics, a higher-order polynomial fit or other more complex models may be required. For example, if the temperature data contains outliers and the overall trend is linear, you can use numpy.polyfit to perform a linear fit. By setting appropriate parameters, such as setting the order of fit to 1, a linear fit is performed. The fitted model will produce a straight line equation that represents the overall trend of the data. Using this model, all outliers are corrected to the closest values ​​in the fitting results, ensuring that these data points better reflect the actual situation and trend. The fitting error can be used to assess the model's accuracy and ensure sufficient accuracy. During the repair process, the error between each data point and the fitted model is calculated, and the model parameters are adjusted accordingly to ensure model stability and accuracy. Multiple adjustments can be made during the repair process, as necessary, until the fitting results meet the predetermined accuracy requirements. The repaired outlier data points are updated with the original data, forming a repair summary. The repaired data will be smoother and more consistent, reducing outliers that do not conform to the overall trend and ensuring the accuracy and reliability of the data.

[0078] Integrate the repair summary and the filled data summary to obtain the inconsistent data repair summary;

[0079] In this example, the padded data summary and the repair summary are merged to ensure that the two data sets have a unified format and consistent data fields. To achieve this, the two datasets can be merged using a union or join operation in the database, or in a programming implementation, using the concat method in Pandas. During the merging process, it is necessary to ensure that the repaired and padded data fields are arranged in chronological order and that there are no duplicate data items. After the merger, the generated "Inconsistent Data Repair Summary" contains all padded and repaired fields, ensuring data consistency and integrity.

[0080] Perform hash calculation on the inconsistent data repair digest to obtain the repair data digest.

[0081] In this embodiment, hash calculation is the process of converting data of arbitrary length into a fixed-length value using a hash function. This process ensures data consistency and integrity. Hash values ​​are unique; identical input data will produce the same hash value, while different input data will produce different hash values. Therefore, hash values ​​can serve as data identifiers for data verification and integrity checks. In specific implementations, the integrated data is first extracted from the repaired data digest obtained in step S15. This data digest is the padded and repaired data, including data for all missing items and data after outlier repair. The integrated data digest is a complex data structure containing multiple fields and values, stored in JSON, XML, or other formats. Next, a suitable hash algorithm is selected to perform a hash calculation on the integrated repaired data digest. Common hash algorithms include SHA-256 and MD5, which have different output lengths and security requirements. The SHA-256 algorithm generates a 256-bit (32-byte) hash value and is typically used in scenarios requiring higher security. The MD5 algorithm, on the other hand, generates a 128-bit (16-byte) hash value. While faster, it offers relatively lower security and is therefore typically used for data integrity verification. In practice, hash calculations can be performed using the hashlib library in Python. The specific steps are as follows: First, import the hashlib library and import hash functions using `import hashlib`. Next, select a hash algorithm. For example, if you're using the SHA-256 algorithm, call `hashlib.sha256()` to generate a SHA-256 hash object. Next, hash the data and pass the combined repair data digest to the hash function for hashing. Since the input data is a string or byte stream, you must first convert the data to byte format. Use the `.encode()` method to encode it (if the data is a string), or pass it directly to the hash function as a byte stream. Finally, call the `.hexdigest()` method to obtain the calculated hash value. This method returns a hexadecimal string representing the hash value of the data. The resulting `hash_value` is the hash value of the repair data digest, a unique, fixed-length string. This hash value can be used for subsequent data verification to ensure data consistency and integrity during the repair process. The resulting repair data digest hash value can be used for data verification. During data transmission or storage, by comparing the hash value of the original data with the hash value of the repaired data, it can be verified whether the data has been altered outside of the repair process. If the hash values ​​are the same, the data has not been tampered with; if the hash values ​​are different, it indicates data inconsistency or corruption. The hash value can also serve as a unique identifier for the repair data digest, facilitating subsequent recordkeeping and auditing. When documenting the data recovery or repair process, the hash value can be used to track the source and history of the repair.The hash algorithm's inherent collision resistance ensures the uniqueness of the repaired data digest even after multiple repair operations, preventing malicious data tampering. This step ensures the integrity and accuracy of inconsistent data during the repair process, providing a reliable basis for subsequent data processing or verification.

[0082] Step S175: Encrypt the repair data summary to generate an encrypted data packet, and store it in the slave database. The encryption process takes no more than 10 seconds when the data packet size is set to 500MB.

[0083] In this embodiment, the repaired data digest is encrypted using the AES-256 symmetric encryption algorithm. Before data encryption, each data digest is split into small blocks no larger than 500MB to ensure that the encrypted data packet size does not exceed 500MB. The encryption process is required to complete within 10 seconds. If the data packet is larger than 500MB, it is encrypted in multiple segments. During the encryption process, hardware acceleration (such as GPU acceleration) is used to increase encryption speed, ensuring that each encrypted packet can be completed within the specified time. Finally, the encrypted data packet is stored in the slave database through the database storage interface according to the specified storage structure, ensuring data security and efficient access.

[0084] Preferably, step S2 is specifically as follows:

[0085] Step S21: using a monitoring tool to obtain disk status data from the database at a sampling period of once per second, including I / O performance data and disk health status data;

[0086] In this embodiment, a monitoring tool is used to obtain disk status data from the database using a sampling cycle of once per second. At this point, the monitoring tool will conduct real-time data acquisition with the server where the database is located, recording I / O performance data and disk health status data. I / O performance data includes indicators such as disk read and write speed, disk I / O queue length, and the number of read and write requests per second. This data is obtained using operating system monitoring interfaces such as iostat or smartctl. Disk health status data includes indicators such as disk temperature, health, and SMART status (Self-Monitoring, Analysis, and Reporting Technology). The acquisition cycle is set to once per second, and data is collected by the real-time monitoring system and stored in the database to ensure the timeliness and accuracy of the data.

[0087] Step S22: performing storage array fault analysis on the I / O performance data to obtain I / O performance fault data;

[0088] In this embodiment, the I / O performance data collected every second is aggregated and stored to form time series data. Then, potential fault conditions are detected by defined thresholds (for example, read and write speeds less than 100MB / s, I / O queue lengths greater than 10, and the number of read and write requests per second less than 500). Data analysis tools (such as the Pandas library in Python or the time series analysis package in R) are used to perform trend analysis on the I / O performance data and calculate I / O performance fluctuations in different time periods. If certain data points exceed the set threshold, it indicates that there is a storage array failure. The algorithm detects whether there are performance bottlenecks, hard drive failures, and other problems, and finally generates I / O performance fault data, recording the specific time period and fault type of the fault.

[0089] Step S23: performing storage array fault analysis on the disk health status data to obtain disk storage array fault data;

[0090] In this embodiment, each parameter of the disk health status data (such as disk temperature, health status, SMART error value, etc.) is input into a preset fault judgment model. The model is based on actual industry standards. For example, a disk temperature exceeding 65°C will cause disk damage, and a SMART status error count exceeding 5 indicates a hardware failure. This data is analyzed using a rule engine (such as the Drools rule engine), and faults are identified based on preset fault judgment criteria (such as excessive temperature, SMART error limit exceeded, etc.). If the health status of a disk is displayed as "bad" or "warning", it is considered that the disk has a storage array failure. In this way, storage array failure data caused by abnormal disk health status is identified and recorded.

[0091] Step S24: Integrate the I / O performance fault data and the disk storage array fault data to generate secondary database storage array fault data.

[0092] In this embodiment, I / O performance failure data and disk health failure data are linked together using a database management system (such as MySQL or PostgreSQL) to generate a consolidated data table containing both information. By linking the data tables, I / O performance data and disk health status data for each time period are combined to form a comprehensive storage array failure data record. The consolidated data table contains information such as the failure occurrence time, disk ID, failure type, failure cause, and related performance indicators. This data is arranged in a specific chronological order and stored in the database to facilitate subsequent failure analysis and report generation.

[0093] Preferably, step S22 is specifically as follows:

[0094] Step S221: When any of the following conditions occurs, the storage array read / write performance is determined to be abnormal, and abnormal data of the storage array read / write performance is obtained: the sequential read / write speed drops by more than 20%, the random access latency exceeds 30% of the normal range, or the I / O queue depth is in a full load state for a long time and lasts for more than 60 minutes;

[0095] In this embodiment, the sequential read and write speeds of the storage array are regularly obtained through the monitoring system and compared with the baseline speed during normal operation. If the sequential read and write speed drops by more than 20%, that is, exceeds the threshold, it is determined that the read and write performance is abnormal. Secondly, the random access delay is monitored, its current value is obtained and compared with the normal range set by the system. If the random access delay exceeds 30% of the normal range, that is, exceeds the threshold, it is also determined that the performance is abnormal. Finally, the I / O queue depth is monitored. If the queue depth is in a full load state for a long time and the duration exceeds 60 minutes, it indicates that the workload of the storage array is overloaded, which is also determined to be a read and write performance abnormality. Under these conditions, the storage array read and write performance abnormality data is collected and integrated to form a performance fault report.

[0096] Step S222: When the following conditions occur simultaneously, it is determined that the storage array controller is faulty and storage array controller fault data is obtained: the controller CPU utilization exceeds 90% for a long period of time and the response time is abnormally prolonged; the cache hit rate drops by more than 25%; the communication error frequency between the controller and the disk continues to increase and cannot be restored to normal within 10 minutes;

[0097] In this embodiment, the controller CPU utilization data is obtained through a monitoring tool. If the CPU utilization continues to exceed 90% and is accompanied by a significant increase in the response time, it indicates that the controller has a performance bottleneck or fault. In addition, the cache hit rate is monitored. When it drops by more than 25%, this is also an indicator of an abnormality in the controller. The communication error frequency between the controller and the disk is monitored. If the error frequency continues to increase, and the error cannot be recovered within 10 minutes, and the system does not automatically repair it, it is clearly determined that the controller has a fault. At this time, data related to the storage array controller fault is collected, including CPU utilization, cache hit rate drop, and communication error frequency, for subsequent processing.

[0098] Step S223: Integrate the storage array read and write performance abnormality data and the storage array controller fault data to obtain I / O performance fault data.

[0099] In this embodiment, I / O performance data is matched and aggregated with controller fault data to ensure that all fault data types are fully recorded and sorted by timestamp. A data analysis system automatically generates a fault data report containing detailed information about both I / O performance and controller faults, including the degree of abnormality for each parameter value and the time of occurrence. This data report provides a basis for subsequent troubleshooting and resolution.

[0100] Preferably, step S23 is specifically as follows:

[0101] Step S231: extracting disk rotation speed features from the disk health status data to obtain disk rotation speed data;

[0102] In this embodiment, the monitoring system needs to connect to the storage array to ensure real-time health data for each disk in the storage array. A storage array typically consists of multiple disks, each equipped with internal monitoring sensors that monitor their operating status, including key indicators such as temperature, rotational speed, and health. When acquiring disk health data, the monitoring system periodically acquires rotational speed information from each disk's sensor. To ensure real-time and accurate data acquisition, the data acquisition cycle is set to once per second. Each disk's rotational speed is directly measured by an internal rotational speed sensor in revolutions per minute (RPM). The sensor continuously monitors the rotation of the disk platter and outputs the current rotational speed value. This rotational speed data is typically transmitted to the monitoring tool via the disk drive's firmware or the storage control program in the operating system. To ensure high data accuracy, the disk rotational speed values ​​are typically subjected to some noise reduction processing to filter out abnormal fluctuations caused by disk hardware failures or external interference. After receiving the rotational speed data collected every second, the monitoring tool stores this data in chronological order, typically in a time series format. This stored data includes the specific rotational speed value for each second and the sampling timestamp. This enables the monitoring system to provide detailed rotational speed history, allowing subsequent analysis to assess and diagnose the disk's operating status at every moment. By filtering rotational speed information from disk health status data, real-time rotational speed data for each disk can be obtained. This data serves as the foundation for disk performance analysis and provides crucial information for subsequent fault diagnosis. Based on this rotational speed data, the monitoring system can assess whether the disks are experiencing rotational speed fluctuations, wear, or potential failures, thereby predicting their remaining lifespan and the need for maintenance. The monitoring system must ensure stable data collection to avoid data loss or delayed collection due to network latency or hardware failure. Monitoring tools can ensure continuous and accurate data collection through regular self-tests and error logging. Furthermore, to cope with the large number of disks in large-scale storage arrays, monitoring systems often use a distributed architecture to parallelize data collection tasks for each disk, improving data collection efficiency and real-time performance.

[0103] Step S232: Obtaining disk standard rotation speed range data using disk technical documentation;

[0104] In this embodiment, this data is typically provided by the disk manufacturer in technical documentation. The standard speed range refers to the speed range that the disk should maintain under normal operating conditions. Typically, for different types of disks (e.g., 5400 RPM, 7200 RPM, 10000 RPM, etc.), the manufacturer will indicate the standard operating speed range for each type of disk in the documentation. The data obtained from the technical documentation will serve as a benchmark for subsequent comparisons to determine whether the actual disk speed falls within the standard range. This data needs to be updated regularly to ensure it reflects the standards of the latest disk models.

[0105] Step S233: identifying abnormal fluctuations in the disk speed data based on the disk standard speed range data. If the speed fluctuation exceeds ±5% of the disk standard speed range, generating abnormal speed fluctuation data.

[0106] In this embodiment, the system obtains the disk's standard speed range data based on the disk model and manufacturer's technical documentation. This data typically includes the disk's rated speed and the speed range under normal operating conditions. The system then compares the real-time speed data (in revolutions per minute) collected by the disk's internal sensor with this standard range data and calculates the deviation between the current speed and the standard range. For example, if the disk's standard speed range is 6000 RPM to 7200 RPM and its real-time speed is 7500 RPM, the deviation is 7500 RPM minus 7200 RPM, which equals 300 RPM. To determine whether abnormal speed fluctuations exist, the system sets a threshold of ±5%. This ±5% range is calculated based on the standard speed range. That is, if the disk's real-time speed exceeds ±5% of the standard speed range, it is considered to have abnormal speed fluctuations. For a disk with a standard speed range of 6000 RPM to 7200 RPM, the ±5% range is 5700 RPM to 7560 RPM. If the real-time speed exceeds this range, the system determines that the speed has fluctuated abnormally. In actual operation, the monitoring system calculates the deviation between the real-time speed and the upper and lower limits of the standard range and determines whether the value exceeds the ±5% range. If it exceeds this range, the system generates abnormal speed fluctuation data. This data includes the speed deviation value, the timestamp of the abnormal fluctuation, and the real-time speed value. This abnormal fluctuation data is stored along with other disk health status data and used for subsequent fault diagnosis and maintenance decisions.

[0107] Step S234: evaluating the physical loss of the disk based on the abnormal rotation speed fluctuation data, thereby obtaining disk loss data;

[0108] In this embodiment, a baseline disk wear assessment model is established. This model typically combines the number, amplitude, and duration of abnormal speed fluctuations to assess disk physical wear. To assess disk physical wear, a wear scoring system must first be established. This system can calculate the extent of disk wear based on data from abnormal speed fluctuations. The assessment process includes the following steps: First, the frequency and amplitude of each abnormal speed fluctuation are recorded, and the amplitude of each fluctuation is calculated. For example, the amplitude can be determined by the difference between the real-time speed and the standard speed range. Next, the duration of the speed fluctuation is assessed based on the actual disk operation time. The duration refers to the duration of the speed fluctuation outside the standard range. If the speed fluctuation persists for more than a certain threshold (for example, 10 minutes), the wear is considered severe. The number and amplitude of abnormal speed fluctuations are then combined with the duration to calculate a comprehensive physical wear score. The wear score calculation formula is based on manufacturer-provided technical standards or industry experience, typically using a weighted average approach. For example, the amplitude and duration of the speed fluctuation can be weighted differently based on their importance, with the amplitude being given a higher weight because larger speed deviations generally indicate a greater risk of physical damage. Assume that each time the rotational speed fluctuation exceeds the standard range by ±5%, different scores are assigned based on the magnitude of the fluctuation (for example, 1 point is added for each ±5% fluctuation). Furthermore, additional points are added for each duration exceeding a certain limit (for example, 10 minutes). Finally, all scores are weighted and summed to obtain a physical wear score. A high score (for example, a score exceeding 70%) indicates that the disk has a high degree of wear and requires replacement or maintenance. A low score (for example, a score below 30%) indicates that the disk has minimal wear and can still be used normally. This physical wear score serves as an important indicator of disk health, helping administrators make subsequent maintenance and replacement decisions.

[0109] Step S235: performing storage array fault determination on the disk health status data based on the disk loss data. When the disk loss score exceeds 60%, disk storage array fault data is obtained.

[0110] In this embodiment, the disk physical wear score calculated in step S234 is compared with a preset wear score threshold (60%). This wear score is calculated using a weighted approach based on factors such as the amplitude, frequency, and duration of disk rotational speed fluctuations. This score reflects the health of the disk. If the disk wear score exceeds 60%, it indicates that the disk's wear has reached a certain level, exceeding the health threshold, impacting the overall performance and data security of the storage array. To determine a failure, a preset threshold is used to determine whether the disk has entered a high-risk state. Once the wear score exceeds 60%, the system automatically determines that the disk's health status does not meet the standard, indicating that the disk is nearing failure and presents a risk of failure. In this case, the system generates disk storage array failure data and marks the disk as a high-failure-risk disk. This failure data is recorded and fed back to operations and maintenance personnel as a basis for disk replacement or repair, preventing data loss or system interruption in the storage array due to disk failure. Through continuous monitoring and regular health checks, potential failure risks can be promptly identified and preventive measures can be taken.

[0111] Preferably, step S234 is specifically as follows:

[0112] Count the frequency of abnormal speed fluctuation data, where the sampling period is set to collect data once per minute to obtain frequency data;

[0113] In this embodiment, the system needs to set a sampling period of one minute, sampling disk speed data once per minute to ensure that the data collection frequency meets real-time requirements. In actual operation, speed data is collected in real time by the disk's internal sensor and transmitted to the monitoring system. At the end of each sampling period, the system records the current speed data point and checks whether this data exceeds the set standard speed range, that is, whether the speed fluctuation exceeds the ±5% threshold. Specifically, for each minute of data, the system compares all speed data within that period one by one to determine whether there are any abnormal fluctuations exceeding the standard range of ±5%. For each sampling period, if the number of abnormal speed fluctuations is greater than 0, the number of occurrences is recorded. For example, if there are three fluctuations exceeding the standard range of ±5% within a minute, the frequency of abnormal speed fluctuations for that minute is 3. The frequency data is calculated using a simple counting method, recording the number of fluctuations within each one-minute period. The frequency data for each minute of abnormal speed fluctuations is recorded and stored in chronological order. This allows the system to conduct further analysis and evaluation based on this frequency data in subsequent processing steps. To ensure data accuracy and timeliness, the sampling frequency is set to once per minute. This frequency is sufficient to capture abnormal rotational speed fluctuations while ensuring the high real-time nature of the calculated frequency data. This frequency data is stored in a database for subsequent use, such as physical wear assessment. In subsequent analysis, frequency data becomes a crucial indicator for assessing disk health and potential failure risks.

[0114] Count the fluctuation time of abnormal speed fluctuation data, where the fluctuation time greater than 5 seconds is set as effective fluctuation, and obtain the fluctuation time data;

[0115] In this embodiment, the system extracts the duration of each abnormal speed fluctuation from the data acquired in step S231. Abnormal speed fluctuations typically manifest as speed data exceeding the standard range by ±5%. Therefore, during the data extraction process, the system records the start and end times of each fluctuation. The duration of each abnormal speed fluctuation is calculated as the time difference between the start and end of the fluctuation. First, the system examines the start and end times of each abnormal speed fluctuation event one by one, calculating the duration of each fluctuation by simply differencing the timestamps. For each abnormal speed fluctuation event, the system calculates the duration of the fluctuation based on the timestamp difference. The system then filters the fluctuation duration based on predefined criteria, defining fluctuations longer than 5 seconds as valid. In other words, if a speed fluctuation lasts longer than 5 seconds, the system considers it a valid fluctuation and records the duration. Fluctuations lasting less than or equal to 5 seconds are excluded from the valid fluctuation data. To ensure data accuracy, the system checks all fluctuation events individually to ensure that the duration of each fluctuation is accurately recorded. The system accumulates and stores the duration of all valid fluctuations, forming a dataset of each valid fluctuation duration. This dataset serves as a key basis for disk health assessment and provides essential input for subsequent wear assessment models. The duration of valid fluctuations helps assess the physical wear of disks, particularly abnormal events with longer fluctuation durations, which can have a greater impact on the long-term stability and reliability of the disks.

[0116] A disk physical loss assessment model is constructed based on frequency data and fluctuation time data, where the weight of the impact of frequency on loss is set to 70% and the weight of the impact of fluctuation time on loss is set to 30%;

[0117] In this embodiment, it is necessary to assign weights to the impact of frequency data and fluctuation time data on disk loss. Based on actual requirements, the weight of frequency on loss is set to 70%, while the weight of fluctuation time on loss is set to 30%. This weighting reflects the dominant role of frequency on loss: the higher the frequency of abnormal speed fluctuations, the greater the risk of disk loss. Fluctuation time, on the other hand, has a more secondary impact on loss and therefore a lower weight. During the evaluation process, each disk's frequency and fluctuation time data must first be normalized to allow comparison of these two parameters within a consistent scale. Common normalization methods include min-max normalization and Z-score normalization. Min-max normalization scales the data to a range between 0 and 1, while Z-score normalization transforms the data into a standard normal distribution by subtracting the mean and dividing by the standard deviation. After normalization, the numerical ranges of the frequency and fluctuation time data are unified, allowing for a comprehensive evaluation using a weighted approach. The disk's physical loss score is then calculated based on the assigned weights using the following formula: Loss score = 0.7 normalized frequency + 0.3 normalized fluctuation time. This formula amplifies the effect of frequency on wear and minimizes the effect of fluctuation duration, ultimately resulting in a comprehensive wear score. The score typically ranges from 0 to 1, representing no wear to maximum wear. This wear score provides a quantitative indicator of drive health, helping to assess the physical wear and tear of a drive during operation.

[0118] The disk physical loss assessment model is used to assess abnormal rotational speed fluctuation data: 31%-50% is considered mild loss, 51%-70% is moderate loss, and 71%-100% is severe loss.

[0119] In this embodiment, the core task of wear assessment is to compare each drive's wear score against a predetermined wear level range to determine the drive's health. Based on the calculated wear score, the score is categorized into three different levels. First, slight wear is defined as a wear score between 31% and 50%. This indicates that, while the drive has experienced abnormal speed fluctuations, the wear is relatively minor and can continue to operate. Next, moderate wear is defined as a wear score between 51% and 70%. This indicates significant physical wear and poses a risk, requiring maintenance or further monitoring. Finally, severe wear is defined as a wear score between 71% and 100%. This indicates that the drive's speed fluctuates frequently and for an extended period, reaching a level of severe damage and requiring immediate replacement or repair. To perform this categorization, each drive's wear score is first calculated and compared with the aforementioned wear level criteria. If the score falls within the slight wear range, the drive is labeled as slight; if the score falls within the moderate wear range, the drive is labeled as moderate; and if the score falls within the severe wear range, the drive is labeled as severe. By classifying the loss level of each disk, a clear basis can be provided for subsequent fault diagnosis and early warning, and data support can be provided for subsequent storage array maintenance and management.

[0120] The disk loss data is obtained by integrating the slight loss data, the moderate loss data and the severe loss data.

[0121] In this embodiment, the wear level of each disk assessed in step S244 is first organized. Each disk's wear score is categorized as minor, moderate, or severe. Based on these categorizations, each disk is labeled and categorized. Then, all disks with different wear levels are aggregated by wear category, listing the number, specific names, and wear scores of disks with minor, moderate, and severe wear, respectively. This aggregation process not only helps administrators visualize the health of each disk at a glance, but also allows them to quickly determine which disks require immediate attention, repair, or replacement based on their wear level. The aggregated disk wear data report includes the following elements: each disk's wear score, wear level (e.g., minor, moderate, severe), specific wear symptoms (e.g., the number and duration of abnormal rotational speed fluctuations), and whether further action (e.g., replacement, maintenance, etc.) is required. This report plays a crucial role in storage array health management, triggering early warning systems, and subsequent maintenance decisions. It can help enterprises or data centers more efficiently manage disk resources and develop appropriate maintenance plans.

[0122] Preferably, step S3 is specifically as follows:

[0123] Step S31: obtaining an image file of the slave database virtual hard disk, and parsing the image file using a disk imaging tool, thereby obtaining the slave database virtual hard disk;

[0124] In this embodiment, a suitable disk imaging tool is selected to extract the image file of the target database virtual hard disk. Commonly used tools include "dd" and "Clonezilla." Taking "dd" as an example, to use this tool, you first need to connect to the target virtual machine or virtual hard disk in the virtualization environment. By specifying the virtual hard disk's device path (e.g., / dev / sda) and the image file save path, execute the following command to perform a block-level copy: "dd if= / dev / sda of= / path / to / output.img bs=4M." This command copies each block of the virtual hard disk one by one, ensuring that the image file contains the complete virtual hard disk data, including the file system structure, operating system, and database files. After the image file is obtained, it is stored on a designated storage device, such as an external hard drive or network storage device, to ensure data integrity and security. During the storage process, the image file should be verified using a reliable file verification method (e.g., SHA256 hashing) to ensure that the file has not been corrupted during the storage process. The verification process involves calculating the image file's hash value and comparing it with the hash value of the original image. If the two match, the data has not been altered. Use tools such as FTK Imager or EnCase to parse the acquired image file. These tools extract file system metadata from disk images, such as file creation and modification times, size, and type, and verify the file contents individually. During the parsing process, the file system structure is reconstructed by analyzing the virtual hard disk's sector information, ensuring that each partition, file, and folder is correctly identified. The tool automatically detects the file system type (such as NTFS, ext4, etc.) and performs parsing based on the file system's characteristics. To ensure data integrity, the parsing tool also performs integrity checks on each file's block data and stores the results in a log file. All parsed data is retained in its original format for easy access. If any problems arise during data parsing, the tool generates an error report to help locate the problem. Finally, during the parsing process, each file's metadata (such as timestamp, file size, and file type) should be recorded for later analysis.

[0125] Step S32: extracting the number of write times and the number of erase cycles of the flash memory based on the storage array fault data to obtain the number of write times and the number of erase cycles of the flash memory;

[0126] In this embodiment, storage array fault data is obtained from the storage array management system, and the number of writes and erase cycles for each flash memory cell is obtained through the storage array control system. This information is typically stored in the health monitoring area of ​​the storage device (such as the SMART attributes). The flash memory controller monitors and records each flash memory cell in real time. To extract this characteristic data, it is necessary to access the device's operating data through the controller's management API, device management interface, or dedicated storage array monitoring tools. Typically, the interfaces provided by the storage array management system (such as RESTful API, SNMP protocol, dedicated command line tools, etc.) support remote query and data extraction. During the data extraction process, the system reads the write count and erase cycle data for each flash memory cell, recording the cumulative number of writes and erase cycles for each flash memory cell. The write count data indicates the number of data blocks written to each flash memory cell, while the erase cycle refers to the number of erase cycles each flash memory cell undergoes during its lifetime. To ensure data integrity and real-time performance, the sampling period is set to once per hour to ensure that the data recorded within each hour can promptly reflect the device status. After data collection, the storage array control system stores the write and erase cycle values ​​in a detailed data table based on the unique identifier of the flash memory chip (such as the physical address or identifier). The data table should include fields such as chip ID, write count, erase cycle, and data extraction timestamp. These data tables are then imported into a database for subsequent analysis.

[0127] Step S33: accumulating the number of write times of the flash memory to obtain the cumulative number of write times data;

[0128] In this embodiment, the write count data for each flash memory particle is accumulated one by one. A period window is set, usually one day or one week, and the write counts in each cycle are calculated based on the write log of the flash memory, and these counts are accumulated. When statistics are taken, ensure that each write operation is accurately recorded and included in the total. By accumulating the number of writes for each flash memory particle within a given time range, the cumulative write count data can be obtained. This data provides the basis for subsequent service life assessment and ensures that the write process of each flash memory particle can be accurately tracked and recorded. For the counting logic, a maximum write count threshold is set. For example, the maximum write count for each flash memory particle can be set to 50,000 times. When this threshold is reached, the particle enters a high-risk state.

[0129] Step S34: Evaluate the service life of the flash memory particles based on the cumulative number of write times to obtain service life data of the flash memory particles;

[0130] In this embodiment, a preset lifespan threshold is set for the flash memory particles. Generally, the preset lifespan of a flash memory particle is 100,000 writes, but the specific threshold can vary depending on the type of flash memory particle (e.g., SLC, MLC, TLC) and the manufacturer's technical standards. Based on this preset lifespan threshold, the system analyzes and evaluates the cumulative number of writes for each flash memory particle. To perform a specific lifespan assessment, it is first necessary to obtain the cumulative number of writes for each flash memory particle. This data is collected by the storage array's health monitoring system (e.g., SMART). The system then compares the cumulative number of writes for each flash memory particle with the preset lifespan threshold. If the cumulative number of writes for a particle approaches or exceeds the preset lifespan threshold, it can be determined that the particle is nearing the end of its service life and entering a higher risk level. Different lifespan intervals can be set based on historical data, the manufacturer's standards for the flash memory particles, or empirical formulas. For example, a percentage range can be used: if the cumulative write count of a flash memory chip is between 0% and 50% of its preset lifespan, it is considered mild wear, indicating a long remaining lifespan; between 50% and 80%, it is moderate wear, indicating the chip is gradually entering a risk phase; and between 80% and 100%, it is severe wear, indicating the chip's lifespan is approaching or exceeding its expected lifespan. Furthermore, in actual operation, these thresholds should be adjusted based on the operating environment. For example, in environments with high temperatures or high loads, the number of writes to flash memory chips will accelerate, necessitating appropriate lowering of the preset lifespan threshold or shortening of the remaining lifespan range. This assessment result can help identify flash memory chips nearing the end of their lifespan, facilitating early replacement or other maintenance measures. By evaluating each chip, flash memory chip lifespan data is generated, which provides important information for subsequent storage system maintenance and fault warnings.

[0131] Step S35: dividing the erase cycle data into flash memory particles. If the erase cycle of the flash memory particle exceeds 1000 times, large-cycle particle data is obtained; if the erase cycle of the flash memory particle is less than 1000 times, small-cycle flash memory particle data is obtained.

[0132] In this embodiment, it is necessary to obtain erase cycle data for each flash memory particle. This data is typically obtained from the storage array's monitoring system or the flash controller's health status log. The data is recorded in units of erase cycles per particle. After obtaining this data, a threshold of 1000 erase cycles is set to classify the flash memory particles. Specifically, the erase cycle data of all flash memory particles is compared against this threshold. If a particle has more than 1000 erase cycles, it is classified as a "high-cycle particle," indicating that it can withstand more erase operations and generally has a longer service life. If a particle has less than 1000 erase cycles, it is classified as a "low-cycle particle," indicating that it reaches its erase cycle limit early during normal use and requires early replacement or maintenance. To ensure accurate classification, the erase cycle data of each particle must be precisely measured. In practice, the erase cycle data of each particle can be queried through the API provided by the flash memory controller or through the storage array's management interface, and the data integrity and accuracy can be verified. During the classification process, special care should be taken to avoid misclassification due to data acquisition errors or storage system anomalies. After the classification is complete, the classification results for each particle are saved as the health status data of the flash memory particle. This data will serve as the basis for subsequent analysis and can be used for tasks such as evaluating the remaining life of the flash memory particles and analyzing erase imbalances. This classification enables more effective management and monitoring of the health status of flash memory particles, allowing particles at risk of failure to be identified in advance and targeted maintenance or replacement.

[0133] Step S36: performing erase imbalance analysis based on the large-cycle flash memory particle data and the small-cycle flash memory particle data to obtain erase imbalance data;

[0134] In this embodiment, erase cycle data for all large-cycle flash memory particles and small-cycle flash memory particles needs to be collected. To analyze erase imbalance, this data is first extracted from the storage array to ensure that the erase cycles of each particle are accurately recorded. Data acquisition can be performed through the flash memory controller's API or management interface to ensure that the data collection process is free of interference and errors. Next, this data is compared and analyzed, focusing on the difference in erase cycles between large-cycle particles and small-cycle particles. During this comparative analysis, a threshold for erase cycle imbalance is set. If the erase cycle of a large-cycle particle is significantly lower than that of a small-cycle particle—that is, the number of erases for a large-cycle particle is significantly less than that for a small-cycle particle—erase imbalance can be determined. To quantify erase imbalance, the standard deviation is used to measure the difference in erase cycles. The specific calculation process is as follows: First, the mean and standard deviation of the erase cycles for large-cycle particles and small-cycle particles are calculated separately. The mean represents the average erase cycle for each particle type, while the standard deviation measures the fluctuation in the erase cycle for each particle type. By comparing the standard deviations of the two types of chips, a significant difference between them indicates significant erase imbalance. To further quantify the extent of erase imbalance, the relative standard deviation of the erase cycle differences between the two types of chips can be calculated. A threshold (for example, 20%) can be set. When the relative standard deviation exceeds this threshold, a severe erase imbalance is considered. Further analysis can be conducted to determine if, for example, some chips experience a heavy workload or are erased more frequently during use, leading to premature endurance. This analysis yields erase imbalance data, helping to assess performance bottlenecks within the storage array. For example, if some chips have significantly fewer erase cycles than others, this indicates premature failure or performance degradation, impacting the lifespan and performance of the entire storage array. This result provides important guidance for subsequent maintenance efforts, allowing adjustments to address the imbalance and adopting appropriate maintenance strategies, such as optimizing load balancing and adjusting write distribution, to ensure the stability and longevity of the array.

[0135] Step S37: Integrate the flash memory particle service life data and the erase imbalance data to generate the flash memory particle wear data.

[0136] In this embodiment, the lifespan assessment results of flash memory particles are combined with the erase imbalance analysis results to generate a detailed flash memory particle wear report for subsequent maintenance and performance monitoring. First, the lifespan data for each flash memory particle is obtained in step S34, and the erase imbalance data for each flash memory particle is obtained in step S36. Lifespan data typically includes the remaining lifespan percentage of each particle, while erase imbalance data reflects the distribution of erase cycles across the particle. When generating the wear data, each flash memory particle is first classified into a wear level. The lifespan classification is based on the particle's remaining lifespan percentage: particles with less than 50% remaining lifespan are labeled "severe wear," particles with a remaining lifespan between 50% and 80% are labeled "moderate wear," and particles with a remaining lifespan above 80% are labeled "slight wear." This classification is typically based on the device manufacturer's recommended values ​​or industry standards. For example, particles with less than 50% remaining lifespan are nearing the end of their lifespan and are prone to failure. Erase imbalance also requires classification. The erase imbalance level of the particle is determined based on the erase cycle difference data obtained in step S36. Chips can be categorized into three categories: low, medium, and high imbalance. For example, a low standard deviation of erase cycles indicates good chip balance and is labeled "low imbalance." A medium standard deviation is labeled "medium imbalance." A high standard deviation, coupled with significant variation in erase cycles, is labeled "high imbalance." By combining the wear level and erase imbalance data for each chip, a comprehensive flash chip wear report is generated through data merging. The report presents each chip's wear level and erase imbalance data in a table format, along with detailed information about each chip (such as its ID, remaining lifespan percentage, erase cycles, and erase imbalance level). The report format should be clear and easy to understand, enabling operators and administrators to quickly understand the health of each chip in the storage array. The generated flash chip wear data report provides valuable support for subsequent storage array maintenance. Administrators can use this report to promptly identify cells with severe wear or erase imbalance, allowing them to take necessary maintenance measures, such as replacing problematic cells, rebalancing storage load, or implementing other optimization measures. This not only ensures the stability and performance of the storage array, but also extends the lifespan of storage devices, reduces failure rates, and avoids system downtime or data loss caused by cell damage.

[0137] Preferably, step S36 is specifically as follows:

[0138] Step S361: performing erasure count statistics based on the large-cycle flash memory particle data to obtain large-cycle erasure count data;

[0139] In this embodiment, the erase cycle data of all large-cycle flash memory particles is obtained from the storage array management system, and flash memory particles with an erase cycle greater than 1000 are screened out. For each large-cycle particle, its erase count is counted. The specific method is to obtain the erase count record of each flash memory particle by reading the log file or health status monitoring data (such as SMART value) provided by the flash memory controller. The erase count of these particles is then accumulated or averaged to obtain the erase count data for the large-cycle particles. This data reflects the durability of the large-cycle particles and their long-term usage. All data should be indexed and stored by particle ID, and the accuracy and integrity of the original data should be maintained.

[0140] Step S362: performing erasure count statistics based on the small-cycle flash memory particle data to obtain small-cycle erasure count data;

[0141] In this embodiment, the erase cycle data for all short-cycle flash memory particles needs to be extracted from the storage array management system. These short-cycle particles are defined as flash memory particles with an erase cycle of less than 1000 times. The core of this step is to filter out particles that meet this condition and obtain their erase count records. By accessing the storage array management system, the storage array's API interface or health status monitoring system (such as SMART values) is used to obtain detailed health data for each flash memory particle. This data includes the erase count of the flash memory cell, which is typically recorded in the flash memory controller's log file. This information is usually stored in the form of a log, containing information such as the erase count and write count of the flash memory particle, and each record has a unique particle ID for easy tracking and management. 1000 times. By writing a filtering program or using a database query statement, filter out the data for all short-cycle particles. The erase count of these particles should be extracted by parsing the data in the log file. The erase count of each particle is the specific value recorded by the flash memory controller. Next, count the erase counts of the short-cycle particles. Specifically, for all eligible short-cycle flash memory cells, the system can choose to sum their erase counts to obtain the total erase count. Alternatively, the system can calculate the average erase count to obtain the average erase count per cell. To further analyze the distribution of cells, the standard deviation of the erase counts can be calculated to assess the distribution of erase counts. To ensure statistical accuracy and consistency, all erase count data should be indexed and stored by cell ID. The cell ID is a unique identifier for each cell, used to associate each cell with its erase count data, ensuring accurate storage and tracking of each cell's data. All collected data should be stored in a database or table, sorted and organized by cell ID, erase count, and other relevant information (such as storage cell location). All the erase count data collected is compiled into a report on the erase count of short-cycle flash memory cells. This report provides foundational data for subsequent analysis, helping to assess the endurance differences of short-cycle flash memory cells and, in turn, estimate their service life and replacement time. It also provides a reference for analyzing the health of the storage array and identifying flash memory cells with potential problems or short lifespans.

[0142] Step S363: Draw a distribution map of the database virtual hard disk based on the large-cycle erasure count data and the small-cycle erasure count data to obtain a cycle erasure distribution map;

[0143] In this embodiment, data on the number of erases per large cycle and the number of erases per small cycle are collected. By extracting this data from a database, the complete and accurate erase count data for each particle is ensured. The data should be categorized and stored by particle type (large cycle or small cycle) so that different types of particles can be distinguished when plotting a distribution graph. To plot the cycle-by-erasure distribution graph, statistical software is required for data visualization. For this step, suitable data processing tools can be used, such as Excel, Matplotlib in Python, or Seaborn. These tools support data processing and graphing. Before plotting, the erase count data is first organized into two datasets based on particle type (large cycle and small cycle). Each dataset should include the erase count and the corresponding particle ID. This step separates the erase count data for different categories for accurate presentation in the distribution graph. Next, the data is visualized using libraries such as Matplotlib or Seaborn in Python. The horizontal axis (X-axis) represents the erase count, and the vertical axis (Y-axis) represents the number or ratio of different flash memory particles. Data points for large-cycle particles should be concentrated in regions with higher erase counts, such as those with more than 1,000 erases. Data points for small-cycle particles should be distributed in regions with lower erase counts, such as those with less than 1,000 erases. Each data point can be resized to represent the actual number or proportion of particles. For example, the data point size can be larger for particles with relatively high erase counts (i.e., large-cycle particles), while the data point size can be smaller for particles with fewer erase counts (i.e., small-cycle particles). This helps clearly distinguish between large-cycle and small-cycle particles on the graph and allows the graph to reflect the distribution of each type of particle across different erase count ranges. To better present the data, color coding can be added to the distribution graph to distinguish large-cycle and small-cycle particles. For example, blue could be used for large-cycle particles and red for small-cycle particles, or different marker symbols could be used to distinguish the two types of data points. This allows the difference in erase counts between large-cycle and small-cycle particles to be clearly displayed on the same graph, helping analysts quickly identify distribution trends between the two. After plotting the erase cycle distribution graph, further analysis should be performed to observe the differences in erase cycle distribution between particles with large and small cycles. This graph can intuitively demonstrate the erase cycle characteristics of each particle in the storage array, helping managers identify potential performance bottlenecks, particle health and lifespan issues, and providing a reference for subsequent maintenance decisions.

[0144] Step S364: performing flash memory particle imbalance calculation on the periodic erase distribution diagram to obtain erase imbalance data.

[0145] In this embodiment, the mean and standard deviation of the erase counts for large-cycle particles and small-cycle particles are calculated separately. Then, by comparing these two standard deviations, the degree of erase imbalance is obtained. If the standard deviation of the erase counts for large-cycle particles is significantly smaller than that for small-cycle particles, it indicates that the erase cycle distribution of the large-cycle particles is relatively balanced. Conversely, if the standard deviation is large, it indicates that there is a more serious erase imbalance. In addition, the overall mean of the erase counts for the two types of particles can be calculated, and by comparing the difference in the mean, the erase cycle distribution of the two types of particles can be further analyzed. The erase imbalance data provides a key indicator for the performance evaluation of the storage array, reflecting the usage status and health distribution of different types of flash memory particles.

[0146] Preferably, this specification also provides a disk-based big data storage system for executing the disk-based big data storage method described above, the disk-based big data storage system comprising:

[0147] The data encryption storage module obtains the storage capacity of the slave database; divides the data into blocks based on the storage capacity of the slave database to generate a first block node to be stored; obtains the big data to be stored, converts the big data to be stored into a request transaction data packet, and sends it to the first block node to be stored; during the sending process, the request transaction data packet is broadcast to other block nodes to be stored; and using a preset consensus node, the data of other block nodes to be stored is compared with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the node consistent data is encrypted and stored in the slave database;

[0148] The storage array fault analysis module uses monitoring tools to obtain the slave database disk status data in real time; performs storage array fault analysis on the slave database disk status during the encryption storage process and generates slave database storage array fault data;

[0149] The flash memory particle loss analysis module uses a disk imaging tool to scan the slave database hard disk; based on the storage array failure data, it performs flash memory particle loss analysis on the slave database hard disk to obtain flash memory particle loss data;

[0150] The thermal protection mechanism setting module sets the thermal protection mechanism for the slave database hard disk based on the flash memory particle loss data and generates thermal protection mechanism data; uploads the thermal protection mechanism data to the slave database, and stores the slave database data in the master database to perform large data storage tasks.

Claims

1. A disk-based big data storage method, characterized in that: The following steps are involved: Step S1: Obtain storage capacity from the database; Divide the data into blocks from the database storage capacity to generate a first block node to be stored; Obtaining the big data to be stored, converting the big data to be stored into a transaction request data packet and sending it to the first block node to be stored; During the sending process, the request transaction data packet is broadcast to other nodes where the block is to be stored; Use the preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the node consistent data is encrypted and stored in the slave database; Step S2: using a monitoring tool to obtain slave database disk status data in real time; performing storage array failure analysis on the slave database disk status during the encryption storage process to generate slave database storage array failure data; Step S3: Scan the slave database hard disk using a disk imaging tool to obtain a slave database virtual hard disk; perform flash memory particle loss analysis on the slave database virtual hard disk based on the storage array failure data to obtain flash memory particle loss data; Step S4: Setting a thermal protection mechanism for the slave database virtual hard disk based on the flash memory particle loss data to generate thermal protection mechanism data; uploading the thermal protection mechanism data to the slave database, and storing the slave database data in the master database to perform the big data storage task.

2. The disk-based big data storage method according to claim 1, characterized in that: Step S1 is specifically as follows: Step S11: Obtain storage capacity from the database; Step S12: vertically partition the storage capacity of the secondary database, wherein the size of the sub-database in the vertical partition is set to not exceed 5TB, and generate the vertical partition data to be stored; Step S13: vertically partitioning the table based on the vertical partitioned database data, wherein the vertical partitioned table size is set to no more than 2TB, and the vertical partitioned table data to be stored is obtained; Step S14: Integrate the vertical sub-library data to be stored and the vertical sub-table data to be stored to obtain the first block node to be stored; Step S15: Using the first block node to be stored, the storage capacity of the slave database is divided into storage block nodes. The storage capacity of the slave database is divided according to a fixed ratio of 1:

1. The capacity of each node is controlled within 2TB-4TB, and other block nodes to be stored are obtained. Step S16: Obtain the big data to be stored, convert it into a transaction request data packet, and send it to the first block node to be stored. The size of each data packet is fixed to within 500MB. During the sending process, a dedicated transmission protocol (TCP / IP) is used to broadcast the transaction request data packet at a fixed broadcast rate of 500MB / s to other block nodes to be stored. Step S17: Use the preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node, and the repair error threshold is set to 0.01% until the data of each node is consistent. The node consistent data is encrypted and stored in the slave database.

3. The disk-based big data storage method according to claim 2, characterized in that: Step S17 is specifically as follows: Step S171: Calculate the data digests of the other block nodes to be stored and the first block node to be stored, wherein the calculation time is set to be controlled within 2 seconds, and obtain the data digests of the other block nodes to be stored and the data digest of the first block node to be stored; Step S172: Integrate the data digests of the other block nodes to be stored and the data digest of the first block node to be stored to obtain a data digest, and broadcast the data digest to the preset consensus node with a timeout of 5 seconds to obtain the consensus node data digest; Step S173: Use the consensus node data digest to determine the consistency of the data digests of other to-be-stored block nodes and the first to-be-stored block node. If the data digest difference exceeds 5%, an inconsistent data digest is obtained, and the first to-be-stored block node is used to initiate a consensus request to the consensus node. The consensus node responds with the voting result within 5 seconds. When at least 3 / 4 of the nodes that meet the protocol requirements agree to the data update, consensus is confirmed and the node consistent data is obtained. Step S174: using the node consistent data to repair the inconsistent data digest to obtain a repaired data digest; Step S175: Encrypt the repair data summary to generate an encrypted data packet, and store it in the slave database. The encryption process takes no more than 10 seconds when the data packet size is set to 500MB.

4. The disk-based big data storage method according to claim 1, characterized in that: Step S2 is specifically as follows: Step S21: using a monitoring tool to obtain disk status data from the database at a sampling period of once per second, including I / O performance data and disk health status data; Step S22: performing storage array fault analysis on the I / O performance data to obtain I / O performance fault data; Step S23: performing storage array fault analysis on the disk health status data to obtain disk storage array fault data; Step S24: Integrate the I / O performance fault data and the disk storage array fault data to generate secondary database storage array fault data.

5. The disk-based big data storage method according to claim 4, characterized in that: Step S22 is specifically as follows: Step S221: When any of the following situations occurs, it is determined that the storage array read / write performance is abnormal, and storage array read / write performance abnormality data is obtained: the sequential read / write speed drops by more than 20%, the random access latency exceeds 30% of the normal range, and the I / O queue depth is in a full load state for a long time and lasts for more than 60 minutes; Step S222: When the following conditions occur simultaneously, it is determined that the storage array controller is faulty and storage array controller fault data is obtained: the controller CPU utilization exceeds 90% for a long period of time and the response time is abnormally prolonged; the cache hit rate drops by more than 25%; the communication error frequency between the controller and the disk continues to increase and cannot be restored to normal within 10 minutes; Step S223: Integrate the storage array read and write performance abnormality data and the storage array controller fault data to obtain I / O performance fault data.

6. The disk-based big data storage method according to claim 4, characterized in that: Step S23 is specifically as follows: Step S231: extracting disk rotation speed features from the disk health status data to obtain disk rotation speed data; Step S232: Obtaining disk standard rotation speed range data using disk technical documentation; Step S233: identifying abnormal fluctuations in the disk rotation speed data based on the disk standard rotation speed range data. If the rotation speed fluctuation exceeds ±5% of the disk standard rotation speed range, generating abnormal rotation speed fluctuation data. Step S234: evaluating the physical loss of the disk based on the abnormal rotation speed fluctuation data, thereby obtaining disk loss data; Step S235: performing storage array fault determination on the disk health status data based on the disk loss data. When the disk loss score exceeds 60%, disk storage array fault data is obtained.

7. The disk-based big data storage method according to claim 6, characterized in that: Step S234 is specifically as follows: Count the frequency of abnormal speed fluctuation data, where the sampling period is set to collect data once per minute to obtain frequency data; Count the fluctuation time of abnormal speed fluctuation data, where the fluctuation time greater than 5 seconds is set as effective fluctuation, and obtain the fluctuation time data; A disk physical loss assessment model is constructed based on frequency data and fluctuation time data, where the weight of the impact of frequency on loss is set to 70% and the weight of the impact of fluctuation time on loss is set to 30%; The disk physical loss assessment model is used to assess the loss of abnormal rotation speed fluctuation data: 31%-50% is considered mild loss data, 51%-70% is considered moderate loss data, and 71%-100% is considered severe loss data. The disk loss data is obtained by integrating the slight loss data, the moderate loss data and the severe loss data.

8. The disk-based big data storage method according to claim 1, characterized in that: Step S3 is specifically as follows: Step S31: obtaining an image file of the slave database virtual hard disk, and parsing the image file using a disk imaging tool, thereby obtaining the slave database virtual hard disk; Step S32: extracting the number of write times and the number of erase cycles of the flash memory based on the storage array fault data to obtain the number of write times and the number of erase cycles of the flash memory; Step S33: accumulating the number of write times of the flash memory to obtain the cumulative number of write times data; Step S34: Evaluate the service life of the flash memory particles based on the cumulative number of write times to obtain service life data of the flash memory particles; Step S35: dividing the erase cycle data into flash memory particles. If the erase cycle of the flash memory particle exceeds 1000 times, large-cycle particle data is obtained; if the erase cycle of the flash memory particle is less than 1000 times, small-cycle flash memory particle data is obtained. Step S36: performing erase imbalance analysis based on the large-cycle flash memory particle data and the small-cycle flash memory particle data to obtain erase imbalance data; Step S37: Integrate the flash memory particle service life data and the erase imbalance data to generate the flash memory particle wear data.

9. The disk-based big data storage method according to claim 8, characterized in that: Step S36 is specifically as follows: Step S361: performing erasure count statistics based on the large-cycle flash memory particle data to obtain large-cycle erasure count data; Step S362: performing erasure count statistics based on the small-cycle flash memory particle data to obtain small-cycle erasure count data; Step S363: Draw a distribution map of the database virtual hard disk based on the large-cycle erasure count data and the small-cycle erasure count data to obtain a cycle erasure distribution map; Step S364: performing flash memory particle imbalance calculation on the periodic erase distribution diagram to obtain erase imbalance data.

10. A disk-based big data storage system, characterized in that: For executing the disk-based big data storage method according to claim 1, the disk-based big data storage system comprises: The data encryption storage module obtains the storage capacity of the slave database; divides the data into blocks based on the storage capacity of the slave database to generate a first block node to be stored; obtains the big data to be stored, converts the big data to be stored into a request transaction data packet, and sends it to the first block node to be stored; during the sending process, the request transaction data packet is broadcast to other block nodes to be stored; and using a preset consensus node, the data of other block nodes to be stored is compared with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the node consistent data is encrypted and stored in the slave database; The storage array fault analysis module uses monitoring tools to obtain the slave database disk status data in real time; performs storage array fault analysis on the slave database disk status during the encryption storage process and generates slave database storage array fault data; The flash memory particle loss analysis module uses a disk imaging tool to scan the slave database hard disk; based on the storage array failure data, it performs flash memory particle loss analysis on the slave database hard disk to obtain flash memory particle loss data; The thermal protection mechanism setting module sets the thermal protection mechanism for the slave database hard disk based on the flash memory particle loss data and generates thermal protection mechanism data; uploads the thermal protection mechanism data to the slave database, and stores the slave database data in the master database to perform large data storage tasks.

Citation Information

Patent Citations

  • A block chain-based quality data processing method and device based on a block chain

    CN109598505A

  • Rebuilding portions of virtual segments based on power

    US20240330305A1