Big data storage method and system based on disk

By performing data chunking and consensus processing in the storage system, data consistency is ensured, and thermal protection mechanism is set in combination with real-time monitoring and loss analysis, the thermal protection mechanism is solved, and the problems of limited capacity expansion, degraded performance and insufficient thermal management in traditional storage methods are achieved, thereby achieving efficient, reliable and secure data storage.

CN120144584AActive Publication Date: 2025-06-13邓志敏

Patent Information

Application Number
CN202510156415.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-13
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

Traditional storage methods are difficult to flexibly cope with the rapid growth of data volume, resulting in limited capacity expansion capabilities of storage systems, and ignore the loss and erasing cycles of flash memory particles, resulting in performance degradation and storage media damage, and lack of effective thermal management solutions.

Method used

By obtaining the storage capacity from the database and chunking data, generating block nodes to be stored, using consensus nodes to ensure data consistency, and encrypting storage. Monitor disk status in real time, perform storage array failure analysis and flash memory particle loss analysis, set thermal protection mechanism, and optimize the temperature management of the storage system.

Benefits of technology

It improves the reliability and stability of data storage, enhances the fault tolerance and data security of the storage system, extends the service life of the storage device, and avoids problems caused by limited capacity expansion, performance degradation and overheating.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144584A_ABST
    Figure CN120144584A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data storage, in particular to a disk-based big data storage method and system. The method comprises the following steps: acquiring the storage capacity of a slave database; performing data partitioning on the storage capacity of the slave database to generate a first to-be-stored block node; obtaining to-be-stored big data, converting the to-be-stored big data into a request transaction data packet, and sending the request transaction data packet to the first to-be-stored block node; the request transaction data packet is broadcasted to other to-be-stored block nodes in the sending process; performing data comparison on other to-be-stored block nodes and the first to-be-stored block node by using a preset consensus node, when the data is inconsistent, performing consensus processing based on the consensus node until the data of each node is consistent, and encrypting and storing the node consistent data to a slave database; and acquiring disk state data of the slave database in real time by using a monitoring tool. The big data storage efficiency is optimized based on the big data storage technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data storage, and particularly relates to a disk-based big data storage method and system. Background Art

[0002] Traditional storage methods usually adopt fixed storage structures and capacities, making it difficult to flexibly cope with the rapid growth of data volume, resulting in limited capacity expansion capabilities of the storage system. When the data volume increases sharply, traditional systems often struggle to provide sufficient storage space, leading to insufficient storage system resources and affecting data processing efficiency and stability. In systems using flash memory storage media, traditional methods usually ignore key factors such as the wear and tear of flash memory particles and the erase cycle. The service life of flash memory particles is limited, and traditional methods do not have effective wear assessment and maintenance mechanisms, which causes the performance of the data storage system to decline after long-term operation, and even the storage media to be damaged, affecting the reliability of data storage. Traditional storage systems often do not provide effective solutions for the thermal management of storage devices. Especially in the case of high-load and big data storage, the lack of a dynamic thermal protection mechanism easily leads to overheating of the storage device, which in turn damages the hard disk or storage chip, affecting the operation efficiency and stability of the entire data storage system. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a disk-based big data storage method and system to solve at least one of the above technical problems.

[0004] To achieve the above object, a disk-based big data storage method includes the following steps: Step S1: Obtain the storage capacity of the slave database; perform data chunking on the storage capacity of the slave database to generate the first block node to be stored; obtain the big data to be stored, convert the big data to be stored into a request transaction data packet and send it to the first block node to be stored; broadcast the request transaction data packet to other block nodes to be stored during the sending process; use a preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the consistent data of the nodes is encrypted and stored in the slave database; Step S2: Use a monitoring tool to obtain the disk status data of the slave database in real time; perform storage array fault analysis on the disk status of the slave database during the encryption storage process to generate slave database storage array fault data; Step S3: Use a disk imaging tool to scan the hard disk of the slave database to obtain a virtual hard disk of the slave database; perform flash memory particle wear analysis on the virtual hard disk of the slave database according to the storage array fault data to obtain flash memory particle wear data; Step S4: Set the thermal protection mechanism for the virtual hard disk of the slave database based on the flash memory particle loss data to generate thermal protection mechanism data; upload the thermal protection mechanism data to the slave database, and store the slave database data in the master database to execute the big data storage task.

[0005] By obtaining the storage capacity of the slave database and performing data chunking, the present invention can convert the big data to be stored into request transaction data packets, and ensure data consistency through the broadcast and consensus mechanisms of multiple nodes, thereby effectively improving the reliability and stability of data storage. By presetting the consensus nodes to handle data inconsistencies between different nodes, it can ensure that the finally stored data is highly consistent among multiple nodes, effectively reducing the risks brought by data loss or inconsistency, and enhancing the fault tolerance of the storage system. The encrypted storage process can not only enhance the security of data, but also ensure the privacy of sensitive data, avoiding the risk of data leakage. By real-time monitoring the disk status data of the slave database and combining with fault analysis, potential hardware failures can be identified and predicted in advance, and preventive measures can be taken in time, thereby reducing the risk of data loss or service interruption caused by storage array failures. Analyzing the flash memory particle loss of the virtual hard disk helps to understand the health status of the storage medium and predict the hard disk life, which can provide an effective basis for subsequent hardware maintenance and avoid performance degradation caused by excessive particle loss. By setting the thermal protection mechanism, it can not only optimize the temperature management of the storage system, avoid overheating damage to the hard disk or storage chip, but also dynamically adjust the operating conditions of the storage device, extending the service life of the storage device. By uploading the thermal protection mechanism data to the slave database and synchronizing it to the master database, it ensures that the monitoring data of the system can be effectively stored and used for subsequent analysis, thus supporting more efficient big data storage and processing. This method can fully guarantee the stability, reliability and efficiency of the storage system in big data storage and processing tasks, avoiding problems such as limited capacity expansion, performance degradation and overheating faced by traditional storage methods, and providing a flexible, efficient and secure storage solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Other features, objects and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings: Figure 1 It is a schematic flow chart of the steps of the big data storage method based on disk of the present invention; Figure 2 It is a detailed schematic flow chart of step S2 in the present invention; The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0007] The technical method of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention.

[0008] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.

[0009] It should be understood that although terms such as "first" and "second" may be used here to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be referred to as the second unit, and similarly the second unit can be referred to as the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed associated items.

[0010] To achieve the above object, please refer to Figures 1 to 2 , the present invention provides a method for storing big data based on a disk. The method includes the following steps: Step S1: Obtain the storage capacity of the slave database; perform data chunking on the storage capacity of the slave database to generate the first block node to be stored; obtain the big data to be stored, convert the big data to be stored into a request transaction data packet and send it to the first block node to be stored; broadcast the request transaction data packet to other block nodes to be stored during the sending process; use a preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the consistent data of the nodes is encrypted and stored in the slave database; In this embodiment, storage capacity data is obtained from the database to determine the total storage space size of the database. The unit of this storage space is GB. Then, the storage capacity is divided according to certain rules (for example, the size of each data block is set to 1 GB). Specifically, the entire storage space is divided into multiple storage blocks of a fixed size, and these storage blocks are called "block nodes". Each block node represents a storage unit and will be used to store the data to be stored. When the large data to be stored arrives, data splitting technology is used to split the data into multiple smaller data packets, ensuring that the size of each data packet does not exceed the size of a single data block (such as 1 GB). The core of this step is to perform reasonable splitting according to the size of the data packet to avoid a data packet exceeding the predetermined storage block size, thereby ensuring the efficient storage of data. Each data packet will be marked according to the block node it belongs to after splitting, ensuring that each data packet can accurately identify its target storage location. All data packets will be simultaneously sent to multiple storage block nodes to be stored through broadcast technology. These block nodes form a multi-node network, and the reliability and fault tolerance of storage are improved through distributed storage. Broadcast technology ensures that all data packets can be quickly and parallelly spread to multiple nodes, thereby realizing data distribution in a multi-node system. To ensure data consistency and correctness, configured consensus nodes are used to coordinate the consistency verification of data packets among block nodes. When a node receives a data packet, it will compare it with the data of other nodes to ensure that the content stored in all nodes is consistent. If it is found that the content of some data packets is inconsistent (such as data loss or transmission error) during the comparison process, a consensus process will be carried out through an algorithm based on a preset consensus mechanism. Commonly used consensus algorithms include the Paxos protocol or the Raft protocol. These algorithms ensure consistency through the voting and confirmation processes among multiple nodes, thereby avoiding data divergence. Once the consensus process is completed, all block nodes will encrypt and store the data. To ensure data security, a strong encryption algorithm AES-256 (Advanced Encryption Standard 256-bit key) is used to encrypt the data. AES-256 is a relatively secure and efficient encryption standard at present, and through it, the stored data can be effectively protected from unauthorized access. After being encrypted, the data will be stored in the storage area of the slave database to ensure the secure storage of the data and have a certain anti-tampering ability.

[0011] Step S2: Use a monitoring tool to obtain the disk status data of the slave database in real time; perform a storage array failure analysis on the disk status of the slave database during the encryption storage process to generate slave database storage array failure data; In this embodiment, a real-time monitoring tool is used to monitor the disk status data of the slave database. The monitoring tool is connected to the disk controller interface to obtain in real time data such as the hard disk health status, temperature, read / write speed, power consumption, etc. of the slave database. Specifically, the monitoring tool reads the SMART (Self-Monitoring, Analysis, and Reporting Technology) information of each hard disk to obtain indicators such as the operating temperature, read / write error rate, remaining life, etc. of the hard disk. Based on these indicators, a storage array fault analysis is performed to determine whether there are signs of an impending hard disk failure. If the monitoring tool detects an abnormality in a certain hard disk (such as the temperature exceeding 65°C or multiple read / write errors occurring), a fault data record is generated, including information such as the ID of the faulty hard disk, the time of the fault occurrence, and the type of the fault. These fault data records will be used to further analyze the potential problems of the storage array and provide a data basis for subsequent flash particle wear analysis.

[0012] Step S3: Use a disk imaging tool to scan the hard disk of the slave database to obtain the virtual hard disk of the slave database; perform flash particle wear analysis on the virtual hard disk of the slave database according to the storage array fault data to obtain flash particle wear data; In this embodiment, a disk imaging tool (such as Clonezilla) is used to scan the hard disk of the slave database to generate a virtual image file of the hard disk. The disk imaging tool will comprehensively scan all partitions of the hard disk of the slave database, record the logical partition table of the hard disk and the data content of each partition, and generate a virtual hard disk copy of the slave database. During the specific operation process, the disk imaging tool will obtain the physical address and data content of the hard disk from the controller of the hard disk, and read the data block by block according to the block size of the hard disk (such as 512 bytes or 4KB) to generate a virtual image file containing all the hard disk content. According to the storage array fault data of the slave database, the wear condition of the flash particles is analyzed. Here, a dedicated hard disk analysis tool (such as SSD Life) is used to read information such as the Write Amplification Factor (WAF), the number of erase / write cycles, and the number of bad blocks of the hard disk to evaluate the wear condition of the flash particles. By comparing the usage conditions of different hard disks, the flash particle wear data of each hard disk is calculated, and a wear analysis report is generated. These wear data provide a basis for subsequent setting of the thermal protection mechanism.

[0013] Step S4: Set a thermal protection mechanism for the virtual hard disk of the slave database based on the flash particle wear data to generate thermal protection mechanism data; upload the thermal protection mechanism data to the slave database, and store the slave database data in the master database to perform the big data storage task.

[0014] In this embodiment, a temperature threshold of the hard disk is set according to the wear data of the flash memory particles. For example, it is set that when the hard disk temperature exceeds 70 °C, the thermal protection mechanism is immediately activated. This temperature threshold is based on the technical specifications provided by the hard disk manufacturer to ensure that the hard disk can be effectively protected when the temperature exceeds a certain range, preventing accelerated wear of the flash memory particles due to excessive temperature. The thermal protection mechanism includes dynamically adjusting the workload of the hard disk, reducing the write frequency, or activating a cooling system (such as an external fan or coolant) to control the hard disk temperature. The temperature data of the hard disk is obtained in real time through a monitoring tool and compared with the set threshold. When the temperature reaches the preset threshold, the protection measures are automatically activated. The generated thermal protection mechanism data is uploaded to the slave database to store the detailed parameters of the mechanism (such as the temperature threshold, protection mode, etc.). Then, the data from the slave database is stored in the master database according to the encryption protocol to perform the big data storage task. Specifically, a dual-encryption storage method is adopted. On the one hand, the data is encrypted by AES-256, and on the other hand, a hash algorithm (such as SHA-256) is used for data integrity verification to ensure the security and reliability of the data during storage.

[0015] Particularly importantly, step S4 includes the following steps: Step S41: Obtain historical flash memory particle wear data; In this embodiment, the health data of the storage device is obtained through a dedicated tool (such as a SMART monitoring tool or an API provided by the hard disk manufacturer). These data include but are not limited to the number of erase / write cycles, the erase / write cycles of each particle, temperature records, error counts, read / write counts, etc. The number of erase / write cycles data is recorded by the hard disk controller, and the number of erase / write cycles of each flash memory particle is usually read through the interface between the log record of the flash memory management unit and the monitoring software. In actual operation, the smartctl command is used to extract historical data from the S.M.A.R.T. log of the hard disk. By comparing the number of erase / write cycles of each particle with the current state, the accuracy of the data is ensured. The key parameters of this step include the "number of erase / write cycles threshold" (for example, when the number of erase / write cycles is greater than 10,000 times, it indicates approaching the end of life) and the "temperature threshold" (for example, when the temperature exceeds 70 °C, it starts to enter the attention state).

[0016] Step S42: Construct a wear model based on the historical flash memory particle wear data and the flash memory particle wear data, thereby obtaining the wear model; In this embodiment, historical data is first sorted out, and the number of erase / write cycles, temperature data, and particle status of each flash memory particle are extracted. Then, linear regression or multiple regression analysis methods are used to fit the historical data to predict the wear trend of the flash memory particles. To construct the wear model, key factors related to wear need to be selected, such as the number of erase / write cycles, operating temperature, power consumption, etc. Assume that for every additional 1,000 erase / write cycles, the particle lifespan decreases by 10%. In this way, the relationship between the wear factor and the particle lifespan is set. The wear model also includes a standard formula for calculating the remaining lifespan, such as L = (Lmax - E) * (1 - T / Tmax), where Lmax is the maximum lifespan, E is the consumed lifespan, T is the current temperature, and Tmax is the maximum operating temperature. The output of the wear model will provide predictions of the remaining lifespan for each particle.

[0017] Step S43: Predict the lifespan of the flash memory particle wear data according to the wear model to obtain lifespan data; In this embodiment, parameters such as the number of erase / write cycles, operating temperature, and power consumption of each particle are input and calculated through the model formula. For each particle, first, its consumed lifespan (such as the wear caused by the number of erase / write cycles and operating temperature) is compared with the preset parameters in the model, and then its remaining lifespan is calculated according to the model. For example, if a certain particle has been erased / written 8,000 times and the current temperature is 65°C, assuming the model parameter is that the lifespan decreases by 1% for every additional 1,000 erase / write cycles and decreases by 0.1% for every 1°C increase in temperature, then according to the calculation, the lifespan will decrease by 8% + 0.5% (temperature wear). Finally, the remaining lifespan data for each particle is calculated. The lifespan data will be used as the basis for subsequent judgment of the hard disk temperature control strategy and write control.

[0018] Step S44: Monitor the temperature of the virtual hard disk from the database to obtain temperature data; In this embodiment, temperature data needs to be obtained through the hard disk controller or a dedicated temperature sensor. The temperature sensor data provided by the hard disk controller (usually part of the S.M.A.R.T. data) can be obtained through the smartctl command to read the operating temperature of the hard disk in real time. In addition, if a virtualization platform such as VMware or Hyper-V is used, relevant temperature data can also be obtained through the virtual hard disk control interface. To ensure data accuracy, the temperature sampling frequency needs to be set, for example, the temperature data is collected once a minute, and the maximum, minimum, and average temperatures are recorded. The key parameters at this time are the temperature upper limit threshold (for example, 80°C is the high-temperature warning threshold), and the temperature standard error of each particle (for example, ±2°C). The monitoring system should set an automatic alarm function to send an alarm signal when the temperature exceeds the set threshold.

[0019] Step S45: Perform statistics based on the temperature data and the life data respectively to obtain high-temperature data and low-life data, and calculate the time overlap degree of the high-temperature data and the low-life data to obtain high-risk overlap time data; In this embodiment, based on the temperature data obtained from step S44 and the life data of step S43, high-temperature data exceeding the set threshold and data with a remaining life lower than the set threshold are respectively counted. For the temperature data, if the temperature of a certain particle exceeds 80°C (for example, the set high-temperature threshold), it is recorded as high-temperature data. For the life data, if the remaining life of the particle is lower than 20% (set low-life threshold), it is recorded as low-life data. Then, based on the overlap of these two types of data within a time period, the time overlap degree is calculated to find the overlap time of the particles with too high temperature and approaching end of life. For example, if a certain particle has a temperature exceeding 80°C and a remaining life lower than 20% within the past 30 minutes, then these 30 minutes are recorded as high-risk overlap time data. The calculation method of the overlap degree is: high-risk overlap time = temperature over-standard time + time with life lower than the threshold (unit: minute).

[0020] Step S46: Based on the high-risk overlap time data, start the cooling mechanism for the database virtual hard disk to obtain cooling mechanism data; In this embodiment, a cooling start threshold is set. For example, if the high-risk overlap time exceeds 10 minutes, the cooling mechanism is immediately started. The specific implementation of the cooling mechanism is to reduce the temperature of the hard disk through a temperature control system, such as fan adjustment, air conditioning equipment, liquid cooling system, etc. The start and stop of the cooling equipment are controlled through a monitoring tool. The cooling equipment should adjust its working state in real time according to the temperature data, and record information such as the start and stop time and duration of each cooling, which is used as the basis for storing and analyzing the cooling mechanism data.

[0021] Step S47: Based on the high-risk overlap time data, start the low data write frequency mechanism for the database virtual hard disk to obtain low data write frequency mechanism data; In this embodiment, for the high-risk overlap time data, if the high-risk overlap time exceeds the set threshold (for example, more than 10 minutes), the low data write frequency mechanism is started. Specifically, when implementing, the write rate of the virtual hard disk is adjusted through the hard disk controller or the operating system. For example, the number of write requests per second is limited, or the concurrency of data writing is reduced. This method can reduce the burden on the flash particles, thereby slowing down the temperature rise and loss caused by frequent writing. During implementation, a low write frequency threshold needs to be set, usually limited to 50% or lower of the normal write rate. This strategy is implemented through the file system of the operating system or the firmware of the hard disk controller, and the time period when the low data write frequency is enabled and the write rate are recorded.

[0022] Step S48: Integrate the data of the low data write frequency mechanism and the cooling mechanism data to obtain the thermal protection mechanism data; In this embodiment, the start time, duration, and cooling effect of the cooling mechanism are merged with the start time, rate, etc. of the low write frequency through a data aggregation tool (such as a database or a log analysis system). By analyzing these data, a unified thermal protection mechanism strategy can be formed, which includes two means: cooling and write control. The thermal protection mechanism data will provide the system administrator with the real-time protection status and can be used for subsequent optimization and adjustment.

[0023] Step S49: Upload the thermal protection mechanism data to the slave database and store the slave database data in the master database to perform the big data storage task.

[0024] In this embodiment, it is necessary to ensure that the thermal protection mechanism data has been generated in step S48 and is available for uploading. During the implementation process, the thermal protection mechanism data is transmitted to the slave database through a network interface (such as the TCP / IP protocol). At this time, the data transmission protocol used should ensure the integrity of the data and the stability of the transmission. Usually, a secure encryption transmission protocol (such as HTTPS or SSL) can be adopted to ensure that the data is not tampered with or lost during the transmission process. The transmitted data should include information such as the enabling time of the cooling mechanism, the duration, the adjustment time of the writing frequency, and the change of the writing rate. An appropriate verification mechanism should be set during data transmission, such as packet verification (e.g., CRC or MD5), to confirm that the data has not been lost or damaged during the transmission process. Once the thermal protection data is successfully uploaded to the slave database, the next operation is to synchronize the data in the slave database to the master database. This step requires setting an appropriate synchronization strategy according to a specific database management system (such as MySQL, PostgreSQL, etc.). Generally speaking, database synchronization is divided into two methods: real-time synchronization and periodic synchronization. For real-time synchronization, the replication function of the database can be used to keep the slave database and the master database in real-time consistency; while periodic synchronization usually uses a scheduled task (such as a cron job) to synchronize data in batches regularly. The synchronization process will adopt an incremental update method to reduce the transmission volume and improve efficiency. The incremental synchronization mechanism detects which data has changed by comparing the timestamps or version numbers of the data in the source database and the target database, and only synchronizes these changed data, thereby reducing unnecessary data transmission. When the data in the slave database changes, the synchronization mechanism of the master database will be automatically triggered to upload the data to the master database. The master database usually sets a data storage strategy to ensure sufficient storage space and high availability. Adopting a distributed storage architecture for data storage can better adapt to the management of large amounts of data and ensure the reliable storage of data. The data stored in the master database can also support subsequent big data analysis, real-time query, report generation, data mining and other tasks. Appropriate backup and recovery strategies should be configured during the data storage process to prevent data loss or damage and ensure the integrity and security of the data.

[0025] Preferably, step S1 is specifically as follows: Step S11: Obtain the storage capacity of the slave database; In this embodiment, the storage capacity of the slave database is obtained. To obtain the accurate storage capacity of the database, the system tables of the database management system (such as the INFORMATION_SCHEMA table in MySQL and the sys.master_files in SQL Server) can be used for query. The query process is carried out by executing the SQL command SELECT SUM(data_length) FROM information_schema.tables, and this command will return the total storage space size of all data tables in the current database. Suppose the query result shows that the storage capacity of the slave database is 20TB. At this time, this storage capacity data is recorded and used as the basis for subsequent database sharding and data partitioning. After the storage capacity data is obtained in this step, it will be passed to the subsequent database sharding and table partitioning processing procedures.

[0026] Step S12: Vertically shard the slave database storage capacity, where the size of each sharded database is set not to exceed 5TB, and generate the vertical sharded database data to be stored; In this embodiment, according to the storage capacity of the previous step (such as 20TB), the slave database is split into multiple sharded databases. The maximum capacity of each sharded database is 5TB, aiming to avoid the performance degradation or increased management complexity caused by the excessive capacity of each database. Using the database partitioning technology (such as the PARTITION function in MySQL or the storage engine function provided in the database management system), the database is divided into multiple logical sharded databases according to the storage capacity. For example, the 20TB storage capacity is divided into 4 sharded databases, and the capacity of each sharded database is 5TB. The generation of each sharded database can be configured in the database management system to ensure that each sharded database has an independent physical storage path and the read-write load balancing of each sharded database is configured according to business requirements. At this time, the 4 generated sharded databases will provide the basis for subsequent data storage and table partitioning operations.

[0027] Step S13: Vertically partition the table based on the vertically sharded database data, where the size of each table is set not to exceed 2TB, and obtain the vertically partitioned table data to be stored; In this embodiment, vertical table partitioning refers to splitting a database table based on fields, with each table containing a portion of the data fields. The goal of this step is to ensure that the size of each table does not exceed 2TB to avoid performance bottlenecks or decreased query efficiency caused by an overly large single table. During the specific implementation process, it is necessary to analyze the data table structure of each sub-database to determine the division of fields and data for each table. For example, if a table contains 100 fields and the table data reaches 10TB, these fields can be divided into several sub-tables according to their relevance. The data volume of each sub-table does not exceed 2TB. Therefore, if 10TB of data can be divided into 5 tables, the capacity of each table is 2TB. Use the table partitioning function provided by the database management system (such as the CREATE TABLE command and PARTITION BY clause in MySQL, or perform table partitioning through Sharding technology) to perform vertical table partitioning operations in each sub-database to ensure that the data size and table fields after each table partition meet the requirements. At this time, each generated sub-table will be divided according to fields and the data volume does not exceed 2TB, providing a basis for the subsequent block node partitioning.

[0028] Step S14: Integrate the vertically partitioned database data to be stored and the vertically partitioned table data to be stored to obtain the first block node to be stored; In this embodiment, the vertically partitioned database data to be stored and the vertically partitioned table data are integrated. The purpose of the integration is to generate a complete dataset to be stored for subsequent storage block node partitioning. Integrate the sub-databases generated in step S12 and the table partition data generated in step S13 through database merge operations. Use the SQL INSERT INTO SELECT operation to insert the data after table partitioning into the sub-database. For example, for the multiple 2TB table partition data contained in a sub-database, these data are merged into the block data to be stored through database transaction management. At this time, the merged dataset meets the requirements of storage node partitioning, and the generated first block node to be stored contains the integrated data, providing a basis for subsequent partitioning operations and consensus node data verification.

[0029] Step S15: Use the first block node to be stored to perform storage block node partitioning on the slave database storage capacity. Divide the slave database storage capacity according to a fixed ratio of 1:1, with the capacity of each node controlled within 2TB - 4TB, to obtain other block nodes to be stored; In this embodiment, it is divided according to the storage capacity of the database (such as 20TB), and is split according to a fixed ratio of 1:1. The capacity of each storage block node is set between 2TB and 4TB. Therefore, the total storage capacity of 20TB needs to be divided into multiple block nodes, and the capacity of each node does not exceed 4TB. To ensure the reasonable capacity of each block node, in actual operation, the storage capacity is divided into 5 nodes, each with a capacity of 4TB, and the remaining 5TB capacity is allocated to other nodes according to the ratio. At this time, the capacity of all nodes is between 2TB and 4TB, meeting the storage requirements. This operation is carried out using the partition table function of the database or a custom storage partition management tool to ensure that each storage node has an independent storage path and can process data read and write requests in parallel.

[0030] Step S16: Obtain the large data to be stored, convert the large data to be stored into a request transaction data packet and send it to the first block node to be stored. The size of each data packet is fixed within 500MB. During the sending process, the request transaction data packet is broadcast to other block nodes to be stored at a fixed broadcast rate of 500MB / s using a dedicated transmission protocol (TCP / IP). In this embodiment, the large data to be stored is split according to a specified packet size (such as 500MB) to ensure that the size of each request data packet does not exceed 500MB. To efficiently transmit these data packets, a dedicated transmission protocol (such as the TCP / IP protocol) is adopted, and a fixed broadcast rate of 500MB / s is set to ensure that the data packets are not lost during transmission and can be quickly spread to other storage nodes. During the transmission process, the sending window size and retransmission mechanism of the TCP / IP protocol are set to ensure effective recovery of transmission in the case of unstable network. In addition, the data packets are encrypted during the sending process to prevent data leakage during transmission. During the entire transmission process, it is ensured that each node can receive the corresponding data packet within a predetermined time and perform storage processing on the received data packet.

[0031] Step S17: Use a preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node. The repair error threshold is set to 0.01% until the data of each node is consistent, and the consistent data of the nodes is encrypted and stored in the slave database.

[0032] In this embodiment, when a data packet arrives at each block node to be stored, data comparison is first performed through a consensus node to ensure data consistency among nodes. The consensus node will compare the data received by each block node. If data inconsistency is found, a consensus algorithm (such as the Raft protocol or the Paxos protocol) will be started for data repair. During this process, the threshold for repair error is set to 0.01%, that is, when the data error between nodes exceeds 0.01%, the repair mechanism is started. The consensus processing process includes steps such as data verification, data synchronization, and conflict resolution. After the consensus node finishes processing, the data of all nodes will be encrypted and stored in the slave database. The encryption operation uses the AES-256 encryption algorithm to ensure the security of data storage. At this time, the data consistency of all storage nodes is guaranteed, and the data has been encrypted and stored, ensuring that subsequent operations can be carried out in an efficient and secure environment.

[0033] Preferably, step S17 is specifically as follows: Step S171: Calculate the data digests of other block nodes to be stored and the first block node to be stored, where the calculation time is set to be within 2 seconds, to obtain the data digests of other block nodes to be stored and the data digest of the first block node to be stored; In this embodiment, hash algorithms such as SHA-256 or SHA-512 are used to calculate the data of each 2TB block node. To ensure that the calculation time is within 2 seconds, the data of the block node will be split into multiple small pieces, and the size of each small piece is 64MB. The data digests of each small piece are processed using parallel computing. After all small pieces are calculated, the digests of these small pieces are merged to generate a complete data digest. All calculation steps use a multi-core processor or GPU for accelerated calculation to ensure that the calculation time of each data digest does not exceed 2 seconds. When calculating, the digest of each data block is completed in parallel by multiple processing units, thereby improving the overall calculation efficiency and ensuring efficient time control.

[0034] Step S172: Integrate the data digests of other block nodes to be stored and the data digest of the first block node to be stored to obtain a data digest, and broadcast the data digest to a preset consensus node, set the timeout time to 5 seconds, to obtain the data digest of the consensus node; In this embodiment, the data digests from each block node are sequentially merged to form a large data digest. After the data digest merging is completed, it is broadcast to the preset consensus nodes through a private network using the TCP / IP protocol. At this time, the broadcast operation needs to ensure stable data transmission. All consensus nodes should receive the data digest within a timeout of 5 seconds. If a response is not received in time, a retry mechanism will be triggered. After each consensus node receives the data digest, it will verify it. The key to the entire broadcast process is to ensure the integrity and reliability of data transmission and ensure that the timeout is controlled within 5 seconds.

[0035] Step S173: Use the data digest of the consensus node to determine the data digest consistency between other block nodes to be stored and the first block node to be stored. If the difference in data digests exceeds 5%, obtain inconsistent data digests, and use the first block node to be stored to initiate a consensus request to the consensus node. The consensus node will reply with the voting result within 5 seconds. When at least 3 / 4 of the nodes that meet the protocol requirements agree to the data update, it is confirmed that consensus is reached, and node-consistent data is obtained. In this embodiment, the consensus node uses the hash algorithm to compare the data digests of each node to determine whether the data digests are consistent. If the difference in data digests exceeds 5%, it is determined to be inconsistent. In this case, the first block node to be stored initiates a consensus request to all consensus nodes. According to distributed protocols such as the Paxos protocol or the Raft protocol, at least 3 / 4 of the consensus nodes need to make a voting decision on data consistency within 5 seconds. If the protocol requirements are met and 3 / 4 of the nodes agree to the data update, data consistency is determined, and the update of the data digest is confirmed. The core of the determination process is to quickly confirm the correctness of the data through the comparison of hash values and combine the consistency protocol to ensure that no data deviation occurs during the execution of the update operation.

[0036] Step S174: Use the node-consistent data to repair the inconsistent data digest to obtain a repaired data digest. In this embodiment, for those digests where differences are found, after confirming the consistent data with the first block node to be stored and other consensus nodes, the inconsistent data is repaired. The repair operation will depend on the voting result of the consensus node. The key to the repair is to obtain the support of at least 3 / 4 of the consensus nodes during the repair process to ensure that the data repair process is recognized by most nodes. The repaired data digest will be recalculated using the hash algorithm to ensure its consistency and integrity. The repair operation needs to coordinate the correct data of each node to ensure that the repaired data no longer differs from the digests of other nodes, completing the consistency repair.

[0037] Particularly importantly, step S174 includes the following steps: Identify inconsistent data by using node-consistent data to perform inconsistent data identification on inconsistent data summaries, thereby obtaining inconsistent data summaries; In this embodiment, it is necessary to obtain a confirmed consistent data set from multiple nodes, and this data is usually extracted from each node through a database or API. Through programming languages, such as the pandas library in Python, this data is loaded into a DataFrame for subsequent operations. The extracted data usually includes multiple fields, such as timestamps, device status, sensor values, etc. Next, set comparison rules to determine which data is "consistent". For numerical data, an error range can be set (for example, ±0.01), and data outside this range is considered inconsistent. For string data, directly compare whether the field contents are exactly the same. When comparing data, use functions such as merge or join in pandas to align data from different nodes according to common fields (such as timestamps, device IDs, etc.) to ensure that the fields can correspond one by one. Then, compare these fields item by item. If the difference in numerical fields exceeds the set error range, the data is considered inconsistent; if it is a string field, directly compare whether its content is the same. After identifying the inconsistent data, mark it as "inconsistent data". Finally, the inconsistent data will be output as an inconsistent data summary, including inconsistent data items, source nodes, field names, and inconsistent numerical values or identifiers, providing a basis for subsequent data repair and analysis.

[0038] Perform missing item analysis on the inconsistent data summary to obtain the missing item target of the data summary; In this embodiment, the missing items usually appear as null values (NaN) or incomplete field contents, which are caused by data transmission errors, record omissions, or node synchronization problems. By checking each field in the inconsistent data summary, the data items that are not completely recorded are marked. In the specific operation, first, the inconsistent data summary is cleaned using the pandas library to remove outliers and null data other than the valid data, and then each field is checked to confirm which data item values are missing. If the value of a certain field is null or invalid (such as NaN, None, empty string, etc.), that field is marked as a missing item. To further analyze the missing items, missing value detection methods can be used, such as the SimpleImputer class in scikit-learn. This tool provides various missing value filling strategies, such as mean filling, median filling, most frequent value filling, etc., which are applicable to different types of data. During the analysis of missing items, first determine the data type of the missing field (such as numeric, character, time, etc.), and then use the filling strategy suitable for this data type for analysis. For example, for a numeric field, the mean or median of the field can be calculated as the filling value; for a character field, the most frequent value or a specified default value can be used for filling. In addition to the automatic filling strategy, the context of each missing field can be analyzed by comparing with the node-consistent data to determine whether there are systematic omissions. During this process, supplementation can also be carried out through the association relationship between fields. For example, the omission of a certain field will affect the data judgment of other related fields. In this step, the missing items are finally classified, all the missing data items that need to be supplemented are listed, and a target is provided for the subsequent filling process. The missing item target of the data summary will include the field name, data type, the number of missing items, and the corresponding filling strategy, providing a reference for the subsequent filling steps.

[0039] Based on the node-consistent data, data filling is performed on the missing item target of the data summary to obtain the filled data summary; In this embodiment, field data with similar characteristics is extracted from the stored node-consistent data, or missing items are filled from the same data source. For example, if the missing item is the operating temperature data of a certain device, it can be filled according to the known data of other devices or the same device type. The filling method can be based on regression analysis, mean filling, or median filling. For example, for numeric data, if there are many missing items in the field, the mean of adjacent nodes or data blocks can be selected for filling; for categorical data, the most common category is selected as the filling value. To maintain the accuracy of the data, the sklearn.impute.SimpleImputer toolkit can be used during the filling process, and the filling strategy can be adjusted according to the data distribution. After filling, the missing items are effectively supplemented, and finally the filled data summary is obtained.

[0040] Perform least squares fitting repair on the inconsistent data summary to obtain a repaired summary; In this embodiment, numerical fields are extracted from the inconsistent data summary obtained in step S12, such as temperature, pressure, humidity, etc. These fields contain some outliers, which usually appear as extreme values far from the mean and are caused by factors such as sensor errors, data transmission problems, or recording errors. After extracting these numerical data, it is first necessary to identify the outliers among them. Statistical methods can be used, such as calculating the mean and standard deviation of each field and setting a reasonable threshold to determine which data are abnormal. For example, if a data point deviates from the mean by more than three times the standard deviation, it can be considered an outlier. In addition, visualization methods such as box plots can be used to assist in identifying outliers. By these means, all outlier data points are marked and ready to enter the fitting repair process. The least squares method is used to repair the abnormal data. The least squares method is an optimization method whose goal is to minimize the sum of the squares of the errors between the fitted data and the actual data, thereby finding an optimal fitting model. When implementing, toolkits such as numpy.polyfit or scipy.optimize.curve_fit can be selected to perform least squares fitting. Specifically, first select a suitable fitting model type, such as linear fitting, quadratic fitting, or more complex polynomial fitting, and decide which model to choose according to the trend of the data. If the data changes relatively smoothly, linear fitting is applicable; if the data changes have non-linear characteristics, higher-order polynomial fitting or other more complex models need to be used. For example, assuming that there are outliers in the temperature data and the overall trend is linear, numpy.polyfit can be used for linear fitting. By setting appropriate parameters, such as the order of fitting being 1, indicating linear fitting. The fitted model will give a straight-line equation representing the overall trend of the data. Through this model, all outliers are corrected to the closest values in the fitting result, ensuring that these data points are more in line with the actual situation and trend. The accuracy of the model can be evaluated according to the fitting error to ensure that the fitting result is accurate enough. During the repair process, calculate the error between each data point and the fitting model, and adjust the parameters of the fitting model according to the error to ensure the stability and accuracy of the model. During the repair process, multiple adjustments can be made if necessary until the fitting result meets the predetermined accuracy requirements. The repaired outlier data points are updated to the original data to form a repaired summary. These repaired data will be smoother and more consistent, reducing the outliers that do not conform to the overall trend and ensuring the accuracy and reliability of the data.

[0041] Integrate the repaired summary and the filled data summary to obtain an inconsistent data repair summary; In this embodiment, the filled data summary and the repaired summary are merged to ensure that the two parts of data have a unified format and consistent data fields. To this end, the union or join operation in the database can be used to merge the two data sets, or the concat method in pandas can be used to merge the two data sets in programming implementation. During the merging process, it is necessary to ensure that the data fields after repair and filling are arranged in chronological order and that there are no duplicate data items. After integration, the generated "inconsistent data repair summary" contains all filled and repaired fields, ensuring data consistency and integrity.

[0042] Hash calculation is performed on the inconsistent data repair summary to obtain the repaired data summary.

[0043] In this embodiment, hash calculation is a process of converting data of any length into a value of a fixed length through a hash function. This process ensures the consistency and integrity of the data. The hash value is unique, that is, the same input data will get the same hash value, while different input data will produce different hash values. Therefore, the hash value can be used as an identifier of the data for data verification and integrity check. Specifically, during implementation, first extract the integrated data from the repaired data summary obtained in step S15. This data summary is the data that has been filled and repaired, including the filled data for all missing items and the data after outlier repair. The integrated data summary is a complex data structure that contains multiple fields and values and is stored in JSON, XML, or other formats. Next, select a suitable hash algorithm to perform hash calculation on the integrated repaired data summary. Commonly used hash algorithms include SHA-256 and MD5, which have different output lengths and security requirements. The SHA-256 algorithm generates a 256-bit (32-byte) hash value and is usually used in scenarios that require higher security; while the MD5 algorithm generates a 128-bit (16-byte) hash value. Although it is faster, its security is relatively low, so it is usually used for data integrity verification. In actual operation, the hashlib library in Python can be used to implement hash calculation. The specific operations are as follows: First, import the hashlib library and import the functions related to hash calculation through import hashlib; then, select the hash algorithm. Assuming the SHA-256 algorithm is used, the hashlib.sha256() function can be called to generate a SHA-256 hash object; next, perform hash calculation on the data. Pass the integrated repaired data summary into the hash function for hash calculation. Since the input data is a string or a byte stream, the data needs to be converted to byte format first. Use the.encode() method for encoding (if the data is a string), or directly pass it into the hash algorithm in the form of a byte stream; finally, obtain the calculated hash value by calling the.hexdigest() method. This method returns a hexadecimal string representing the hash value of the data. Through the above steps, the obtained hash_value is the hash value of the repaired data summary, which is a unique and fixed-length string. This hash value can be used for subsequent data verification to ensure the consistency and integrity of the data during the repair process. Finally, the obtained hash value of the repaired data summary can be used for data verification. During data transmission or storage, by comparing the hash value of the original data with the hash value of the repaired data, it can be verified whether the data has changed outside the repair process. If the hash values are the same, the data has not been tampered with; if the hash values are different, it means that the data is inconsistent or damaged. The hash value can also be used as the unique identifier of the repaired data summary, which is convenient for subsequent recording and auditing. When recording the data recovery or repair process, the source and history of the repair can be traced through the hash value.The hash algorithm itself has collision resistance, ensuring that the uniqueness of the repaired data digest can still be guaranteed even after multiple repair operations, preventing data from being maliciously tampered with. By implementing this step, the integrity and accuracy of inconsistent data during the repair process can be ensured, and a reliable basis can be provided for subsequent data processing or verification.

[0044] Step S175: Encrypt the repaired data digest to generate an encrypted data packet and store it in the slave database, where the encryption process is set to not exceed 10 seconds when the data packet size is 500MB.

[0045] In this embodiment, the repaired data digest is encrypted using the AES-256 symmetric encryption algorithm. Before data encryption, each data digest is split into small blocks not exceeding 500MB to ensure that the encrypted data packet size does not exceed 500MB. The encryption process is required to be completed within 10 seconds. If the data packet is larger than 500MB, it will be encrypted in multiple segments. During the encryption process, hardware acceleration (such as GPU acceleration) is used to improve the encryption speed to ensure that each encrypted packet can be completed within the specified time. Finally, the encrypted data packet will be stored in the slave database through the database storage interface according to the specified storage structure to ensure data security and efficient access.

[0046] Preferably, step S2 is specifically as follows: Step S21: Use a monitoring tool to obtain the disk status data of the slave database with a sampling period of once per second, including I / O performance data and disk health status data; In this embodiment, the disk status data of the slave database is obtained through a monitoring tool with a sampling period of once per second. At this time, the monitoring tool will perform real-time data collection with the server where the database is located, recording I / O performance data and disk health status data. The I / O performance data includes indicators such as the disk read and write speed, disk I / O queue length, and the number of read and write requests per second. These data are obtained using the monitoring interfaces of the operating system such as iostat or smartctl. The disk health status data includes indicators such as the disk temperature, health, and S.M.A.R.T. status (Self-Monitoring, Analysis and Reporting Technology). The collection period is set to once per second, and the data is collected and saved to the database through a real-time monitoring system to ensure the timeliness and accuracy of the data.

[0047] Step S22: Perform storage array failure analysis on the I / O performance data to obtain I / O performance failure data; In this embodiment, the I / O performance data collected per second is summarized and stored to form time series data. Then, potential fault conditions are detected by defined thresholds (such as read / write speed below 100 MB / s, I / O queue length greater than 10, number of read / write requests per second below 500). Using data analysis tools (such as the Pandas library in Python or the time series analysis package in R language), trend analysis is performed on the I / O performance data to calculate the I / O performance fluctuations within different time periods. If some data points exceed the set thresholds, it indicates a fault in the storage array. Whether problems such as performance bottlenecks and hard disk drive failures occur is detected through algorithms, and finally I / O performance fault data is generated, recording the specific time period when the fault occurs and the type of fault.

[0048] Step S23: Perform storage array fault analysis on the disk health status data to obtain disk storage array fault data; In this embodiment, each parameter of the disk health status data (such as disk temperature, health status, S.M.A.R.T. error value, etc.) is input into a preset fault determination model. The model is based on actual industry standards. For example, when the disk temperature exceeds 65 °C, the disk is damaged, and when the number of errors reported in the S.M.A.R.T. status is greater than 5 times, it indicates a hardware fault. These data are analyzed using a rule engine (such as the Drools rule engine), and faults are identified according to the preset fault determination criteria (such as too high temperature, S.M.A.R.T. error exceeding the limit, etc.). If the health status of a certain disk shows "bad" or "warning", it is considered that there is a storage array fault for this disk. In this way, the storage array fault data caused by abnormal disk health status is identified and recorded.

[0049] Step S24: Integrate the I / O performance fault data and the disk storage array fault data to generate database storage array fault data.

[0050] In this embodiment, the I / O performance fault data is associated with the disk health status fault data through a database management system (such as MySQL, PostgreSQL, etc.) to generate an integrated data table containing information of both. Through data table connection, the I / O performance data and the disk health status data within each time period are merged to form a comprehensive storage array fault data record. The integrated data table contains information such as the fault occurrence time, disk ID, fault type, fault cause, relevant performance indicators, etc. These data will be arranged in a certain time sequence and stored in the database for subsequent fault analysis and report generation.

[0051] Preferably, step S22 is specifically: Step S221: When any of the following situations occurs, it is determined that the read / write performance of the storage array is abnormal, and the abnormal data of the read / write performance of the storage array is obtained: the sequential read / write speed drops by more than 20%, the random access latency exceeds 30% of the normal range, the I / O queue depth is in a full-load state for a long time and the duration exceeds 60 minutes; In this embodiment, the monitoring system periodically obtains the sequential read / write speed of the storage array and compares it with the benchmark speed during normal operation. If the sequential read / write speed drops by more than 20%, that is, exceeds the threshold, it is determined that the read / write performance is abnormal. Secondly, the random access latency is monitored, its current value is obtained and compared with the normal range set by the system. If the random access latency exceeds 30% of the normal range, that is, exceeds the threshold, it is also determined that the performance is abnormal. Finally, the I / O queue depth is monitored. If the queue depth is in a full-load state for a long time and the duration exceeds 60 minutes, it indicates that the workload of the storage array is overloaded, and it is also determined that the read / write performance is abnormal. Under these conditions, the abnormal data of the read / write performance of the storage array is collected and integrated to form a performance failure report.

[0052] Step S222: When the following situations occur simultaneously, it is determined that the storage array controller fails, and the failure data of the storage array controller is obtained: the CPU utilization rate of the controller exceeds 90% for a long time and the response time is abnormally extended, the cache hit rate drops by more than 25%, the communication error frequency between the controller and the disk continues to increase, and it cannot return to normal within 10 minutes; In this embodiment, the utilization rate data of the controller CPU is obtained through a monitoring tool. If the CPU utilization rate continuously exceeds 90% and is accompanied by a significant extension of the response time, it indicates that there is a performance bottleneck or failure in the controller. In addition, the cache hit rate is monitored. When it drops by more than 25%, this is also an indication signal of an abnormality in the controller. Monitor the communication error frequency between the controller and the disk. If the error frequency continues to increase, the error fails to be recovered within 10 minutes, and no automatic repair by the system is seen, it is clearly determined that the controller fails. At this time, the data related to the failure of the storage array controller is collected, including the CPU utilization rate, the drop in the cache hit rate, and the communication error frequency, for subsequent processing.

[0053] Step S223: Integrate the abnormal data of the read / write performance of the storage array and the failure data of the storage array controller to obtain the I / O performance failure data.

[0054] In this embodiment, the I / O performance data is matched and summarized with the controller failure data to ensure that each type of failure data is fully recorded, and the data is sorted according to the time stamp. Through the data analysis system, a failure data report is automatically generated. This report contains all the detailed information of the I / O performance failure and the controller failure, including the abnormal degree of each parameter value and the time node when it occurs. Through this data report, it can provide a basis for subsequent troubleshooting and handling.

[0055] Preferably, step S23 is specifically as follows: Step S231: Extract the disk rotation speed characteristics from the disk health status data to obtain the disk rotation speed data; In this embodiment, the monitoring system needs to be connected to the storage array to ensure that the health status data of each disk in the storage array can be obtained in real time. The storage array usually consists of multiple disks, and these disks are equipped with internal monitoring sensors to monitor their operating status, including key indicators such as temperature, rotation speed, and health status. When obtaining the disk health status data, the monitoring system will regularly obtain the rotation speed information from the sensors of each disk. To ensure the real-time and accuracy of data collection, the data collection period is set to once per second. The rotation speed information of each disk is directly measured by a built-in rotation speed sensor, and the unit is "revolutions per minute" (RPM). The sensor continuously monitors the rotation of the disk platter and outputs the current rotation speed value. These rotation speed data are usually transmitted to the monitoring tool through the firmware of the disk drive or the storage control program in the operating system. To ensure the high precision of the data, the disk rotation speed value is usually denoised to filter out abnormal fluctuations caused by disk hardware failures or external interferences. After receiving the rotation speed data collected every second, the monitoring tool stores these data in chronological order, usually in the form of a time series. The stored data includes the specific rotation speed value and the sampling timestamp per second. In this way, the monitoring system can provide a detailed rotation speed history record and, in subsequent analysis, evaluate and diagnose the operating status of the disk at each moment. By filtering out the rotation speed information from the disk health status data, the real-time rotation speed data of each disk can be obtained. These data will serve as the basis for disk performance analysis and provide important evidence for subsequent fault diagnosis. Based on these rotation speed data, the monitoring system can evaluate whether there are problems such as rotation speed fluctuations, wear, or potential failures in the disk, and then predict the remaining service life of the disk or whether maintenance is required. The monitoring system needs to ensure the stability of data collection and avoid data loss or untimely collection caused by network latency or hardware failures. The monitoring tool can ensure the continuity and accuracy of data collection through regular self-checks and error log records. In addition, to cope with the large number of disks in a large-scale storage array, the monitoring system usually uses a distributed architecture to parallel process the data collection tasks of each disk, thereby improving the data collection efficiency and real-time performance.

[0056] Step S232: Obtain the disk standard rotation speed range data by using the disk technical documentation; In this embodiment, this data is usually provided by disk manufacturers in technical documents. The standard rotational speed range data refers to the rotational speed range that the disk should maintain under normal operating conditions. Usually, for different types of disks (such as 5400 RPM, 7200 RPM, 10000 RPM, etc.), the manufacturer will mark the standard operating rotational speed range of each type of disk in the document. The data obtained from the technical document will be used as the benchmark for subsequent comparison to determine whether the actual disk rotational speed meets the normal standard range. This data needs to be updated regularly to ensure that it reflects the standards of the latest disk models.

[0057] Step S233: Identify abnormal fluctuations in the disk rotational speed data based on the disk standard rotational speed range data. If the rotational speed fluctuation exceeds ±5% of the disk standard rotational speed range, generate rotational speed abnormal fluctuation data. In this embodiment, the system obtains the disk standard rotational speed range data according to the disk model and the technical document provided by the manufacturer. This data usually includes the rated rotational speed of the disk and the rotational speed range under normal operating conditions. The system compares the real-time rotational speed data (in revolutions per minute, RPM) collected from the internal sensors of the disk with this standard range data and calculates the deviation value of the current rotational speed from the standard range. For example, if the standard rotational speed range of the disk is from 6000 RPM to 7200 RPM and the real-time rotational speed is 7500 RPM, the deviation value is 7500 RPM minus 7200 RPM, which equals 300 RPM. To determine whether there are abnormal fluctuations in the rotational speed, the system sets a threshold, that is, ±5%. This ±5% range is calculated based on the standard rotational speed range. That is, if the real-time rotational speed of the disk exceeds ±5% of the standard rotational speed range, it is considered that there are abnormal fluctuations in the rotational speed. For a disk with a standard rotational speed range of 6000 RPM to 7200 RPM, the ±5% range is from 5700 RPM to 7560 RPM. If the real-time rotational speed exceeds this range, the system will determine that there is an abnormal fluctuation in the rotational speed. In actual operation, the monitoring system will calculate the deviation value between the real-time rotational speed and the upper and lower limits of the standard range and determine whether this value exceeds the ±5% range. If it exceeds this range, the system will generate rotational speed abnormal fluctuation data. This data includes the rotational speed deviation value, the timestamp when the abnormal fluctuation occurs, and the real-time rotational speed value. These abnormal fluctuation data will be stored together with other health status data of the disk and used for subsequent fault diagnosis and maintenance decisions.

[0058] Step S234: Evaluate the physical wear of the disk based on the rotational speed abnormal fluctuation data to obtain disk wear data. In this embodiment, a benchmark disk loss assessment model needs to be established, which usually evaluates by combining the number, amplitude, and duration of abnormal fluctuations in the rotation speed. To evaluate the physical loss of the disk, a loss scoring system needs to be set up first, which can calculate the loss degree of the disk according to the data of abnormal fluctuations in the rotation speed. The evaluation process includes the following steps: First, record the occurrence frequency and fluctuation amplitude of each abnormal rotation speed fluctuation, and calculate the amplitude of each fluctuation. For example, the fluctuation amplitude can be determined by the difference between the real-time rotation speed and the standard rotation speed range. Next, combine the actual running time of the disk to evaluate the duration of the rotation speed fluctuation. The duration refers to the time when the rotation speed fluctuation lasts outside the standard range. If the rotation speed fluctuation lasts for more than a certain threshold (e.g., 10 minutes), it is considered that the loss is relatively serious. Then, combine the number and amplitude of the abnormal rotation speed fluctuations with the duration to calculate a comprehensive physical loss score. The calculation formula of the loss score will be based on the technical standards provided by the manufacturer or industry experience, and usually uses the weighted average method. For example, different weight coefficients can be assigned to the amplitude and duration of the rotation speed fluctuation according to different importance. The amplitude will have a higher weight because a larger rotation speed deviation usually means a greater risk of physical damage. Assume that when the amplitude of each rotation speed fluctuation exceeds ±5% of the standard range, different scores are assigned according to the size of the amplitude (e.g., add 1 point for each ±5% exceedance), and at the same time, an additional score is added for each exceedance of a certain duration (e.g., add 5 points for each 10-minute exceedance). Finally, sum up all the scores after weighting to obtain a physical loss score. If the score is relatively high (e.g., the score exceeds 70%), it means that the loss degree of the disk is relatively high and needs to be replaced or maintained. If the score is relatively low (e.g., the score is lower than 30%), it indicates that the loss of the disk is small and it can still be used normally. This physical loss score will be used as an important indicator of the disk health status to help the administrator make subsequent maintenance and replacement decisions.

[0059] Step S235: Determine the storage array failure based on the disk health status data according to the disk loss data. When the disk loss score exceeds 60%, the disk storage array failure data is obtained.

[0060] In this embodiment, the disk physical loss score calculated in step S234 is compared with a preset loss score threshold (60%). This loss score is obtained by weighted calculation based on factors such as the amplitude, frequency, and duration of disk rotation speed fluctuations. Through this score, the health status of the disk can be reflected. If the disk loss score exceeds 60%, it means that the loss of the disk has reached a certain level, exceeding the healthy threshold, which will affect the overall performance and data security of the storage array. To determine a fault, it is necessary to judge whether the disk enters a high-risk state according to the preset threshold. Once the loss score exceeds 60%, the system will automatically determine that the health status of the disk does not meet the standard, indicating that the disk is close to failure and there is a risk of failure. In this case, the system will generate disk storage array fault data and mark the disk as a high-fault-risk disk. These fault data will be recorded and fed back to the operation and maintenance personnel as the basis for disk replacement or repair, avoiding data loss or system interruption of the storage array due to disk failure. Through continuous monitoring and regular health status checks, potential fault risks can be detected in a timely manner and corresponding preventive measures can be taken.

[0061] Preferably, step S234 is specifically as follows: Count the frequency of abnormal rotation speed fluctuation data, where the sampling period is set to be collected once per minute to obtain frequency data; In this embodiment, the system needs to set one minute as a sampling period, and sample the disk rotation speed data once per minute to ensure that the data acquisition frequency meets the real-time requirements. In actual operation, the rotation speed data is collected in real time through the internal sensor of the disk, and the collected data is transmitted to the monitoring system. At the end of each sampling period, the system records the data points of the current rotation speed and checks whether these data exceed the set standard rotation speed range, that is, whether the rotation speed fluctuation exceeds the threshold of ±5%. Specifically, for the data of each minute, the system will compare all the rotation speed data in this time period one by one to determine whether there is an abnormal fluctuation exceeding the standard range of ±5%. For each sampling period, if the number of abnormal rotation speed fluctuations is greater than 0, the occurrence times of the abnormal rotation speed fluctuations are recorded. Suppose there are 3 fluctuation events exceeding the standard range of ±5% in one minute, then the abnormal rotation speed fluctuation frequency of this minute is 3 times. The calculation of the frequency data is completed by a simple counting method, recording the number of fluctuations in each one-minute period. The abnormal rotation speed fluctuation frequency data of each minute will be recorded and stored in chronological order. In this way, the system can perform further analysis and evaluation based on these frequency data in subsequent processing steps. To ensure the accuracy and timeliness of the data, the frequency of the sampling period is set to once per minute. This frequency is sufficient to capture the changes in abnormal rotation speed fluctuations and can also ensure that the calculated frequency data has high real-time performance. These frequency data will be stored in the database for use in subsequent steps (such as physical loss assessment). In subsequent analysis, the frequency data will become an important indicator to help evaluate the health status of the disk and potential failure risks.

[0062] Statistical fluctuation time of abnormal rotation speed fluctuation data, where it is set that the fluctuation time greater than 5 seconds is an effective fluctuation to obtain the fluctuation time data; In this embodiment, from the abnormal rotational speed fluctuation data obtained in step S231, the system needs to extract the duration of each rotational speed fluctuation. Abnormal rotational speed fluctuations usually manifest as the rotational speed data exceeding the standard range of ±5%. Therefore, during the data extraction process, the system will record the start time and end time of each fluctuation. The calculation method for the duration of each abnormal rotational speed fluctuation is: the time difference between the start moment and the end moment of the fluctuation. First, the system will check the start and end times of each abnormal rotational speed fluctuation event one by one, and obtain the duration of each fluctuation through simple timestamp difference calculation. For each abnormal rotational speed fluctuation event, the system will calculate the fluctuation duration of the event based on the timestamp difference. The system will screen the fluctuation time according to the set conditions, and set the fluctuation time greater than 5 seconds as an effective fluctuation. That is to say, if the duration of a certain rotational speed fluctuation exceeds 5 seconds, the system will regard it as an effective fluctuation and record the fluctuation time. For those fluctuation events with a duration less than or equal to 5 seconds, the system will exclude them and not include them in the effective fluctuation data. To ensure the accuracy of the data, the system will check all fluctuation events one by one to ensure that the time of each fluctuation is accurately recorded. The system will accumulate and store the duration data of all effective fluctuations to form a dataset on the duration of each effective fluctuation. This dataset will serve as an important basis for disk health assessment and provide necessary input information for the subsequent loss assessment model. The time data of effective fluctuations will help evaluate the physical loss of the disk. In particular, abnormal events with longer fluctuation durations will have a greater impact on the long-term stability and reliability of the disk.

[0063] Construct a disk physical loss assessment model based on frequency data and fluctuation time data, where the weight of the influence of frequency on loss is set to 70%, and the weight of the influence of fluctuation time on loss is set to 30%; In this embodiment, it is necessary to set the influence weights of frequency data and fluctuation time data on disk loss. According to actual requirements, the influence weight of frequency on loss is set to 70%, while the influence weight of fluctuation time on loss is set to 30%. This weight distribution reflects the dominant role of frequency in loss, that is, the higher the frequency of abnormal rotational speed fluctuations, the greater the loss risk of the disk, while the influence of fluctuation time on loss is relatively auxiliary, so the weight is lower. During the evaluation process, it is first necessary to normalize the frequency data and fluctuation time data of each disk so that these two parameters can be compared within a unified scale range. Common normalization methods include min-max normalization and Z-score standardization. Through min-max normalization, the data is scaled proportionally between 0 and 1, while Z-score standardization transforms the data into a standard normal distribution by subtracting the mean and dividing by the standard deviation. After normalization, the numerical ranges of the frequency data and fluctuation time data are unified, enabling them to be comprehensively evaluated through weighting. Then, according to the set weights, the following formula is used to calculate the physical loss score of the disk: Loss score = 0.7 * normalized frequency + 0.3 * normalized fluctuation time. Through this formula, the influence of frequency on loss is amplified, while the influence of fluctuation time is reduced, and finally a comprehensive loss score is obtained. The score range is usually from 0 to 1, indicating no loss to maximum loss. This loss score provides a quantitative indicator for the disk health status, helping to evaluate the physical loss situation of the disk during operation.

[0064] Perform loss evaluation on the abnormal rotational speed fluctuation data according to the disk physical loss evaluation model: an evaluation of 31%-50% is slightly lossy data, an evaluation of 51%-70% is moderately lossy data, and an evaluation of 71%-100% is severely lossy data; In this embodiment, the core task of loss assessment is to compare the loss score of each disk with a predetermined loss level range to determine the health status of the disk. According to the calculation result of the loss score, the scores are classified into three different grade intervals. First, the data with minor loss is defined as the loss score between 31% and 50%, indicating that although there are abnormal fluctuations in the rotation speed of the disk, the loss is relatively minor and the disk can still continue to operate; then, the data with moderate loss is defined as the loss score between 51% and 70%, which means that the physical loss of the disk is relatively obvious and there is a certain risk, and maintenance or further monitoring is required; finally, the data with severe loss is defined as the loss score between 71% and 100%, indicating that the rotation speed of the disk fluctuates frequently and lasts for a long time, and has reached the level of severe damage, and immediate replacement or repair is required. To classify, first calculate the loss score of each disk and compare the calculation result with the above loss level standard. If the score is in the minor loss interval, it is marked as minor loss; if the score is in the moderate loss interval, it is marked as moderate loss; if the score is in the severe loss interval, it is marked as severe loss. By classifying the loss level of each disk, it can provide a clear basis for subsequent fault diagnosis and early warning, and provide data support for subsequent storage array maintenance and management.

[0065] Integrate the minor loss data, moderate loss data, and severe loss data to obtain the disk loss data.

[0066] In this embodiment, first, sort out the loss level of each disk evaluated in step S244. The loss score of each disk has been classified as minor loss, moderate loss, or severe loss. According to these classification results, label and classify each disk. Then, summarize all disks with loss levels according to the loss category, and list the number of disks with minor loss, moderate loss, and severe loss, their specific names, and their loss scores respectively. This summarization work can not only help managers clearly see the health status of each disk, but also quickly judge which disks need to be focused on or repaired or replaced as soon as possible according to the loss level. The integrated disk loss data report will include the following elements: the loss score of each disk, the loss level (such as minor, moderate, severe), the specific manifestation of the loss (such as the number and duration of abnormal rotation speed fluctuations), and whether further measures need to be taken (such as replacement, maintenance, etc.). This report is crucial for the health management of the storage array, the triggering of the early warning system, and subsequent repair decisions, and can help enterprises or data centers more efficiently manage disk resources and formulate reasonable repair plans during maintenance work.

[0067] Preferably, step S3 is specifically: Step S31: Obtain the image file of the database virtual hard disk, and use a disk imaging tool to parse the image file, so as to obtain the database virtual hard disk; In this embodiment, a suitable disk imaging tool is selected to extract the image file of the target database virtual hard disk. Commonly used tools include "dd" and "Clonezilla". Taking "dd" as an example, when using this tool, it is first necessary to connect to the target virtual machine or virtual hard disk in the virtualization environment. By specifying the device path of the virtual hard disk (such as / dev / sda) and the save path of the image file, execute the following command for block-level replication: "dd if= / dev / sda of= / path / to / output.img bs=4M". This command copies each block of the virtual hard disk one by one to ensure that the image file contains the complete virtual hard disk data, including the file system structure, operating system, and database files, etc. After the image file is obtained, it is stored on a specified storage device, such as an external hard disk, network storage device, etc., to ensure the integrity and security of the data. During the saving process, a reliable file verification method (such as SHA256 hash verification) should be used to verify the image file to ensure that the file has not been damaged during the saving process. The verification process involves calculating the hash value of the image file and comparing it with the hash value of the original image. If the two are the same, it means that the data has not changed. Use tools such as "FTK Imager" or "EnCase" to parse the obtained image file. These tools can extract the metadata of the file system from the disk image, such as information about the creation time, modification time, size, type, etc. of the file, and verify each file content one by one. During the parsing process, the file system structure is reconstructed by analyzing the sector information of the virtual hard disk to ensure that each partition, file, and folder is correctly identified. The tool will automatically detect the file system type (such as NTFS, ext4, etc.) and parse it according to the characteristics of the file system. To ensure the integrity of the data, the parsing tool will also perform an integrity check on the block data of each file and store the check result as a log file. All parsed data will be retained in the original format for subsequent operations. If there are problems during data parsing, the tool will generate an error report to help locate the problem. Finally, during the parsing process, the metadata of each file (such as timestamp, file size, file type, etc.) should also be recorded for subsequent analysis.

[0068] Step S32: Extract the flash write count feature and the erase cycle feature according to the storage array failure data, and obtain the flash write count data and the erase cycle data; In this embodiment, failure data of the storage array is obtained from the storage array management system, and the write count and erase cycle of each flash cell are obtained through the storage array control system. This information is usually stored in the health monitoring area of the storage device (such as in SMART attributes), and the flash controller will perform real-time monitoring and recording of each flash memory chip. To extract these characteristic data, it is necessary to access the operating data of the device through the management API provided by the controller, the device management interface, or a dedicated storage array monitoring tool. Usually, the interfaces provided by the storage array management system (such as RESTful API, SNMP protocol, dedicated command-line tools, etc.) can support remote query and data extraction. During the data extraction process, the system will read the write count data and erase cycle data of each flash memory chip, and record the cumulative write count and erase cycle count of each flash memory chip. The write count data represents the number of data blocks written to each flash cell, and the erase cycle refers to the number of erasures experienced by each flash cell during its life cycle. To ensure the integrity and real-time nature of the data, the sampling period is set to once per hour to ensure that the data recorded within each hour can promptly reflect the device status. After data collection, the storage array control system is used to store the values of the write count and erase cycle according to the unique identifier of the flash memory chip (such as the physical address or identifier) as a detailed data table. The data table should include fields such as chip ID, write count, erase cycle, and data extraction timestamp, etc. These data tables are then imported into the database for subsequent analysis.

[0069] Step S33: Perform cumulative write count statistics on the flash write count data to obtain cumulative write count data; In this embodiment, for the write count data of each flash memory chip, it is accumulated one by one. A period window is set, usually one day or one week. According to the write log of the flash memory, the write count within each period is calculated and these counts are accumulated. During the statistics, ensure that each write operation is accurately recorded and included in the total count. By accumulating the write counts of each flash memory chip within a given time range, cumulative write count data can be obtained. This data provides a basis for subsequent service life assessment, ensuring that the write process of each flash memory chip can be accurately traced and recorded. For the counting logic, a maximum write count threshold is set. For example, the maximum write count of each flash memory chip can be set to 50,000 times. When this threshold is reached, the chip enters a high-risk state.

[0070] Step S34: Perform flash memory chip service life assessment based on the cumulative write count data to obtain flash memory chip service life data; In this embodiment, a preset life threshold of the flash memory particles is set. Generally, the preset life of the flash memory particles is 100,000 write operations, but the specific threshold may vary according to different types of flash memory particles (such as SLC, MLC, TLC) and the technical standards of the manufacturers. Based on this preset life threshold, the system will analyze and evaluate the cumulative write count of each flash memory particle. To perform the specific life evaluation operation, first, the cumulative write count data of each flash memory particle needs to be obtained, and these data are collected through the health monitoring system (such as SMART) of the storage array. Then, the system compares the cumulative write count of each flash memory particle with the preset life threshold. If the cumulative write count of a certain particle is close to or exceeds the preset life threshold, it can be determined that the service life of this particle is approaching the end, and it enters a higher risk level. Different life intervals can be set according to historical data, the manufacturer's standards of the flash memory particles, or empirical formulas. For example, it can be divided according to percentage intervals: if the cumulative write count of the flash memory particle is 0%-50% of the preset life, it is regarded as mild wear, indicating that its remaining life is relatively long; if it is 50%-80%, it is moderate wear, indicating that the particle is gradually entering the risk stage; if it is 80%-100%, it is severe wear, indicating that the life of this particle has approached or exceeded its predetermined service life. In addition, in actual operation, these thresholds should also be adjusted according to the working environment. For example, in environments such as high temperature and high load, the write count of the flash memory particles will accelerate, and it is necessary to appropriately lower the preset life threshold or shorten the remaining life interval. The evaluation result can help identify the flash memory particles whose service life is about to be reached, facilitating early replacement or other maintenance measures. By evaluating each particle, the service life data of the flash memory particles are finally generated, and these data will provide an important basis for the subsequent maintenance of the storage system and fault warning.

[0071] Step S35: Divide the flash memory particles according to the erase cycle data. If the erase cycle of the flash memory particle exceeds 1000 times, the large-cycle particle data is obtained; if the erase cycle of the flash memory particle is less than 1000 times, the small-cycle flash memory particle data is obtained; In this embodiment, it is necessary to obtain the erase cycle data of each flash memory cell. These data are usually obtained through the monitoring system of the storage array or the health status log of the flash memory controller, and the unit of data recording is the number of erasures per cell. After obtaining the data, a threshold of 1000 erase cycles is set to classify the flash memory cells. The specific operation is to compare the erase cycles of all flash memory cells with this threshold: if the erase cycle of a certain cell is greater than 1000 times, it is classified as a "large-cycle cell", indicating that this cell can withstand more erase operations and usually has a longer service life; if the erase cycle of a certain cell is less than 1000 times, it is classified as a "small-cycle cell", indicating that this cell reaches the upper limit of the erase cycle earlier during normal use and needs to be replaced or maintained in advance. To ensure the accuracy of the classification, the erase cycle of each cell needs to be accurately measured. In actual operation, the erase cycle data of each cell can be queried through the API provided by the flash memory controller or through the management interface of the storage array, and the integrity and accuracy of the data are verified. During the classification process, special attention should be paid to avoiding misclassification caused by data acquisition errors or storage system abnormalities. After the classification is completed, the classification results of each cell are saved as the health status data of the flash memory cells. These data will serve as the basis for subsequent analysis and can be used for tasks such as evaluating the remaining service life of the flash memory cells and analyzing erase imbalance. Through this classification, the health status of the flash memory cells can be managed and monitored more effectively, and cells at risk of failure can be detected in advance for targeted maintenance or replacement.

[0072] Step S36: Perform erase imbalance analysis based on the large-cycle flash memory cell data and the small-cycle flash memory cell data to obtain erase imbalance data; In this embodiment, it is necessary to collect the erasure cycle data of all large-cycle flash memory particles and small-cycle flash memory particles. To perform the erasure imbalance analysis, these data are first extracted from the storage array to ensure that the erasure cycles of each particle are accurately recorded. The data can be obtained through the API or management interface of the flash memory controller to ensure that the data collection process is not interfered with and error-free. Next, a comparative analysis of these data is carried out, with a focus on the difference in erasure cycles between large-cycle particles and small-cycle particles. When performing the comparative analysis, a threshold for erasure cycle imbalance is set. If the erasure cycle of large-cycle particles is significantly lower than that of small-cycle particles, that is, the number of erasure times of large-cycle particles is significantly less than that of small-cycle particles, it can be determined that there is erasure imbalance. To quantify the erasure imbalance, the standard deviation is used to measure the difference in erasure cycles. The specific calculation process is as follows: First, calculate the mean and standard deviation of the erasure cycles of large-cycle particles and small-cycle particles respectively. The mean represents the average erasure cycle of each type of particle, and the standard deviation measures the fluctuation of the erasure cycles of each type of particle. By comparing the standard deviation differences between the two types of particles, if the standard deviations between them differ significantly, it indicates an obvious erasure imbalance phenomenon. To further quantify the degree of erasure imbalance, the relative standard deviation of the erasure cycle differences between the two types of particles can be calculated, and a certain threshold (such as 20%) is set. When the relative standard deviation exceeds this threshold, it can be considered that there is a serious erasure imbalance problem. At this time, the reasons can be further analyzed. For example, some particles have a heavier workload or some particles are erased more frequently during use, resulting in reaching their lifespan limit in advance. Through this analysis, erasure imbalance data can be obtained to help evaluate whether there are performance bottlenecks in the storage array. For example, if the erasure cycles of some particles are significantly less than those of other particles, it means that these particles will fail in advance or experience performance degradation, thus affecting the service life and performance of the entire storage array. This result provides an important basis for subsequent maintenance work, and adjustments can be made for the imbalance problem, and reasonable maintenance strategies can be adopted, such as optimizing load balancing and adjusting write distribution, to ensure the stability and durability of the array.

[0073] Step S37: Integrate the service life data of flash memory particles and the erasure imbalance data to generate flash memory particle loss data.

[0074] In this embodiment, the service life evaluation results of the flash memory particles are combined with the erase imbalance analysis results to generate a detailed flash memory particle wear report for subsequent maintenance and performance monitoring. First, obtain the service life data of each flash memory particle from step S34, and obtain the erase imbalance data of each flash memory particle from step S36. The service life data usually includes the remaining life percentage of each particle, while the erase imbalance data reflects the erase cycle distribution of the particles. When generating the wear data, first classify the wear level of each flash memory particle. When classifying the service life, it is divided according to the remaining life percentage of the particle: particles with a remaining life of less than 50% are marked as "severely worn", particles with a remaining life between 50% and 80% are marked as "moderately worn", and particles with a remaining life of more than 80% are marked as "slightly worn". The criteria for this distinction usually come from the recommended values of the device manufacturer or industry standards. For example, particles with a remaining life of less than 50% are approaching the end of their service life and are prone to failure. The erase imbalance also needs to be classified. According to the erase cycle difference data obtained in step S36, judge the erase imbalance level of the particle. The particles can be divided into three categories: low, medium imbalance, and high. For example, when the standard deviation of the erase cycle is low, it can be considered that the particle has good balance and is marked as "low imbalance"; when the standard deviation is in the medium range, it is marked as "medium imbalance"; when the standard deviation is high and there are significant differences in the erase cycles of the particles, it is marked as "high imbalance". Combine the wear level and erase imbalance data of each particle, and generate a complete flash memory particle wear report through the method of data merging. In the report, the wear level and erase imbalance data of each particle will be presented in tabular form, and the detailed information of each particle (such as ID, remaining life percentage, erase cycle, erase imbalance level, etc.) will be listed. The format of the report should be clear and easy to understand, facilitating the operation and maintenance personnel and administrators to quickly understand the health status of each particle in the storage array. The generated flash memory particle wear data report provides important support for the subsequent maintenance work of the storage array. The administrator can identify particles with severe wear or erase imbalance problems in a timely manner through this report, and thus take necessary maintenance measures, such as replacing the problematic particles, readjusting the storage load distribution, or taking other optimization measures. This can not only ensure the stability and performance of the storage array, but also extend the service life of the storage device, reduce the failure rate, and avoid system downtime or data loss caused by particle damage.

[0075] Preferably, step S36 is specifically as follows: Step S361: Count the number of erase times according to the large-cycle flash memory particle data to obtain the large-cycle erase times data; In this embodiment, the erasure cycle data of all large-cycle flash memory particles are obtained from the storage array management system, and the flash memory particles with an erasure cycle greater than 1000 times are screened out. For each large-cycle particle, its erasure times are counted. The specific method is to obtain the erasure times record of each flash memory particle by reading the log file or health status monitoring data (such as SMART value) provided by the flash memory controller. Then, the erasure times of these particles are accumulated or averaged to obtain the erasure times data of the large-cycle particles. These data reflect the durability of the large-cycle particles and their long-term usage conditions. All the data should be indexed and stored according to the particle ID, and the accuracy and integrity of the original data should be maintained.

[0076] Step S362: Count the erasure times according to the small-cycle flash memory particle data to obtain the small-cycle erasure times data; In this embodiment, it is necessary to extract the erase cycle data of all small-cycle flash memory particles from the storage array management system. The definition of these small-cycle particles is flash memory particles with an erase cycle less than 1000 times. The core of this step is to screen out the particles that meet this condition and obtain their erase count records. By accessing the management system of the storage array, using the API interface of the storage array or the health status monitoring system (such as SMART values) to obtain the detailed health data of each flash memory particle. These data include the erase count of the flash memory cells, which is usually recorded in the log file of the flash memory controller. This information is usually stored in the form of a log, including information such as the erase count and write count of the flash memory particle, and each record has a unique particle ID for easy tracking and management. Particles with 1000 times. By writing a filtering program or using a database query statement, filter out the data of all small-cycle particles. The erase count of these particles should be extracted by parsing the data in the log file, and the erase count of each particle is the specific value recorded by the flash memory controller. Then, count the erase count of the small-cycle particles. Specifically, for all eligible small-cycle particles, the system can choose to sum their erase counts to obtain the total sum of the overall erase count; or calculate the average value of their erase counts to get the average erase count of each particle. If further analysis of the particle distribution is required, the standard deviation of the erase count can also be calculated to evaluate the distribution of the erase count. To ensure the accuracy and consistency of the statistical data, all erase count data should be indexed and stored by particle ID. The particle ID is the unique identifier of each particle, used to associate each particle with its erase count data, ensuring that the data of each particle can be accurately stored and traced. All the collected data should be stored in a database or table, sorted and organized according to the particle ID, erase count, and other relevant information (such as the storage unit location, etc.). Integrate all the statistically obtained erase count data into an erase count report of small-cycle flash memory particles. This report provides the basic data for subsequent analysis, which can help evaluate the durability differences of small-cycle particles, and then infer their service life and replacement time. In addition, it can also provide a reference for the health status analysis of the storage array, identifying potential problems or flash memory particles with short life spans.

[0077] Step S363: Draw a distribution map of the database virtual hard disk based on the large-cycle erase count data and the small-cycle erase count data to obtain a cycle erase distribution map; In this embodiment, the data of the large-cycle erase count and the small-cycle erase count are collected. By extracting this data from the database, it is ensured that the erase count data of each particle is complete and accurate. The data should be classified and stored according to the type of particle (large cycle or small cycle) so that different types of particles can be distinguished when drawing the distribution map later. To draw the cycle erase distribution map, statistical software is needed for data visualization. In this step, appropriate data processing tools can be selected, such as Excel, Matplotlib or Seaborn in Python, etc. These tools can all support data processing and graph drawing. Before drawing, first sort the erase count data into two data sets according to the particle type (large cycle and small cycle). Each data set should include the erase count and the corresponding particle ID. Through this step, different types of erase data are separated so that they can be accurately displayed in the distribution map. Then, libraries such as Matplotlib or Seaborn in Python are used to visualize this data. The horizontal axis (X-axis) represents the erase count, and the vertical axis (Y-axis) represents the number or proportion of different flash particles. For the data points of large-cycle particles, they should be concentrated in the area of higher erase counts, such as the area with more than 1000 erases; while the data points of small-cycle particles should be distributed in the area of lower erase counts, such as the area with less than 1000 erases. The size of each data point can be adjusted to represent the actual number or proportion of the particles. For example, for particles with relatively more erase counts (i.e., large-cycle particles), the size of the data point can be relatively large; while for particles with fewer erase counts (i.e., small-cycle particles), the size of the data point can be smaller. This helps to clearly show the difference between large-cycle and small-cycle particles on the graph and enables the graph to reflect the distribution of each type of particle in different erase count intervals. To better display the data, different color codings can be added to the distribution map to distinguish large-cycle particles and small-cycle particles. For example, blue can be used to represent large-cycle particles, red for small-cycle particles, or different marker symbols can be used to distinguish these two types of data points. This can clearly show the difference in erase counts between large-cycle and small-cycle particles in the same chart and help analysts quickly identify the distribution trends between the two. After drawing the cycle erase distribution map, the graph should be further analyzed to observe the difference in the erase cycle distribution between large-cycle particles and small-cycle particles. This graph can intuitively display the erase cycle characteristics of each particle in the storage array, help managers identify potential performance bottlenecks, the health status and lifespan issues of the particles, and provide a reference for subsequent maintenance decisions.

[0078] Step S364: Calculate the flash particle imbalance of the cycle erase distribution map to obtain the erase imbalance data.

[0079] In this embodiment, the average and standard deviation of the erasure counts of large-cycle particles and small-cycle particles are calculated respectively. Then, by comparing these two standard deviations, the degree of erasure imbalance is obtained. If the standard deviation of the erasure counts of large-cycle particles is significantly smaller than that of small-cycle particles, it indicates that the erasure cycle distribution of large-cycle particles is more balanced. On the contrary, if the standard deviation is larger, it means that there is a more serious erasure imbalance phenomenon. In addition, the overall average of the erasure counts of the two types of particles can also be calculated, and by comparing the differences in the averages, the erasure cycle distribution of the two types of particles can be further analyzed. The erasure imbalance data provides a key indicator for the performance evaluation of the storage array, reflecting the usage status and healthy distribution of different types of flash memory particles.

[0080] Preferably, this specification also provides a disk-based big data storage system for executing a disk-based big data storage method as described above. The disk-based big data storage system includes: A data encryption storage module that obtains the storage capacity of the slave database; performs data chunking on the storage capacity of the slave database to generate the first block node to be stored; obtains the big data to be stored, converts the big data to be stored into a request transaction data packet and sends it to the first block node to be stored; broadcasts the request transaction data packet to other block nodes to be stored during the sending process; uses a preset consensus node to compare the data of other block nodes to be stored with the data of the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the consistent data of the nodes is encrypted and stored in the slave database; A storage array failure analysis module that uses a monitoring tool to obtain the disk status data of the slave database in real time; performs storage array failure analysis on the disk status of the slave database during the encryption storage process to generate the storage array failure data of the slave database; A flash memory particle wear analysis module that scans the hard disk of the slave database using a disk imaging tool; performs flash memory particle wear analysis on the hard disk of the slave database according to the storage array failure data to obtain the flash memory particle wear data; A thermal protection mechanism setting module that sets the thermal protection mechanism for the hard disk of the slave database based on the flash memory particle wear data to generate the thermal protection mechanism data; uploads the thermal protection mechanism data to the slave database and stores the slave database data in the master database to execute the big data storage task.

Claims

1. A disk-based big data storage method, characterized in that: The following steps are involved: Step S1: Obtain storage capacity from the database; Divide the data into blocks from the database storage capacity to generate a first block node to be stored; Obtaining the big data to be stored, converting the big data to be stored into a request transaction data packet and sending it to the first block node to be stored; During the sending process, the request transaction data packet is broadcast to other block nodes to be stored; Use the preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the node consistent data is encrypted and stored in the slave database; Step S2: using a monitoring tool to obtain the slave database disk status data in real time; performing storage array failure analysis on the slave database disk status during the encryption storage process to generate slave database storage array failure data; Step S3: Scan the slave database hard disk using a disk imaging tool to obtain a slave database virtual hard disk; perform flash memory particle loss analysis on the slave database virtual hard disk according to the storage array failure data to obtain flash memory particle loss data; Step S4: Setting a thermal protection mechanism for the slave database virtual hard disk based on the flash memory particle loss data to generate thermal protection mechanism data; uploading the thermal protection mechanism data to the slave database, and storing the slave database data to the master database to perform the big data storage task.

2. The disk-based big data storage method according to claim 1, characterized in that: Step S1 is specifically as follows: Step S11: Obtain storage capacity from the database; Step S12: vertically divide the storage capacity of the slave database, wherein the size of the sub-database is set to not exceed 5TB, and the vertical sub-database data to be stored is generated; Step S13: vertically partitioning the table based on the vertical partitioned database data, wherein the vertical partitioned table size is set to be no more than 2TB, and the vertical partitioned table data to be stored is obtained; Step S14: Integrate the vertical sub-library data to be stored and the vertical sub-table data to be stored to obtain the first block node to be stored; Step S15: using the first block node to be stored to divide the storage capacity of the slave database into storage block nodes, dividing the storage capacity of the slave database according to a fixed ratio of 1:1, controlling the capacity of each node to be 2TB-4TB, and obtaining other block nodes to be stored; Step S16: Obtain the big data to be stored, convert the big data to be stored into a request transaction data packet and send it to the first block node to be stored. The size of each data packet is fixed to within 500MB. During the sending process, a dedicated transmission protocol (TCP / IP) is used to broadcast the request transaction data packet at a fixed broadcast rate of 500MB / s to other block nodes to be stored; Step S17: Use the preset consensus node to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node, and the repair error threshold is set to 0.01% until the data of each node is consistent, and the node consistent data is encrypted and stored in the slave database.

3. The disk-based big data storage method according to claim 2, characterized in that: Step S17 is specifically as follows: Step S171: Calculate the data digests of other block nodes to be stored and the first block node to be stored, wherein the calculation time is set to be controlled within 2 seconds, and obtain the data digests of other block nodes to be stored and the data digest of the first block node to be stored; Step S172: Integrate the data summaries of other block nodes to be stored and the data summaries of the first block node to be stored to obtain a data summary, and broadcast the data summary to a preset consensus node, setting a timeout time of 5 seconds to obtain a consensus node data summary; Step S173: Use the consensus node data summary to determine the consistency of the data summary of other to-be-stored block nodes and the first to-be-stored block node. If the data summary difference exceeds 5%, an inconsistent data summary is obtained, and the first to-be-stored block node is used to initiate a consensus request to the consensus node. The consensus node replies with the voting result within 5 seconds. When at least 3 / 4 of the nodes that meet the protocol requirements agree to the data update, it is confirmed that a consensus has been reached and the node consistent data is obtained; Step S174: using the node consistent data to perform data repair on the inconsistent data summary to obtain a repaired data summary; Step S175: Encrypt the repair data summary to generate an encrypted data packet, and store it in the slave database, wherein the encryption process is set to take no more than 10 seconds when the data packet size is 500MB.

4. The disk-based big data storage method according to claim 1, characterized in that: Step S2 is specifically as follows: Step S21: using a monitoring tool to obtain disk status data from the database at a sampling period of once per second, including I / O performance data and disk health status data; Step S22: performing storage array fault analysis on the I / O performance data to obtain I / O performance fault data; Step S23: performing storage array fault analysis on the disk health status data to obtain disk storage array fault data; Step S24: Integrate the I / O performance fault data and the disk storage array fault data to generate secondary database storage array fault data.

5. The disk-based big data storage method according to claim 4, characterized in that: Step S22 is specifically as follows: Step S221: When any of the following situations occurs, it is determined that the storage array read / write performance is abnormal, and the storage array read / write performance abnormality data is obtained: the sequential read / write speed drops by more than 20%, the random access delay exceeds 30% of the normal range, and the I / O queue depth is in a full load state for a long time and the duration exceeds 60 minutes; Step S222: When the following conditions occur at the same time, it is determined that the storage array controller is faulty, and the storage array controller fault data is obtained: the controller CPU utilization rate exceeds 90% for a long time and the response time is abnormally prolonged, the cache hit rate decreases by more than 25%, and the communication error frequency between the controller and the disk continues to increase and cannot be restored to normal within 10 minutes; Step S223: Integrate the storage array read and write performance abnormality data and the storage array controller failure data to obtain I / O performance failure data.

6. The disk-based big data storage method according to claim 4, characterized in that: Step S23 is specifically as follows: Step S231: extracting disk rotation speed features from the disk health status data, thereby obtaining disk rotation speed data; Step S232: using the disk technical document to obtain the disk standard rotation speed range data; Step S233: identifying abnormal fluctuations of the disk rotation speed data according to the disk standard rotation speed range data, and generating abnormal rotation speed fluctuation data if the rotation speed fluctuation exceeds ±5% of the disk standard rotation speed range; Step S234: evaluating the physical loss of the disk according to the abnormal rotation speed fluctuation data, thereby obtaining disk loss data; Step S235: Perform storage array fault determination on the disk health status data based on the disk loss data. When the disk loss score exceeds 60%, disk storage array fault data is obtained.

7. The disk-based big data storage method according to claim 6, characterized in that: Step S234 is specifically as follows: Count the frequency of abnormal speed fluctuation data, where the sampling period is set to collect data once per minute to obtain frequency data; Count the fluctuation time of abnormal speed fluctuation data, where the fluctuation time greater than 5 seconds is set as effective fluctuation, and obtain the fluctuation time data; A disk physical loss assessment model is constructed based on frequency data and fluctuation time data, where the weight of the frequency effect on loss is set to 70% and the weight of the fluctuation time effect on loss is set to 30%; According to the disk physical loss assessment model, the abnormal rotation speed fluctuation data is evaluated for loss: 31%-50% is evaluated as slight loss data, 51%-70% is evaluated as moderate loss data, and 71%-100% is evaluated as severe loss data; The disk loss data is obtained by integrating the slight loss data, the moderate loss data and the severe loss data.

8. The disk-based big data storage method according to claim 1, characterized in that: Step S3 is specifically as follows: Step S31: obtaining an image file of the slave database virtual hard disk, and parsing the image file using a disk imaging tool, thereby obtaining the slave database virtual hard disk; Step S32: extracting the number of flash memory write features and the number of erase cycles according to the storage array fault data, and obtaining the number of flash memory write data and the number of erase cycles data; Step S33: Count the cumulative number of write times of the flash memory to obtain the cumulative number of write times; Step S34: evaluating the service life of the flash memory particles based on the accumulated write times data to obtain the service life data of the flash memory particles; Step S35: dividing the erase cycle data into flash memory particles, if the erase cycle of the flash memory particles exceeds 1000 times, obtaining large cycle particle data; if the erase cycle of the flash memory particles is less than 1000 times, obtaining small cycle flash memory particle data; Step S36: performing erase imbalance analysis according to the large-cycle flash memory particle data and the small-cycle flash memory particle data to obtain erase imbalance data; Step S37: Integrate the flash memory particle service life data and the erase imbalance data to generate the flash memory particle wear data.

9. The disk-based big data storage method according to claim 8, characterized in that: Step S36 is specifically as follows: Step S361: performing erasure count according to the large-cycle flash memory particle data to obtain large-cycle erasure count data; Step S362: performing erasure count according to the small-cycle flash memory particle data to obtain small-cycle erasure count data; Step S363: Draw a distribution map of the database virtual hard disk based on the large-cycle erasure number data and the small-cycle erasure number data to obtain a periodic erasure distribution map; Step S364: Calculate the flash memory particle imbalance on the periodic erase distribution diagram to obtain erase imbalance data.

10. A disk-based big data storage system, characterized in that: Used to execute the disk-based big data storage method according to claim 1, the disk-based big data storage system comprises: The data encryption storage module obtains the storage capacity of the slave database; divides the storage capacity of the slave database into data blocks to generate the first block node to be stored; obtains the big data to be stored, converts the big data to be stored into a request transaction data packet and sends it to the first block node to be stored; during the sending process, the request transaction data packet is broadcast to other block nodes to be stored; the preset consensus node is used to compare the data of other block nodes to be stored with the first block node to be stored. When the data is inconsistent, consensus processing is performed based on the consensus node until the data of each node is consistent, and the node consistent data is encrypted and stored in the slave database; The storage array fault analysis module uses monitoring tools to obtain the slave database disk status data in real time; performs storage array fault analysis on the slave database disk status during the encrypted storage process, and generates slave database storage array fault data; The flash memory particle loss analysis module uses a disk imaging tool to scan the slave database hard disk; performs flash memory particle loss analysis on the slave database hard disk according to the storage array failure data to obtain flash memory particle loss data; The thermal protection mechanism setting module sets the thermal protection mechanism for the slave database hard disk based on the flash memory particle loss data and generates thermal protection mechanism data; uploads the thermal protection mechanism data to the slave database, and stores the slave database data in the master database to perform big data storage tasks.

Citation Information

Patent Citations

  • A block chain-based quality data processing method and device based on a block chain

    CN109598505A

  • Rebuilding portions of virtual segments based on power

    US20240330305A1

Cited By

  • Method and device for offline storage of real-time sensing data based on Flash ROM in single-chip microcomputer

    CN121455423A

  • Air conditioning system database management method, system, medium and equipment

    CN121478752A