Enterprise-level solid-state storage method for big data processing

Through the data temperature index and cluster analysis method, combined with CRC checking and Merkle tree hash pre-checking, the hot and cold data storage strategy of solid state hard disk is optimized, the performance bottlenecks and life problems of solid state hard disks in big data processing are solved, and efficient and reliable enterprise-level storage solutions are realized.

CN120371224AInactive Publication Date: 2025-07-25HYUNDAI DIGITAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510863916.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional mechanical hard disks have high random access latency, solid-state hard disk write amplification effect and limited erase life affect big data processing performance. Existing storage management technologies cannot efficiently respond to the needs of high concurrent data access, and fail to optimize the storage strategies of hot and cold data in a targeted manner, resulting in waste of storage resources and degradation of system performance.

Method used

The data temperature index is introduced to separate hot and cold data, and cluster analysis is performed by combining time window constraints and DBSCAN algorithm. In-block CRC verification and cluster-level Merkle tree hash pre-checking are used to optimize the data storage order through the log structure merging and writing strategy, and a hybrid index structure is constructed based on the B+ tree to manage metadata. The hot data is stored in the SLC area and the cold data is stored in the QLC area.

Benefits of technology

It improves storage resource utilization, improves data access efficiency and system performance, ensures data integrity and security, extends the service life of the equipment, and optimizes the stability and reliability of enterprise-level storage systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371224A_ABST
    Figure CN120371224A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data processing, in particular to an enterprise-level solid-state storage method oriented to big data processing. The method comprises the following steps: firstly, acquiring enterprise data, calculating a data temperature index, and dividing the enterprise data into cold data and hot data through a temperature threshold value; performing clustering analysis on the hot data in combination with time window constraint and a DBSCAN algorithm, aligning a clustering result with the size of a physical block of the solid state disk, and performing decomposition to obtain a hot data block; performing in-block CRC (Cyclic Redundancy Check) and cluster-level Merkle tree Hash pre-check on the hot data block to generate a two-level check code; a write-in sequence of the hot data blocks is optimized through a log structure merging write-in strategy, and a mixed index structure is constructed based on a B + tree to manage metadata of the hot data blocks; and storing the metadata into a memory, storing the hot data blocks into an SLC (Single Level Cell) area of the solid state disk, and storing the cold data into a QLC (Quality Level Cell) area of the solid state disk. According to the method, the storage efficiency, the data integrity and the data security of enterprise-level big data processing are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing, and specifically to an enterprise-level solid-state storage method for big data processing. Background Art

[0002] With the rapid growth of the data scale of enterprise-level applications, data storage and management face severe challenges. Due to the relatively high random access latency of traditional hard disk drives (HDDs), it is difficult to meet the enterprise's demand for high-performance data access. In contrast, solid-state drives (SSDs) have gradually become the mainstream choice for enterprise-level storage due to their advantages of high throughput and low latency. However, the write amplification effect, limited erase-write lifespan, and data management complexity of SSDs still restrict their application in big data processing.

[0003] In the prior art, although SSDs are superior to traditional hard disk drives in terms of performance, some of their drawbacks still affect big data processing. First of all, the write amplification effect of SSDs is a major bottleneck. Since data writing inside SSDs is performed in units of pages rather than individual data items, frequent small data writes will result in a large amount of invalid data being written, increasing the number of write amplifications, thereby accelerating device wear and reducing its service life. Secondly, the limited erase-write lifespan of SSDs means that frequent write operations, especially in large-scale data processing scenarios, may lead to premature failure of storage devices, affecting the long-term stable storage of data. Moreover, as the amount of data continues to increase, existing storage management technologies have deficiencies in data indexing, verification, and sorting, and are unable to efficiently handle the access requirements of massive data. Especially during high-concurrency data reading and writing, performance bottlenecks are likely to occur. In addition, traditional data storage methods do not fully consider the difference in data heat, and often fail to optimize the storage strategies for hot data and cold data specifically, resulting in waste of storage resources and degradation of system performance.

[0004] Therefore, an enterprise-level solid-state storage method for big data processing is proposed. Summary of the Invention

[0005] The purpose of the present invention is to provide an enterprise-level solid-state storage method for big data processing. First of all, a data temperature index is introduced to dynamically evaluate data heat based on access frequency, storage time, and priority, and cold data and hot data are divided through a temperature threshold to achieve intelligent hierarchical storage and improve the utilization rate of storage resources. Secondly, combining time window constraints with the DBSCAN clustering algorithm, only recently hot data is classified, the KD tree is used to efficiently search for neighborhood data, and the data storage layout is optimized based on Euclidean distance and time constraints to improve access efficiency. Finally, block-level CRC verification and cluster-level Merkle tree hash pre-verification are adopted to achieve fast error detection at the data block level and cluster-level integrity verification, improving the reliability and security of data storage.

[0006] To achieve the above object, the present invention provides the following technical solutions: An enterprise-level solid-state storage method for big data processing, comprising: Obtaining enterprise data and calculating a data temperature index, and classifying the enterprise data into cold data and hot data through a temperature threshold; Performing clustering analysis on the hot data in combination with a time window constraint and a DBSCAN algorithm, aligning and decomposing the clustering result with the physical block size of the solid-state drive to obtain hot data blocks; Performing in-block CRC check and cluster-level Merkle tree hash pre-check on the hot data blocks to generate two-level check codes; Optimizing the writing order of the hot data blocks through a log-structured merge writing strategy, and constructing a hybrid index structure based on a B+ tree to manage the metadata of the hot data blocks; Storing the metadata in memory, storing the hot data blocks in the SLC area of the solid-state drive, and storing the cold data in the QLC area of the solid-state drive.

[0007] Further, the calculation formula of the data temperature index is: ; wherein, represents the access frequency, represents the data storage time, represents the current time, represents the most recent access time, represents the visitor priority, , and respectively represent the frequency weight, the time weight, and the priority weight.

[0008] Further, the steps of performing clustering analysis on the hot data are as follows: Constructing a KD tree according to the data temperature index of the hot data; Randomly selecting an unlabeled hot data as a seed point and labeling it, searching for all neighborhood points within the neighborhood radius of the seed point according to the KD tree, and checking whether each neighborhood point meets the composite determination condition of the Euclidean distance and the time constraint; if the neighborhood point meets the composite determination condition, the number of neighborhood points of the seed point is incremented; If the number of neighborhood points of the seed point is greater than the lowest neighbor point threshold, marking the seed point as a core point and forming a new cluster, and continuing to add neighborhood points that meet the composite determination condition to the cluster until no further expansion is possible; Traversing all hot data and deleting isolated hot data to form a complete clustering structure.

[0009] Further, the steps of the in-block CRC check are: Calculate the CRC32 checksum for each hot data block; Append the checksum to the end of the hot data block to form a checksum structure in the format of "data + CRC".

[0010] Furthermore, the steps of the in-block CRC check also include: When reading data, recalculate the CRC checksum of the hot data block and compare it with the CRC checksum at the end of the checksum structure. If they are the same, it indicates that the data is complete. If they are different, reread the data or perform a repair operation.

[0011] Furthermore, the steps of the cluster-level Merkle tree hash pre-check are as follows: Arrange all the hot data blocks within the same cluster in sequence and calculate the hash value to generate a Merkle tree; Generate a Merkle hash for each cluster and store it in an independent secure area.

[0012] Furthermore, the steps of the cluster-level Merkle tree hash pre-check also include: When reading data, use the Merkle hash to verify the data integrity. If the hash values match, the data is valid; otherwise, reread the data or perform a repair operation.

[0013] Furthermore, the steps of the log-structured merge write policy are as follows: Temporarily store the hot data block in memory and establish a write-ahead log; Cache the random write requests of the hot data block in memory and merge them into sequential write requests according to the clustering label; When the cached data reaches the threshold, align the cached data with the physical page size of the solid-state drive.

[0014] Furthermore, the metadata includes logical address, physical address, cluster label, CRC checksum, and Merkle hash.

[0015] The beneficial effects of the present invention are as follows: 1. The present invention introduces a data temperature index to calculate enterprise data, classifies the enterprise data into cold data and hot data through a temperature threshold, effectively realizes data hierarchical storage, preferentially stores the frequently accessed hot data into a high-performance storage medium to improve the read and write efficiency, and at the same time stores the infrequently accessed cold data into a large-capacity low-cost storage medium, optimizes the utilization of storage resources, reduces the storage cost, and improves the performance, reliability, and management efficiency of the enterprise-level storage system.

[0016] 2. The present invention combines time - window constraints and only performs clustering analysis on hot data within the most recent time range. By using an improved DBSCAN algorithm, it efficiently finds neighborhood data points through a KD - tree, determines data attribution based on Euclidean distance and time constraints, forms a reasonable data - cluster structure, and eliminates isolated data points simultaneously, ensuring the rationality and consistency of data distribution. This method can accurately identify highly correlated hot data, improve the organization and query efficiency of data management, reduce ineffective storage and computing overhead, optimize the read - write performance of solid - state drives, and enhance the stability and overall performance of enterprise - level storage systems.

[0017] 3. The present invention adopts in - block CRC check and cluster - level Merkle - tree hash pre - check. Through CRC check, it quickly detects errors during data transmission and storage at the data - block level to ensure the integrity of data blocks; meanwhile, Merkle - tree hash pre - check provides a higher - level data - integrity verification at the cluster level, supporting efficient data - consistency checking and anti - tampering mechanisms. This method double - guarantees data reliability, can quickly detect and repair damaged data during storage and reading, improves the system's error - code resistance and data security, and ensures the high reliability and efficient operation of enterprise - level storage systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of an enterprise - level solid - state storage method for big - data processing provided by the present invention; Figure 2 is a flowchart of hot - data clustering analysis provided by the present invention; Figure 3 is a flowchart of a log - structured merge - write strategy provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only for explaining and understanding the present invention, and are not used to limit the present invention.

[0020] Embodiment 1 An enterprise - level solid - state storage method for big - data processing, as Figure 1 shown, includes: S100: Obtain enterprise data and calculate the data - temperature index, and divide the enterprise data into cold data and hot data through a temperature threshold; Further, the calculation formula of the data - temperature index is: ; Wherein, Indicates the access frequency, Indicates the data storage time, Indicates the current time, Indicates the most recent access time, Indicates the visitor priority, 、 and respectively represent the frequency weight, time weight, and priority weight.

[0021] By calculating the data temperature index, the access activity of data can be dynamically evaluated, and the data can be classified into cold data and hot data according to the temperature threshold, so as to achieve hierarchical storage optimization. Among them, by combining factors such as access frequency, storage time, most recent access time, and visitor priority, high-frequency access data can be more accurately identified, ensuring that hot data is preferentially stored in high-performance storage media and cold data is stored in low-cost storage media. This method not only improves the read and write efficiency of the storage system, but also reduces the storage cost and extends the service life of the device.

[0022] S200: Combine the time window constraint and the DBSCAN algorithm to perform clustering analysis on the hot data, align and decompose the clustering result with the physical block size of the solid-state drive to obtain hot data blocks; Furthermore, the steps for performing clustering analysis on the hot data are as Figure 2 shown, including: Construct a KD tree according to the data temperature index of the hot data; Randomly select an unlabeled hot data as a seed point and label it. Search all neighborhood points within the neighborhood radius of the seed point according to the KD tree, and check whether each neighborhood point meets the composite determination condition of Euclidean distance and time constraint; if the neighborhood point meets the composite determination condition, the number of neighborhood points of the seed point increments; If the number of neighborhood points of the seed point is greater than the minimum neighbor point threshold, mark the seed point as a core point and form a new cluster, and continue to add neighborhood points that meet the composite determination condition to the cluster until it can no longer be expanded; Traverse all hot data and delete isolated hot data to form a complete clustering structure.

[0023] Specifically, select the data temperature index of the hot data as the splitting dimension, sort the data according to this dimension, select the median as the root node, recursively construct the left subtree with the left half of the data and the right subtree with the right half of the data until all data points are constructed; traverse all hot data points, randomly select an unlabeled data point as a seed point and mark the seed point as the "visited" state. With as the center, search all neighborhood points in the KD tree whose distance is less than the neighborhood radius for each neighborhood point Calculate whether the Euclidean distance between two points is less than the neighborhood radius The calculation formula can be expressed as: ; Wherein, represents the Euclidean distance, represents the seed point, represents the neighborhood point, represents the number of dimensions of the feature. In this embodiment , which are the access frequency, data storage time, recent access time, and visitor priority respectively, represents the dimension index of the feature, represents the neighborhood radius; the time constraint calculation formula can be expressed as: ; Wherein, represents the recent access time of the seed point, represents the recent access time of the neighborhood point, represents the time constraint value, and check whether each neighborhood point meets the composite determination condition of Euclidean distance and time constraint; if the neighborhood point meets the composite determination condition, the number of neighborhood points of the seed point is incremented; when the number of neighborhood points of the seed point is greater than the minimum neighbor point threshold, mark the seed point as a core point and form a new cluster, and continue to add neighborhood points that meet the composite determination condition to the cluster until it can no longer be expanded; traverse all hot data to form a complete clustering structure.

[0024] By constructing a KD tree based on the data temperature index, perform efficient spatial indexing on the hot data, accelerate neighborhood search, and thus optimize density-based dynamic clustering. Introduce the composite determination condition of Euclidean distance and time constraint to ensure that the clustering result is more accurate and can adapt to the access characteristics of the data. In addition, remove isolated hot data, improve the utilization rate of storage space, and avoid low-correlation data affecting query performance. This method improves the clustering calculation efficiency, optimizes the data storage structure, and enhances the retrieval ability of enterprise-level storage systems in the big data environment.

[0025] S300: Perform in-block CRC check and cluster-level Merkle tree hash pre-check on the hot data block to generate two-level check codes; Furthermore, the steps of the in-block CRC check are: Calculate the CRC32 check code for each hot data block; Append the check code to the end of the hot data block to form a check structure in the "data + CRC" format.

[0026] Specifically, the specific process of calculating the CRC32 check code is: Select a standard CRC32 polynomial (such as 0x04C11DB7 used by IEEE802.3), initialize the CRC register to 0xFFFFFFFF, perform an exclusive OR operation and a bit shift on the hot data block byte by byte, and then obtain the CRC32 checksum; append the checksum to the end of the hot data block to form a check structure in the "data + CRC" format.

[0027] By performing CRC32 check on the hot data block, calculating and appending the checksum before data storage to form the "data + CRC" format, the data integrity and reliability are effectively improved. This method has low calculation overhead and small storage space occupation, can quickly detect data corruption during the SSD read and write process, and avoid storage errors from affecting data availability.

[0028] Furthermore, the steps of the in-block CRC check further include: When reading data, recalculate the CRC checksum of the hot data block and compare it with the CRC checksum at the end of the check structure. If they are the same, it means the data is complete; if they are different, reread the data or perform a repair operation.

[0029] By recalculating the CRC checksum when reading data and comparing it with the stored checksum, the data integrity can be quickly detected, effectively preventing data anomalies caused by bit flips, write errors, or media damage during storage. If the check fails, the data reread or repair mechanism can be triggered to ensure the correctness and availability of the data, improve the reliability and fault tolerance of the enterprise-level storage system, and reduce the risk of damage during data migration.

[0030] Furthermore, the steps of the cluster-level Merkle tree hash pre-check are as follows: Arrange all the hot data blocks within the same cluster in sequence and calculate the hash value to generate a Merkle tree; Generate a Merkle hash for each cluster and store it in an independent secure area.

[0031] Specifically, for the same clustering cluster , sort the hot data blocks in the decomposition order and use the hot data blocks as the leaf nodes of the Merkle tree; calculate the hash value of each hot data block using the SHA-256 hash function , where represents the hash value of the hot data block , represents the hash function, and then obtain the leaf node hash value sequence ; pair adjacent hash values in pairs and calculate the parent node hash , where represents the hot data block and The hash of the parent node; repeat this process to calculate the upper-level hash values layer by layer until a unique Merkle hash is generated. Store the Merkle hash corresponding to each cluster in an independent secure area.

[0032] Through the Merkle tree hash pre-verification, calculate the hash of the hot data blocks within the same cluster and generate a unique root hash to be stored in an independent secure area, which can efficiently verify data integrity, prevent data tampering and corruption. The hierarchical structure of the Merkle tree makes data verification faster and more efficient. Only by comparing the root hash can the consistency check of the entire cluster of data be completed, significantly reducing the computational overhead and improving the security, reliability and verification efficiency of the enterprise-level storage system.

[0033] Furthermore, the steps of the cluster-level Merkle tree hash pre-verification further include: When reading data, use the Merkle hash to verify the data integrity. If the hash values match, the data is valid; otherwise, re-read the data or perform a repair operation.

[0034] Through the cluster-level Merkle tree hash pre-verification, the data integrity can be quickly verified when reading data. If the hash values match, it ensures that the data has not been tampered with and improves the reading reliability; if the hash values do not match, the data corruption can be detected in a timely manner, and the data availability can be guaranteed through the re-reading or repair mechanism. This method improves the security and fault tolerance of data storage.

[0035] S400: Optimize the writing order of the hot data blocks through the log-structured merge writing strategy, and build a hybrid index structure based on the B+ tree to manage the metadata of the hot data blocks; Furthermore, the steps of the log-structured merge writing strategy are as Figure 3 shown, including: Temporarily store the hot data blocks in the memory and establish a write-ahead log. Cache the random write requests of the hot data blocks in the memory and merge them into sequential write requests according to the clustering labels. When the cached data reaches the threshold, align the cached data with the physical page size of the solid-state drive.

[0036] Specifically, temporarily store the hot data blocks in the memory buffer to avoid frequent direct writes to the SSD, reduce write amplification, generate a write-ahead log to record the metadata of the data blocks and the write requests, ensuring data consistency and recoverability; cache the randomly written hot data blocks in the memory to avoid fragmentation and performance loss caused by direct writes to the SSD, merge the data blocks according to the clustering labels (based on the DBSCAN clustering results), and convert them into sequential write requests; monitor the cached data volume, when the data reaches the set write threshold, prepare for batch writing to the SSD, and adjust the data blocks according to the SSD physical page size (usually 4KB or 8KB) to ensure that the data writing conforms to the SSD physical structure and reduce the additional overhead caused by cross-page writing.

[0037] Through the log-structured merge write strategy, convert random writes into sequential writes, reduce SSD write amplification, and improve write performance and storage life. Adopt the write-ahead log and data alignment mechanism to ensure data consistency, reduce the internal garbage collection pressure of the SSD, and improve the stability and reliability of the enterprise-level storage system.

[0038] S500: Store the metadata in the memory, store the hot data blocks in the SLC area of the solid-state drive, and store the cold data in the QLC area of the solid-state drive.

[0039] Further, the metadata includes a logical address, a physical address, a cluster label, a CRC check code, and a Merkle hash.

[0040] Store the hot data blocks in the SLC area to improve read and write performance by using its high throughput and low latency characteristics; store the cold data in the QLC area to make full use of its large capacity and low cost advantages to improve storage efficiency. At the same time, the metadata contains a logical address, a physical address, a cluster label, a CRC check code, and a Merkle hash to ensure data integrity, fast retrieval, and efficient management, and improve the reliability and performance of the enterprise-level storage system.

[0041] Embodiment 2 A certain company adopted an enterprise-level solid-state storage method for big data processing proposed by the present invention to achieve efficient and secure data migration, including: Obtain enterprise data and calculate the data temperature index, and divide the enterprise data into cold data and hot data through a temperature threshold; Perform clustering analysis on the hot data by combining time window constraints and the DBSCAN algorithm, align and decompose the clustering results with the solid-state drive physical block size to obtain hot data blocks; Perform in-block CRC check and cluster-level Merkle tree hash pre-check on the hot data blocks to generate two-level check codes; Optimize the writing order of the hot data blocks through a log-structured merge writing strategy, and construct a hybrid index structure based on a B+ tree to manage the metadata of the hot data blocks; Store the metadata in memory, store the hot data blocks in the SLC area of the solid-state drive, and store the cold data in the QLC area of the solid-state drive.

[0042] During the data migration process, the access frequency of a certain piece of data is 100 times per minute, the creation time is 1 hour ago, the last access time is 10 seconds ago, and the visitor permission is guest, with the corresponding visitor priority of 5. The frequency weight, time weight, and priority weight corresponding to this data are 0.6, 0.3, and 0.1 respectively. According to the data temperature index calculation formula, we can get , since the temperature threshold is set to 50, this data is classified as hot data.

[0043] To achieve efficient writing of hot data to the SSD, clustering analysis is performed on enterprise data. During the process of building a KD tree for a group of data, the median of the data temperature index is 75. Select the median as the root node, recursively construct the left subtree with the left half of the data and the right subtree with the right half of the data until all data points are included in the KD tree; during the clustering process, a randomly selected seed point has a neighborhood radius = 100, and the time constraint value = 15s seconds. Under this constraint, the number of its neighborhood points that meet the composite judgment condition is 732, which is greater than the minimum neighbor point threshold of 500. Finally, a new cluster with a space size of 89KB is formed with this point as the core point. , after aligning with the physical block size of the solid-state drive, it is decomposed into 23 4KB hot data blocks.

[0044] Before and after writing, in-block CRC check and cluster-level Merkle tree hash pre-check are required. Calculate the CRC32 checksum for each 4KB hot data block in cluster C1 and append the CRC32 checksum to the end of the block. The CRC32 checksum of a certain hot data block is 0x3A7B9C, and the CRC32 checksum of this hot data block obtained during the read test is 0x5D8E2F. This hot data block is repaired during this test; at the same time, use the 23 hot data blocks in cluster C1 as leaf nodes, calculate the SHA-256 hash layer by layer, generate the Merkle hash 0x8E2F3A and store it in the independent security area (Block0, Page0) of the SSD. During the read test process, if the Merkle hashes before and after writing are consistent, it proves that the data in this cluster is valid.

[0045] During the write process, three random write requests are cached in memory (Cluster C1 - Data A, Cluster C2 - Data B, Cluster C1 - Data C). To achieve efficient writing, they are merged into a sequential operation according to the cluster tags (first write Data A and Data C in Cluster C1, then write Data B in Cluster C2). After merging, the total size of Cluster C1 is 8KB, aligned to 8KB (2 × 4KB pages), and written in batches to Block 5 in the SLC area.

[0046] During the implementation of the above steps, the performance comparison between the present invention and the traditional method is shown in Table 1, where the traditional method is the direct write from the database to the hard disk. It can be seen from the table that the present invention is significantly superior to the traditional method in multiple key storage performance indicators. First, the write throughput of the traditional method is only 50K IOPS, while the present invention reaches 120K IOPS, an increase of 140%. This is because of the separation of hot and cold data storage, storing frequently accessed data in the SLC area of the SSD, reducing the performance loss caused by random writes, and improving the sequential write efficiency through the log-structured merge write strategy. Second, the read latency of the present invention is only 1ms, a reduction of 80% compared with 5ms of the traditional method. This benefits from the clustering alignment strategy based on KD-tree and DBSCAN, which optimizes data locality, reduces the query overhead caused by the scattered storage of physical blocks, and combines with the B+ tree index to speed up data location. Finally, in terms of data integrity guarantee, the data repair time of the traditional method is 10ms, while the present invention only needs 3ms, and the repair efficiency is increased by 233%. This is mainly attributed to the two-level verification mechanism, that is, the intra-block CRC verification quickly detects errors, and the cluster-level Merkle tree hash pre-verification efficiently repairs, ensuring data integrity while significantly reducing the repair cost.

[0047] In summary, through optimization strategies such as hot and cold separation, clustering alignment, and two-level verification, the present invention effectively improves the storage performance of the SSD, enhances the reliability of the system and the data management efficiency, and provides an efficient, low-latency, and high-reliability storage solution for enterprise-level big data processing.

[0048] Table 1 Performance Comparison between the Present Invention and the Traditional Method

[0049]

[0050] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An enterprise-level solid-state storage method for big data processing, characterized in that, Including: Obtain enterprise data and calculate the data temperature index, and classify the enterprise data into cold data and hot data through a temperature threshold; Combine time window constraints and the DBSCAN algorithm to perform clustering analysis on the hot data, align and decompose the clustering results with the physical block size of the solid-state drive to obtain hot data blocks; Perform in-block CRC check and cluster-level Merkle tree hash pre-check on the hot data blocks to generate two-level check codes; Optimize the write order of the hot data blocks through a log-structured merge write policy, and build a hybrid index structure based on the B+ tree to manage the metadata of the hot data blocks; Store the metadata in memory, store the hot data blocks in the SLC area of the solid-state drive, and store the cold data in the QLC area of the solid-state drive.

2. The enterprise-level solid-state storage method for big data processing according to claim 1, characterized in that, The calculation formula of the data temperature index is: ; Among them, represents the access frequency, represents the data storage time, represents the current time, represents the most recent access time, represents the visitor priority, 、 and represent the frequency weight, the time weight, and the priority weight respectively.

3. An enterprise-level solid-state storage method for big data processing according to claim 1, characterized in that, The steps for performing clustering analysis on the hot data are: Construct a KD tree based on the data temperature index of the hot data; Randomly select an unlabeled hot data as a seed point and label it, search for all neighborhood points within the neighborhood radius of the seed point according to the KD tree, and check whether each neighborhood point meets the composite decision condition of Euclidean distance and time constraint; if the neighborhood point meets the composite decision condition, the number of neighborhood points of the seed point is incremented; If the number of neighborhood points of the seed point is greater than the minimum neighbor point threshold, mark the seed point as a core point and form a new cluster, and continue to add neighborhood points that meet the composite decision condition to the cluster until no further expansion is possible; Traverse all hot data and delete isolated hot data to form a complete clustering structure.

4. An enterprise-level solid-state storage method for big data processing according to claim 1, characterized in that, The steps of the in-block CRC check are: Calculate the CRC32 check code for each hot data block; Append the check code to the end of the hot data block to form a check structure in the format of "data + CRC".

5. An enterprise-level solid-state storage method for big data processing according to claim 4, characterized in that The steps of the in-block CRC check further include: When reading data, recalculate the CRC check code of the hot data block and compare it with the CRC check code at the end of the check structure. If they are the same, it means the data is complete. If they are different, reread the data or perform a repair operation.

6. The enterprise-level solid-state storage method for big data processing according to claim 1, wherein The steps of the cluster-level Merkle tree hash pre-check are: Arrange all hot data blocks within the same cluster in order and calculate the hash value to generate a Merkle tree; Generate a Merkle hash for each cluster and store it in an independent secure area.

7. An enterprise-level solid-state storage method for big data processing according to claim 6, characterized in that, The steps of the cluster-level Merkle tree hash pre-check further include: When reading data, use the Merkle hash to verify the data integrity. If the hash values match, the data is valid. Otherwise, reread the data or perform a repair operation.

8. An enterprise-level solid-state storage method for big data processing according to claim 1, characterized in that The steps of the log-structured merge write policy are: Temporarily store the hot data blocks in memory and establish a write-ahead log; Cache the random write requests of the hot data blocks in memory and merge them into sequential write requests according to the clustering labels; When the cached data reaches the threshold, align the cached data with the physical page size of the solid-state drive.

9. An enterprise-level solid-state storage method for big data processing according to claim 1, characterized in that, The metadata includes logical address, physical address, cluster label, CRC check code, and Merkle hash.