Hierarchical storage method and device for data, storage medium and electronic equipment
By deploying storage controllers in storage devices and dividing data storage layers according to data volume and storage granularity, the problem of low efficiency in traditional data tiered storage is solved, and more efficient data access and storage utilization are achieved.
Patent Information
- Application Number
- CN202511250840.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Traditional tiered data storage methods suffer from accelerated memory wear due to frequent data migration and limited data access bandwidth, resulting in low storage efficiency.
By deploying storage controllers in storage devices, data is divided into different storage layers based on the amount of target data and the granularity of storage. Data can then be accessed directly through the storage controllers, thus achieving tiered data storage.
It improves data access speed, avoids memory wear caused by frequent data migration, and enhances the efficiency of data tiered storage.
Smart Images

Figure CN120743203B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to methods and apparatuses for hierarchical data storage, storage media and electronic devices. Background Technology
[0002] With the advent of the information age, data is experiencing explosive growth, making traditional data storage and management increasingly difficult and less cost-effective. To address this data storage demand, a tiered data storage scheme has been proposed, dividing data into different tiers based on access frequency and storing them on different storage media. These tiers utilize differentiated storage media such as solid-state storage (SSDs) and mechanical storage (MH). SSDs, with their superior read / write performance, are used to store frequently accessed data, while MH, with its larger storage capacity, is used to store less frequently accessed data. Data is then migrated between SSDs and MH according to access frequency. However, this tiered data storage method suffers from frequent data migration, accelerating memory wear, and its reliance on SSDs for data access limits bandwidth. Summary of the Invention
[0003] This application provides a method and apparatus for hierarchical data storage, a storage medium and an electronic device, to at least solve the technical problem of low efficiency in hierarchical data storage in related technologies.
[0004] This application provides a hierarchical data storage method applied to a target storage controller deployed in a target storage device, comprising: receiving a data write request, wherein the data write request is used to request writing target data into the target storage device, the target storage device including multiple storage layers for storing data, the target storage controller being used to access data in the storage layers, and different storage layers having different data storage granularities; dividing the target storage layer into which the target data falls according to the target data volume and the data storage granularity; and writing the target data into the target storage layer.
[0005] This application also provides a storage device, including: a storage controller and a plurality of storage layers, the storage controller and the storage layers being connected, the different storage layers having different data storage granularities, the storage controller being configured to receive a data write request, wherein the data write request is used to request that target data be written into the storage device; to classify the target storage layer into which the target data falls according to the target data volume and the data storage granularity; and to write the target data into the target storage layer.
[0006] This application also provides a tiered data storage device, applied to target storage control deployed in a target storage device, comprising: a first receiving module, configured to receive a data write request, wherein the data write request is used to request that target data be written into the target storage device, the target storage device including multiple storage layers for storing data, the target storage controller being used to access data in the storage layers, and different storage layers having different data storage granularities; a partitioning module, configured to partition the target storage layer into which the target data falls based on the target data volume and the data storage granularity; and a first writing module, configured to write the target data into the target storage layer.
[0007] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the hierarchical data storage method described above when executing the computer program.
[0008] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of any of the above-described hierarchical data storage methods.
[0009] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described hierarchical data storage methods.
[0010] This application utilizes a storage controller deployed in a storage device. Upon receiving a data write request to write target data to the storage device, the controller divides the target storage layer according to the target data volume and data storage granularity. The controller then directly accesses each storage layer, achieving layered data storage based on the matching relationship between data volume and storage granularity. This enables efficient utilization of storage space and persistent data storage. On one hand, it improves data access speed; on the other hand, it avoids memory wear caused by frequent data migration during storage. Therefore, it solves the technical problem of low efficiency in related technologies for layered data storage, achieving the technical effect of improving the efficiency of layered data storage. Attached Figure Description
[0011] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1This is a hardware structure block diagram of the hierarchical data storage method according to an embodiment of this application;
[0013] Figure 2 This is a flowchart of a hierarchical data storage method according to an embodiment of this application;
[0014] Figure 3 This is a schematic diagram of an optional edge storage system according to an embodiment of this application;
[0015] Figure 4 This is an optional edge node architecture diagram according to an embodiment of this application;
[0016] Figure 5 This is a schematic diagram of an optional hybrid media tiered storage according to an embodiment of this application;
[0017] Figure 6 This is a schematic diagram of data storage for an optional storage engine according to an embodiment of this application;
[0018] Figure 7 This is a schematic diagram of a storage device structure for optional data tiered storage according to an embodiment of this application;
[0019] Figure 8 This is a structural block diagram of a hierarchical data storage device according to an embodiment of this application;
[0020] The above figures include the following reference numerals:
[0021] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. Detailed Implementation
[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0023] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0024] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] The specific application environment architecture or specific hardware architecture on which the implementation of the hierarchical storage method for combined data depends is described here.
[0026] The methods and embodiments provided in this application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure block diagram of a hierarchical data storage method according to an embodiment of this application. For example... Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the image. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The server device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the server equipment described above. For example, the server equipment may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0027] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the hierarchical data storage method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to server devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0028] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the server device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0029] The embodiments of this application provide a hierarchical data storage method, and the method is described in detail in conjunction with the execution flow of the hierarchical data storage method.
[0030] The following explains the technical terms used in this application:
[0031] Distributed storage systems: Distributed storage systems improve storage reliability, availability, and access efficiency by distributing data across multiple independent devices and utilizing a scalable architecture, thus solving the performance bottleneck problem of traditional centralized storage.
[0032] TLC / QLC Solid State Drives: TLC (Triple-Level Cell) and QLC (Quad-Level Cell) are solid state drive (SSD) storage technologies. TLC stores 3 bits of data per cell, offering better performance but a shorter lifespan; QLC stores 4 bits per cell, providing larger capacity but lower durability. TLC media offers higher read and write performance, while QLC's small IO write performance is about one-third that of TLC, while their read performance is essentially the same.
[0033] NVMe over Fabric: NVMe over Fabric (NVMe-oF) is a network protocol that extends the NVMe storage protocol to remote servers through transmission media such as fiber optics and Ethernet, enabling high-performance, low-latency distributed storage access.
[0034] RDMA: RDMA (Remote Direct Memory Access) is a network technology that allows a computer to directly access the memory of another computer without CPU intervention, significantly reducing latency and improving data transfer efficiency. It is commonly used in high-performance computing and storage systems.
[0035] FC: FC (Fibre Channel) is a high-speed network technology designed for Storage Area Networks (SANs). It provides high-bandwidth, low-latency data transmission, supports both fiber optic and copper cable media, and is widely used in enterprise-level storage systems.
[0036] TCP / IP: TCP / IP is the core communication protocol suite of the Internet, which includes Transmission Control Protocol (TCP) and Internet Protocol (IP). It is responsible for data segmentation, addressing, routing and reliable transmission, and forms the basic architecture for global network interconnection.
[0037] Performance pool: A storage pool with high performance such as bandwidth or latency. This invention consists of several TLC SSDs.
[0038] Capacity pool: Compared to the performance pool, it has lower bandwidth, latency and other performance characteristics, but higher capacity density. This invention consists of several QLCSSDs.
[0039] Metadata pool: A storage device specifically for storing metadata. Since metadata changes frequently and occupies relatively little space compared to the data itself, in this invention it is composed of several TLC SSDs.
[0040] IO scheduling: Differentiate IO characteristics and adopt different response strategies for IO. In this invention, IO requests are responded to by TLC SSD or QLCSSD.
[0041] Data tiering: Data tiering is a strategy that categorizes data according to access frequency, value, or performance requirements and stores it on different media (such as SSDs, HDDs, and tapes) to optimize storage costs and access efficiency. A typical example is storing hot data in a high-speed tier and cold data in a low-cost tier.
[0042] Data sharing: In this invention, it refers to data sharing between edge nodes. Geographically dispersed edge computing nodes directly exchange data through a collaborative protocol, reducing cloud backhaul latency and improving real-time performance. This is commonly used in IoT, CDN, and other scenarios.
[0043] This embodiment provides a hierarchical data storage method. Figure 2 This is a flowchart of a hierarchical data storage method according to an embodiment of this application, applied to a target storage controller deployed in a target storage device, such as... Figure 2 As shown, the method includes the following steps:
[0044] Step S202: Receive a data write request, wherein the data write request is used to request that target data be written into the target storage device, the target storage device includes multiple storage layers for storing data, the target storage controller is used to access the storage layers, and different storage layers have different data storage granularities;
[0045] Step S204: Divide the target storage layer into which the target data falls based on the target data volume and the data storage granularity;
[0046] Step S206: Write the target data into the target storage layer.
[0047] Through the above steps, by deploying a storage controller in the storage device, upon receiving a data write request to write target data to the storage device, the target storage layer to which the target data falls is divided according to the target data volume and data storage granularity. The storage controller directly accesses each storage layer, realizing hierarchical storage of data based on the matching relationship between data volume and storage granularity. This achieves effective utilization of storage space and persistent data storage, improving data access speed on the one hand and avoiding memory wear caused by frequent data migration during storage on the other. Therefore, it can solve the technical problem of low efficiency in data hierarchical storage in related technologies and achieve the technical effect of improving the efficiency of data hierarchical storage.
[0048] In the embodiment provided in step S202, the storage controller is an intermediate layer deployed between the data access party and the storage layer in the memory. The storage controller interfaces with the data access party and can receive data access requests and respond to data access requests to read data from the data access party. On the other hand, the storage controller can also directly access each storage layer and directly perform data access operations on each storage layer, thereby improving the data access bandwidth of the storage device.
[0049] Optionally, in this embodiment of the application, the data storage granularity is the data storage capacity of the smallest data storage unit of the memory in the storage layer. The smallest data storage unit is the basic unit for writing data in the memory. When it is necessary to modify a certain character stored in the smallest data storage unit, it is necessary to delete all the contents stored in the smallest data storage unit and perform a data rewrite operation on the data storage unit.
[0050] Optionally, in this embodiment, the storage layer is obtained by classifying multiple memories included in the target storage device according to the data storage granularity of the memory cells. A storage layer may include, but is not limited to, one or more memories. By dividing memories with the same data storage granularity into one storage layer, the storage device is layered according to the data storage granularity, and data is stored in layers according to the storage granularity. This achieves a reasonable match between the data to be written and the storage cell capacity of the memory, avoiding the problems of write amplification and short lifespan caused by small data blocks occupying large capacity storage cells. In this embodiment, the memories in the storage layer may be, but are not limited to, memories with high data access rates, such as TLC (Triple-Level Cell) solid-state memories and QLC (Quad-Level Cell) solid-state memories.
[0051] In the embodiment provided in step S204, the core of the target storage layer to which the target data falls in this application is to improve the matching relationship between the data and the data storage capacity (data storage granularity) of the occupied storage unit, so as to avoid write amplification caused by the mismatch between the data source and the data storage capacity of the storage unit, thereby causing the unoccupied storage resources to be repeatedly rewritten during data rewriting, resulting in the waste of storage resource life. Therefore, the target storage controller can match the target data volume of the target data with the data storage granularity of each storage layer to find the first storage layer whose data storage granularity matches the target data volume. Then, it checks the remaining storage resources in the first storage layer. If the remaining storage resources in the first storage layer are less than the first resource capacity, it searches upwards for a second storage layer with a data storage granularity greater than the first storage layer. The first resource capacity is the minimum remaining storage resource required for the normal operation of the first storage layer. If the remaining storage resources in the second storage layer are greater than or equal to the second resource capacity, the target data is aggregated with the first data currently to be stored on the target storage device to obtain aggregated second data whose data storage granularity matches that of the second storage layer. This second data includes the target data, and the second storage layer is then designated as the target storage layer for the second data. In this embodiment, by aggregating the target data and the first data, the data volume of the aggregated second data matches the data storage granularity of the second storage layer, thus avoiding waste of storage resources. Further, in this embodiment of the application, the method of aggregating the target data with the first data to be stored on the target storage device can be as follows: Obtain the initial data access information of all third data to be written on the target storage device. This data access information is used to indicate the operation execution cycle and execution sequence of the external device performing data rewriting operations on each third data. Then, match the target data access information of the target data with the initial data access information of the third data, thereby filtering out fourth data whose initial data access information matches the target data access information from the third data. Then, based on the data storage granularity of the second storage layer and the data volume of the target data, filter out the first data from the fourth data, and aggregate the first data and the target data to obtain the second data. In this way, the first data to be aggregated is found for the target data according to the execution cycle and execution sequence of the data rewriting operation, thereby synchronizing the rewriting operations of the target data and the first data. This avoids frequent data rewriting operations caused by using the same storage unit to store the same data, thus avoiding the waste of storage resources caused by repeated data rewriting.
[0052] Optionally, in this embodiment, the method of dividing the target storage layer into which the target data falls based on the target data volume and data storage granularity can also match the target data with the data storage granularity, thereby selecting the storage layer whose data storage granularity matches the target data volume from multiple storage layers as the target storage layer. In this embodiment, when the data volume is equal to the data storage granularity or the data volume is less than the target threshold for data storage granularity, it can be determined that the data volume of the current data matches the storage granularity of the current storage layer.
[0053] In the embodiment provided in step S206, the target data is directly written to the target storage layer by the storage controller. After the target storage controller completes the data writing operation of writing the target data to the target storage layer, in order to facilitate subsequent data retrieval, metadata representing the data writing information of the target data in the target storage layer can be generated and the target metadata of the target data can be stored in the target storage device.
[0054] As an optional embodiment, the step of dividing the target storage layer into which the target data falls based on the data volume of the target data and the data storage granularity includes:
[0055] Select an initial storage layer from the plurality of storage layers that matches the target data volume of the target data;
[0056] Based on the first data to be stored in the initial storage layer, the target data is aggregated to obtain aggregated candidate data.
[0057] The target storage layer is selected from the plurality of storage layers to match the amount of candidate data that the candidate data has.
[0058] Optionally, in this embodiment of the application, when the amount of data is equal to the data storage granularity or the amount of data is less than the target threshold for data storage granularity, it can be determined that the amount of data of the current data matches the storage granularity of the current storage layer.
[0059] Optionally, in this embodiment of the application, the aggregation process is to merge and store the target data and the first data, that is, to store the target data and the first data in the same data storage unit.
[0060] Optionally, in this embodiment, when new target data needs to be written, an initial storage layer matching the target data volume is first selected. Then, the first data currently to be stored in that layer is checked. If the access attributes (such as access frequency and access pattern) of the first data are similar to the access attributes of the target data, the first data and the target data are aggregated to form candidate data. The aggregated candidate data is then re-evaluated to determine the optimal storage layer to which it should be written.
[0061] Based on the above, technically, this embodiment optimizes storage space utilization by merging data with similar characteristics through data aggregation processing. In principle, by analyzing the first data in the initial storage layer to determine whether it has similar data access needs to the target data, it decides whether to perform aggregation processing, which helps improve the flexibility and efficiency of data storage. In terms of effect, the aggregated candidate data can more effectively utilize the storage resources of the target storage layer, reduce data fragmentation, and improve overall storage performance. In other embodiments, the storage layer allocation strategy can be dynamically adjusted to address the problem of storage layer inapplicability caused by changes in data access patterns.
[0062] As an optional embodiment, the step of performing data aggregation processing on the target data based on the first parameter data currently to be stored in the initial storage layer to obtain aggregated candidate data includes:
[0063] Obtain reference data access attributes for the first data, wherein the reference data access attributes are used to indicate the data access requirements of the first data;
[0064] Select second data from the multiple reference data that satisfies the target similarity condition between the access attributes of the reference data and the target data access attributes of the target data;
[0065] The second data and the target data are aggregated into the candidate data.
[0066] Optionally, in this embodiment, the data access attributes may include, but are not limited to, the data access type and the execution sequence, execution frequency, etc., of the data access operations corresponding to the data access type. This comparison does not impose any limitations. By matching the first execution sequence of the data rewriting operation performed on the reference data with the second execution sequence of the data rewriting operation performed on the target data, if the first execution sequence and the second execution sequence match, it can be considered that the reference data access data and the target data access data satisfy the target similarity condition.
[0067] Based on the above, technically, this embodiment improves data access efficiency by meticulously analyzing data access attributes to ensure that aggregated data exhibits similar access patterns. In principle, by comparing the similarity of data access attributes, it intelligently selects second data suitable for aggregation with the target data, which helps enhance the intelligence of data storage. In terms of effectiveness, the aggregated data better adapts to the characteristics of the target storage layer, reducing unnecessary data migration and extending the lifespan of storage devices.
[0068] As an optional embodiment, the step of dividing the target storage layer into which the target data falls based on the data volume of the target data and the data storage granularity includes:
[0069] Select an initial storage layer from the plurality of storage layers that matches the target data volume of the target data;
[0070] The initial storage layer is determined as the target storage layer for writing the target data.
[0071] Through the above, this embodiment simplifies the data storage layer selection process, directly determining the initial storage layer that matches the target data volume as the target storage layer, thereby improving the efficiency of data writing.
[0072] As an optional embodiment, the storage layer includes multiple memories with the same storage granularity, and these memories are used for distributed data storage.
[0073] Writing the target data into the target storage layer includes:
[0074] The target data is split into multiple target sub-data based on the number of memories included in the target storage layer;
[0075] The target sub-data is written into the multiple memories included in the target storage layer, wherein each memory stores one target sub-data.
[0076] Generate first metadata for each of the target sub-data, wherein the first metadata is used to indicate the write status of the corresponding target sub-data in the target storage device;
[0077] The first metadata of the target data is written into the target memory, wherein the target memory is a memory with metadata storage function included in the target storage device.
[0078] Optionally, in this embodiment, the target storage layer includes multiple memories, which can be used to perform distributed storage of data. That is, the target data to be stored is split into multiple parts and stored on different memories. This can improve the data writing speed by enabling multiple memories to write data in parallel, and can also maintain the stability of data storage in the storage system.
[0079] Optionally, in this embodiment, the target memory is a memory configured in the target storage device that has metadata storage function. It may be a storage device specifically deployed in the target storage device for storing metadata, or it may be a memory included in the storage layer of the target storage device. This solution does not limit this.
[0080] Through the above, this embodiment improves data writing speed and storage system stability by using data partitioning and distributed storage. In principle, it utilizes multiple storage devices to write target sub-data in parallel, while simultaneously recording the writing status of each sub-data to ensure data integrity. In terms of effectiveness, this method significantly increases data writing throughput, and the recording of metadata facilitates subsequent data management and maintenance.
[0081] As an optional embodiment, writing the target data into the target storage layer includes:
[0082] Candidate memories for writing the target data are selected from the plurality of memories included in the target storage layer;
[0083] Write the target data into the candidate memory;
[0084] Generate second metadata of the target data based on the write status of the target data in the candidate memory;
[0085] The second metadata of the target data is written into the target memory, wherein the target memory is a memory with metadata storage function included in the target storage device.
[0086] Optionally, in this embodiment, by screening candidate storage devices, precise data location and storage are achieved, avoiding resource waste. Based on the characteristics of the target data, the most suitable storage device is selected for data writing, and the writing status is recorded to facilitate subsequent data retrieval and updates. This method can improve the accuracy and efficiency of data storage, reduce unnecessary storage space occupation, and improve the utilization rate of storage resources.
[0087] As an optional embodiment, after writing the target data into the target storage layer, the method further includes:
[0088] Receive an update request for the target data, wherein the update request for the target data is used to request an update of the target data;
[0089] Obtain the update data carried in the update request for updating the target data;
[0090] The updated data is appended to the target data in the target storage layer.
[0091] Optionally, in this embodiment of the application, the append writing of data is not done by directly using the updated data to update the target data when the target data needs to be updated, but by writing the updated data of the target data on top of the target data stored in the target storage layer, thereby combining the target data and the updated data into a whole.
[0092] Through the above methods, efficient data updates are achieved by appending updated data, avoiding complete data rewriting, reducing the complexity of storage operations, and promptly obtaining updated data by detecting update requests for target data and appending it to the target storage layer. This helps maintain the timeliness and accuracy of the data. This method can significantly reduce the cost of data updates, improve the efficiency of data updates, and at the same time reduce wear and tear on storage devices and extend their lifespan.
[0093] As an optional embodiment, appending the updated data to the target data in the target storage layer includes:
[0094] The updated data is written to the first unoccupied storage location in the target storage layer;
[0095] Establish a binding relationship between the data stored in the first storage location and the data stored in the second storage location, wherein the second storage location is the storage location in the target storage layer used to store the target data.
[0096] Optionally, in this embodiment, when appending updated data, it is not written to the original storage location where the target data was originally stored. Therefore, this requires performing a data erasure operation on the first storage location where the target data was originally stored, which affects the lifespan of the memory. By writing updated data to the first storage location and establishing a binding relationship between the first and second storage locations, the target data and updated data are merged into a whole.
[0097] By writing updated data to unoccupied storage locations and establishing a binding relationship with the existing data, orderly data updates are achieved, avoiding data chaos. By locating unoccupied storage locations to ensure the correct writing of updated data, and by maintaining data continuity and consistency through binding relationships, this method improves the accuracy and security of data updates, reduces the risk of data loss, and facilitates subsequent data retrieval and management.
[0098] As an optional embodiment, the step of constructing the binding relationship between the data stored in the first storage location and the data stored in the second storage location includes:
[0099] The updated data is used to update the third metadata of the target data stored in the target storage device using the data write information in the target storage layer.
[0100] Optionally, in this embodiment of the application, the binding relationship between the updated data stored in the first storage location and the target data stored in the second storage location can also be constructed by adding metadata of the updated data to the storage device and constructing a binding relationship between the metadata of the updated data and the metadata of the target data. That is, an index tree structure of metadata can be created in the target storage device, and key-value pairs can be constructed in the index tree structure with the data identifier of the target data and the metadata under the target data. The metadata of the updated data is added to the index tree structure of the target data, and the version tag of the metadata is configured in the index tree structure. Then, the relationship of the data corresponding to the metadata can be determined according to the version tag.
[0101] Through the above, by updating third-party metadata, data update information is recorded and managed, improving the transparency of data updates. By modifying the target data's metadata and recording the write information of updated data, this facilitates subsequent data queries and maintenance. This method improves data management efficiency, reduces the uncertainty brought about by data updates, and facilitates data tracking and auditing.
[0102] As an optional embodiment, after establishing the binding relationship between the data stored in the first storage location and the data stored in the second storage location, the method further includes:
[0103] Detect the data update cycle of the target data in the target storage layer;
[0104] If the data update cycle is greater than or equal to a preset duration, the target data and the updated data are aggregated to obtain aggregated data;
[0105] Candidate storage layers whose data storage granularity matches the data volume of the aggregated data are selected from the reference storage layers, wherein the reference storage layer is a storage layer other than the target storage layer among the multiple storage layers included in the target storage device;
[0106] The aggregated data is written into the candidate storage layer.
[0107] As an optional embodiment, the target storage device is a target edge storage device among a plurality of edge storage devices included in the edge storage system, and the storage controllers in the edge storage devices are interconnected.
[0108] After writing the target data into the target storage layer, the method further includes:
[0109] The target data and the target metadata of the target data are transmitted to a reference storage controller deployed on a reference edge storage device. The target metadata is used to indicate the data writing method of the target data in the target edge storage device, and the reference storage controller is used to write the target data into the reference edge storage device according to the data writing method indicated by the target metadata.
[0110] This embodiment achieves redundant data backup through cross-edge storage device data synchronization, improving data reliability and availability. By communicating between storage controllers, target data and metadata are transmitted to the reference edge storage device, ensuring that data can be stored consistently across multiple devices. This method enhances data disaster recovery capabilities, ensuring normal access and use of data even if one edge storage device fails.
[0111] As an optional embodiment, the storage controllers are connected via a target communication bus, which is used to implement data access functions between the storage controllers.
[0112] The step of transmitting the target data and the target metadata of the target data to a reference storage controller deployed on a reference edge storage device includes:
[0113] The target communication bus is invoked to send a data update instruction to the reference storage controller, wherein the data update instruction is used to indicate that the target data has been updated in the target edge storage device. The reference storage controller is used to respond to the data update instruction by invoking the target communication bus to read the target metadata stored in the target edge storage device, and to read the target data in the target storage layer of the target edge storage device according to the data writing method indicated by the target metadata.
[0114] Optionally, in this embodiment, the target communication bus may be, but is not limited to, the NVMe over Fabric network bus protocol. This communication bus can extend the NVMe storage protocol to remote servers through transmission media such as fiber optics and Ethernet, enabling high-performance, low-latency distributed storage access. A novel high-performance storage engine is designed based on NoF technology, and a streamlined software architecture for efficient data access is built based on the NVMe over Fabric (NoF) protocol. This architecture breaks the limitations of traditional multi-layer protocol stacks, enabling direct access from local clients to remote NVMe storage devices through a new high-speed network link. This bypasses intermediate protocol conversion stages such as TCP / iSCSI, significantly shortening the data transmission path, reducing software stack processing latency, and decreasing CPU resource consumption in protocol parsing and context switching, thereby improving the overall efficiency of the storage system. Edge storage devices directly carry protocol layer parsing and performance layer read / write logic, enabling direct access to remote NVMe storage devices through high-speed network links such as RoCE / InfiniBand, bypassing intermediate forwarding nodes of TCP / iSCSI in traditional storage architectures, thus improving data transmission efficiency. The unified data engine is responsible for the space pooling management of the performance layer TLC SSD and the capacity layer QLC SSD, integrating read and write operation acceleration and multi-replica / erasure coding consistency guarantee. Through direct client connection and intelligent engine scheduling, it builds a high-performance storage architecture with low latency and high reliability.
[0115] This application provides a possible application scenario for a high-performance edge storage system for novel computing applications. Figure 3 This is a schematic diagram of an optional edge storage system according to an embodiment of this application, such as... Figure 3 As shown, the edge computing system consists of m edge storage nodes, n terminal devices, and one cloud data center. The edge storage nodes are connected to the terminal devices, the cloud data center, and other edge storage nodes via a network (Ethernet, Internet, LAN, or WAN, etc.). In this system, the edge storage nodes act as an intermediate layer, serving data requests from the cloud and terminals, as well as data requests from other nearby edge storage nodes.
[0116] To address the demands of massive amounts of data from terminals for large-capacity storage, the requirements of cloud data centers for high-bandwidth read / write performance, the need for low-latency response from terminals, and the need for efficient data sharing between edge storage nodes.
[0117] This invention employs a heterogeneous storage solution using TLC SSDs and QLC SSDs to construct a performance layer and a capacity layer (collectively referred to as the upper-level storage layer). The performance layer, built with TLC SSDs, leverages their superior read / write performance and lifespan to handle small data block writes, efficiently handling operations such as user small data block writes and metadata writes. The capacity layer, relying on QLC SSDs, utilizes their large capacity and low cost to focus on large data block write requests. Furthermore, an NVMe over Fabric technology is used to design a data storage engine, aggregating and flushing small I / O operations from the performance layer to the storage pool, achieving rapid data persistence. RDMA technology is also combined to enable efficient data sharing between edge nodes. In metadata management, a distributed KV metadata engine is designed to linearly improve the system's metadata performance. Through these technical solutions, the performance, storage efficiency, and scalability of the edge storage system are improved.
[0118] 1. Edge storage node system design.
[0119] In this embodiment, a novel architecture is designed for edge storage nodes. Figure 4 This is an optional edge node architecture diagram according to an embodiment of this application, which consists of the following parts:
[0120] 1) Smart NIC: In this embodiment, the smart NIC is responsible for classifying I / O and interacting with the edge switch to enable communication with other edge storage nodes;
[0121] 2) Performance Pool: Composed of TLC SSDs, it is primarily responsible for small IO read / write operations, small IO aggregation and flushing, and metadata storage and management. Small IO requests in this specification mainly originate from terminal requests and other edge storage nodes;
[0122] 3) Capacity Pool: Composed of QLC SSDs, primarily responsible for large I / O read and write operations. By bypassing the software stack overhead of traditional read and write paths, it improves system read and write bandwidth performance. These large I / O read and write operations can originate from the cloud, terminals, and other edge storage nodes.
[0123] 4) Metadata pool: Composed of TLC SSDs. Since the metadata itself is relatively small and read / write operations are frequent, the metadata is placed in TLC media.
[0124] 5) Data Engine / Metadata Engine: Based on NoF technology, this invention designs a data engine and a metadata engine to decouple computing and storage resources, realize the pooling of storage resources, and enhance the scalability of the system.
[0125] 6) Edge Switch: Data sharing between edge storage nodes can be achieved through the switch. This invention adopts a data access protocol based on RDMA.
[0126] Figure 4 This document describes the I / O request processing of storage nodes and presents the distributed storage system architecture of this project. It mainly consists of a performance layer, a capacity layer, a data engine, and a metadata engine, along with modules for space management and metadata management. The performance layer primarily manages IOPs and a high-bandwidth, low-latency TLC SSD performance pool, enabling high-performance access to small I / Os (i.e., small data in the diagram) and metadata. The capacity layer manages a capacity pool composed of QLC SSDs, handling high-speed read and write requests for large I / Os (i.e., large data in the diagram). During the response to read and write requests, the client determines the operation type and I / O (data) size based on the request initiator and the read / write data size threshold. Small I / Os (small data) follow a 1-4 process: passing through the protocol layer, the performance layer manages space and metadata, the data engine performs persistence via NoF, and finally, the metadata is updated. Large I / Os follow a-c process: passing through the metadata engine, the data is persisted to the capacity layer, and the metadata is updated.
[0127] This application embodiment mainly achieves efficient hierarchical data storage through the following three aspects:
[0128] 1) Data tiering technology based on heterogeneous storage media.
[0129] This storage system innovatively adopts a hybrid TLC / QLC media tiered architecture to precisely match different storage needs. Figure 5 This is a schematic diagram of an optional hybrid media tiered storage according to an embodiment of this application, such as... Figure 5 As shown, the performance layer is built using TLCSSD, leveraging its excellent read / write performance and lifespan to handle small data block write scenarios, such as user small data block writes and metadata writes, all efficiently. The capacity layer relies on QLC SSDs, utilizing their large capacity and low cost to focus on large data block write requests. Through an IO classification and scheduling mechanism, small data block writes are handled by the performance layer TLC, avoiding write amplification and reduced lifespan issues caused by small data block writes in QLC, ensuring improved lifespan metrics for QLC SSDs, and achieving an optimal balance between storage performance and cost.
[0130] In addition, append-only write methods are used for both the performance layer and the storage layer to achieve fast data writing. When the capacity of the performance layer reaches a predetermined threshold, the aggregated data from small I / O operations in the performance layer is flushed down to the capacity layer.
[0131] 2) High-performance storage engine based on NoF.
[0132] The project designs a novel high-performance storage engine based on NoF technology, and builds a streamlined software architecture for efficient data access based on the NVMe over Fabric (NoF) protocol. This architecture breaks through the limitations of traditional multi-layered protocol stacks, enabling local clients to directly access remote NVMe storage devices through a new high-speed network link. Bypassing intermediate protocol conversion links such as TCP / iSCSI, the data transmission path is significantly shortened, software stack processing latency is significantly reduced, and CPU resource consumption in protocol parsing and context switching is reduced, thereby improving the overall efficiency of the storage system.
[0133] Figure 6 This is a schematic diagram of data storage for an optional storage engine according to an embodiment of this application, as shown in the attached diagram. Figure 6 As shown, based on the NoF protocol, the client directly carries the protocol layer parsing and performance layer read / write logic. Through high-speed network links such as RoCE / InfiniBand, it achieves direct access to remote NVMe storage devices, bypassing the intermediate forwarding nodes of TCP / iSCSI in traditional storage architectures, thus improving data transmission efficiency. The unified data engine is responsible for the space pooling management of performance-layer TLC SSDs and capacity-layer QLC SSDs, integrating read / write operation acceleration and multi-replica / erasure coding consistency guarantees. Through direct client connection and intelligent engine scheduling, it constructs a low-latency, highly reliable, high-performance storage architecture.
[0134] Furthermore, a three-replica storage method is used for data at the performance layer, and an 8+2 erasure storage scheme is used for data at the capacity layer. This further improves the storage efficiency of edge storage nodes.
[0135] 3) High-performance metadata service based on distributed key-value pairs.
[0136] This invention designs a high-performance metadata service based on distributed key-value pairs. It employs a Multi-Raft distributed architecture (corresponding to the target communication bus mentioned earlier) to achieve horizontal data scaling, while combining a storage engine to optimize single-machine performance, thus achieving the dual advantages of distributed elastic scaling and extreme single-node performance. Intelligent sharding ensures globally uniform data distribution, multi-replica replication guarantees high availability, and dynamic load balancing automatically adjusts resource allocation, ensuring seamless system scaling with data growth. Furthermore, an adaptive scheduling strategy monitors cluster status in real time and optimizes request routing, improving throughput and reducing latency while maintaining strong consistency. This makes the system highly scalable, high-performance, and highly reliable, suitable for massive data storage and high-concurrency access scenarios.
[0137] Based on the above, this invention employs a heterogeneous media storage solution combining TLC and QLC solid-state drives (SSDs), with TLC forming the performance layer and QLC forming the capacity layer, to address the performance issues of HDD storage solutions. Furthermore, by designing an IO scheduling scheme, system performance is improved, and the lifespan of QLCs is extended. Then, NVMe over Fabric (including but not limited to RDMA, FC-NVMe, and CXL) technology is introduced to significantly improve the performance of data transmission between edge storage nodes vertically (cloud and terminal) and horizontally (adjacent edge storage nodes) for data sharing and access, reducing access latency. Moreover, addressing the performance bottlenecks of traditional metadata architectures, a high-performance metadata service architecture based on distributed key-value pairs is designed, using a Multi-Raft distributed architecture to achieve horizontal data scaling. Finally, a 3-replica storage scheme is used in the performance layer, and an erasure propagation storage scheme is used in the capacity layer, thereby solving the problem of low storage efficiency.
[0138] This invention achieves a balance between performance and cost through a heterogeneous storage architecture combining TLC and QLC, extending the lifespan of QLCs with intelligent IO scheduling; it utilizes NVMe over Fabric (supporting RDMA / FC-NVMe / CXL) technology to achieve microsecond-level low-latency transmission between edge nodes; it employs a distributed KV metadata service and a Multi-Raft architecture to enhance scalability; and it significantly improves storage efficiency while ensuring high availability through a hybrid redundancy strategy of 3 replicas at the performance layer and erasure coding at the capacity layer. The overall solution significantly optimizes storage performance, network latency, and resource utilization in edge computing scenarios.
[0139] The above embodiments of this application have the following beneficial effects: This invention proposes an innovative solution to the core problems of performance bottlenecks, low storage efficiency, and high data sharing latency in edge storage systems. Traditional solutions adopt a TLCSSD+HDD layered architecture, which faces shortcomings such as poor HDD performance (random IOPS less than 200), limited TLC capacity (8TB per disk), and short QLC lifespan (1000 P / E cycles). At the same time, the reliance on complex protocol stacks (such as TCP / IP+iSCSI) results in cross-node latency exceeding 500ms, and storage efficiency is only 30%.
[0140] Advantages of using three key technologies in edge storage systems:
[0141] 1) Heterogeneous storage tiering: TLC SSD (performance tier) handles small I / O and metadata, while QLC SSD (capacity tier) handles large I / O. Intelligent scheduling reduces QLC write amplification and significantly improves lifespan.
[0142] 2) NoF storage engine: Based on NVMe over Fabric (RDMA / FC-NVMe), it enables direct connection between nodes, reducing protocol interaction steps by 70% and latency to 50μs;
[0143] 3) Distributed metadata service: The Multi-Raft architecture supports linear scaling, with metadata throughput reaching 1M OPS and fault recovery completed in seconds.
[0144] This invention can significantly improve the performance and efficiency of edge storage systems. In terms of storage performance, TLC random read / write reaches 100K IOPS, and QLC sequential bandwidth reaches 2.8GB / s, which is 3 times higher than traditional solutions. In terms of resource utilization, the performance layer has 3 replicas + the capacity layer has 8+2 erasure coding, which improves storage efficiency from 30% to 80%. In terms of network optimization, RDMA achieves a data locality rate of 90% and improves bandwidth utilization by 40%.
[0145] Furthermore, this invention is applicable to a wider range of application scenarios and boasts strong system scalability. It is suitable for low-latency scenarios such as Industrial IoT and Vehicle-to-Everything (V2X) and supports exabyte-level (EB) data expansion. Dynamic load balancing technology improves data migration speed by 50% during expansion, avoiding hotspot skew. Intelligent network interface cards (NICs) collaborate with edge switches to achieve cross-node data sharing without cloud relay.
[0146] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0147] Embodiments of this application also provide a storage device for hierarchical data storage. Figure 7 This is a schematic diagram of a storage device structure for optional data tiered storage according to an embodiment of this application, such as... Figure 7 As shown, the storage device includes:
[0148] The system includes a storage controller and multiple storage tiers, with the storage controller and storage tiers connected together. Different storage tiers have different data storage granularities.
[0149] The storage controller is configured to receive a data write request, wherein the data write request is used to request that target data be written into the storage device; divide the target storage layer into which the target data falls according to the target data volume and the data storage granularity; and write the target data into the target storage layer.
[0150] Based on the above, by deploying a storage controller in the storage device, upon receiving a data write request to write target data to the storage device, the target storage layer to which the target data falls is divided according to the target data volume and data storage granularity. The storage controller directly accesses each storage layer, achieving hierarchical storage of data based on the matching relationship between data volume and storage granularity. This enables efficient utilization of storage space and persistent data storage. On the one hand, it improves the data access speed; on the other hand, it avoids memory wear caused by frequent data migration during storage. Therefore, it can solve the technical problem of low efficiency in data hierarchical storage in related technologies and achieve the technical effect of improving the efficiency of data hierarchical storage.
[0151] As an optional embodiment, the storage device is a target edge storage device among a plurality of edge storage devices included in the edge storage system, and the storage controllers in the edge storage devices are interconnected.
[0152] The target storage controller in the target edge storage device is used to transmit the target data and the target metadata of the target data to the reference storage controller deployed on the reference edge storage device, wherein the target metadata is used to indicate the data writing method of the target data in the target edge storage device;
[0153] The reference storage controller is used to write the target data to the reference edge storage device according to the data writing method indicated by the target metadata.
[0154] As an optional embodiment, the storage controllers are connected via a target communication bus, which is used to implement data access functions between the storage controllers;
[0155] The target edge storage device is configured to invoke the target communication bus to send a data update instruction to the reference storage controller, wherein the data update instruction is used to indicate that the target data has been updated in the target edge storage device;
[0156] The reference edge storage device is configured to respond to the data update instruction by calling the target communication bus to read the target metadata stored in the target edge storage device, and to read the target data in the target storage layer of the target edge storage device according to the data writing method indicated by the target metadata.
[0157] As an optional embodiment, the storage layer includes multiple memories with the same storage granularity, and the storage controller is connected to the multiple memories.
[0158] Multiple memory units are used for distributed data storage;
[0159] The storage controller is configured to split the target data into multiple target sub-data according to the number of memories included in the target storage layer; write the multiple target sub-data onto the multiple memories included in the target storage layer, wherein each memory stores one target sub-data; generate first metadata for each target sub-data, wherein the first metadata is used to indicate the write status of the corresponding target sub-data in the target storage device; and write the first metadata of the target data into the target memory, wherein the target memory is a memory with metadata storage function included in the target storage device.
[0160] Embodiments of this application also provide a hierarchical data storage device. Figure 8 This is a structural block diagram of a data tiered storage device according to an embodiment of this application, applied to a target storage controller deployed in a target storage device, such as... Figure 8 As shown, the device includes:
[0161] The first receiving module is used to receive a data write request, wherein the data write request is used to request that target data be written into the target storage device, the target storage device includes multiple storage layers for storing data, and the target storage controller is used to access the data of the storage layers, and the different storage layers have different data storage granularities.
[0162] The partitioning module is used to partition the target storage layer into which the target data falls based on the target data volume and the data storage granularity.
[0163] The first writing module is used to write the target data into the target storage layer.
[0164] The above-described device, by deploying a storage controller in the storage device, upon receiving a data write request to write target data to the storage device, divides the target storage layer according to the target data volume and data storage granularity, and directly accesses each storage layer through the storage controller. This achieves hierarchical storage of data based on the matching relationship between data volume and storage granularity, thereby realizing efficient utilization of storage space and persistent data storage. On the one hand, it improves the data access speed, and on the other hand, it avoids memory wear caused by frequent data migration during storage. Therefore, it can solve the technical problem of low efficiency in data hierarchical storage in related technologies and achieve the technical effect of improving the efficiency of data hierarchical storage.
[0165] Optionally, the partitioning module includes:
[0166] The first filtering unit is used to filter out an initial storage layer from the plurality of storage layers that matches the target data volume of the target data;
[0167] The processing unit is configured to perform data aggregation processing on the target data based on the first data to be stored in the initial storage layer, and obtain aggregated candidate data;
[0168] The second filtering unit is used to filter out the target storage layer from the plurality of storage layers that matches the amount of candidate data that the candidate data has.
[0169] Optional, processing unit, used for:
[0170] Obtain reference data access attributes for the first data, wherein the reference data access attributes are used to indicate the data access requirements of the first data;
[0171] Select second data from the multiple reference data that satisfies the target similarity condition between the access attributes of the reference data and the target data access attributes of the target data;
[0172] The second data and the target data are aggregated into the candidate data.
[0173] Optionally, the partitioning module includes:
[0174] The third filtering unit is used to filter out an initial storage layer from the plurality of storage layers that matches the target data volume of the target data;
[0175] A determining unit is configured to determine the initial storage layer as the target storage layer for writing the target data.
[0176] Optionally, the storage layer includes multiple memories with the same storage granularity, and these memories are used for distributed data storage.
[0177] The first writing module includes:
[0178] The splitting unit is used to split the target data according to the number of memories included in the target storage layer to obtain multiple target sub-data;
[0179] The first writing unit is used to write multiple target sub-data into multiple memories included in the target storage layer, wherein each memory stores one target sub-data.
[0180] A first generation unit is configured to generate first metadata for each of the target sub-data, wherein the first metadata is used to indicate the write status of the corresponding target sub-data in the target storage device;
[0181] The second writing unit is used to write the first metadata of the target data into the target memory, wherein the target memory is a memory with metadata storage function included in the target storage device.
[0182] Optionally, the first write module includes:
[0183] The fourth filtering unit is used to filter out candidate memories for writing the target data from the plurality of memories included in the target storage layer;
[0184] The third writing unit is used to write the target data into the candidate memory;
[0185] The second generation unit is used to generate second metadata of the target data according to the write status of the target data in the candidate memory;
[0186] The fourth writing unit is used to write the second metadata of the target data into the target memory, wherein the target memory is a memory with metadata storage function included in the target storage device.
[0187] Optionally, the device further includes:
[0188] The second receiving module is configured to receive an update request for the target data after writing the target data into the target storage layer, wherein the update request for the target data is used to request an update of the target data;
[0189] The acquisition module is used to acquire the update data carried in the update request for updating the target data;
[0190] The second write module is used to append the updated data to the target data in the target storage layer.
[0191] Optionally, the second write module includes:
[0192] The fifth writing unit is used to write the updated data to a first storage location in the target storage layer that is in an unoccupied state;
[0193] A construction unit is used to construct a binding relationship between data stored in the first storage location and data stored in the second storage location, wherein the second storage location is a storage location in the target storage layer used to store the target data.
[0194] Optionally, the building unit is used for:
[0195] The updated data is used to update the third metadata of the target data stored in the target storage device using the data write information in the target storage layer.
[0196] Optionally, the device further includes:
[0197] The detection module is used to detect the data update cycle of the target data in the target storage layer after the binding relationship between the data stored in the first storage location and the data stored in the second storage location is established;
[0198] The aggregation module is used to aggregate the target data and the updated data to obtain aggregated data when the data update cycle is greater than or equal to a preset duration.
[0199] A filtering module is used to filter out candidate storage layers from the reference storage layer whose data storage granularity matches the data volume of the aggregated data, wherein the reference storage layer is a storage layer other than the target storage layer among the multiple storage layers included in the target storage device;
[0200] The third writing module is used to write the aggregated data into the candidate storage layer.
[0201] Optionally, the target storage device is a target edge storage device among multiple edge storage devices included in the edge storage system, and the storage controllers in the edge storage devices are interconnected.
[0202] The device further includes:
[0203] A transmission module is configured to, after writing the target data into the target storage layer, transmit the target data and the target metadata of the target data to a reference storage controller deployed on a reference edge storage device, wherein the target metadata is used to indicate the data writing method of the target data in the target edge storage device, and the reference storage controller is configured to write the target data into the reference edge storage device according to the data writing method indicated by the target metadata.
[0204] Optionally, the storage controllers are connected via a target communication bus, which is used to implement data access functions between the storage controllers.
[0205] The transmission module includes:
[0206] A sending unit is configured to invoke the target communication bus to send a data update instruction to the reference storage controller, wherein the data update instruction is configured to indicate that the target data has been updated in the target edge storage device, and the reference storage controller is configured to respond to the data update instruction by invoking the target communication bus to read the target metadata stored in the target edge storage device, and read the target data in the target storage layer of the target edge storage device according to the data writing method indicated by the target metadata.
[0207] For a description of the features in the embodiment corresponding to the hierarchical data storage device, please refer to the relevant description in the embodiment corresponding to the hierarchical data storage method, which will not be repeated here.
[0208] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the hierarchical storage method for data.
[0209] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described hierarchical storage method embodiments of data at runtime.
[0210] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0211] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described hierarchical data storage method embodiments.
[0212] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described hierarchical data storage method embodiments.
[0213] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0214] The foregoing has provided a detailed description of a hierarchical data storage method, apparatus, storage medium, and electronic device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to aid in understanding the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A hierarchical data storage method, characterized in that, The target storage controller used in the target storage device includes: A data write request is received, wherein the data write request is used to request that target data be written into the target storage device, the target storage device includes multiple storage layers for storing data, the target storage controller is used to access the storage layers, and different storage layers have different data storage granularities; The target storage layer to which the target data falls is determined based on the target data volume and the data storage granularity. Write the target data into the target storage layer; The target storage device is a target edge storage device among multiple edge storage devices included in the edge storage system. The storage controllers in the edge storage devices are interconnected. After writing the target data into the target storage layer, the method further includes: transferring the target data and the target metadata of the target data to a reference storage controller deployed on the reference edge storage device. The target metadata is used to indicate the data writing method of the target data in the target edge storage device, and the reference storage controller is used to write the target data into the reference edge storage device according to the data writing method indicated by the target metadata.
2. The method according to claim 1, characterized in that, The step of dividing the target storage layer into which the target data falls based on the data volume of the target data and the data storage granularity includes: Select an initial storage layer from the plurality of storage layers that matches the target data volume of the target data; Based on the first data to be stored in the initial storage layer, the target data is aggregated to obtain aggregated candidate data. The target storage layer is selected from the plurality of storage layers to match the amount of candidate data that the candidate data has.
3. The method according to claim 2, characterized in that, The step of performing data aggregation processing on the target data based on the first data to be stored in the initial storage layer to obtain aggregated candidate data includes: Obtain reference data access attributes for the first data, wherein the reference data access attributes are used to indicate the data access requirements of the first data; Select second data from the multiple reference data that satisfies the target similarity condition between the access attributes of the reference data and the target data access attributes of the target data; The second data and the target data are aggregated into the candidate data.
4. The method according to claim 1, characterized in that, The step of dividing the target storage layer into which the target data falls based on the data volume of the target data and the data storage granularity includes: Select an initial storage layer from the plurality of storage layers that matches the target data volume of the target data; The initial storage layer is determined as the target storage layer for writing the target data.
5. The method according to claim 1, characterized in that, The storage layer includes multiple memories with the same storage granularity, which are used for distributed data storage. Writing the target data into the target storage layer includes: The target data is split into multiple target sub-data based on the number of memories included in the target storage layer; The target sub-data is written into the multiple memories included in the target storage layer, wherein each memory stores one target sub-data. Generate first metadata for each of the target sub-data, wherein the first metadata is used to indicate the write status of the corresponding target sub-data in the target storage device; The first metadata of the target data is written into the target memory, wherein the target memory is a memory with metadata storage function included in the target storage device.
6. The method according to claim 1, characterized in that, Writing the target data into the target storage layer includes: Candidate memories for writing the target data are selected from the plurality of memories included in the target storage layer; Write the target data into the candidate memory; Generate second metadata of the target data based on the write status of the target data in the candidate memory; The second metadata of the target data is written into the target memory, wherein the target memory is a memory with metadata storage function included in the target storage device.
7. The method according to claim 1, characterized in that, After writing the target data into the target storage layer, the method further includes: Receive an update request for the target data, wherein the update request for the target data is used to request an update of the target data; Obtain the update data carried in the update request for updating the target data; The updated data is appended to the target data in the target storage layer.
8. The method according to claim 7, characterized in that, The step of appending the updated data to the target data in the target storage layer includes: The updated data is written to the first unoccupied storage location in the target storage layer; Establish a binding relationship between the data stored in the first storage location and the data stored in the second storage location, wherein the second storage location is the storage location in the target storage layer used to store the target data.
9. The method according to claim 8, characterized in that, The process of establishing the binding relationship between the data stored in the first storage location and the data stored in the second storage location includes: The updated data is used to update the third metadata of the target data stored in the target storage device using the data write information in the target storage layer.
10. The method according to claim 9, characterized in that, After establishing the binding relationship between the data stored in the first storage location and the data stored in the second storage location, the method further includes: Detect the data update cycle of the target data in the target storage layer; If the data update cycle is greater than or equal to a preset duration, the target data and the updated data are aggregated to obtain aggregated data; Candidate storage layers whose data storage granularity matches the data volume of the aggregated data are selected from the reference storage layers, wherein the reference storage layer is a storage layer other than the target storage layer among the multiple storage layers included in the target storage device; The aggregated data is written into the candidate storage layer.
11. The method according to claim 1, characterized in that, The storage controllers are connected via a target communication bus, which is used to enable data access between the storage controllers. The step of transmitting the target data and the target metadata of the target data to a reference storage controller deployed on a reference edge storage device includes: The target communication bus is invoked to send a data update instruction to the reference storage controller, wherein the data update instruction is used to indicate that the target data has been updated in the target edge storage device. The reference storage controller is used to respond to the data update instruction by invoking the target communication bus to read the target metadata stored in the target edge storage device, and to read the target data in the target storage layer of the target edge storage device according to the data writing method indicated by the target metadata.
12. A data tiered storage device, characterized in that, include: The system includes a storage controller and multiple storage tiers, with the storage controller and storage tiers connected together. Different storage tiers have different data storage granularities. The storage controller is configured to receive a data write request, wherein the data write request is used to request that target data be written into the storage device; divide the target storage layer into which the target data falls according to the target data volume and the data storage granularity; and write the target data into the target storage layer. The storage device is a target edge storage device among multiple edge storage devices included in the edge storage system. The storage controllers in the edge storage devices are interconnected. After writing the target data into the target storage layer, the storage controller is further configured to: transfer the target data and the target metadata of the target data to a reference storage controller deployed on a reference edge storage device. The target metadata is used to indicate the data writing method of the target data in the target edge storage device, and the reference storage controller is used to write the target data into the reference edge storage device according to the data writing method indicated by the target metadata.
13. The storage device according to claim 12, characterized in that, The storage device is a target edge storage device among multiple edge storage devices included in the edge storage system, and the storage controllers in the edge storage devices are interconnected. The target storage controller in the target edge storage device is used to transmit the target data and the target metadata of the target data to the reference storage controller deployed on the reference edge storage device, wherein the target metadata is used to indicate the data writing method of the target data in the target edge storage device; The reference storage controller is used to write the target data to the reference edge storage device according to the data writing method indicated by the target metadata.
14. The storage device according to claim 13, characterized in that, The storage controllers are connected via a target communication bus, which is used to enable data access between the storage controllers. The target edge storage device is configured to invoke the target communication bus to send a data update instruction to the reference storage controller, wherein the data update instruction is used to indicate that the target data has been updated in the target edge storage device; The reference edge storage device is configured to respond to the data update instruction by calling the target communication bus to read the target metadata stored in the target edge storage device, and to read the target data in the target storage layer of the target edge storage device according to the data writing method indicated by the target metadata.
15. The storage device according to claim 12, characterized in that, The storage layer includes multiple memories with the same storage granularity, and the storage controller is connected to the multiple memories. Multiple memory units are used for distributed data storage; The storage controller is configured to split the target data into multiple target sub-data according to the number of memories included in the target storage layer; write the multiple target sub-data onto the multiple memories included in the target storage layer, wherein each memory stores one target sub-data; generate first metadata for each target sub-data, wherein the first metadata is used to indicate the write status of the corresponding target sub-data in the target storage device; and write the first metadata of the target data into the target memory, wherein the target memory is a memory with metadata storage function included in the target storage device.
16. A hierarchical data storage device, characterized in that, The target storage controller used in the target storage device includes: The first receiving module is used to receive a data write request, wherein the data write request is used to request that target data be written into the target storage device, the target storage device includes multiple storage layers for storing data, and the target storage controller is used to access the data of the storage layers, and the different storage layers have different data storage granularities. The partitioning module is used to partition the target storage layer into which the target data falls based on the target data volume and the data storage granularity. The first writing module is used to write the target data into the target storage layer; The target storage device is a target edge storage device among multiple edge storage devices included in the edge storage system. The storage controllers in the edge storage devices are interconnected. The device further includes: a transmission module, used to transmit the target data and the target metadata of the target data to a reference storage controller deployed on the reference edge storage device after writing the target data into the target storage layer. The target metadata is used to indicate the data writing method of the target data in the target edge storage device, and the reference storage controller is used to write the target data into the reference edge storage device according to the data writing method indicated by the target metadata.
17. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the hierarchical storage method for data as described in any one of claims 1 to 11 when executing the computer program.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the hierarchical storage method for data as described in any one of claims 1 to 11.
19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the hierarchical storage method for data as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Data storage method and device
CN106406759A