Data management method and apparatus
By distinguishing and managing data blocks based on data heat attributes in solid-state drives, the problem of low data management efficiency is solved, and more efficient data management and utilization is achieved.
Patent Information
- Application Number
- CN202510341115.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-21
AI Technical Summary
In the prior art, the problem of low data management efficiency is mainly due to the high probability of spam data generation due to the indiscriminate management of data of different popularity, which affects the overall efficiency.
By obtaining the data to be written to the solid state hard disk and its heat attributes, the data is written to data blocks of different heat conditions according to the heat attributes, and data with different heat conditions are distinguished and managed.
It reduces the probability of spam data generation, improves the efficiency and data utilization of data management, and ensures efficient update of hot data and the effectiveness of cold data.
Smart Images

Figure CN119861879B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computers, and more particularly, to a data management method and apparatus. Background Art
[0002] In the data management scenario, data with different heat levels is usually shuffled and mixed into the same data block of a solid-state drive. Since hot data has a high update frequency and a high probability of becoming garbage data, while cold data has a low update frequency and may still be valid data after long-term writing, this indiscriminate data management method affects the overall efficiency of data management, thus leading to the problem of low data management efficiency.
[0003] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0004] Embodiments of the present application provide a data management method and apparatus to at least solve the technical problem of low data management efficiency in related technologies.
[0005] According to one aspect of the embodiments of the present application, a data management method is provided, including: obtaining a plurality of data to be written into a solid-state drive; obtaining the heat level attribute corresponding to each of the plurality of data; and writing each of the data into each data block in the solid-state drive according to the heat level attribute, where different data blocks correspond to different heat level attribute conditions.
[0006] According to another aspect of the embodiments of the present application, a data management apparatus is further provided, including: a first obtaining unit configured to obtain a plurality of data to be written into a solid-state drive; a second obtaining unit configured to obtain the heat level attribute corresponding to each of the plurality of data; and a writing unit configured to write each of the data into each data block in the solid-state drive according to the heat level attribute, where different data blocks correspond to different heat level attribute conditions.
[0007] As an optional solution, the second obtaining unit includes: a first obtaining module configured to obtain the data type corresponding to each of the data; and a second obtaining module configured to obtain the heat level attribute according to the data type.
[0008] As an optional solution, the second obtaining module includes: a first obtaining sub-module configured to obtain the heat level corresponding to each of the data according to the data type, where different data types correspond to different heat levels; and a second obtaining sub-module configured to obtain the heat level attribute according to the heat level, where the different heat level attribute conditions are different heat level intervals.
[0009] As an alternative solution, the above-mentioned first acquisition module includes at least one of the following: a third acquisition sub-module for acquiring a specified data type corresponding to the first data among the above-mentioned multiple data; a fourth acquisition sub-module for acquiring a stored data type corresponding to the second data among the above-mentioned multiple data; a fifth acquisition sub-module for acquiring a service data type corresponding to the third data among the above-mentioned multiple data.
[0010] As an alternative solution, the above-mentioned first acquisition sub-module includes: a first determination subunit for determining that the heat corresponding to the data of the specified data type is greater than the heat corresponding to the data of the stored data type when the data of the specified data type and the data of the stored data type are acquired.
[0011] As an alternative solution, the above-mentioned first acquisition sub-module includes: a second determination subunit for determining that the heat corresponding to the data of the service data type is greater than the heat corresponding to the data of the stored data type when the data of the service data type and the data of the stored data type are acquired.
[0012] As an alternative solution, the above-mentioned fourth acquisition sub-module includes: a third determination subunit for determining the stored data type corresponding to the second data as the cache pool storage type when the second data is data stored in the cache pool; a fourth determination subunit for determining the stored data type corresponding to the second data as the normal pool storage type when the second data is data stored in the normal pool, where the heat corresponding to the cache pool storage type is greater than the heat corresponding to the normal pool storage type; a fifth determination subunit for determining the stored data type corresponding to the second data as the data pool storage type when the second data is data stored in the data pool, where the heat corresponding to the normal pool storage type is greater than the heat corresponding to the data pool storage type.
[0013] As an alternative solution, the above-mentioned third acquisition sub-module includes: a sixth determination subunit for determining the specified data type corresponding to the first data as the metadata type when the first data is metadata; a seventh determination subunit for determining the specified data type corresponding to the first data as the WAL data type when the first data is WAL data, where the heat corresponding to the metadata type is greater than the heat corresponding to the WAL data type.
[0014] As an alternative solution, the above-mentioned fifth acquisition sub-module includes: an eighth determination subunit for determining the data that conforms to the service logic among the above-mentioned multiple data as the third data; a ninth determination subunit for determining the service type corresponding to the service logic as the service data type.
[0015] As an alternative, the above-mentioned first acquisition sub-module includes: a first acquisition sub-unit, configured to acquire at least two types of popularity-related parameters associated with the third data when the data type is the business data type; an allocation sub-unit, configured to, when the data type is the business data type, allocate calculation weights to each of the at least two types based on the business logic; a second acquisition sub-unit, configured to, when the data type is the business data type, use the calculation weights and combine the popularity-related parameters of the at least two types to acquire the popularity of the third data.
[0016] As an alternative, the above-mentioned first acquisition sub-unit includes at least two of the following: a first sub-acquisition module, configured to acquire the data access frequency associated with the third data; a second sub-acquisition module, configured to acquire the data update frequency associated with the third data; a third sub-acquisition module, configured to acquire the data access times associated with the third data; a fourth sub-acquisition module, configured to acquire the data update times associated with the third data; a fifth sub-acquisition module, configured to acquire the data access timestamp associated with the third data; a sixth sub-acquisition module, configured to acquire the data update timestamp associated with the third data.
[0017] As an alternative, the above-mentioned writing unit includes: an allocation module, configured to allocate virtual flow identifiers to each of the data according to the popularity attribute, where the virtual flow identifier is used to represent the popularity of each of the data; a first writing module, configured to write each of the data into each of the data blocks according to the virtual flow identifier, where different data blocks correspond to different virtual flow identifiers, and different virtual flow identifiers correspond to different popularity attribute conditions.
[0018] As an alternative, the above-mentioned allocation module includes: a first allocation sub-module, configured to allocate a first virtual flow identifier to the fourth data based on the popularity information when the popularity attribute corresponding to the fourth data among the multiple data carries the popularity information of the fourth data.
[0019] As an alternative, the above-mentioned allocation module includes: a sixth acquisition sub-module, configured to acquire the popularity information of the storage pool to which the fifth data belongs when the popularity attribute corresponding to the fifth data among the multiple data lacks the popularity information of the fifth data; a second allocation sub-module, configured to allocate a second virtual flow identifier to the fifth data based on the popularity information of the storage pool to which it belongs.
[0020] As an alternative solution, the above-mentioned second acquisition unit includes: a setting module, configured to set the heat attribute corresponding to the preset heat condition as the heat attribute corresponding to the sixth data when the sixth data among the above-mentioned multiple data satisfies the preset heat condition.
[0021] As an alternative solution, the above-mentioned writing unit includes: a second writing module, configured to write the data that meets the high heat condition among the above-mentioned multiple data into the first data block in the above-mentioned solid-state drive according to the above-mentioned heat attribute, where the above-mentioned heat attribute condition includes the above-mentioned high heat condition; a third writing module, configured to write the data that meets the low heat condition among the above-mentioned multiple data into the second data block in the above-mentioned solid-state drive according to the above-mentioned heat attribute, where the storage space of the above-mentioned first data block is larger than that of the above-mentioned second data block, and the above-mentioned heat attribute condition includes the above-mentioned low heat condition.
[0022] As an alternative solution, the above-mentioned device further includes: an updating unit, configured to, after writing each of the above-mentioned data into each data block in the above-mentioned solid-state drive according to the above-mentioned heat attribute, when the heat attribute corresponding to the seventh data written into the above-mentioned data block changes, update the position of the seventh data in the above-mentioned solid-state drive according to the changed heat attribute.
[0023] According to another embodiment of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0024] According to another embodiment of the present application, there is also provided an electronic device, including a memory and a processor, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0025] Through the present application, by adopting the method of clearly distinguishing and separately managing data with different heats, the probability of generating garbage data is reduced. Because hot data is centrally managed, its update and replacement are more efficient, reducing the accumulation of invalid data. At the same time, since cold data is stored separately, it can still maintain its validity after being written for a long time, which can solve the technical problem of low data management efficiency, and thus achieve the technical effect of improving data management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic diagram of an application environment of an alternative data management method according to an embodiment of the present application;
[0027] Figure 2It is a schematic diagram of the process of an optional data management method according to an embodiment of the present application;
[0028] Figure 3 It is a schematic diagram of an optional data management method according to an embodiment of the present application;
[0029] Figure 4 It is a schematic diagram of another optional data management method according to an embodiment of the present application;
[0030] Figure 5 It is a schematic diagram of an optional data management device according to an embodiment of the present application; Detailed implementation manners
[0031] In the following, embodiments of the present application will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence.
[0033] The method embodiments provided in the embodiments of the present application can be executed on a server device or a similar computing device. Taking running on a server device as an example, Figure 1 It is a hardware structure block diagram of a server device of a data management method according to an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic, and it does not limit the structure of the above-mentioned server device. For example, the server device may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0034] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the data management method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0035] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0036] In this embodiment, a data management method is provided. Figure 2 It is a flowchart of the data management method according to the embodiments of the present application, as Figure 2 shown, and the process includes the following steps:
[0037] S202, obtain a plurality of data to be written into the solid-state drive;
[0038] This step refers to obtaining a set of data that is about to be written into the solid-state drive (SSD) from a data source (such as memory, CPU, or other storage devices). This data may come from user application operations, system tasks, or other data generation processes. Accurately obtaining the data to be written is the basis for the subsequent data management process, ensuring the correctness and efficiency of data writing.
[0039] Before the data is actually written to the SSD, the system needs to first determine which data needs to be written. This may involve the management of the data buffer to ensure the integrity and order of the data. For example, when a user saves a document or uploads a video file, the file data is the data to be written to the SSD.
[0040] S204, obtain the heat attribute corresponding to each data in the plurality of data;
[0041] "Heat attribute" refers to the frequency at which data is accessed or modified. In this step, the system analyzes and determines the heat value of each data item to be written, which is usually based on the historical access patterns of the data, the expected usage frequency, or other relevant metrics. Accurately obtaining the heat attribute of data is the key to achieving data classification and optimizing storage, which helps improve data access speed and storage efficiency.
[0042] Different data has different access and update patterns. For example, some data may be frequently read and modified (high heat), while other data may remain unchanged for a long time (low heat). By identifying the heat attribute of the data, the system can manage this data more effectively.
[0043] For example, on a social media platform, the user's frequently updated profile information can be regarded as high-heat data, while the old posts that the user occasionally accesses may be low-heat data.
[0044] S206, according to the heat attribute, write each data into each data block in the solid-state drive, where different data blocks correspond to different heat attribute conditions.
[0045] This step involves classifying the data according to its heat attribute and writing them into the corresponding data blocks in the SSD. Each data block is designed to store data with similar heat attributes. The data writing strategy based on the heat attribute optimizes the data storage layout, improves the locality of data access, thereby reducing the read / write latency and wear of the SSD, and extending the service life of the device. At the same time, it also enhances the overall data processing capacity and response speed of the system.
[0046] By clustering data with similar heat together, the system can manage the storage resources of the SSD more effectively. High-heat data blocks may adopt faster storage media or more optimized data layouts to support higher access speeds. On the contrary, low-heat data blocks may focus more on data persistence and cost-effectiveness.
[0047] For example, in a database system, the table data that is frequently queried may be written into high-heat data blocks, while the archived old data may be stored in low-heat data blocks.
[0048] In an alternative embodiment, multiple data to be written to the solid-state drive: refers to the set of data that has not been written to the solid-state drive but is ready for storage. It contains various types of information, such as documents, pictures, videos, etc. These information are called "data to be written" before being permanently saved to the SSD.
[0049] In an alternative embodiment, the heat property: is an indicator used to measure the frequency of data access or modification. It reflects the activity level of the data and is usually determined based on the historical access records of the data and the expected usage frequency. High-heat data means that it is frequently accessed or modified, while low-heat data is the opposite.
[0050] In an alternative embodiment, a data block: in a solid-state drive, a data block is a physical or logical unit for storing data. Each data block has a certain storage capacity, and according to this scheme, different data blocks are designed to store data with different heat properties.
[0051] It should be noted that the solution of this embodiment is a solid-state drive (SSD) write policy based on the heat property of data. Its core idea is to intelligently write data into different data blocks of the SSD according to the heat property of the data, that is, the frequency of data access or update. In this way, the solution aims to improve data management efficiency, optimize the use of storage resources, and enhance data access speed.
[0052] For further illustration, assume that a large enterprise needs to store a large amount of business data, including frequently updated sales records and less accessed archived files. According to this scheme, the system will first analyze the heat properties of these data. Since the sales records are frequently queried and updated, they are marked as high-heat data; while the archived files are marked as low-heat data due to their low access frequency. Subsequently, the system will write the sales records into the data blocks with higher performance in the SSD to ensure fast access and update; while the archived files are written into the data blocks with slightly lower performance but more optimized costs.
[0053] Through the embodiments provided in this application, obtain multiple data to be written into the solid-state drive; obtain the heat property corresponding to each data among the multiple data; according to the heat property, write each data into each data block in the solid-state drive, where different data blocks correspond to different heat property conditions.
[0054] By clearly distinguishing and separately managing different heat data, the probability of generating garbage data is reduced, because the hot data is centrally managed, and its update and replacement are more efficient, reducing the accumulation of invalid data. At the same time, since the cold data is stored separately, it can still maintain its validity after long-term writing, thereby achieving the purpose of improving data utilization and persistence, and thus realizing the technical effect of improving data management efficiency.
[0055] As an alternative scheme, obtaining the heat property corresponding to each data among the multiple data includes:
[0056] S1-1, obtain the data type corresponding to each data;
[0057] This step refers to the system's need to identify and determine the data type of each data item to be processed. The data type is a basic attribute of data, which defines the structure, representation, and operations that can be performed on the data.
[0058] S1-2. Obtain the heat attribute according to the data type.
[0059] After determining the data type, the system will assign or calculate a heat attribute for each data type based on preset rules or algorithms, as well as possible historical data. This heat attribute reflects the frequency with which data of that type is accessed or updated.
[0060] In an alternative embodiment, the "data type" refers to the classification or category of data, such as text, numbers, images, videos, etc., which stipulates the format and storage method of the data.
[0061] In an alternative embodiment, the "heat attribute" specifically refers to an indicator of data activity level determined based on the data type, which is used to measure and distinguish the importance or frequency of different data types during storage and access.
[0062] It should be noted that this embodiment describes a specific method for obtaining the heat attribute corresponding to each data among multiple data. This process is divided into two main steps: First, obtain the data type of each data; Second, determine or obtain the corresponding heat attribute according to these data types.
[0063] For further illustration, for example, in a social media application, the data generated by users may include text comments, picture sharing, and video uploads, etc. The system will first identify the data types of these data, such as text, images, and videos. Then, based on historical data and user behavior analysis, the system may find that video data is usually more popular than text and picture data, and thus is accessed and shared more frequently. Based on this discovery, the system will mark the video data type as having a higher heat attribute.
[0064] Through the embodiments provided in this application, different types of data can be understood and managed more precisely, and optimized storage and access can be performed according to their heat attributes. This can not only improve the efficiency of data processing, but also enhance the user experience, because more popular data types (i.e., data with higher heat) will be processed and stored preferentially, thus responding to user requests faster. At the same time, this method of assigning heat attributes based on data types also helps to reasonably allocate and utilize system resources, reducing storage and computing costs.
[0065] As an alternative solution, obtaining the heat attribute according to the data type includes:
[0066] S2-1. Obtain the popularity corresponding to each piece of data according to the data type, where different data types correspond to different popularities.
[0067] The system will allocate or calculate a popularity value according to the type of data (such as text, picture, video, etc.). This popularity value reflects the popularity or the frequency of access of this type of data in the system.
[0068] S2-2. Obtain the popularity attribute according to the popularity, where different popularity attribute conditions are different popularity ranges.
[0069] Once the system determines the popularity value of each piece of data, it will match these popularity values with the preset popularity ranges, so as to classify the data into different popularity attribute conditions. These popularity ranges are divided according to the activity level of the data.
[0070] In an alternative embodiment, "popularity" here refers to a quantitative indicator used to represent the activity level or the degree of attention of a specific data type in the system. Different data types may have different popularity values.
[0071] In an alternative embodiment, "popularity attribute" is an attribute label determined based on the popularity value of the data, which reflects the belonging of the data in the popularity range. Different popularity ranges correspond to different popularity attribute conditions.
[0072] It should be noted that this embodiment further details the process of obtaining the popularity attribute of data according to the data type. It includes two core steps: First, determine the popularity corresponding to each piece of data according to the data type; Second, based on these popularity values, classify the data into different popularity attribute conditions or popularity ranges.
[0073] For further illustration, for example, in an online news platform, articles, pictures and videos are common data types. The system may find according to historical access data that video content is usually more popular among users than articles and pictures, so videos will be assigned a higher popularity value. Then, the system sets several popularity ranges, such as "high popularity", "medium popularity" and "low popularity". When the popularity value of a certain video falls within the "high popularity" range, it is given the corresponding popularity attribute.
[0074] Through the embodiments provided by this application, it is possible to more accurately understand and manage the popularity attributes of different types of data. This helps to optimize the data storage and access strategies, improve the system response speed and user experience. For example, high-popularity data can be preferentially stored in a storage medium with higher performance to ensure fast access; while low-popularity data can be stored in a storage medium with lower cost and moderate performance, so as to achieve reasonable allocation and utilization of storage resources.
[0075] As an alternative solution, obtain the data types corresponding to each data, including at least one of the following:
[0076] Obtain the specified data type corresponding to the first data among multiple data;
[0077] Obtain the stored data type corresponding to the second data among multiple data;
[0078] Obtain the business data type corresponding to the third data among multiple data.
[0079] In an alternative embodiment, the specified data type: refers to the data type that has been clearly specified when the data is processed or created. It is the type of data that is clearly marked or defined, usually directly specified by the creator or processor of the data.
[0080] In an alternative embodiment, the stored data type: refers to the data type recorded by the data in the storage system or database, usually related to the physical storage format and method of the data. It is the type of data recorded in the storage medium or system, which reflects the format and characteristics of the data when it is stored.
[0081] In an alternative embodiment, the business data type: refers to the type of data in a specific business domain or application scenario, which is closely related to the logic and rules of the business. It is the type of data classified in a specific business environment and can be associated with the actual needs and processes of the business.
[0082] It should be noted that this embodiment describes different ways and categories when obtaining data types. The data type here refers to the specific classification or attribute of the data, which can help understand the essence and use of the data. This content lists three ways to obtain data types, namely: specified data type, stored data type, and business data type.
[0083] For further illustration, for example, in an e-commerce system, the first data may be the search keywords entered by the user, and its "specified data type" is text; the second data may be the product picture, and its "stored data type" is image file; the third data may be the purchase record of the user, and its "business data type" is transaction data.
[0084] By clearly and meticulously obtaining different types of data, the system can more precisely understand the essence and characteristics of the data, and then select the most appropriate processing methods and analysis strategies. This can not only improve the efficiency and accuracy of data processing, but also help to explore the potential value of the data, providing more powerful data support for the enterprise's decision-making. At the same time, clarifying the data type also helps to optimize the data storage and management strategies, ensuring the security and accessibility of the data.
[0085] As an alternative, according to the data type, obtain the popularity corresponding to each data, including:
[0086] When the data of the specified data type and the data of the stored data type are obtained, determine that the popularity corresponding to the data of the specified data type is greater than the popularity corresponding to the data of the stored data type.
[0087] In an alternative embodiment, the popularity corresponding to the data of the specified data type: refers to the popularity value associated with the data that is clearly specified in type (such as user input, data generated by a specific application, etc.). The popularity of these data is usually relatively high because they directly reflect the user's intention or the key operations of the system.
[0088] In an alternative embodiment, the popularity corresponding to the data of the stored data type: refers to the popularity value associated with the type recorded for the data in the storage system or database. The popularity of these data may be relatively low because they are more related to the physical storage state of the data and do not necessarily directly reflect user behavior or business requirements.
[0089] It should be noted that this embodiment describes how to determine the popularity of data according to the data type. In particular, it compares the differences in popularity determination between the specified data type and the stored data type, and points out that the data of the specified data type usually has a higher popularity than the data of the stored data type.
[0090] For further illustration, for example, in an online document editing system, a document (specified data type) explicitly uploaded and marked as "important" by the user will have a higher popularity than the version of the document automatically saved by the system (stored data type). This is because the "important" document marked by the user is more likely to be accessed and edited frequently, while the automatically saved version is mainly used for backup and recovery, and the access frequency is relatively low.
[0091] By distinguishing the differences in popularity determination between the specified data type and the stored data type, the system can more accurately identify and manage the data that is more important to the user or business. This helps to optimize the data storage strategy to ensure that high-popularity data can be accessed and processed more quickly and reliably. At the same time, this popularity determination method also helps to improve the overall performance and user experience of the system because it enables the system to more effectively respond to the user's needs and prioritize the processing of critical data.
[0092] As an alternative, according to the data type, obtain the popularity corresponding to each data, including:
[0093] When the data of the business data type and the data of the stored data type are obtained, determine that the popularity corresponding to the data of the business data type is greater than the popularity corresponding to the data of the stored data type.
[0094] In an alternative embodiment, the heat value corresponding to data of the business data type: This refers to the heat value assigned to data that is closely related to a specific business logic or process. Since this type of data is directly associated with the core activities and decisions of the business, it usually has a higher heat value.
[0095] In an alternative embodiment, the heat value corresponding to data of the storage data type: This refers to the heat value corresponding to the type to which the data is classified at the storage level. This type of data pays more attention to the physical storage and format of the data, and generally has a lower heat value compared to direct business activities.
[0096] It should be noted that this embodiment describes how to determine the heat value of data according to the data type, especially in the case of business data types and storage data types. It points out that when comparing data of the business data type and data of the storage data type, data of the business data type is usually given a higher heat value.
[0097] For further illustration, optionally, for example, in a customer relationship management system of a bank, the transaction records of customers (business data type) will have a higher heat value than the storage format of customer account information (storage data type). This is because transaction records directly reflect the financial activities and business needs of customers and are the basis for important business decisions such as customer analysis and risk assessment by the bank.
[0098] By clearly distinguishing the differences in heat value determination between business data types and storage data types, the system can more effectively identify and utilize data that is crucial to the business. This can not only improve the accuracy and timeliness of business decisions, but also optimize data storage and management strategies to ensure that key data can be processed and protected with priority. At the same time, this heat value allocation mechanism also helps to improve the overall performance and response speed of the system, thus providing users with a smoother and more efficient service experience.
[0099] As an alternative solution, obtaining the storage data type corresponding to the second data among multiple data includes:
[0100] S3-1, when the second data is data stored in the cache pool, determining the storage data type corresponding to the second data as the cache pool storage type;
[0101] S3-2, when the second data is data stored in the normal pool, determining the storage data type corresponding to the second data as the normal pool storage type, where the heat value corresponding to the cache pool storage type is greater than the heat value corresponding to the normal pool storage type;
[0102] S3-3. When the second data is the data stored in the data pool, determine the storage data type corresponding to the second data as the data pool storage type. The popularity corresponding to the normal pool storage type is greater than the popularity corresponding to the data pool storage type.
[0103] In an optional embodiment, the cache pool storage type refers to the data type when the data is stored in the cache pool. The cache pool is usually used to store the data that is frequently accessed to provide faster reading speed.
[0104] In an optional embodiment, the normal pool storage type refers to the data type when the data is stored in the normal storage pool. The performance and access speed of the normal pool may be lower than those of the cache pool, but it is suitable for general data storage requirements.
[0105] In an optional embodiment, the data pool storage type refers to the data type when the data is stored in a more general or archival data pool. The data pool may be used for long-term storage or backup, and its access frequency is relatively low.
[0106] It should be noted that this embodiment describes how to obtain the storage data type corresponding to the second data among multiple data, and determine its corresponding popularity according to the different pools (cache pool, normal pool, data pool) where the data is stored. The storage data type here refers to the specific method or technical classification adopted when the data is stored, and different storage pools represent different data storage strategies and performance characteristics.
[0107] For further illustration, for example, in a large online video platform, the cache pool may be used to store the recently popular and frequently watched video data by users, and these data are marked as "cache pool storage type"; the normal pool may be used to store most of the regular video data, and these data are marked as "normal pool storage type"; while the data pool may be used to store the old and infrequently accessed video data or backup files, and these data are marked as "data pool storage type".
[0108] By clearly distinguishing the data types in different storage pools and assigning different popularity values according to their characteristics, the system can more effectively manage and optimize the storage and access of data. High-popularity data (such as the data in the cache pool) can enjoy faster access speed and higher availability, thus improving the user experience and system performance. At the same time, this classification method also helps to reduce the storage cost because different types of storage media and strategies can be flexibly configured and adjusted according to actual needs.
[0109] As an optional solution, obtaining the specified data type corresponding to the first data among multiple data includes:
[0110] S4-1. When the first data is metadata, determine the specified data type corresponding to the first data as the metadata type;
[0111] S4-2. When the first data is WAL data, determine the specified data type corresponding to the first data as the WAL data type, where the heat corresponding to the metadata type is greater than the heat corresponding to the WAL data type.
[0112] In an optional embodiment, the metadata type refers to the data that describes other data, that is, the data about the structure, attributes, relationships, etc. of the data. Metadata is usually used to help understand, manage, and use other data.
[0113] In an optional embodiment, the WAL data type: WAL stands for Write-Ahead Logging, which is a technology used in database management systems to ensure data integrity and persistence. WAL data records the logs of database changes for data recovery after system crashes or other failures.
[0114] It should be noted that this embodiment describes how to determine the specified data type corresponding to the first data according to the characteristics of the first data, and compares the differences in heat among different types of data. The specified data type refers to the data type clearly specified according to the specific content and use of the data.
[0115] For further illustration, for example, in a library management system, information such as the title, author, and publication date of a book can be regarded as metadata, which describes the basic attributes and characteristics of the book. And the log records generated when the library system performs operations such as borrowing and returning books can be regarded as WAL data for restoring the data state when the system has problems.
[0116] By clearly distinguishing the metadata type and the WAL data type, and assigning different heat values according to their characteristics and uses, the system can more effectively manage and optimize the access and processing of data. High-heat metadata can be accessed faster and have higher availability, thus enhancing the user experience and system performance. At the same time, this classification method also helps to ensure the integrity and reliability of WAL data, providing strong support for data recovery in case of system failures. By reasonably managing the heat of different types of data, the system can respond to user needs more efficiently and ensure the security and availability of data.
[0117] As an optional solution, obtaining the business data type corresponding to the third data among multiple data includes:
[0118] S5-1. Among the multiple data, determine the data that conforms to the business logic as the third data;
[0119] S5-2. Determine the business data type based on the business type corresponding to the business logic.
[0120] In an alternative embodiment, the third data refers to data related to specific business logic, which is usually directly associated with the core business activities or processes of an enterprise.
[0121] In an alternative embodiment, business logic refers to specific rules, processes, and strategies in enterprise operations that define the execution manner and sequence of business activities.
[0122] In an alternative embodiment, the business data type is the data type determined according to the business logic, which reflects the specific role and meaning of the data in business activities.
[0123] It should be noted that this embodiment describes how to obtain the business data type corresponding to the third data from multiple data. Here, the "business data type" refers to the data type related to specific business logic or business processes. This type of data is crucial for business analysis, decision-making, and the implementation of system functions.
[0124] For further illustration, for example, on an e-commerce platform, the third data may include users' purchase records, browsing histories, search keywords, etc. These data are closely related to business logics such as sales, recommendations, and market popularity on the e-commerce platform. For example, the business data type corresponding to the user's purchase record data can be "transaction data", and the business data type corresponding to the user's search keywords can be "user demand data".
[0125] By accurately obtaining and identifying the business data type corresponding to the third data, an enterprise can gain a deeper understanding of the key data and processes in its business operations. This helps the enterprise make more informed decisions, optimize business processes, improve customer satisfaction, and drive business growth. At the same time, clarifying the business data type also helps the enterprise manage and utilize its data resources more effectively, improving the quality and availability of the data.
[0126] As an alternative solution, when the data type is the business data type, according to the data type, obtain the popularity of each data, including:
[0127] S6-1. Obtain at least two types of popularity-related parameters associated with the third data;
[0128] S6-2. Based on the business logic, assign calculation weights to each type in at least two types;
[0129] S6-3. Use the calculation weights to combine the popularity-related parameters of at least two types to obtain the popularity of the third data.
[0130] In an alternative embodiment, the heat-related parameter refers to a parameter directly related to the calculation of data heat, which may include the access frequency of data, the update frequency, the number of user interactions, etc. These parameters can reflect the activity and importance of the data.
[0131] In an alternative embodiment, the calculation weight: When calculating the data heat, different types of heat-related parameters may contribute differently to the final heat. Therefore, different weights need to be assigned to these parameters to reflect their relative importance in influencing the heat.
[0132] In an alternative embodiment, the heat of the third data refers to an indicator calculated based on the business logic and heat-related parameters, reflecting the importance and activity of the third data in the business.
[0133] It should be noted that this embodiment describes how to obtain the heat corresponding to each data (specifically the third data here) according to the data type when the data type is the business data type. It involves the acquisition of heat-related parameters, the assignment of weights, and the calculation of the final heat.
[0134] For further illustration by example, optionally, in a news recommendation system, the third data may represent different news articles. The system can obtain the click-through rate, the number of comments, the number of shares, etc. of each article as heat-related parameters. According to the business logic of news recommendation, the click-through rate may be considered more important than the number of comments, so it will be assigned a higher calculation weight. Finally, by combining these parameters and weights, the heat of each article can be calculated to determine the recommendation priority.
[0135] By obtaining the heat-related parameters associated with the third data and combining the business logic to assign calculation weights, it is finally possible to obtain the heat value reflecting the importance and activity of the third data in the business. This enables the system to more accurately identify and utilize the data that is crucial to the business, optimizing data processing, storage, and recommendation strategies. At the same time, this heat calculation method also helps to improve the overall performance and user experience of the system, because it enables the system to more effectively respond to user needs and prioritize the processing of key data.
[0136] As an alternative solution, at least two types of heat-related parameters associated with the third data are obtained, including at least two of the following:
[0137] Obtain the data access frequency associated with the third data;
[0138] Obtain the data update frequency associated with the third data;
[0139] Obtain the number of data accesses associated with the third data;
[0140] Obtain the number of data updates associated with the third data;
[0141] Obtain the data access timestamp associated with the third data;
[0142] Obtain the data update timestamp associated with the third data.
[0143] In an optional embodiment, data access frequency: refers to the number of times data is accessed per unit time, reflecting the popularity and usage frequency of the data.
[0144] In an optional embodiment, data update frequency: refers to the number of times data is updated per unit time, reflecting the timeliness and activity level of the data.
[0145] In an optional embodiment, data access count: refers to the total number of times data is accessed, which is an intuitive indicator for evaluating the popularity of the data.
[0146] In an optional embodiment, data update count: refers to the total number of times data is modified or updated, which can reflect the change situation and maintenance activity level of the data.
[0147] In an optional embodiment, data access timestamp: records the specific time when data is accessed each time, which helps to analyze the access pattern of the data and user behavior.
[0148] In an optional embodiment, data update timestamp: records the specific time when data is updated each time, which is very useful for tracking the change history of the data and analyzing the data update pattern.
[0149] It should be noted that this embodiment describes the specific method for obtaining the heat-related parameters associated with the third data. The heat-related parameters are important indicators for evaluating and calculating the heat of the data, and they can directly or indirectly reflect the activity level, popularity and usage frequency of the data. The parameters mentioned here include data access frequency, data update frequency, data access count, data update count, data access timestamp and data update timestamp.
[0150] For further illustration by way of example, optionally, taking an online forum as an example, the third data may represent a specific post. The data access frequency of this post can be the number of times it is viewed per day; the data update frequency can be the frequency of modification of the post content or comments; the data access count is the total number of times the post has been viewed since its publication; the data update count is the total number of times the post or comments have been edited; the data access timestamp records the specific time when the user views the post each time; and the data update timestamp records the time when the post or comments are modified each time.
[0151] By obtaining these detailed heat-related parameters, the system can more accurately evaluate the heat and importance of the third data. This helps optimize data storage, retrieval, and recommendation strategies, improving the system's performance and user experience. For example, for data with high access frequency, the system can preferentially cache it in memory to speed up access; for data with frequent updates, the system can set a higher synchronization frequency to ensure data timeliness. In short, these heat-related parameters provide rich data characteristics for the system, facilitating more refined data management and services.
[0152] As an alternative solution, according to the heat attribute, write each data into each data block in the solid-state drive, including:
[0153] S7-1, according to the heat attribute, assign a virtual flow identifier to each data, where the virtual flow identifier is used to represent the heat of each data;
[0154] S7-2, write each data into each data block according to the virtual flow identifier, where different data blocks correspond to different virtual flow identifiers, and different virtual flow identifiers correspond to different heat attribute conditions.
[0155] In an alternative embodiment, the virtual flow identifier: is an identifier used to represent the heat attribute of each data. This identifier can help the system identify the heat level of the data and determine which data block the data should be written into according to this level.
[0156] In an alternative embodiment, the data block: is a physical or logical unit in the solid-state drive for storing data. Different data blocks can correspond to different virtual flow identifiers, so as to store data with different heat attributes.
[0157] It should be noted that this embodiment describes how to write data into different data blocks of the solid-state drive according to the heat attribute of the data. This process involves assigning virtual flow identifiers to the data and using these identifiers to guide the data writing operation.
[0158] For further illustration, assume that there is a solid-state drive, which is divided into multiple data blocks, and each data block can store data with specific heat attributes. Now there is a set of data to be written into this hard drive, and this set of data includes high-heat data, medium-heat data, and low-heat data. The system will first assign different virtual flow identifiers to these data. For example, high-heat data is assigned the identifier "H", medium-heat data is assigned the identifier "M", and low-heat data is assigned the identifier "L". Then, the system will write the data into the corresponding data blocks according to these identifiers. For example, the data with the identifier "H" is written into the data block dedicated to storing high-heat data.
[0159] By assigning virtual stream identifiers to data and writing the data into different data blocks according to these identifiers, the system can achieve more efficient and intelligent data storage management. The use of virtual stream identifiers provides the system with greater flexibility, allowing it to easily adjust the data storage strategy to adapt to different application scenarios and changing requirements.
[0160] As an alternative solution, virtual stream identifiers are assigned to each data according to the popularity attribute, including:
[0161] In the case where the popularity attribute corresponding to the fourth data among multiple data carries the popularity information of the fourth data, a first virtual stream identifier is assigned to the fourth data based on the popularity information.
[0162] In an alternative embodiment, the fourth data: refers to a specific data item among multiple data, which has a popularity attribute associated with it.
[0163] In an alternative embodiment, the popularity information: is a specific value or level regarding the popularity of the data, which reflects the popularity, usage frequency, or importance of the data.
[0164] In an alternative embodiment, the first virtual stream identifier: is a specific identifier assigned to the fourth data according to the popularity information of the fourth data, used to represent the popularity level or classification of the data.
[0165] It should be noted that this embodiment describes a specific scenario of how to assign virtual stream identifiers to data according to the popularity attribute of the data. Here, when the fourth data among multiple data has a popularity attribute associated with it, and this popularity attribute carries the popularity information of the fourth data, the system will assign a specific virtual stream identifier, namely the first virtual stream identifier, to the fourth data based on this popularity information.
[0166] For further illustration, it is optionally assumed that in a video streaming media platform, the fourth data represents a specific video file. This video file has a popularity attribute, which carries the popularity information about the popularity of the video, such as the number of views, the number of likes, etc. Based on this information, the system assigns a first virtual stream identifier, such as "popular video", to this video file in order to distinguish it from other less popular videos.
[0167] By allocating a first virtual flow identifier to the fourth data based on heat information, the system can more accurately identify and classify data with different heat levels. This classification helps the system optimize data storage, retrieval, and recommendation strategies. For example, data with a higher heat identifier can be preferentially stored on high-performance storage devices to respond more quickly to user requests; at the same time, this data can also be recommended to users more frequently to increase user satisfaction and platform activity. Generally speaking, this method of allocating virtual flow identifiers improves the efficiency and intelligence level of the system in processing data.
[0168] As an alternative solution, according to the heat attribute, virtual flow identifiers are allocated to each data, including:
[0169] S8-1, in the case where the heat attribute corresponding to the fifth data among multiple data lacks the heat information of the fifth data, obtain the heat information of the storage pool to which the fifth data belongs;
[0170] S8-2, allocate a second virtual flow identifier to the fifth data based on the heat information of the storage pool to which it belongs.
[0171] In an alternative embodiment, the fifth data: refers to a specific data item among multiple data, whose direct heat information is missing or unavailable.
[0172] In an alternative embodiment, the storage pool to which it belongs: is a logical or physical container in which the fifth data is stored. It may contain multiple similar data items, and this storage pool itself has a measurable heat attribute.
[0173] In an alternative embodiment, the second virtual flow identifier: is an identifier allocated to the fifth data based on the heat information of the storage pool to which it belongs when the direct heat information of the fifth data is unavailable, and is used to indirectly represent the heat of the data.
[0174] It should be noted that this embodiment describes how to allocate a virtual flow identifier to a certain data (specifically the fifth data here) when the heat attribute of the data lacks direct heat information. Specifically, the system will instead obtain the heat information of the storage pool to which the data belongs, and allocate a second virtual flow identifier to the fifth data based on the heat information of this storage pool.
[0175] For further illustration, assume that in a cloud storage system, the fifth data is a newly uploaded file that has not had enough user interaction to generate direct heat information. However, this file is stored in a storage pool called "Popular File Pool", which has a clear heat information based on the average download volume and access frequency of the files in it. Based on the heat information of this storage pool, the system allocates a second virtual flow identifier, such as "Potentially Popular File", to the newly uploaded file so as to give it corresponding priority in subsequent processing.
[0176] By obtaining the heat information of the storage pool to which the fifth data belongs and allocating a second virtual flow identifier for the fifth data based on this information, the system can still perform effective heat classification and management of the data in the absence of direct heat information. This method improves the flexibility and accuracy of the system's perception of data heat, helps optimize the data storage strategy, improve access performance, and enhance the user experience. For example, data marked as "potentially popular files" can be cached by the system in advance to high-speed storage media to cope with possible future access peaks.
[0177] As an alternative solution, obtain the heat attributes corresponding to each of the multiple data, including:
[0178] When the sixth data among the multiple data meets the preset heat condition, set the heat attribute corresponding to the preset heat condition as the heat attribute corresponding to the sixth data.
[0179] In an alternative embodiment, the sixth data: refers to a specific data item in a multiple data set.
[0180] In an alternative embodiment, the preset heat condition: is a set of criteria or thresholds set in advance for determining whether the data reaches a specific heat level.
[0181] In an alternative embodiment, the heat attribute: is an identifier or label used to represent the popularity, usage frequency, or other heat-related characteristics of the data.
[0182] It should be noted that this embodiment describes the process of obtaining and setting the heat attribute for a specific data (here referring to the sixth data) when processing multiple data. Specifically, when the sixth data meets a certain preset heat condition, the system assigns the heat attribute corresponding to this preset heat condition to the sixth data as the identifier of its heat attribute.
[0183] For further illustration by example, optionally, for example, in an online news platform, the sixth data can be a newly released news report. The platform sets a preset heat condition, that is, the news report is read more than 100,000 times within 24 hours after its release. If this new report indeed reaches this number of readings within 24 hours, the system will automatically assign the heat attribute of "hot news" to this report as its heat identifier.
[0184] By setting corresponding heat attributes for the sixth data that meets the preset heat conditions, the system can quickly identify and classify the heat status of the data. This mechanism helps improve data processing efficiency, optimize resource allocation, and provide more personalized and accurate content recommendations for users. For example, in a news platform, reports marked as "hot news" can be preferentially displayed on the home page or in the recommendation list, thereby increasing their exposure rate and user interaction opportunities.
[0185] As an alternative solution, according to the heat attributes, write each data into each data block in the solid-state drive, including:
[0186] S9-1, according to the heat attributes, write the data that meets the high heat condition among multiple data into the first data block in the solid-state drive, where the heat attribute condition includes the high heat condition;
[0187] S9-2, according to the heat attributes, write the data that meets the low heat condition among multiple data into the second data block in the solid-state drive, where the storage space of the first data block is larger than that of the second data block, and the heat attribute condition includes the low heat condition.
[0188] In an alternative embodiment, the high heat condition: refers to the condition in the data heat attribute that indicates the data is very popular, has a high usage frequency, or is important.
[0189] In an alternative embodiment, the low heat condition: is opposite to the high heat condition, indicating the condition that the data is relatively less popular, has a low usage frequency, or has a lower importance.
[0190] In an alternative embodiment, the first data block: the area in the solid-state drive used to store data that meets the high heat condition.
[0191] In an alternative embodiment, the second data block: the area in the solid-state drive used to store data that meets the low heat condition.
[0192] It should be noted that this embodiment elaborates on how to write data into different data blocks of the solid-state drive according to the heat attributes of the data. Specifically, it describes how to write data that meets the high heat condition into the first data block and how to write data that meets the low heat condition into the second data block. In addition, it is also mentioned that the storage space of the first data block is larger than that of the second data block.
[0193] For further illustration by example, optionally, for example, in a video streaming service, high-heat videos (such as popular movies and TV series) will be written into the first data block because these contents are frequently accessed by users and require faster read and write speeds and larger storage space to handle high-concurrency access. While those less-watched and relatively unpopular videos will be written into the second data block.
[0194] By writing data into the first data block and the second data block respectively according to the heat attribute, the system can achieve hierarchical storage and management of data.
[0195] As an alternative solution, after writing each data into each data block in the solid-state drive according to the heat attribute, the method further includes:
[0196] In the case where the heat attribute of the seventh data corresponding to the written data block changes, update the position of the seventh data in the solid-state drive according to the changed heat attribute.
[0197] In an alternative embodiment, the seventh data refers to a specific data item that has been previously written into the data block of the solid-state drive.
[0198] In an alternative embodiment, the change in the heat attribute means that the original heat attribute of the seventh data (such as high heat, low heat, etc.) has changed due to certain reasons (such as increased user interaction, time lapse, etc.).
[0199] In an alternative embodiment, updating the position means moving the data from the current storage position to the data block in the solid-state drive that matches its new heat attribute according to the changed heat attribute.
[0200] It should be noted that this embodiment describes how the system responds to the change when the heat attribute of a certain data (specifically the seventh data here) changes after writing the data into the data block of the solid-state drive according to the heat attribute. Specifically, the system will update the storage position of this data in the solid-state drive according to the changed heat attribute.
[0201] For further illustration, optionally assume that in a music streaming application, the seventh data was originally a relatively unpopular song and was stored in the low-heat data block. However, over time and with continuous discovery and sharing by users, this song suddenly becomes very popular. The system detects a significant change in the heat attribute of this song and automatically moves it from the low-heat data block to the high-heat data block to ensure faster access speed and better user experience.
[0202] By updating the position of the seventh data in the solid-state drive according to the change in the heat attribute, the system can ensure that the data storage strategy always matches its actual heat.
[0203] As an alternative solution, for the convenience of understanding, the above data management method is applied to the distributed storage scenario, aiming to solve problems such as a certain percentage decrease in the steady-state performance of the cluster compared to the initial performance and a shortened service life of the cluster disks after the cluster has been continuously used for a period of time in a distributed all-flash storage system, so as to improve the steady-state performance of the cluster in the distributed storage system under long-term use, extend the service life of the cluster, and improve the service quality of the distributed cluster.
[0204] Specifically, this embodiment proposes an innovative multi-stream write shunting method for a distributed storage system. This method writes the data that has undergone active shunting processing (i.e., the data for which hot and cold data have been distinguished and different stream IDs have been marked respectively) into a multi-stream solid-state disk through a refined data shunting strategy. This strategy can not only make full use of the performance characteristics of the multi-stream solid-state disk, but also significantly improve the overall performance stability of the distributed storage system, while extending the service life of the disk, thereby improving the service quality of the entire storage system.
[0205] It should be noted that this embodiment innovatively introduces a data shunting method in the distributed storage architecture, that is, refined distinction based on data heat. Throughout the process from the data receiving operation to the final writing into the solid-state drive (SSD) disk, we classify the data according to its heat attribute. When the data is written into the SSD disk, data with different heats will carry specific stream identifiers (stream id), and the multi-stream solid-state disk will write the data into different data blocks (block) according to these identifiers, thereby significantly reducing the migration amount of cold data during the garbage collection (GC) process, improving the GC efficiency, and reducing the write amplification phenomenon.
[0206] In a distributed storage system, the diversity of storage pools provides flexibility for data management. For example, the cache pool is designed specifically to improve the performance of the cluster, and it stores data with high read and write frequencies, that is, high-heat data. In contrast, the large-capacity data pool is used for persistent storage of user data, and its data update frequency is relatively low. In addition, in specific application scenarios, the storage pool will be further divided into three types: hot, warm, and cold, to more precisely manage data with different heats.
[0207] In addition to the data in the storage pool, the distributed storage system also involves the update of metadata and WAL (Write-Ahead Logging) data. Metadata records all the modification histories of data objects, so its update frequency is extremely high. And WAL data plays an important role in the scenario of random writes of small objects, improving I / O performance by writing logs first and then flushing the disk, and its update is also quite frequent.
[0208] Based on the above analysis, in this embodiment, the data heat in distributed storage is divided into the following levels from high to low: metadata, WAL data, cache pool data, specific storage pool data (such as hot, warm, and cold storage pool data), and data pool data. This division method fully considers the access frequency and update characteristics of the data, providing a solid foundation for subsequent data shunting operations.
[0209] Optionally, this embodiment also supports data shunting and writing on the business side. This means that the business layer can distinguish the data heat according to its own data heat and cold algorithm, and submit the distinguished data to the OSD (Object Storage Daemon) for disk writing. In this mode, the heat priority of business data is higher than the heat setting of the storage pool.
[0210] When business data is written, if the data itself does not carry heat information, the OSD will assign a specific heat value (i.e., virtual flow ID) to it according to the heat of the storage pool to which the data belongs. When the data is submitted to the multi-stream solid-state disk, this virtual flow ID will be passed through, and the multi-stream solid-state disk will write the data into the corresponding data block according to this ID. In this way, we have realized the multi-stream write operation in distributed storage, which not only improves the overall performance of the system but also extends the service life of the solid-state disk.
[0211] For further illustration, the specific implementation steps of the above shunting method are as follows:
[0212] S10-1, Data heat virtual flow ID assignment: First, assign a unique virtual flow ID to each type of data according to the data heat. The metadata is given the lowest heat value of 1, and then the WAL data, cache pool data, common pool data, and data pool data increase the heat value in turn to ensure the priority and processing efficiency of the data in the storage system.
[0213] S10-2, Storage pool heat configuration function: The Monitor module integrates the storage pool data heat assignment function, and through a standardized external interface / command, allows users to flexibly configure and modify the heat information (pool_hotness) of the storage pool. This function enhances the flexibility and adaptability of the system and meets the requirements of different application scenarios.
[0214] S10-3, Storage pool Hotness setting strategy: During the creation or subsequent maintenance stage of the storage pool, according to the specific usage scenario of the storage pool, assign a Hotness value to the storage pool through active configuration or system automatic recognition. The setting range is limited between 3 and n, where the metadata and WAL data occupy the first two heat levels to ensure the high-priority processing of key data.
[0215] S10-4, OSDMAP Update and Synchronization: Once the Monitor receives a change request for the storage pool heat value, it immediately updates the OSDMAP and pushes the updated OSDMAP to all OSD nodes. This step ensures the real-time synchronization of the storage pool heat information among all nodes in the storage system, providing an accurate basis for data shunting.
[0216] S10-5, Client Business OP Data Heat Processing: When the client business OP writes to the OSD, the system performs special processing on the data part. During the Transaction encapsulation phase, the op_hotness variable is introduced to record the heat value of the storage pool to which this OP belongs. The pool_hotness corresponding to the pool_id is obtained by querying the OSDMAP and assigned to op_hotness, and then it is submitted to bluestore along with the Transaction. When bluestore submits the data, op_hotness is passed to the MSSD as the virtual flow ID to achieve data shunting according to the heat.
[0217] S10-6, Metadata and WAL Data Heat Management: For the metadata and WAL data of the OP, the system adopts an independent processing flow. The data is first submitted to Rocksdb, and the data types are distinguished by the prefix information. The db_hotness variable is added to the FileWriter structure, and db_hotness is set to 2 or 1 (metadata) according to the prefix value (L represents WAL data). When writing to the disk in BlueFS, the data is submitted to the MSSD according to db_hotness as the virtual flow ID to ensure that the metadata and WAL data are stored according to the heat.
[0218] S10-7, Special Scenario Data Heat Processing: For the special case where the data OP write occurs before the storage pool heat distribution, the system defaults the heat of this data OP to the lowest (i.e., the data pool data heat), and processes it through the preset op_hotness value to ensure that the data can still be stored according to the rules in special cases.
[0219] S10-8, Data Reconstruction and Heat Maintenance in Fault Scenarios: During the data reconstruction process caused by OSD failures, the system follows the processing flows of steps S10-5 and S10-6 to ensure that the reconstructed data is shunted and written to the MSSD according to the heat value of the original storage pool. This mechanism guarantees the consistency and integrity of the data during the fault recovery process, and at the same time maintains the efficient management of the data in the storage system.
[0220] In an alternative embodiment, the above shunting method involves two aspects: the storage pool heat management module and the data write shunting module. Among them, the optimized design of the storage pool heat management module includes expanding the OSDMAP attribute. Given that the OSDMAP (OSD mapping) maintained by the Monitor already contains the attribute information of some storage pools, and this information can be synchronously shared with each OSD (Object Storage Device Daemon), this embodiment adds a key attribute - data heat in the OSDMAP. The maintenance of this attribute is the responsibility of the Monitor, and at the same time, the Monitor also provides an interface that can be actively configured and updated externally to ensure the flexibility and accuracy of the data heat attribute.
[0221] The optimized design of the storage pool heat management module also includes the storage pool heat configuration. When creating a storage pool during the deployment of the storage system, according to the specific usage mode and business requirements of the storage pool, the administrator can set different data heat values for the storage pool through the management software interface or by executing operation and maintenance commands. In addition, the system also supports automatically identifying and setting the heat value internally in the code, which is especially suitable for the automatic partitioning mode of hot, warm, cold, etc. storage pools. After receiving these configuration information, the Monitor will publish a new OSDMAP to each OSD to ensure that all OSDs can obtain the latest storage pool heat information.
[0222] The optimized design of the storage pool heat management module also includes virtual flow ID assignment and data writing. When the OSD receives the updated OSDMAP, it will assign a corresponding virtual flow ID to the operation request (OP) issued by the client according to the storage pool information to which the request belongs. This virtual flow ID is used as the identifier for data writing and will subsequently be submitted to the Device layer. After receiving the data carrying the virtual flow ID, the Multi-Stream Solid State Disk (MSSD) will write the data into the corresponding Block according to these ID information. This design ensures that data with different heats can be effectively isolated and managed separately, thereby improving the GC (Garbage Collection) efficiency, reducing the write amplification, and extending the service life of the MSSD.
[0223] For further illustration, the process is as Figure 3 shown as follows:
[0224] Storage pool creation and heat partitioning: In the user interface (UI) or the operation and maintenance node, when a new storage pool is created, the system will perform heat partitioning by the performance module according to the functions and expected usage scenarios it provides. This heat information is then sent to the MON (monitoring node).
[0225] The MON (monitor daemon of the cluster) processes storage: After receiving the heat mark of the storage pool, the MON saves and integrates it into the OSD MAP (Object Storage Device Map). The OSD MAP is a key data structure in the storage system used to describe the status, layout, and configuration of OSDs (Object Storage Devices).
[0226] OSD MAP publication and OSD operations: The updated OSD MAP is published to the OSDs in the system. The OSDs make corresponding configuration and status adjustments according to the received MAP (mapping) information to ensure that the heat information of the storage pool is correctly applied.
[0227] Client operations and application of heat information: When the client initiates data operation requests (such as writes), these operations are optimized according to the heat information of the storage pool. For example, high-heat storage pools may prioritize the processing of critical business data to ensure system performance and response speed.
[0228] In an alternative embodiment, the optimization design of the data write shunt module includes the following steps:
[0229] S11-1, Obtain the data heat value of the storage pool:
[0230] After the Monitor publishes a new OSD MAP, each OSD will be able to synchronously obtain the data heat value of each storage pool. This step ensures that when the OSD performs subsequent data write operations, it can perform reasonable shunt processing based on the latest heat information of the storage pool;
[0231] S11-2, Distribution of client write operations:
[0232] The write operation (OP) initiated by the client to the storage pool, after reaching the OSD (and finally being allocated to the corresponding PG for data submission), is split into three main parts: metadata, WAL data (for small object writes), and the actual data part. This split helps to perform more refined management and shunt for different types of data;
[0233] S11-3, Heat distribution and disk writing of the data part:
[0234] After the data part is encapsulated into a Transaction, it will be submitted to Bluestore for disk writing. During this process, the OSD will assign a corresponding data heat value to it according to the storage pool to which the data object belongs, and embed it into the Transaction. Subsequently, when Bluestore performs data disk writing, it will assign a virtual stream ID to the data according to this heat value and submit it to the multi-stream solid-state disk (MSSD). This design ensures that data with different heats can be effectively isolated and managed separately;
[0235] S11-4, Heat processing and disk writing of metadata and WAL data:
[0236] Before the metadata and WAL data are submitted to RocksDB, different prefixes will be assigned according to their types (for example, the prefix of WAL data is "L", while the metadata contains multiple prefixes such as "O, M, P", etc.). This embodiment uses this prefix information to assign different heat values to the data submitted to BlueFS, mainly divided into two categories: metadata and WAL. When BlueFS performs disk writing operations, it will also assign corresponding virtual stream IDs to the data according to these heat values and submit them to the MSSD. This method not only simplifies the complexity of heat management but also improves the flexibility and efficiency of data writing;
[0237] S11-5, Write processing of fault reconstruction data:
[0238] For the data reconstruction situation caused by OSD failure, the submission process of its data writing is the same as steps S11-3 and S11-4 above. This design ensures data consistency during the fault recovery process and the continuity of heat management;
[0239] S11-6, Heat allocation strategy for historical data:
[0240] For historical data that has been written to the OSD before the storage pool data heat configuration, this embodiment adopts a conservative strategy: the processing of its metadata and WAL data is consistent with the above rules, while for the data part, it is uniformly assigned according to the heat value of the coldest data. This strategy ensures that when the heat of these data becomes higher later, they can be written into a new data stream, thus avoiding performance losses caused by repeated migrations during garbage collection (GC).
[0241] For further illustration with examples, the process is as Figure 4 shown as follows:
[0242] S12-1, Data sending: The client can choose to directly send the data to the storage pool (OF (write operation) is issued to the storage pool) or first send it to the PG (placement group) (OF is issued to the PG).
[0243] S12-2, Direct storage process: The data directly flows to the OSD (Object Storage Device) / PG, and then enters the Bluefs (Blue File System) storage system. Bluefs allocates the data to different sIDs (such as unique identifier 1) according to the policy, and finally stores it in the SSD Device (Solid State Drive).
[0244] S12-3, Storage process through PG: The data is first preprocessed through the WAL (Write-Ahead Log) to ensure data consistency. A part of the processed data continues to flow to Bluefs, and is allocated and stored in the SSD Device through the sID (unique identifier) (such as unique identifier 2).
[0245] Another part of the data directly enters the Transactions module, which allocates Hotness according to the storage pool to which the OP belongs and the OSDMAP (Object Storage Device Mapping) information. The data after allocating Hotness flows to Bluestore, and then Bluestore allocates and stores it in the SSD Device (Solid State Drive) according to the sID (such as unique identifier n).
[0246] S12-4, WAL and Hotness processing: When the op is submitted, Hotness is allocated according to the prefix information (O / M / P and L) to ensure that the data is stored according to the priority.
[0247] Through the embodiments provided by the present application, the data flow is finely managed, significantly improving the utilization efficiency of the solid-state disk (SSD) in the distributed all-flash storage system, and further greatly enhancing the garbage collection (GC) efficiency of the SSD, bringing a significant improvement to the overall performance of the storage cluster. Compared with the prior art, the present embodiment shows the following technical advantages:
[0248] Performance improvement: The present embodiment uses the shunt technology to intelligently classify the data according to the hotness and allocate it to different storage streams. This effectively improves the GC efficiency of the SSD disk because the aggregation of the same type of data optimizes the erasure and writing patterns. This optimization is directly reflected in the significant improvement of the overall performance of the storage system, especially in the application scenarios with high concurrency and large data volume, where the performance improvement is particularly obvious.
[0249] Function Enhancement and Lifespan Extension: By implementing an effective shunting strategy, this embodiment reduces the write amplification phenomenon of the SSD, that is, the number of repeated erasures of the same storage unit is effectively controlled. This not only extends the service life of the SSD but also reduces the risk of storage particle damage caused by frequent erasures. At the same time, it reduces the need for disk replacement and the disk replacement cost, directly enhancing the reliability and economy of the storage system.
[0250] Operation and Maintenance Simplification and Space Utilization Optimization: Traditional storage deployment methods often require dividing a single disk into multiple partitions (such as database partitions, WAL partitions, etc.), which not only occupies additional disk frame space but also increases the complexity of operation and maintenance management. This embodiment utilizes the characteristics of multi-streams within the disk to intelligently write different types of data into different blocks, thereby saving valuable disk frame space and increasing the density of data storage. This improvement not only simplifies the operation and maintenance process but also makes the management of storage resources more efficient and flexible.
[0251] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0252] In this embodiment, a data management device is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0253] According to another aspect of the embodiments of this application, a data management device for implementing the above data management method is also provided. As Figure 5 shown, this device includes:
[0254] A first acquisition unit 502, configured to acquire a plurality of data to be written into the solid-state drive;
[0255] A second acquisition unit 504, configured to acquire the heat attribute corresponding to each data among the plurality of data;
[0256] A writing unit 506 for writing each piece of data into each data block in the solid-state drive according to the heat property, where different data blocks correspond to different heat property conditions.
[0257] For specific embodiments, reference can be made to the examples shown in the above data management method, and details are not described herein again.
[0258] As an alternative solution, the second acquisition unit 504 includes:
[0259] A first acquisition module for acquiring the data types corresponding to the respective pieces of data;
[0260] A second acquisition module for acquiring the heat property according to the data type.
[0261] For specific embodiments, reference can be made to the examples shown in the above data management method, and details are not described herein again.
[0262] As an alternative solution, the second acquisition module includes:
[0263] A first acquisition sub-module for acquiring the heat corresponding to the respective pieces of data according to the data type, where different data types correspond to different heats;
[0264] A second acquisition sub-module for acquiring the heat property according to the heat, where the different heat property conditions are different heat ranges.
[0265] For specific embodiments, reference can be made to the examples shown in the above data management method, and details are not described herein again.
[0266] As an alternative solution, the first acquisition module includes at least one of the following:
[0267] A third acquisition sub-module for acquiring the specified data type corresponding to the first data among the multiple pieces of data;
[0268] A fourth acquisition sub-module for acquiring the stored data type corresponding to the second data among the multiple pieces of data;
[0269] A fifth acquisition sub-module for acquiring the service data type corresponding to the third data among the multiple pieces of data.
[0270] For specific embodiments, reference can be made to the examples shown in the above data management method, and details are not described herein again.
[0271] As an alternative solution, the first acquisition sub-module includes:
[0272] A first determination subunit, configured to determine that the popularity corresponding to the data of the specified data type is greater than the popularity corresponding to the data of the stored data type when the data of the specified data type and the data of the stored data type are obtained.
[0273] For specific embodiments, reference may be made to the examples shown in the above data management method, and details are not described herein again in this example.
[0274] As an alternative solution, the first acquisition sub-module includes:
[0275] A second determination subunit, configured to determine that the popularity corresponding to the data of the service data type is greater than the popularity corresponding to the data of the stored data type when the data of the service data type and the data of the stored data type are obtained.
[0276] For specific embodiments, reference may be made to the examples shown in the above data management method, and details are not described herein again in this example.
[0277] As an alternative solution, the fourth acquisition sub-module includes:
[0278] A third determination subunit, configured to determine the storage data type corresponding to the second data as the cache pool storage type when the second data is the data stored in the cache pool;
[0279] A fourth determination subunit, configured to determine the storage data type corresponding to the second data as the general pool storage type when the second data is the data stored in the general pool, where the popularity corresponding to the cache pool storage type is greater than the popularity corresponding to the general pool storage type;
[0280] A fifth determination subunit, configured to determine the storage data type corresponding to the second data as the data pool storage type when the second data is the data stored in the data pool, where the popularity corresponding to the general pool storage type is greater than the popularity corresponding to the data pool storage type.
[0281] For specific embodiments, reference may be made to the examples shown in the above data management method, and details are not described herein again in this example.
[0282] As an alternative solution, the third acquisition sub-module includes:
[0283] A sixth determination subunit, configured to determine the specified data type corresponding to the first data as the metadata type when the first data is metadata;
[0284] A seventh determination subunit, configured to, when the first data is WAL data, determine the specified data type corresponding to the first data as the WAL data type, where the popularity of the metadata type is greater than the popularity of the WAL data type.
[0285] For specific embodiments, reference may be made to the examples shown in the above data management method, and details are not described herein again in this example.
[0286] As an alternative solution, the fifth acquisition sub-module includes:
[0287] An eighth determination subunit, configured to determine the data that conforms to the business logic among the multiple data as the third data;
[0288] A ninth determination subunit, configured to determine the business data type based on the business type corresponding to the business logic.
[0289] For specific embodiments, reference may be made to the examples shown in the above data management method, and details are not described herein again in this example.
[0290] As an alternative solution, the first acquisition sub-module includes:
[0291] A first acquisition subunit, configured to, when the data type is the business data type, acquire at least two types of popularity-related parameters associated with the third data;
[0292] An allocation subunit, configured to, when the data type is the business data type, allocate calculation weights to each of the at least two types based on the business logic;
[0293] A second acquisition subunit, configured to, when the data type is the business data type, use the calculation weights and combine with the popularity-related parameters of the at least two types to acquire the popularity of the third data.
[0294] For specific embodiments, reference may be made to the examples shown in the above data management method, and details are not described herein again in this example.
[0295] As an alternative solution, the first acquisition subunit includes at least two of the following:
[0296] A first sub-acquisition module, configured to acquire the data access frequency associated with the third data;
[0297] A second sub-acquisition module, configured to acquire the data update frequency associated with the third data;
[0298] A third sub-acquisition module, configured to acquire the data access times associated with the third data;
[0299] A fourth sub-acquisition module, configured to acquire the number of data updates associated with the third data;
[0300] A fifth sub-acquisition module, configured to acquire the data access timestamp associated with the third data;
[0301] A sixth sub-acquisition module, configured to acquire the data update timestamp associated with the third data.
[0302] For specific embodiments, reference may be made to the examples shown in the above data management method, and details thereof are not described herein again.
[0303] As an alternative solution, the writing unit 506 includes:
[0304] An allocation module, configured to allocate virtual stream identifiers for the respective data according to the popularity attribute, where the virtual stream identifier is used to represent the popularity of the respective data;
[0305] A first writing module, configured to write the respective data into the respective data blocks according to the virtual stream identifier, where different data blocks correspond to different virtual stream identifiers, and different virtual stream identifiers correspond to different popularity attribute conditions.
[0306] For specific embodiments, reference may be made to the examples shown in the above data management method, and details thereof are not described herein again.
[0307] As an alternative solution, the allocation module includes:
[0308] A first allocation sub-module, configured to, when the popularity attribute corresponding to the fourth data among the multiple data carries the popularity information of the fourth data, allocate a first virtual stream identifier for the fourth data based on the popularity information.
[0309] For specific embodiments, reference may be made to the examples shown in the above data management method, and details thereof are not described herein again.
[0310] As an alternative solution, the allocation module includes:
[0311] A sixth acquisition sub-module, configured to, when the popularity attribute corresponding to the fifth data among the multiple data lacks the popularity information of the fifth data, acquire the popularity information of the storage pool to which the fifth data belongs;
[0312] A second allocation sub-module, configured to allocate a second virtual stream identifier for the fifth data based on the popularity information of the storage pool to which it belongs.
[0313] For specific embodiments, reference may be made to the examples shown in the above data management method, and details thereof are not described herein again.
[0314] As an alternative, the second acquisition unit 504 includes:
[0315] A setting module, configured to, when the sixth data among the multiple data satisfies a preset popularity condition, set the popularity attribute corresponding to the preset popularity condition as the popularity attribute corresponding to the sixth data.
[0316] For specific embodiments, reference may be made to the examples shown in the above data management method, and details are not described herein again in this example.
[0317] As an alternative, the writing unit 506 includes:
[0318] A second writing module, configured to write the data among the multiple data that meets the high popularity condition into the first data block in the solid-state drive according to the popularity attribute, where the popularity attribute condition includes the high popularity condition;
[0319] A third writing module, configured to write the data among the multiple data that meets the low popularity condition into the second data block in the solid-state drive according to the popularity attribute, where the storage space of the first data block is larger than that of the second data block, and the popularity attribute condition includes the low popularity condition.
[0320] For specific embodiments, reference may be made to the examples shown in the above data management method, and details are not described herein again in this example.
[0321] As an alternative, the device further includes:
[0322] An updating unit, configured to, after writing each of the data into each data block in the solid-state drive according to the popularity attribute, when the popularity attribute corresponding to the seventh data written in the data block changes, update the position of the seventh data in the solid-state drive according to the changed popularity attribute.
[0323] For specific embodiments, reference may be made to the examples shown in the above data management method, and details are not described herein again in this example.
[0324] It should be noted that the above-mentioned various virtual devices (modules, units, sub-modules, sub-units, components, etc.) can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: the above-mentioned virtual devices are all located in the same processor; or, the above-mentioned various virtual devices are respectively located in different processors in any combination form.
[0325] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0326] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media capable of storing computer programs such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0327] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0328] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0329] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be elaborated herein.
[0330] Obviously, those skilled in the art should understand that the above virtual devices or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. Thus, the present application is not limited to any specific combination of hardware and software.
[0331] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included in the protection scope of the present application.
Claims
1. A data management method, characterized in that, Including: Obtain a plurality of data to be written into a solid-state drive; Determine the data that conforms to the business logic among the plurality of data as the third data; Determine the business data type corresponding to the business logic, where the business data type is a data type related to a specific business logic or business process; Obtain a heat attribute according to the data type; when the data type is the business data type, obtain at least two types of heat-related parameters associated with the third data; Based on the business logic, assign a calculation weight to each of the at least two types, where the calculation weight is used to represent the importance of the heat-related parameters; Utilize the calculation weight to combine the heat-related parameters of the at least two types to obtain the heat of the third data, where the heat of the third data is used to represent the importance and activity of the third data in the business; Write each of the data into each data block in the solid-state drive according to the heat attribute, where different data blocks correspond to different heat attribute conditions.
2. The method according to claim 1, characterized in that, The obtaining the heat attribute according to the data type includes: Obtain the heat corresponding to each of the data according to the data type, where different data types correspond to different heats; Obtain the heat attribute according to the heat, where the different heat attribute conditions are different heat ranges.
3. The method according to claim 2, wherein The obtaining the data type corresponding to each of the data includes at least one of the following: Obtain the specified data type corresponding to the first data among the plurality of data; Obtain the storage data type corresponding to the second data among the plurality of data.
4. The method according to claim 3, characterized in that The obtaining the heat corresponding to each of the data according to the data type includes: When the data of the specified data type and the data of the storage data type are obtained, determine that the heat corresponding to the data of the specified data type is greater than the heat corresponding to the data of the storage data type.
5. The method according to claim 3, wherein The obtaining the heat corresponding to each of the data according to the data type includes: When the data of the business data type and the data of the storage data type are obtained, determine that the heat corresponding to the data of the business data type is greater than the heat corresponding to the data of the storage data type.
6. The method according to claim 3, wherein The obtaining the storage data type corresponding to the second data among the plurality of data includes: When the second data is data stored in a cache pool, determine the storage data type corresponding to the second data as the cache pool storage type; When the second data is data stored in a normal pool, determine the storage data type corresponding to the second data as the normal pool storage type, where the heat corresponding to the cache pool storage type is greater than the heat corresponding to the normal pool storage type; When the second data is data stored in a data pool, determine the storage data type corresponding to the second data as the data pool storage type, where the heat corresponding to the normal pool storage type is greater than the heat corresponding to the data pool storage type.
7. The method according to claim 3, wherein The obtaining the specified data type corresponding to the first data among the plurality of data includes: When the first data is metadata, determine the specified data type corresponding to the first data as the metadata type; When the first data is WAL data, determine the specified data type corresponding to the first data as the WAL data type, where the popularity of the metadata type is greater than the popularity of the WAL data type.
8. The method according to claim 1, wherein The obtaining of at least two heat-related parameters associated with the third data includes at least two of the following: Obtain the data access frequency associated with the third data; Obtain the data update frequency associated with the third data; Obtain the number of data accesses associated with the third data; Obtain the number of data updates associated with the third data; Obtain the data access timestamp associated with the third data; Obtain the data update timestamp associated with the third data.
9. The method according to claim 1, wherein The writing of each data into each data block in the solid-state drive according to the heat attribute includes: According to the heat attribute, assign a virtual stream identifier to each data, where the virtual stream identifier is used to represent the heat of each data; Write each data into each data block according to the virtual stream identifier, where different data blocks correspond to different virtual stream identifiers, and different virtual stream identifiers correspond to different heat attribute conditions.
10. The method according to claim 9, wherein The assigning of a virtual stream identifier to each data according to the heat attribute includes: When the heat attribute corresponding to the fourth data among the multiple data carries the heat information of the fourth data, assign a first virtual stream identifier to the fourth data based on the heat information.
11. The method according to claim 9, wherein The assigning of a virtual stream identifier to each data according to the heat attribute includes: When the heat attribute corresponding to the fifth data among the multiple data lacks the heat information of the fifth data, obtain the heat information of the storage pool to which the fifth data belongs; Assign a second virtual stream identifier to the fifth data based on the heat information of the storage pool to which it belongs.
12. The method according to any one of claims 1 to 11, characterized in that, The obtaining of the heat attribute corresponding to each data among the multiple data includes: When the sixth data among the multiple data meets a preset heat condition, set the heat attribute corresponding to the preset heat condition as the heat attribute corresponding to the sixth data.
13. The method according to any one of claims 1 to 11, characterized in that, The writing of each data into each data block in the solid-state drive according to the heat attribute includes: According to the heat attribute, write the data that meets the high heat condition among the multiple data into the first data block in the solid-state drive, where the heat attribute condition includes the high heat condition; According to the heat attribute, write the data that meets the low heat condition among the multiple data into the second data block in the solid-state drive, where the storage space of the first data block is larger than the storage space of the second data block, and the heat attribute condition includes the low heat condition.
14. The method according to any one of claims 1 to 11, characterized in that, After writing each data into each data block in the solid-state drive according to the heat attribute, the method further includes: In the case where the heat property corresponding to the seventh data written in the data block changes, update the position of the seventh data in the solid-state drive according to the changed heat property.
15. A data management device, characterized in that, Including: A first acquisition unit for acquiring a plurality of data to be written into a solid-state drive; A second acquisition unit for acquiring the heat property corresponding to each data in the plurality of data; A writing unit for writing each data into each data block in the solid-state drive according to the heat property, where different data blocks correspond to different heat property conditions; The device is further configured to: Determine the third data from the plurality of data that conforms to the business logic; Determine the service data type according to the service type corresponding to the business logic, where the service data type is a data type related to a specific business logic or business process; Obtain the heat property according to the data type; in the case where the data type is the service data type, obtain at least two types of heat-related parameters associated with the third data; Based on the business logic, assign a calculation weight to each type in the at least two types, where the calculation weight is used to represent the importance of the heat-related parameters; Utilize the calculation weight to combine the heat-related parameters of the at least two types to obtain the heat of the third data, where the heat of the third data is used to represent the importance and activity of the third data in the business.
16. A computer-readable storage medium, characterized in that A computer program is stored in the computer-readable storage medium, where when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 14 are implemented.
17. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the computer program, the steps of the method described in any one of claims 1 to 14 are implemented.
Citation Information
Patent Citations
Data processing method, device and equipment based on solid-state disk array and storage medium
CN111880745A
Multi-storage-pool data classified storage method and system and electronic equipment
CN115470190A