Partition identification method, storage system, electronic equipment and storage medium

By collecting and analyzing access characteristic data of data partitions in the table storage system, generating access hot data to identify hot spot partitions, the problem of inaccurate identification of hot spot partitions in the prior art is solved and the access performance of the system is improved.

CN120104037APending Publication Date: 2025-06-06HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311650510.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the table storage system, when dynamically solving hotspot partitioning problems, it is key to accurately identify hotspot partitioning, but the existing technology is difficult to effectively implement.

Method used

By collecting access characteristic data generated by multiple data partitions in a distributed storage system, and generating access hot data under each access operation type based on these data and supported access operation types, the target data partition that meets the requirements is identified.

Benefits of technology

It improves the accurate recognition rate of hotspot partitions, dynamically adjusts partitions to reduce data skew problems, and improves the access performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104037A_ABST
    Figure CN120104037A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a partition recognition method, a storage system, electronic equipment and a storage medium. In the embodiment of the invention, access operation types supported by a plurality of data partitions in the distributed storage system and a plurality of pieces of generated access feature data are considered at the same time; afterwards, according to a plurality of pieces of access feature data generated by the plurality of collected data partitions and supported access operation types, generating access heat data of the plurality of data partitions under each access operation type, and based on the access heat data, selecting a target data partition meeting the access heat requirement; the access feature data and the access operation type are comprehensively considered when the hotspot partition is identified, so that the accuracy of identifying the hotspot partition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storage technology, and in particular to a partition identification method, a storage system, an electronic device and a storage medium. Background Art

[0002] The emergence of non-relational (Not Only SQL, NoSQL) databases solves the challenges brought by large-scale data sets and multiple data types. Table storage, as a type of NoSQL database, can provide table storage services for structured data in scenarios such as massive bills, instant messaging, the Internet of Things, Internet of Vehicles, risk control, and recommendations.

[0003] In table storage, data is organized in tables, and a table is divided into multiple partitions using data sharding technology. These partitions are then dispatched to different storage devices to provide external services using load balancing technology. If you want better access performance for the table, the read and write access volume and data volume on these partitions should be distributed as evenly as possible. In actual applications, there is often a data skew problem where the access popularity of some partitions is much higher than that of other partitions due to various reasons, resulting in poor access performance.

[0004] In some existing solutions, data skew caused by hotspot partitions is prevented by properly designing table structures or partitioning tables. However, due to the uneven level of table building by users and the large fluctuation of access, it is difficult to prevent hotspot partition problems through table design and pre-partitioning, which makes it important to dynamically solve the hotspot partition problem during the access process. However, in the process of dynamically solving the hotspot partition problem, accurate identification of hotspot partitions is a key issue. Summary of the invention

[0005] Multiple aspects of the present application provide a partition identification method, a storage system, an electronic device and a storage medium, which are used to improve the accuracy of identifying hotspot partitions in the process of dynamically solving hotspot partition problems.

[0006] An embodiment of the present application provides a partition identification method, including: collecting multiple access feature data generated by each of multiple data partitions in a distributed storage system, where the multiple access feature data are feature data generated by access operations performed on the data partitions; generating access heat data for the multiple data partitions under each access operation type based on the multiple access feature data generated by each of the multiple data partitions and the supported access operation types; and identifying a target data partition whose access heat meets the requirements based on the access heat data for the multiple data partitions under each access operation type.

[0007] An embodiment of the present application also provides a storage system, including at least one storage node, on which at least one data partition from at least one data table is stored; the storage system also includes: a partition identification node; the partition identification node is used to collect multiple access feature data generated by each of multiple data partitions in the distributed storage system, and the multiple access feature data are feature data generated by access operations performed on the data partitions; based on the multiple access feature data generated by each of the multiple data partitions and the supported access operation types, access heat data of the multiple data partitions under each access operation type is generated; based on the access heat data of the multiple data partitions under each access operation type, a target data partition whose access heat meets the requirements is identified.

[0008] An embodiment of the present application also provides an electronic device, including: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to implement the steps in the above method.

[0009] An embodiment of the present application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor implements the steps in the above method.

[0010] The technical solution provided by the embodiment of the present application collects access feature data generated by each of the multiple data partitions in the distributed storage system, generates access heat data of the multiple data partitions under each access operation type according to the multiple access feature data generated by each of the collected data partitions and the access operation types supported, and based on this, based on the access heat data of the multiple data partitions under each access operation type, selects the target data partition that meets the access heat requirements, and comprehensively considers the access feature data and the access operation type when identifying the hotspot partitions, thereby improving the accuracy of identifying the hotspot partitions. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0012] Figure 1a A schematic diagram of the structure of a distributed storage system provided in an embodiment of the present application;

[0013] Figure 1b A schematic diagram of the structure of another distributed storage system provided in an embodiment of the present application;

[0014] Figure 1c A schematic diagram of the structure of another distributed storage system provided in an embodiment of the present application;

[0015] Figure 2A schematic diagram of table partitioning provided in an embodiment of the present application;

[0016] Figure 3 A schematic diagram of a hotspot partitioning process provided in an embodiment of the present application;

[0017] Figure 4 A schematic diagram of a flow chart of a partition identification method provided in an embodiment of the present application;

[0018] Figure 5 A flowchart of another partition identification method provided in an embodiment of the present application;

[0019] Figure 6 A schematic diagram of the structure of a partition identification device provided in an embodiment of the present application;

[0020] Figure 7 A schematic diagram of the structure of a storage system provided in an embodiment of the present application;

[0021] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0023] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0024] In the process of dynamically solving the hotspot partition problem, how to accurately identify the hotspot partition becomes a key issue that needs to be solved urgently. In the embodiment of the present application, the access operation types supported by multiple data partitions in the distributed storage system and the access feature data generated are considered at the same time; by collecting the access feature data generated by each of the multiple data partitions, according to the multiple access feature data generated by each of the collected multiple data partitions and the access operation types supported, the access heat data of multiple data partitions under each access operation type is generated, based on which, the target data partition that meets the access heat requirements is selected, and the access feature data and the access operation type are comprehensively considered when identifying the hotspot partition, thereby improving the accuracy of identifying the hotspot data partition.

[0025] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0026] Figure 1a-Figure 1c The framework diagram of the distributed storage system provided by the embodiment of the present application is shown. In the embodiment of the present application, the product implementation form of the distributed storage system is not limited. For example, it can be implemented as various non-relational database products, and further can be implemented as but not limited to: a distributed table storage system, a remote dictionary service (RemoteDictionary Server, Redis) system or a column-oriented database system, etc. Figure 1a-Figure 1c As shown, no matter which storage product is specifically implemented, the distributed storage system may include an agent node 11 , a master node 12 and a storage node 13 .

[0027] The proxy node 11 serves as an interactive interface between the user and the distributed storage system, such as an application programming interface (API) or a software development kit (SDK) of the distributed storage system, and is responsible for receiving various user requirements for the distributed storage system, such as receiving user requests to create a data table, and receiving user requirements to access the data table, and providing the user's various requirements to the master node 12 in the distributed storage system, so that the master node 12 responds to the user's various requirements and provides corresponding services to the user, such as creating a data table for the user and providing the user with read and write services for the data table, etc. The user in this embodiment refers to the user of the distributed storage system, such as an individual user, a company or various institutions, etc., and the user of the distributed storage system may be one or some of the user's applications, application systems or a functional module in the application system, etc.

[0028] It should be noted that the proxy node 11 is an optional component, and the distributed storage system may not include the proxy node 11, but the user directly interacts with the master node 12.

[0029] The master node 12 is a node responsible for scheduling in the distributed storage system. On the one hand, it is responsible for receiving various user requirements for the distributed storage system provided by the client 11 or directly provided by the user, such as the requirement to create a data table and the read and write requests for the data table; on the other hand, it is responsible for responding to various user requirements and scheduling the corresponding storage node 13 to provide data storage and read and write services for the user. For example, when the master node 12 receives the requirement to create a data table, it can create a data table for the user, and the data table has a unique identifier. For another example, when the master node 12 receives a read and write request for a data table, it can provide the read and write request to the corresponding storage node 13, and the storage node 13 performs the data read and write operation and returns the data read and write result to the user. In actual applications, one or more master nodes 12 can be deployed in a distributed storage system. For example, when multiple master nodes 12 are deployed, 3 can be deployed, but it is not limited to this.

[0030] The storage node 13 is a node in the distributed storage system responsible for data storage. On the one hand, it is responsible for storing data, and on the other hand, it can respond to the user's read and write requests, perform data read and write operations according to the read and write requests, and return the read and write results. The storage node 13 can also be called a worker node. Compared with the master node 12, the number of storage nodes 13 is relatively large. The specific data depends on the scale of the distributed storage system. It can be dozens, hundreds or even thousands, and there is no limitation on this.

[0031] In the distributed storage system provided in this embodiment, partition technology can be used, that is, the overall data can be divided into multiple data partitions (referred to as partitions for short), and these partitions are dispersed and stored on different storage nodes 13 under the scheduling of the master node 12, which is conducive to improving the concurrent access capability of the distributed storage system, and thus improving the overall service capability of the distributed storage system. According to the different product implementation forms of the distributed storage system, the implementation form of the overall data will also be different. For example, in a table storage system or a column-oriented database product, data is stored in the form of a data table, and the overall data that can be partitioned can specifically be a data table, that is, a data table can be divided into multiple partitions. In this embodiment, the partitioning operation can be performed by the master node 12, and the master node 12 can divide the overall data into multiple partitions according to the data range, and schedule these partitions to different storage nodes 13 for storage according to a certain scheduling strategy.

[0032] Further, if Figure 1a-Figure 1cAs shown, the distributed storage system also includes: a load balancing node 14, which is used to assist the master node 12 in distributing the workload in the distributed storage system to different storage nodes, so as to improve the service performance and reliability of the distributed storage system. For example, the load balancing node 14 can distribute the multiple partitions cut out by the master node 12 to different storage nodes 13 for storage, or can distribute the user's read and write requests to different storage nodes 13, so that when a storage node 13 fails, the data distributed on other storage nodes can still be accessed normally, which is conducive to improving the service performance and reliability of the distributed storage system.

[0033] According to different product implementation forms of the distributed storage system, the master node 12 may divide the overall data into multiple partitions in different ways. Figure 2 As shown in Table Storage, each data table includes multiple rows and multiple columns, and each data table includes one or more primary key columns and common columns. In the embodiment of the present application, a column of the primary key column is defined as a partition key. Figure 2 In the example, the first column in the primary key is defined as the shard key. When partitioning, you can partition according to the shard key. Data with the same shard key will be partitioned in the same partition. Of course, the shard keys of data in the same partition can be different. Figure 2 As shown in the figure, the data table is divided into three partitions: Partition A, Partition B, and Partition C.

[0034] In actual applications, it is often encountered that user access requests are concentrated on a certain data range, that is, the data requested by users are concentrated on some shard keys. Specifically, the user's access requests are concentrated on one partition or a few partitions. If the access volume is large enough, it is easy to exceed the capacity limit of a single partition or a single storage node, resulting in access errors. This phenomenon is called the hot spot access phenomenon, the partition being accessed is called the hot spot partition, and the access imbalance problem caused by the hot spot partition is called the data skew problem. For example, in the e-commerce field, the number of visits to products participating in flash sales will be relatively large, which will lead to a large number of visits to the partition where the product information is located. A large number of users request to access this partition, and this partition will become a hot spot partition and then become the performance bottleneck of the entire system, affecting system performance.

[0035] In order to solve the data skew problem caused by hotspot partitions, in the embodiment of the present application, at least one of the following methods can be used:

[0036] Method 1: During data table design, reasonably design the table structure of the data table, such as reasonably designing the primary key column and shard key to reduce the probability of hot spot partitions.

[0037] Method 2: After creating the table, partition the data table reasonably and try to distribute popular data to different partitions to reduce the probability of hotspot partitions.

[0038] Method 3: During dynamic access, hot partitions are dynamically identified and split, so that hot partitions can be discovered and processed in a timely manner, reducing the data skew problem caused by hot partitions.

[0039] The above methods 1-3 can be implemented at different stages. In actual applications, one can be selected for use, or two or more can be selected for use in combination, without limitation. Among them, considering that the construction of the data table is done by the user, the user may not have strong professional knowledge, and method 1 can be flexibly selected according to the user's professional knowledge and ability. For method 2, the popularity of the data can only be estimated in advance and reasonable partitioning can be performed as much as possible, but since the user's access situation will change dynamically with the application requirements and is unpredictable, the probability of hotspot partitions can only be reduced to a certain extent. Method 3 can dynamically and timely discover hotspot partitions, and process hotspot partitions in a timely manner, which can become the main means to solve the hotspot partition problem.

[0040] In this embodiment, a hotspot partition detection service can be deployed in the distributed storage system, and the hotspot partition detection service can solve the hotspot partition problem during the dynamic access process of the distributed storage system. In this embodiment, the deployment implementation method of the hotspot partition detection service is not limited, such as Figure 1a As shown in Figure 2, the hotspot partition detection service can be deployed on the master node, or as Figure 1b As shown in Figure 2, the hotspot partition detection service can also be deployed on a load balancing node, or as Figure 1c As shown, the hotspot partition detection service can also be deployed independently as a new functional node. Regardless of the system architecture, the principle and process of the hotspot partition detection service when performing hotspot partition identification are the same. In the following embodiments, Figure 3 Describes the process of solving the hotspot partition problem by using the hotspot partition detection service.

[0041] like Figure 3As shown, no matter how the hotspot partition detection service is deployed and implemented, when solving the hotspot partition problem, multiple access feature data generated by each data partition distributed on each storage node can be collected. These access feature data are feature data generated by access operations on data partitions; then, at least identify the hotspot partitions based on these access feature data; then, give corresponding splitting or migration plans for the identified hotspot partitions and perform splitting or migration operations according to the splitting or migration plans. Among them, how to accurately identify hotspot partitions is the key to solving the hotspot partition problem in the dynamic access process. In an embodiment of the present application, in order to improve the accuracy of identifying hotspot partitions, a hotspot partition identification scheme based on comprehensive indicators is proposed, which can efficiently and accurately identify hotspot partitions, and provide a basis for the subsequent automatic or manual splitting of hotspot partitions, assisting the downstream splitting operation to be carried out in a timely and accurate manner, and can effectively reduce the errors and performance losses caused by hotspot partition problems, and cope with the challenges brought by hotspot partition problems. In the following embodiments, the hotspot partition identification process will be described in detail in conjunction with the accompanying drawings.

[0042] Figure 4 This is a flow chart of a partition identification method provided by an embodiment of the present application. Figure 4 As shown, the method includes:

[0043] 401. Collect multiple access feature data generated by multiple data partitions in a distributed storage system, where the multiple access feature data are feature data generated by access operations performed on the data partitions;

[0044] 402. Generate access popularity data of the multiple data partitions under each access operation type according to the multiple access feature data generated by each of the multiple data partitions and the access operation types supported;

[0045] 403. According to the access heat data of the plurality of data partitions under each access operation type, a target data partition whose access heat meets the requirements is identified.

[0046] In this embodiment, the distributed storage system includes at least one storage node, and one storage node can store at least one data partition from at least one data table. In other words, one storage node stores one or more data partitions, and these data partitions can come from the same data table or from multiple different data tables; further, these data tables can come from the same user or from different users.

[0047] In this embodiment, multiple data partitions in the distributed storage system can be identified, and a target data partition whose access popularity meets the requirements can be identified. In this embodiment, a data partition with higher popularity can be identified as a target data partition, and a data partition with lower popularity can also be identified as a target data partition. In short, both hot data partitions and unpopular data partitions can be identified, depending on the partition identification requirements.

[0048] In this embodiment, the data partitions involved in partition identification are not limited. For example, it can be all data partitions in the distributed storage system, or it can be part of the data partitions, which can be flexibly set according to application requirements. In an optional embodiment, partition identification requirement information can be obtained, and multiple data partitions that need to be partition identified in the distributed storage system can be determined based on the partition identification requirement information. Depending on the different partition identification requirements, the multiple data partitions determined will also be different.

[0049] For example, in some application scenarios, it is necessary to identify the data partitions on each storage node. In this case, the data partitions distributed on the same storage node can be used as multiple data partitions that need to be identified, and then from the dimension of each storage node, the target data partitions whose access heat meets the requirements on each storage node can be identified.

[0050] In other application scenarios, it is necessary to identify data partitions on specific storage nodes, and the data distributed on the specific storage nodes can be grouped as multiple data partitions that need to be partitioned, and then the target data partitions on the specific storage nodes that meet the access heat requirements can be identified from the dimension of the specific storage nodes. Among them, the specific storage nodes can be flexibly determined according to application requirements, for example, they can be storage nodes with heavy loads, or storage nodes that store certain user data, or storage nodes distributed in a certain availability zone, etc.

[0051] In some other application scenarios, it is necessary to identify data partitions in the same data table. The data partitions from the same data table can be used as multiple data partitions that need to be identified, and then the target data partitions in the same data table whose access popularity meets the requirements can be identified from the dimension of the data table.

[0052] In some other application scenarios, it is necessary to identify data partitions in a data table for one or more specific tenants. The data partitions from the data table of one or more specific tenants can be used as multiple data partitions that need to be identified, and then from the dimension of the specific tenant, the target data partitions whose access popularity meets the requirements in the data table of the specific tenant are identified.

[0053] In this embodiment, accessing data partitions generates various access characteristic data. These access characteristic data are characteristic data generated by accessing data partitions, including but not limited to: the number of read operations on the data partitions, the amount of data read each time, the sum of the amount of data read by all read operations, and the number of write operations on the data partitions, the amount of data written each time, the sum of the amount of data written by all write operations, etc. Among them, accessing data partitions can generate a large amount of access characteristic data. In this embodiment, multiple access characteristic data suitable for partition identification can be selected from them, and then the selected multiple access characteristic data suitable for partition identification are collected. Among them, the access characteristic data suitable for partition identification are some characteristic data that can reflect the consumption or occupation of the resources of the storage node when accessing the data partition, mainly including access characteristic data related to the number of reads and writes and access characteristic data related to the amount of read and write data. According to the different product forms of the distributed storage system, these access characteristic data will have different names or definitions, which will be illustrated in the subsequent embodiments.

[0054] In this embodiment, multiple access feature data may change with the change of time. In order to ensure the timeliness of multiple access feature data of multiple data partitions, multiple access feature data generated by multiple data partitions in the distributed system can be collected. In this embodiment, the implementation method of collecting multiple access feature data generated by multiple data partitions is not limited. For example, a plug-in, API or SDK with data collection function can be deployed on each storage node, and multiple access feature data generated by the corresponding data partition on the storage node can be collected through the plug-in, API or SDK. In this embodiment, multiple access feature data generated by multiple data partitions in the distributed storage system can be periodically collected, and the length of the collection period is not limited. The collection period is generally in seconds or finer granularity, for example, it can be 1s or 0.5s, etc., which can be flexibly determined according to application requirements. Of course, multiple access feature data generated by multiple data partitions in the distributed storage system can also be continuously collected. Further, no matter which collection method is adopted, the collected access feature data can also be stored.

[0055] In this embodiment, multiple data partitions can support multiple access operation types, and the multiple access operation types include at least two types: read operation and write operation; further optionally, the read operation and the write operation can be subdivided according to application requirements to obtain more fine-grained access operation types. There is no limitation on the way to subdivide the read operation and the write operation. In an optional embodiment, the read operation can be further divided according to the attribute information of the read operation, and the attribute information of the read operation includes but is not limited to the object of reading, the time of reading, and the way of reading. Correspondingly, the write operation can also be further divided according to the attribute information of the write operation, and the attribute information of the write operation includes but is not limited to the object of writing, the time of writing, and the way of writing. Among them, according to the attribute information of the read operation, the read operation can be further divided into read operation 1, read operation 2, read operation 3, or more types of read operations; similarly, according to the attribute information of the write operation, the write operation can be further divided into write operation 1, write operation 2, write operation 3, or more types of write operations.

[0056] For example, write operations can be further subdivided into batch write operations and non-batch write operations according to the write operation mode. Batch write operations refer to the situation of writing multiple data to a data partition at the same time; non-batch write operations refer to the situation of writing one data to a data partition at a time. Correspondingly, read operations can be further subdivided into cell-granularity read operations and row-granularity read operations according to the object of the read operation. Cell-granularity read operations refer to the situation of reading data in one or more cells; row-granularity read operations refer to the situation of reading one or more rows of data.

[0057] Wherein, when performing access operations on data partitions, the data partitions will generate multiple access characteristic data under various access operation types according to different access operation types. Wherein, the types of access characteristic data generated by the same data partition under different access operation types may be the same or different, and different data partitions will generate the same type of access characteristic data under the same access operation. In an optional embodiment, the access characteristic data generated by performing access operations on data partitions can be divided into two categories, one is read characteristic data, and the other is write characteristic data. Wherein, the read characteristic data is characteristic data generated by performing at least one read operation on the data partition, and the write characteristic data is characteristic data generated by performing at least one write operation on the data partition. Based on this, an implementation method of collecting multiple access characteristic data generated by multiple data partitions in a distributed system includes: collecting multiple read characteristic data and multiple write characteristic data generated by each of the multiple data partitions.

[0058] In an embodiment of the present application, when collecting multiple access feature data generated by each of multiple data partitions in a distributed storage system, access heat data of multiple data partitions under each access operation type can be generated according to the multiple access feature data generated by each of the multiple data partitions and the access operation types supported. Among them, for the same data partition, by analyzing the access heat data of the data partition under each access operation type, the access heat of the data partition can be more accurately identified, which is conducive to more accurately identifying the target data partition whose access heat meets the requirements.

[0059] In an embodiment of the present application, access heat data of multiple data partitions under each access operation type can be periodically generated according to multiple access feature data and supported access operation types generated by each of the multiple data partitions, and the access heat data of the data partition under each access operation type can be analyzed to accurately identify the target data partition whose access heat meets the requirements, which is referred to as periodically executing the identification operation of the target data partition. This embodiment does not limit the period length of the periodic execution of the identification operation of the target data partition. The execution period is generally minute-level or second-level, for example, the identification operation of the target data partition can be executed once every 10 minutes, 10s or 1s. Of course, in addition to periodically executing the identification operation of the target data partition, the user can also issue an instruction to indicate the identification of the target data partition when the target data partition identification is required according to the partition identification requirements, and execute the identification operation of the target data partition according to the instruction; or, a trigger event for triggering the execution of the identification operation of the target data partition can be pre-set, and the identification operation of the target data partition can be executed when the trigger event occurs.

[0060] Optionally, whether the target data partition identification operation is performed periodically, or the target data partition identification operation is performed when the user's instruction is received, or the target data partition identification operation is performed when a preset trigger event occurs, each time the target data partition identification operation is performed, the access feature data within the specified time period can be obtained from the access feature data that has been collected and stored, and the access heat data of multiple data partitions under each access operation type are generated according to the access feature data within the specified time period and the supported access operation types, and the access heat data of the data partition under each access operation type are analyzed to accurately identify the target data partition whose access heat meets the requirements. The embodiment of the present application does not limit the length of the specified time period. Compared with the collection cycle and the execution cycle, the time granularity of the specified time period is much larger, for example, it can be 1 week, 10 days, 15 days, 1 month or longer in the near future.

[0061] Further optionally, considering that the specified time period is relatively long, the access feature data collected during this period will be relatively large. In order to reduce the number of access feature data processed when identifying the target data partition, the access feature data within the specified time period can be downsampled, and the downsampled access feature data can be used to identify the target hotspot partition. Based on this, in an embodiment of the present application, the access feature data used in the process of generating the access heat data of multiple data partitions under each access operation type can be the access feature data of each data partition within the specified time period, or the access feature data after downsampling the access feature data of each data partition within the specified time period. In an optional embodiment, according to the multiple access feature data generated by each of the multiple data partitions and the access operation types supported, the access heat data of the multiple data partitions under each access operation type is generated, including: for any access operation type, from the multiple access feature data generated by each of the multiple data partitions, at least one access feature data generated by each of the multiple data partitions under any access operation type is obtained; further, according to the at least one access feature data generated by each of the multiple data partitions under any access operation type, the access heat data of the multiple data partitions under any access operation type is generated.

[0062] In this embodiment, considering that the connection between each access operation type is relatively weak, each access operation type can be processed separately, and the access heat data of the data partition under each access operation type can be calculated. By reflecting the access heat of the data partition from the perspective of different access operations, the access heat of the data partition can be reflected more comprehensively, objectively and accurately.

[0063] Optionally, for any access operation type, access heat data of multiple data partitions under any access operation type is generated based on at least one access feature data generated by each of the multiple data partitions under any access operation type, including: for any data partition, calculating at least one resource consumption ratio corresponding to any data partition under any access operation type based on at least one access feature data generated by any data partition and other data partitions under any access operation type; generating access heat data of any data partition under any access operation type based on at least one resource consumption ratio corresponding to any data partition under any access operation type.

[0064] In this embodiment, the implementation method of calculating at least one resource consumption ratio corresponding to any data partition under any access operation type according to at least one access characteristic data generated by any data partition and other data partitions under any access operation type is not limited. In an optional embodiment, at least one resource consumption ratio corresponding to any data partition under any access operation type is calculated according to at least one access characteristic data generated by any data partition and other data partitions under any access operation type, including: for any access characteristic data, calculating the mean data of any access characteristic data generated by other data partitions under any access operation type; calculating the ratio of any access characteristic data generated by any data partition under any access operation type to the mean data as a resource consumption ratio corresponding to any data partition under any access operation type. Among them, the number of resource consumption ratios corresponding to each data partition under each access operation type can be determined according to the number of access characteristic data corresponding to the data partition under the access operation type, and one resource consumption ratio can be calculated for one access characteristic data.

[0065] Take a read operation type C1 as an example. A data partition A1 will generate multiple read feature data under the read operation type C1. Correspondingly, other data partitions will also generate the same type of read feature data. Then, for any read feature data B1, the mean value of the feature data B1 generated by other data partitions except data partition A1 can be calculated. The ratio is calculated based on the read feature data B1 generated by data partition A1 and the mean value. The ratio can be used as a resource consumption ratio corresponding to data partition A1 under the read operation type C1. Similarly, for read feature data B2, B3...Bn, the same method can be used to obtain a resource consumption ratio corresponding to data partition A1 under the read operation type. Thus, at least one resource consumption ratio corresponding to data partition A1 under the read operation type C1 is obtained. For other data partitions, a method similar to that of data partition A1 can also be used to calculate at least one resource consumption ratio corresponding to other data partitions under the read operation type C1. Similarly, for other read operation types and various write operation types, a method similar to the read operation type C1 can be used to obtain at least one resource consumption ratio corresponding to each data partition under each access operation type. Of course, the above method of calculating a resource consumption ratio corresponding to data partition A1 based on the average data of the read feature data B1 generated by data partition A1 and the read feature data B1 generated by other data partitions is only an example and is not limited to this. Any method that can calculate a resource consumption ratio corresponding to data partition A1 based on the read feature data B1 generated by data partition A1 and the read feature data B1 generated by other data partitions is applicable to the embodiments of the present application.

[0066] In an optional embodiment, according to at least one resource consumption ratio corresponding to any data partition under any access operation type, the access heat data of any data partition under any access operation type is generated, including: according to at least one resource consumption ratio corresponding to any data partition under any access operation type, a resource ratio that meets the requirements is selected from at least one resource consumption ratio as the access heat data of the current data partition under the current access operation type. Alternatively, according to at least one resource consumption ratio corresponding to any data partition under any access operation type, the resource ratio with the largest resource ratio is selected from at least one resource consumption ratio as the access heat data of the current data partition under the current access operation type. Alternatively, according to at least one resource consumption ratio corresponding to any data partition under any access operation type, the mean of at least one resource ratio is calculated, and the mean is used as the access heat data of the current data partition under the current access operation type. Alternatively, according to at least one resource consumption ratio corresponding to any data partition under any access operation type, the weighted average of at least one resource ratio is calculated, and the weighted average is used as the access heat data of the current data partition under the current access operation type.

[0067] Continuing with the above example, still taking a read operation type C1 and data partition A1 as an example, assuming that when a read operation corresponding to the read operation type C1 is performed on the data partition AI, 4 corresponding read feature data can be generated, and then the 4 resource consumption ratios corresponding to the data partition A1 can be calculated based on the 4 read feature data. Based on this, the ratio that meets the resource ratio requirements can be selected from the 4 resource consumption ratios as the access heat data of the data partition A1 under the read operation type C1. Optionally, the largest resource consumption ratio can be selected as the access heat data of the data partition A1 under the read operation type C1; or, the average of the 4 resource consumption ratios can be used as the access heat data of the data partition A1 under the read operation type C1; or, the weighted sum of the 4 resource consumption ratios can be used as the access heat data of the data partition A1 under the read operation type C1.

[0068] Furthermore, after obtaining the access heat data of any data partition under any access operation, the target data partition whose access heat meets the requirements can be identified based on the access heat data of multiple data partitions under each access operation type. Optionally, global access heat data of multiple data partitions can be generated based on the access heat data of multiple data partitions under each access operation type; and the target data partition whose global access heat meets the requirements can be identified based on the global access heat data of multiple data partitions. Alternatively, the target access operation type can be determined based on the access heat data of multiple data partitions under each access operation type, and the target data partition whose access heat meets the requirements can be identified based on the access heat data of multiple data partitions under the target access operation type.

[0069] Further optionally, global access heat data of multiple data partitions are generated based on the access heat data of multiple data partitions under each access operation type, including: for any data partition, according to the weight corresponding to each access operation type, weighted summation of the access heat data of any data partition under each access operation type is performed to obtain the global access heat data of any data partition; wherein the weight corresponding to each access operation type is a preset value, or is dynamically determined based on the size relationship between at least part of the access feature data corresponding to each access operation type.

[0070] Optionally, the weight corresponding to each access operation type is dynamically determined based on the size relationship between at least part of the access feature data corresponding to each access operation type, including: when the number of read operations in multiple access feature data is greater than the number of write operations, the weight corresponding to the read operation type is greater than the weight corresponding to the write operation type; or, when the amount of data involved in the read operation in multiple access feature data is greater than or equal to a set multiple of the amount of data involved in the write operation, the weight corresponding to the read operation type is greater than the weight corresponding to the write operation type.

[0071] Furthermore, after obtaining the global access heat data of multiple data partitions, the target data partition whose global access heat meets the requirements can be identified based on the global access heat data of the multiple data partitions. However, for a storage node, there may be a situation where the storage node does not actually contain a hotspot partition, but the access heat data of this embodiment is determined to have a hotspot partition. In order to avoid such a misjudgment, multiple data partitions can be filtered according to some access feature data; then, based on the global access heat data of the remaining data partitions after multiple considerations, the target data partition whose global access heat meets the requirements can be identified. Specifically, before identifying the target data partition based on the global access heat data of multiple data partitions, filter the data partitions whose multiple access feature data do not meet the preset value, and select the target data partition from the retained data partitions, wherein the preset value can be set according to the demand or an empirical threshold.

[0072] Based on the above analysis, in an optional implementation, according to the global access heat data of multiple data partitions, a target data partition whose global access heat meets the requirements is identified, including: selecting respective reference access operation data from multiple access feature data generated by each of the multiple data partitions; filtering the multiple data partitions according to the respective reference access feature data of the multiple data partitions to obtain candidate data partitions; and identifying the target data partition whose global access heat meets the requirements according to the global access heat data of the candidate data partitions. For example, the reference access feature data may be the number of read operations, the number of write operations, the amount of data for read operations and / or the amount of data for write operations, etc. Based on this, data partitions whose number of read operations is less than the corresponding preset value, the amount of data for read operations is less than the corresponding preset value, the number of write operations is less than the corresponding preset value, and / or the amount of data for write operations is less than the corresponding preset value in the multiple data partitions can be filtered out. Among them, the corresponding preset values ​​will be different for different types of access feature data.

[0073] Furthermore, after identifying the target data partition, if the target data partition is a hotspot partition, the target data partition can be split or migrated to reduce the data skew problem caused by the hotspot partition and improve the access efficiency and service performance of the entire system. Among them, the method of splitting the target data partition includes but is not limited to automatic splitting or manual splitting, which is not limited to. Splitting the target data partition refers to splitting the target data partition into two or more data partitions to reduce the access volume of each data partition. Migrating the target data partition refers to migrating part of the data in the target data partition to other storage nodes, or migrating the entire target data partition from the current storage node to a storage node with higher performance and resource specifications, so as to improve the access efficiency of the data partition and the service performance that can be improved by the data partition.

[0074] In this embodiment, the target data partition may be one or more. In the case where there are multiple target data partitions, how to arrange the order of splitting or migrating the multiple target data partitions is a key issue. In this regard, in the case where there are multiple target data partitions, the order of splitting or migrating the multiple target data partitions is determined according to the load growth degree of the multiple target data partitions within the set time period; the multiple target data partitions are split or migrated in order. Preferably, the target data partition with a greater load growth degree can be split or migrated preferentially. The load growth degree can be reflected by the growth ratio of the access volume of the data partition within the set time period. For example, the greater the growth ratio of the access volume of the data partition within the set time period, the greater the load growth degree. The set time period can be a relatively short period, such as within the last 1 minute or 1 hour, and there is no limitation on this. Optionally, the set time period is greater than the collection period and can include one or more collection periods.

[0075] At this point, the partition identification method is completed. In this embodiment, multiple data partitions in the distributed storage system support various access operation types, and multiple data partitions will generate multiple access feature data under each access operation type; then, based on the multiple access feature data generated by each of the collected multiple data partitions and the supported access operation types, multiple access heat data of multiple data partitions under each access operation type are generated, such as the larger the access heat data, the higher the access heat of the corresponding data partition; based on this, from the multiple access heat data, the data partition corresponding to the target access heat data that meets the access heat requirements is selected as the target data partition, such as the data partition corresponding to the maximum access heat is used as the target data partition, thereby improving the accuracy of identifying hotspot partitions.

[0076] For ease of understanding, the following uses a table storage system as an example to illustrate the technical solution provided in the embodiment of the present application. Figure 5 As shown in the figure, the process of identifying hotspot partitions for the table storage system includes the following steps:

[0077] Step 1: Monitor the working status of each storage node in the table storage system to identify the storage node that is suspected to have hot data partition problems.

[0078] For example, if the overall access volume of the storage node is monitored to be large, or the access volume of a data partition on the storage node is monitored to be large, or the access volume of the storage node or a data partition on the storage node increases rapidly in a short period of time, it can be determined that the storage node is suspected of having a hot data partition problem. The number of storage nodes suspected of having a hot data partition problem can be one or more.

[0079] Step 2: Based on user instructions or preset judgment conditions, determine whether it is necessary to process the storage node suspected of having hot data partition problems; if the judgment result is no, the execution process ends; if the judgment result is yes, continue to execute steps 3-9.

[0080] In an optional embodiment, when a storage node suspected of having a hot data partition problem is determined, it can be determined based on a preset judgment condition whether the storage node suspected of having a hot data partition problem needs to be processed. In the embodiment of the present application, the preset judgment condition is not limited and can be flexibly set according to application requirements. For example, the maximum number of storage nodes allowed to be processed simultaneously can be preset. If the number of storage nodes currently being processed does not reach the maximum number, the storage node suspected of having a hot data partition problem can be processed; otherwise, no processing is performed. For another example, a working status threshold that needs to be processed for the storage node can be preset, such as an access volume threshold. If the access volume of the storage node suspected of having a hot data partition problem is greater than the access volume threshold, it means that the storage node has a high risk of having a hot data partition problem, and the storage node can be processed; otherwise, no processing is performed.

[0081] In another optional embodiment, when a storage node suspected of having a hot data partition problem is determined, confirmation information can be output to the user, such as sending an in-application message to the user, or sending an email to the user, or sending a short message to the user, so that the user can confirm whether to process the storage node suspected of having a hot data partition problem; in response to the user's confirmation message that the storage node suspected of having a hot data partition problem is to be processed, the storage node is processed; otherwise, no processing is performed.

[0082] Step 3: Obtain access characteristic data of each data partition in the storage node suspected of having a hot data partition problem, and proceed to step 4.

[0083] In this embodiment, the hotspot data partitions are identified at the storage node granularity. Therefore, when it is determined that a storage node suspected of having a hotspot data partition problem is to be processed, multiple access feature data of each data partition stored on the storage node can be periodically collected. The multiple access feature data here refers to some access feature data that are selected from a large number of access feature data and are suitable for hotspot partition identification.

[0084] In this embodiment, taking table storage as an example, we can focus on some access feature data related to read operations and some access feature data related to write operations. The read operation here includes various subdivided read operations, and correspondingly, the read operation includes various subdivided write operations. Among them, some access characteristic data related to the read operation (referred to as read characteristic data for short) include but are not limited to the following examples: the total number of accesses to the block cache by the read operation during the acquisition cycle, which can be expressed as total_block_cache_io_cnt; the data size of the invalid data column read by the read operation during the acquisition cycle, which can be expressed as invalid_cell_size; the number of invalid data columns read by the read operation during the acquisition cycle, which can be expressed as invalid_cell_cnt; the number of files read by the read operation during the acquisition cycle, which can be expressed as read_file_cnt; the total amount of data scanned by the read operation during the acquisition cycle (where the scanned data may not be read, only part of the data will be read), which can be expressed as scan_data_size; the size of the raw data read and returned to the user during the acquisition cycle, which can be expressed as resp_return_raw_data_size; the number of requests with errors during the acquisition cycle, which can be expressed as fail_req_cnt; the total number of read requests received during the acquisition cycle, which can be expressed as total_req_cnt and other characteristic data. Among them, a read request will be converted into one or more read operations. Among them, some access characteristic data related to write operations (referred to as write characteristic data) include but are not limited to the following examples: the total number of rows involved in write operations during the collection period, which can be expressed as total_row_count; the write throughput involved in write operations during the collection period, which can be expressed as write_size; the read throughput involved in write operations during the collection period, which can be expressed as read_size; the number of rows successfully written by write operations during the collection period, which can be expressed as successful_row_count, etc. Among them, some write operations need to read data first and then write data, so read throughput is involved.

[0085] It should be noted that the read characteristic data and write characteristic data in the above examples are partial characteristic data applied to table storage, and their definitions and names cannot be limited to other storage systems. These characteristic data depend on specific circumstances.

[0086] In this step, multiple access feature data generated by each of the multiple data partitions can be periodically collected according to the access feature data focused on above. This embodiment does not limit the length of the collection period, and the collection period is generally in seconds or finer granularity, for example, it can be 1s or 0.5s. Of course, multiple access feature data generated by each of the multiple data partitions can also be continuously collected. Furthermore, no matter which collection method is used, the collected access feature data can also be stored.

[0087] Step 4: Count the resource usage of each of the multiple data partitions according to the access operation type to obtain the access popularity data of the multiple data partitions under each access operation type, and then proceed to step 5.

[0088] In the embodiment of the present application, the hotspot partition detection service mentioned in the above embodiment can be executed periodically, and the hotspot partition detection service counts the resource occupancy of each of the multiple data partitions according to the access operation type. The period here is generally minute or second level, for example, the identification operation of the hotspot data partition can be executed once every 10 minutes, 10 seconds or 1 second.

[0089] In this embodiment, resource occupancy of multiple data partitions may be counted according to access operation type. The resource occupancy of each data partition reflects the access popularity data of the data partition under the corresponding access operation type to a certain extent.

[0090] In this embodiment, the access operation types are divided into at least two categories: read operations and write operations. Further, the read operations and write operations can be subdivided to obtain multiple types of read operations and multiple types of write operations. Taking table storage as an example, the read operation is further divided into get_range (cell read operation) and get_row (row read operation), and correspondingly, the write operation is further divided into batch_modify (batch modification operation) and modify (modification operation), but it is not limited to this. Among them, get_range means locking the cell range in the data table, and then reading the data in the cells within the cell range; get_row means reading a row of data from the data table. batch_modify means batch modifying the attribute values ​​of fields in the data table; modify means modifying the attribute values ​​of fields in the data table.

[0091] For the access operation type get_range, the resource occupancy of each data partition is counted to obtain the access popularity data of each data partition under the access operation type get_range; for the access operation type get_row, the resource occupancy of each data partition is counted to obtain the access popularity data of each data partition under the access operation type get_row; for the access operation type batch_modify, the resource occupancy of each data partition is counted to obtain the access popularity data of each data partition under the access operation type batch_modify; and for the access operation type modify, the resource occupancy of each data partition is counted to obtain the access popularity data of each data partition under the access operation type modify. Among them, for each access operation type, the process of counting the resource occupancy of each data partition to obtain the access popularity data of each data partition under the access operation type is the same or similar, so the process is described in detail below using the access operation type get_range as an example.

[0092] Specifically, for the access operation type get_range, multiple access feature data collected within a specified period can be obtained from multiple access feature data collected for each data partition, and then the read feature data of each data partition can be obtained from the multiple access feature data collected within the specified period; for each data partition, the resource proportion of the data partition under each read feature data is calculated, and then the resource proportion that meets the requirements is selected from the resource proportion of the data partition under multiple read feature data as the resource proportion of the data partition under the access operation type get_range, and the resource proportion is used as the access heat data of the data partition under the access operation type get_range. The embodiment of the present application does not limit the length of the specified period. Compared with the collection cycle and the execution cycle, the time granularity of the specified period is much larger, for example, it can be 1 week, 10 days, 15 days, 1 month or longer in the recent period. Of course, multiple access feature data collected within the specified period are used here, but it is not limited to this. The latest collected multiple access feature data can also be used, or all the collected access feature data can also be used. It can be determined according to the computing power of the method execution subject.

[0093] Furthermore, for each data partition, when calculating the resource share of the data partition under each read feature data, for each read feature data, the ratio of the read feature data to the mean data of the same type of read feature data of other data partitions can be calculated as the resource share of the data partition under the read feature data. For example, taking the read feature data total_block_cache_io_cnt as an example, for a data partition, the mean data of total_block_cache_io_cnt of other data partitions can be calculated, and then the ratio of total_block_cache_io_cnt of the data partition to the mean data of total_block_cache_io_cnt of other data partitions can be calculated as the resource share of the data partition under total_block_cache_io_cnt.

[0094] Similarly, for the access operation type get_range, you can refer to the above method to calculate the resource proportion of each data partition under invalid_cell_size, invalid_cell_cnt, read_file_cnt, scan_data_size, resp_return_raw_data_size, fail_req_cnt and total_req_cnt. Among them, the resource proportion of a data partition under a certain read characteristic data (such as invalid_cell_size) indicates the resources consumed by the certain read characteristic data (such as invalid_cell_size) when the corresponding access operation type (such as get_range) is executed on the data partition.

[0095] Finally, for the access operation type get_range, for each data partition, the resource proportion that meets the requirements can be selected from the resource proportions of the data partition under total_block_cache_io_cnt, invalid_cell_size, invalid_cell_cnt, read_file_cnt, scan_data_size, resp_return_raw_data_size, fail_req_cnt and total_req_cnt as the resource proportion of the data partition under the access operation type get_range. For example, the largest resource proportion can be selected as the resource proportion of the data partition under the access operation type get_range; or, the resource proportion within the set ratio range can be selected as the resource proportion of the data partition under the access operation type get_range; or, the smallest resource proportion can be selected as the resource proportion of the data partition under the access operation type get_range. There is no limitation on this, as long as the same selection method is used for all data partitions. At this point, the resource proportion of each data partition under the access operation type get_range can be obtained.

[0096] By adopting a similar method as above, the resource proportion of each data partition under the access operation type get_row, the resource proportion of each data partition under the access operation type batch_modify, and the resource proportion of each data partition under the access operation type modify can be obtained. Among them, the resource proportion of each data partition under a certain access operation type indicates the resources consumed when executing a certain access operation type (such as get_range, get_row, batch_modify or modify) on the data partition.

[0097] Step 5: Comprehensively process the resource consumption of multiple data partitions between various access operation types to obtain global access popularity data of multiple data partitions.

[0098] In the above steps, the resource consumption of multiple data partitions is counted according to the access operation type. However, when judging whether a data partition is a hot data partition, the resource consumption of the data partition under various access operation types can be comprehensively considered, and the resource consumption of multiple data partitions can be comprehensively sorted to consider the access popularity of multiple data partitions.

[0099] Optionally, for each data partition, the resource proportion of the data partition under various access operation types can be weighted and summed to obtain the overall resource situation consumed when the data partition is accessed, and the overall resource situation reflects the global access heat of the data partition. For example, the resource proportion of each data partition under the access operation type get_range, the resource proportion under the access operation type get_row, the resource proportion under the access operation type batch_modify, and the resource proportion under the access operation type modify can be weighted and summed to obtain the overall resource situation consumed when the data partition is accessed, which is used as the global access heat data of the data partition.

[0100] Optionally, multiple data partitions may be sorted based on the global access popularity data of each data partition, for example, they may be sorted from large to small according to the global access popularity data, or they may be sorted from small to large according to the global access popularity, and there is no limitation on this.

[0101] Step 6: Filter the false hotspot data partitions among the multiple data partitions to obtain candidate data partitions, and execute step 7.

[0102] In this embodiment, the false hotspot data partitions in the multiple data partitions are filtered, including: selecting respective reference access operation data from the multiple access feature data generated by the multiple data partitions; filtering the multiple data partitions according to the reference access feature data of the multiple data partitions to obtain candidate data partitions. For example, access operation data such as total_block_cache_io_cnt, resp_return_raw_data_size, total_row_count and / or write_size can be selected as reference access operation data, and these reference access operation data are compared with corresponding thresholds to filter out data partitions whose reference access operation data is lower than the corresponding threshold.

[0103] Step 7: Determine the hot data partition based on the global access popularity data of the candidate data partition, and proceed to step 8.

[0104] Step 8: Split or migrate hot data partitions.

[0105] In this embodiment, when there are multiple hot data partitions, the order of splitting or migrating the multiple hot data partitions can be determined according to the load growth degree of the multiple hot data partitions within a set period; and then the multiple hot data partitions are split or migrated according to the order. At this point, the process of identifying hot data partitions for the table storage system is completed.

[0106] It should be noted that the above steps 4 to 8 can be executed periodically, or can be executed on demand according to the instructions for hotspot partition identification required by the user, or can be executed on demand according to event triggers, that is, the above-mentioned hotspot data partition identification process is executed when the set trigger event occurs. The trigger event here refers to the event that will trigger the hotspot data partition identification.

[0107] Figure 6 This is a schematic diagram of the structure of the partition identification device provided in the embodiment of the present application. Figure 6 As shown, the device comprises:

[0108] A data collection module 61 is used to collect a plurality of access characteristic data generated by each of a plurality of data partitions in a distributed storage system, where the plurality of access characteristic data are characteristic data generated by performing access operations on the data partitions;

[0109] The data generating module 62 is used to generate access heat data of the multiple data partitions under each access operation type according to the multiple access feature data generated by each of the multiple data partitions and the access operation types supported;

[0110] The partition identification module 63 is used to identify a target data partition whose access heat meets the requirements according to the access heat data of multiple data partitions under various access operation types.

[0111] In an optional embodiment, when the data acquisition module 61 is used to collect multiple access feature data generated by multiple data partitions in a distributed system, it is specifically used to: collect multiple read feature data and multiple write feature data generated by each of the multiple data partitions; wherein the multiple read feature data are feature data generated by performing at least one read operation on the data partition, and the multiple write feature data are feature data generated by performing at least one write operation on the data partition; wherein the access operation type includes at least one read operation and at least one write operation.

[0112] In an optional embodiment, when the data generation module 62 is used to generate access heat data of multiple data partitions under each access operation type based on multiple access feature data generated by each of the multiple data partitions and the supported access operation types, it is specifically used to: for any access operation type, obtain at least one access feature data generated by each of the multiple data partitions under any access operation type from the multiple access feature data generated by each of the multiple data partitions; generate access heat data of the multiple data partitions under any access operation type based on at least one access feature data generated by each of the multiple data partitions under any access operation type.

[0113] Optionally, when the data generation module 62 is used to generate access heat data of multiple data partitions under any access operation type based on at least one access feature data generated by each of the multiple data partitions under any access operation type, it is specifically used to: for any data partition, calculate at least one resource consumption ratio corresponding to any data partition under any access operation type based on at least one access feature data generated by any data partition and other data partitions under any access operation type; generate access heat data of any data partition under any access operation type based on at least one resource consumption ratio corresponding to any data partition under any access operation type.

[0114] Among them, when the data generation module 62 is used to calculate at least one resource consumption ratio corresponding to any data partition under any access operation type based on at least one access characteristic data generated by any data partition and other data partitions under any access operation type, it is specifically used to: calculate the mean data of any access characteristic data generated by other data partitions under any access operation type for any characteristic access data; calculate the ratio of any access characteristic data generated by any data partition under any access operation type to the mean data, as a resource consumption ratio corresponding to any data partition under any access operation type.

[0115] In this embodiment, the partition identification module 63 is used to identify the target data partition whose access heat meets the requirements according to the access heat data of multiple data partitions under each access operation type. Specifically, it is used to: generate global access heat data of multiple data partitions according to the access heat data of multiple data partitions under each access operation type; identify the target data partition whose global access heat meets the requirements according to the global access heat data of multiple data partitions.

[0116] Optionally, when the partition identification module 63 is used to generate global access heat data of multiple data partitions based on the access heat data of multiple data partitions under each access operation type, it is specifically used to: for any data partition, according to the weight corresponding to each access operation type, perform weighted summation of the access heat data of any data partition under each access operation type to obtain the global access heat data of any data partition; wherein the weight corresponding to each access operation type is a preset value, or is dynamically determined based on the size relationship between at least part of the access feature data corresponding to each access operation type.

[0117] Among them, when the number of read operations in multiple access feature data is greater than the number of write operations, the weight corresponding to the read operation type is greater than the weight corresponding to the write operation type; or, when the amount of data involved in the read operation in multiple access feature data is greater than or equal to a set multiple of the amount of data involved in the write operation, the weight corresponding to the read operation type is greater than the weight corresponding to the write operation type.

[0118] Optionally, when the partition identification module 63 is used to identify a target data partition whose global access heat meets the requirements based on the global access heat data of multiple data partitions, it is specifically used to: select respective reference access operation data from multiple access feature data generated by each of the multiple data partitions; filter the multiple data partitions based on the reference access feature data of each of the multiple data partitions to obtain candidate data partitions; and identify the target data partition whose global access heat meets the requirements based on the global access heat data of the candidate data partitions.

[0119] Furthermore, before the data acquisition module 61 collects multiple access feature data generated by each of the multiple data partitions in the distributed storage system, it also includes: a partition determination module, which is used to determine the multiple data partitions that need to be partitioned in the distributed storage system according to the partition identification requirement information; wherein the multiple data partitions are data partitions distributed on the same storage node in the distributed storage system, or, the multiple data partitions are data partitions distributed on multiple specific storage nodes in the distributed storage system, or, the multiple data partitions are data partitions from the same data table in the distributed storage system, or, the multiple data partitions are data partitions in the data table from one or more specific tenants in the distributed storage system.

[0120] Furthermore, it also includes: a sequence determination module, which is used to determine the order of splitting or migrating multiple target data partitions according to the load growth degree of the multiple target data partitions within a set time period when there are multiple target data partitions; and splitting or migrating the multiple target data partitions in order.

[0121] It should be noted that the detailed implementation and beneficial effects of each module in the device of this embodiment have been described in detail in the aforementioned embodiments and will not be elaborated here.

[0122] Figure 7 This is a schematic diagram of the structure of the storage system provided in this embodiment. Figure 7 As shown, the storage system includes at least one storage node 71 , and at least one data partition from at least one data table is stored on one storage node; the storage system also includes: a partition identification node 72 .

[0123] The partition identification node 72 is used to collect multiple access feature data generated by each of the multiple data partitions in the distributed storage system, where the multiple access feature data are feature data generated by access operations performed on the data partitions; based on the multiple access feature data generated by each of the multiple data partitions and the supported access operation types, generate access heat data for the multiple data partitions under each access operation type; based on the access heat data for the multiple data partitions under each access operation type, identify the target data partition whose access heat meets the requirements.

[0124] In an optional embodiment, when the partition identification node 72 is used to collect multiple access feature data generated by multiple data partitions in a distributed system, it is specifically used to: collect multiple read feature data and multiple write feature data generated by each of the multiple data partitions; wherein the multiple read feature data are feature data generated by performing at least one read operation on the data partition, and the multiple write feature data are feature data generated by performing at least one write operation on the data partition; wherein the access operation type includes at least one read operation and at least one write operation.

[0125] In an optional embodiment, when the partition identification node 72 is used to generate access heat data of multiple data partitions under each access operation type based on multiple access feature data generated by each of the multiple data partitions and the supported access operation types, it is specifically used to: for any access operation type, obtain at least one access feature data generated by each of the multiple data partitions under any access operation type from the multiple access feature data generated by each of the multiple data partitions; generate access heat data of the multiple data partitions under any access operation type based on at least one access feature data generated by each of the multiple data partitions under any access operation type.

[0126] Optionally, when the partition identification node 72 is used to generate access heat data of multiple data partitions under any access operation type based on at least one access feature data generated by each of the multiple data partitions under any access operation type, it is specifically used to: for any data partition, calculate at least one resource consumption ratio corresponding to any data partition under any access operation type based on at least one access feature data generated by any data partition and other data partitions under any access operation type; generate access heat data of any data partition under any access operation type based on at least one resource consumption ratio corresponding to any data partition under any access operation type.

[0127] Among them, when the partition identification node 72 is used to calculate at least one resource consumption ratio corresponding to any data partition under any access operation type based on at least one access characteristic data generated by any data partition and other data partitions under any access operation type, it is specifically used to: calculate the mean data of any access characteristic data generated by other data partitions under any access operation type for any characteristic access data; calculate the ratio of any access characteristic data generated by any data partition under any access operation type to the mean data, as a resource consumption ratio corresponding to any data partition under any access operation type.

[0128] In this embodiment, when the partition identification node 72 is used to identify a target data partition whose access heat meets the requirements based on the access heat data of multiple data partitions under each access operation type, it is specifically used to: generate global access heat data of multiple data partitions based on the access heat data of multiple data partitions under each access operation type; and identify a target data partition whose global access heat meets the requirements based on the global access heat data of multiple data partitions.

[0129] Optionally, when the partition identification node 72 is used to generate global access heat data of multiple data partitions based on the access heat data of multiple data partitions under each access operation type, it is specifically used to: for any data partition, according to the weight corresponding to each access operation type, perform weighted summation of the access heat data of any data partition under each access operation type to obtain the global access heat data of any data partition; wherein the weight corresponding to each access operation type is a preset value, or is dynamically determined based on the size relationship between at least part of the access feature data corresponding to each access operation type.

[0130] Among them, when the number of read operations in multiple access feature data is greater than the number of write operations, the weight corresponding to the read operation type is greater than the weight corresponding to the write operation type; or, when the amount of data involved in the read operation in multiple access feature data is greater than or equal to a set multiple of the amount of data involved in the write operation, the weight corresponding to the read operation type is greater than the weight corresponding to the write operation type.

[0131] Optionally, when the partition identification node 72 is used to identify a target data partition whose global access heat meets the requirements based on the global access heat data of multiple data partitions, it is specifically used to: select respective reference access operation data from multiple access feature data generated by each of the multiple data partitions; filter the multiple data partitions based on the respective reference access feature data of the multiple data partitions to obtain candidate data partitions; and identify the target data partition whose global access heat meets the requirements based on the global access heat data of the candidate data partitions.

[0132] Furthermore, before the partition identification node 72 collects multiple access feature data generated by each of the multiple data partitions in the distributed storage system, it is also used to determine multiple data partitions that need to be partitioned identified in the distributed storage system according to the partition identification requirement information; wherein the multiple data partitions are data partitions distributed on the same storage node in the distributed storage system, or, the multiple data partitions are data partitions distributed on multiple specific storage nodes in the distributed storage system, or, the multiple data partitions are data partitions from the same data table in the distributed storage system, or, the multiple data partitions are data partitions in data tables from one or more specific tenants in the distributed storage system.

[0133] Furthermore, the partition identification node 72 is also used to determine the order of splitting or migrating the multiple target data partitions according to the load growth degree of the multiple target data partitions within a set time period when there are multiple target data partitions; and split or migrate the multiple target data partitions in the order of priority.

[0134] It should be noted that the detailed implementation and beneficial effects of each node and each step in the storage device of this embodiment have been described in detail in the aforementioned embodiments and will not be elaborated here.

[0135] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 8 As shown, the electronic device includes: a memory 80a and a processor 80b; the memory 80a is used to store a computer program; the processor 80b is coupled to the memory 80a and is used to execute the computer program to implement the following steps:

[0136] Collect multiple access feature data generated by each of the multiple data partitions in the distributed storage system, where the multiple access feature data are feature data generated by access operations on the data partitions; generate access heat data for the multiple data partitions under each access operation type based on the multiple access feature data generated by each of the multiple data partitions and the supported access operation types; identify the target data partition whose access heat meets the requirements based on the access heat data for the multiple data partitions under each access operation type.

[0137] In an optional embodiment, when the processor 80b is used to collect multiple access feature data generated by multiple data partitions in a distributed system, it is specifically used to: collect multiple read feature data and multiple write feature data generated by each of the multiple data partitions; wherein the multiple read feature data are feature data generated by performing at least one read operation on the data partition, and the multiple write feature data are feature data generated by performing at least one write operation on the data partition; wherein the access operation type includes at least one read operation and at least one write operation.

[0138] In an optional embodiment, when the processor 80b is used to generate access heat data of multiple data partitions under each access operation type based on multiple access feature data generated by each of the multiple data partitions and the access operation types supported, it is specifically used to: for any access operation type, obtain at least one access feature data generated by each of the multiple data partitions under any access operation type from the multiple access feature data generated by each of the multiple data partitions; generate access heat data of the multiple data partitions under any access operation type based on at least one access feature data generated by each of the multiple data partitions under any access operation type.

[0139] Optionally, when the processor 80b is used to generate access heat data of multiple data partitions under any access operation type based on at least one access characteristic data generated by each of the multiple data partitions under any access operation type, it is specifically used to: for any data partition, calculate at least one resource consumption ratio corresponding to any data partition under any access operation type based on at least one access characteristic data generated by any data partition and other data partitions under any access operation type; generate access heat data of any data partition under any access operation type based on at least one resource consumption ratio corresponding to any data partition under any access operation type.

[0140] Among them, when the processor 80b is used to calculate at least one resource consumption ratio corresponding to any data partition under any access operation type based on at least one access characteristic data generated by any data partition and other data partitions under any access operation type, it is specifically used to: calculate the mean data of any access characteristic data generated by other data partitions under any access operation type for any characteristic access data; calculate the ratio of any access characteristic data generated by any data partition under any access operation type to the mean data, as a resource consumption ratio corresponding to any data partition under any access operation type.

[0141] In this embodiment, when the processor 80b is used to identify a target data partition whose access heat meets the requirements based on the access heat data of multiple data partitions under each access operation type, it is specifically used to: generate global access heat data of multiple data partitions based on the access heat data of multiple data partitions under each access operation type; and identify a target data partition whose global access heat meets the requirements based on the global access heat data of multiple data partitions.

[0142] Optionally, when the processor 80b is used to generate global access heat data of multiple data partitions based on the access heat data of multiple data partitions under each access operation type, it is specifically used to: for any data partition, according to the weight corresponding to each access operation type, perform weighted summation of the access heat data of any data partition under each access operation type to obtain the global access heat data of any data partition; wherein the weight corresponding to each access operation type is a preset value, or is dynamically determined based on the size relationship between at least part of the access feature data corresponding to each access operation type.

[0143] Among them, when the number of read operations in multiple access feature data is greater than the number of write operations, the weight corresponding to the read operation type is greater than the weight corresponding to the write operation type; or, when the amount of data involved in the read operation in multiple access feature data is greater than or equal to a set multiple of the amount of data involved in the write operation, the weight corresponding to the read operation type is greater than the weight corresponding to the write operation type.

[0144] Optionally, when the processor 80b is used to identify a target data partition whose global access heat meets the requirements based on the global access heat data of multiple data partitions, it is specifically used to: select respective reference access operation data from multiple access feature data generated by each of the multiple data partitions; filter the multiple data partitions based on the respective reference access feature data of the multiple data partitions to obtain candidate data partitions; and identify the target data partition whose global access heat meets the requirements based on the global access heat data of the candidate data partitions.

[0145] Furthermore, before the processor 80b collects multiple access feature data generated by each of the multiple data partitions in the distributed storage system, it is also used to determine multiple data partitions that need to be partitioned in the distributed storage system according to the partition identification requirement information; wherein the multiple data partitions are data partitions distributed on the same storage node in the distributed storage system, or, the multiple data partitions are data partitions distributed on multiple specific storage nodes in the distributed storage system, or, the multiple data partitions are data partitions from the same data table in the distributed storage system, or, the multiple data partitions are data partitions in a data table from one or more specific tenants in the distributed storage system.

[0146] Furthermore, the processor 80b is also used to determine the order of splitting or migrating the multiple target data partitions according to the load growth degree of the multiple target data partitions within a set time period when there are multiple target data partitions; and split or migrate the multiple target data partitions in the order of priority.

[0147] It should be noted that the detailed implementation and beneficial effects of each node and each step in the storage device of this embodiment have been described in detail in the aforementioned embodiments and will not be elaborated here.

[0148] Further, if Figure 8 As shown, the electronic device also includes: a communication component 80c, a display 80d, a power component 80e, an audio component 80f and other components. Figure 8 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 8 Components shown.

[0149] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor is caused to implement the steps in the above method.

[0150] It should be noted that the detailed implementation and beneficial effects of each step in the storage medium of this embodiment have been described in detail in the aforementioned embodiments and will not be elaborated here.

[0151] The above-mentioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0152] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology and other technologies.

[0153] The above-mentioned display includes a screen, and the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundary of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0154] The power supply assembly provides power to various components of the device where the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply assembly is located.

[0155] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (Microphone, MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting an audio signal.

[0156] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-readable storage media (including but not limited to disk storage, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical storage, etc.) containing computer-usable program code.

[0157] The present application is described with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, and the combination of the process and / or box in the flowchart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.

[0158] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0159] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0160] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interface, network interface and memory.

[0161] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0162] Computer readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0163] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0164] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A partition identification method, It is characterized in that include: Collecting a plurality of access characteristic data generated by each of a plurality of data partitions in a distributed storage system, wherein the plurality of access characteristic data are characteristic data generated by performing access operations on the data partitions; Generate access heat data of the multiple data partitions under each access operation type according to the multiple access feature data generated by each of the multiple data partitions and the access operation types supported; According to the access heat data of the multiple data partitions under each access operation type, a target data partition whose access heat meets the requirements is identified.

2. The method according to claim 1, It is characterized in that Collect multiple access feature data generated by multiple data partitions in a distributed system, including: Collecting a plurality of read characteristic data and a plurality of write characteristic data generated by each of the plurality of data partitions; wherein the plurality of read characteristic data are characteristic data generated by performing at least one read operation on the data partitions, and the plurality of write characteristic data are characteristic data generated by performing at least one write operation on the data partitions; The access operation type includes at least one read operation and at least one write operation.

3. The method according to claim 1, It is characterized in that Generating access heat data of the multiple data partitions under each access operation type according to the multiple access feature data generated by each of the multiple data partitions and the access operation types supported, including: For any access operation type, obtaining, from the plurality of access feature data respectively generated by the plurality of data partitions, at least one access feature data respectively generated by the plurality of data partitions under the any access operation type; Access popularity data of the multiple data partitions under any access operation type is generated according to at least one access feature data generated by each of the multiple data partitions under any access operation type.

4. The method according to claim 3, It is characterized in that Generating access heat data of the multiple data partitions under any access operation type according to at least one access feature data generated by each of the multiple data partitions under any access operation type, including: For any data partition, according to at least one access characteristic data generated by the any data partition and other data partitions under any access operation type, calculate at least one resource consumption ratio corresponding to the any data partition under any access operation type; According to at least one resource consumption ratio corresponding to any one of the data partitions under any one of the access operation types, access popularity data of any one of the data partitions under any one of the access operation types is generated.

5. The method according to claim 4, It is characterized in that Calculating at least one resource consumption ratio corresponding to any one data partition under any one access operation type according to at least one access feature data generated by any one data partition and other data partitions under any one access operation type, including: For any one type of access characteristic data, calculate mean data of any one type of access characteristic data generated by other data partitions under any one type of access operation; A ratio of any one of the access characteristic data generated by any one of the data partitions under any one of the access operation types to the mean data is calculated as a resource consumption ratio corresponding to any one of the data partitions under any one of the access operation types.

6. The method according to claim 1, It is characterized in that According to the access heat data of the plurality of data partitions under each access operation type, identifying a target data partition whose access heat meets the requirements, including: Generate global access heat data of the multiple data partitions according to the access heat data of the multiple data partitions under each access operation type; According to the global access heat data of the multiple data partitions, a target data partition whose global access heat meets the requirements is identified.

7. The method according to claim 6, It is characterized in that Generating global access heat data of the multiple data partitions according to the access heat data of the multiple data partitions under each access operation type includes: For any data partition, according to the weights corresponding to the respective access operation types, weighted sum is performed on the access heat data of the any data partition under the respective access operation types, so as to obtain the global access heat data of the any data partition; The weight corresponding to each access operation type is a preset value, or is dynamically determined according to the size relationship between at least part of the access feature data corresponding to each access operation type.

8. The method according to claim 7, It is characterized in that When the number of read operations in the multiple access feature data is greater than the number of write operations, the weight corresponding to the read operation type is greater than the weight corresponding to the write operation type; or, when the amount of data involved in the read operation in the multiple access feature data is greater than or equal to a set multiple of the amount of data involved in the write operation, the weight corresponding to the read operation type is greater than the weight corresponding to the write operation type.

9. The method according to claim 6, It is characterized in that According to the global access heat data of the multiple data partitions, identifying a target data partition whose global access heat meets the requirements, including: selecting respective reference access operation data from a plurality of access characteristic data generated by respective ones of the plurality of data partitions; filtering the plurality of data partitions according to the respective reference access characteristic data of the plurality of data partitions to obtain candidate data partitions; According to the global access heat data of the candidate data partitions, a target data partition whose global access heat meets the requirements is identified.

10. The method according to any one of claims 1 to 9, It is characterized in that Before collecting a plurality of access feature data generated by each of a plurality of data partitions in the distributed storage system, the method further includes: Determining, according to the partition identification requirement information, a plurality of data partitions that need to be partition identified in the distributed storage system; Among them, the multiple data partitions are data partitions distributed on the same storage node in the distributed storage system, or, the multiple data partitions are data partitions distributed on multiple specific storage nodes in the distributed storage system, or, the multiple data partitions are data partitions from the same data table in the distributed storage system, or, the multiple data partitions are data partitions in data tables from one or more specific tenants in the distributed storage system.

11. The method according to any one of claims 1 to 9, It is characterized in that Also includes: In the case where there are multiple target data partitions, determining the order of splitting or migrating the multiple target data partitions according to the load growth degree of the multiple target data partitions within a set period of time; The multiple target data partitions are split or migrated in the order.

12. A storage system, It is characterized in that comprising at least one storage node, wherein at least one data partition from at least one data table is stored on one storage node; The storage system further includes: a partition identification node; The partition identification node is used to collect multiple access feature data generated by each of the multiple data partitions in the distributed storage system, where the multiple access feature data are feature data generated by access operations performed on the data partitions; based on the multiple access feature data generated by each of the multiple data partitions and the supported access operation types, generate access heat data of the multiple data partitions under each access operation type; based on the access heat data of the multiple data partitions under each access operation type, identify the target data partition whose access heat meets the requirements.

13. An electronic device, It is characterized in that include: Memory and processor; The memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to implement the steps in the method according to any one of claims 1 to 11.

14. A computer-readable storage medium storing a computer program, It is characterized in that When the computer program is executed by a processor, the processor is caused to implement the steps in the method according to any one of claims 1 to 11.