Data processing method, data processing system, data processing apparatus, computing device, computer readable storage medium, and computer program product

By obtaining operational information in heterogeneous distributed storage systems to cluster storage components and migrate data, the problem of predicting data block access popularity is solved, the data placement strategy is optimized, the efficiency of the storage system is improved, and the cost is reduced.

CN120653184APending Publication Date: 2025-09-16HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202410303995.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In heterogeneous distributed storage systems, it is difficult to predict the access popularity of data blocks, which affects the efficiency of data placement strategies and increases system storage costs and latency.

Method used

By obtaining operational information of storage services, clustering storage components based on the operational information, generating analytical data, migrating data according to preset data processing strategies, and optimizing data placement.

Benefits of technology

Improves data access efficiency, reduces storage costs, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653184A_ABST
    Figure CN120653184A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method, a data processing system, a data processing device, computing equipment, a computer readable storage medium and a computer program product, and the data processing method comprises the steps that operation information of a storage service is acquired, and at least one storage component class corresponding to the storage service is determined, the storage businesses have storage components with corresponding relations; counting access information corresponding to the at least one storage component class based on the operation information, and generating analysis data corresponding to the storage component according to the access information; and performing data migration on the storage data in the storage component based on the analysis data according to a preset data processing strategy to obtain a migrated storage component. According to the method, the purpose of better placing the data according to the data popularity is achieved, the storage service performance is improved, the storage cost is reduced, and meanwhile better storage service can be provided for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a data processing method, a data processing system, and a data processing device. Background Art

[0002] With the continuous development of internet technology, we have entered the era of big data. An increasing number of internet applications are adopting cloud storage for their data storage. Due to the large number of users, high data access volume, and complex network environments, data storage systems face challenges in ensuring the quality of data storage services. In data storage systems, user demand, storage performance, and system resources all influence the quality of data storage services. Therefore, the optimal placement and storage of data is a pressing issue that needs to be addressed. Summary of the Invention

[0003] In view of this, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a data processing apparatus, a data processing system, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0004] According to a first aspect of an embodiment of this specification, there is provided a data processing method, including:

[0005] Acquire operation information of a storage service, and determine at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component;

[0006] collecting access information corresponding to the at least one storage component class based on the operation information, and generating analysis data corresponding to the storage component according to the access information;

[0007] The storage data in the storage component is migrated based on the analysis data according to a preset data processing strategy to obtain a migrated storage component.

[0008] According to a second aspect of the embodiments of this specification, a data processing system is provided, the system comprising a storage service end and an analysis service end, wherein:

[0009] The analysis server is configured to obtain operational information of a storage service for the storage service client, determine at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component; collect access information corresponding to the at least one storage component class based on the operational information, and generate analysis data corresponding to the storage component based on the access information;

[0010] The storage service end is used to migrate the storage data in the storage component based on the analysis data according to a preset data processing strategy to obtain a migrated storage component.

[0011] According to a third aspect of the embodiments of this specification, there is provided a data processing device, including:

[0012] a determination module configured to obtain operation information of a storage service and determine at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component;

[0013] a statistics module configured to collect access information corresponding to the at least one storage component class based on the operation information, and generate analysis data corresponding to the storage component according to the access information;

[0014] The migration module is configured to perform data migration on the storage data in the storage component based on the analysis data according to a preset data processing strategy to obtain a storage component after migration.

[0015] According to a fourth aspect of the embodiments of this specification, there is provided a computing device, including:

[0016] memory and processor;

[0017] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned data processing method are implemented.

[0018] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned data processing method are implemented.

[0019] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program or instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.

[0020] This specification provides a data processing method, including obtaining operational information of a storage business, determining at least one storage component class corresponding to the storage business, wherein the storage business has a corresponding storage component; collecting statistics on access information corresponding to the at least one storage component class based on the operational information, and generating analysis data corresponding to the storage component based on the access information; migrating the storage data in the storage component based on the analysis data in accordance with a preset data processing strategy to obtain a migrated storage component.

[0021] One embodiment of the present specification achieves this by acquiring and analyzing operational information of storage services, clustering storage components of storage services based on the operational information, and determining multiple different storage component classes, facilitating subsequent access information statistics for different types of storage components. After the access information for each storage component class is statistically calculated based on the operational information, analytical data corresponding to the storage component can be generated based on the access information. Since the analytical data is generated based on the operational data combined with the characteristics of the storage service type, the accuracy of the analytical data is improved. This allows for better data placement based on data popularity when migrating storage data based on the analytical data according to a preset data processing strategy, thereby improving storage service performance and reducing storage costs. Furthermore, after the storage data is placed based on data popularity, better storage services can be provided to users, ensuring user data access performance and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a scenario diagram of a data processing method provided by an embodiment of this specification;

[0023] Figure 2 is a flow chart of a data processing method provided by one embodiment of this specification;

[0024] Figure 3 This is a flowchart of a data processing method provided by one embodiment of this specification;

[0025] Figure 4 This is a schematic diagram of the structure of a data processing system provided by one embodiment of this specification;

[0026] Figure 5 This is a schematic diagram of the structure of a data processing device provided by one embodiment of this specification;

[0027] Figure 6 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0028] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0029] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0030] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0031] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0032] First, the terms involved in one or more embodiments of this specification are explained.

[0033] Elastic Block Storage: Elastic Block Storage (EBS) is an advanced storage solution for cloud computing environments, providing users with a scalable, highly available, and stable block-level storage service.

[0034] Data blocks: Data blocks can be understood as storage units in elastic block storage systems. At the underlying layer, storage volumes are divided into multiple smaller, fixed-size logical units, known as data blocks. When applications read and write to storage volumes, they actually operate on data blocks.

[0035] Currently, in elastic block storage systems, due to the extremely limited semantic information of data blocks, it is difficult to predict the access popularity of data blocks, which greatly affects the efficiency of data placement strategies. In other words, in heterogeneous distributed storage systems, it is impossible to determine the placement location of data in a data block, which affects the data access performance of the entire system and increases the system storage cost.

[0036] Based on this, this specification provides a data processing method, and this specification also involves a data processing device, a data processing system, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0037] See also Figure 1 , Figure 1 A schematic diagram of a data processing method according to an embodiment of the present specification is shown, wherein in a heterogeneous distributed storage system, multiple different storage services can be deployed on each storage node. Figure 1 Only storage business A is used for illustration. When performing data migration, the operation information corresponding to each storage business can be obtained first, and all storage businesses can be classified based on the operation information to obtain multiple storage component classes. The storage components in the storage component class are data blocks of the same type. Then, the access information of each storage component class can be calculated based on the operation information, and the analysis data of the storage business can be determined based on the access information of each storage component class. The analysis data includes the popularity information corresponding to each storage component, that is, each data block. In this way, it can be determined which data blocks are frequently accessed by users, and data migration can be performed based on the analysis data according to the preset data processing strategy, such as Figure 1 In the example, data blocks a1, a2, and a3 of storage service A deployed on storage node 1 are migrated to storage node 2, so that users can access each data block in storage service A through storage node 2, reducing access latency and improving access efficiency.

[0038] See also Figure 2 , Figure 2 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0039] Step 202: Acquire operation information of a storage service, and determine at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component.

[0040] In practice, cloud service providers offer users an online storage service, known as a cloud disk storage service. This service allows users to store user data, such as documents, images, videos, and audio files, on a remote server cluster via the network. When users store data online in cloud storage, they can set up different service types for different cloud disks, such as one for storing log data and another for storing system files. This creates multiple storage services with different service types. The business data corresponding to each storage service may be divided into several storage components, known as data blocks, for storage. The storage components of the same storage service may be located on the same storage node or on different storage nodes.

[0041] Among them, the storage business is the cloud disk storage business. The number of storage businesses can include multiple, and the business type of each storage business is different. The storage business can be classified according to certain standards. Operation information can be understood as the data collected by the cloud service provider in the process of providing cloud disk storage services, such as business, user and other dimensions. Operation information can include user-dimensional information such as user-defined tag information, user business information, user BPS (Bits Per Second, bit rate per second) timing information, etc., and can also include cloud disk-dimensional information such as cloud disk I / OTrace (input / output tracking) information, cloud disk size, cloud disk creation time and other information. By analyzing based on operation information, it is possible to subsequently implement coarse-grained classification of user dimensions for different storage businesses, and classify storage businesses according to business type, so as to facilitate the subsequent statistics of access characteristics of different types of storage businesses. In one embodiment of this specification, the storage components of a storage service can be clustered based on operational information to determine multiple storage component classes. A storage component can be understood as a data block corresponding to a cloud disk. By clustering the data blocks corresponding to a cloud disk storage service based on operational information, the multiple data blocks under the cloud disk storage service can be divided into multiple data block classes, facilitating subsequent access information statistics for data blocks of the same class. In actual applications, storage services can differentiate storage components based on their own needs, such as classifying storage components on storage nodes far from user addresses. This specification uses the classification of storage components based on operational information as an example.

[0042] During specific implementation, operational information can be obtained from the operational system, which is responsible for recording data related to storage services, such as system logs, user information, cloud disk information, load information, etc. The acquisition module can be used to collect and store information or data from the operational system to facilitate subsequent processing and analysis of the collected operational information. When clustering the storage components corresponding to the storage services based on operational information, it is actually possible to perform fine-grained clustering of the storage services through operational information. Since each storage service corresponds to its own storage component, and the number of storage components is large, in order to achieve the purpose of clustering the storage components corresponding to the storage services and reduce the cost and amount of clustering implementation, it can be implemented from the dimension of the storage services. That is, the storage services are fine-grainedly clustered through operational information, so that the data blocks of the storage services belonging to the same class after clustering are divided into the same data block class.

[0043] Based on this, in order to be able to reasonably place block storage data, the data processing method provided in this specification adopts a method based on operational information analysis to count the hot access information of the data blocks. In order to facilitate statistics, it is necessary to classify the services for different storage services, and cluster the data blocks for different types of storage services, and then perform heat statistics on the data blocks of the same type, so as to improve the accuracy of statistics. Therefore, in a specific embodiment of this specification, by obtaining the operational information of all storage services in the cloud disk storage system, clustering the storage components of the storage services according to the operational information, and determining multiple different types of storage component classes, the purpose of clustering data blocks based on operational analysis is achieved, which is convenient for subsequent heat statistics on data blocks of the same type.

[0044] Furthermore, since the operational information collected from the operational system may contain some noise data or information that is of little significance for training guidance, preprocessing and data cleaning operations can be performed on the collected operational information to reduce the amount of subsequent calculations. Specifically, obtaining the operational information of the storage business includes: collecting initial operational information of the storage business; screening the initial operational information according to a preset data cleaning strategy to determine target operational information; and eliminating the target operational information from the initial operational information to obtain operational information.

[0045] In actual applications, the operation system will record all information during the operation of the cloud disk storage service. After the operation information is collected from the operation system, there may be information in the operation information that needs to be filtered. Therefore, the collected information data needs to be preprocessed and cleaned to remove outliers and noise, and necessary feature selection and conversion are performed to ensure the data quality and availability of the operation information. During the specific operation process, when the operation system records the operation information, there may be outliers and noise in the collected operation information due to problems with the data printing itself or abnormalities in the data collection end or network. Therefore, this part of the data needs to be eliminated. In the collected operation information, there is also some data that is not very meaningful for subsequent heat statistics, such as the feature information of the user's location. This part of the data will be filtered. Based on this, by selecting representative and important feature information from the original data, the amount of calculation for subsequent heat statistics is reduced.

[0046] Among them, the initial operation information can be understood as the original data recorded in the operation information. When the acquisition module collects information from the operation information, the information obtained is the initial operation information. There may be outliers and noise in the initial operation information, as well as some unnecessary feature information, and these data need to be eliminated and filtered. Therefore, the initial operation information can be screened according to the preset data cleaning strategy, that is, the data can be cleaned. The preset data cleaning strategy is a strategy for determining the data that needs to be eliminated from the initial operation information. The preset data cleaning strategy may include judgment conditions for noise and outlier data, and unnecessary feature identifiers. Therefore, the initial operation information can be screened according to the preset data cleaning strategy to determine the target operation information. The target operation information is the data that needs to be eliminated, such as noise data and data with little training guidance significance. After removing these target operation information from the initial operation information, the preprocessed operation information can be obtained.

[0047] In a specific embodiment of the present specification, the initial operation information of each storage business is collected from the operation information through the collection module. The initial operation information may include user identification information, user business information, cloud disk load characteristics and other information. The initial operation information is preprocessed, including eliminating outliers and noise information in the initial operation information, and filtering out non-essential feature information in the initial operation information, such as user geographic location information.

[0048] Based on this, by cleaning and preprocessing the operational information, representative and important features can be screened out, which can reduce the amount of statistical calculations and improve the performance of the statistical model when performing subsequent heat statistics.

[0049] Furthermore, since each cloud disk business corresponds to multiple data blocks, the cloud disk business can be coarsely classified first, and then the data blocks of each type of cloud disk business can be fine-grainedly clustered to facilitate subsequent statistical access popularity. Specifically, at least one storage component class corresponding to the storage business is determined, including: classifying the storage business according to the business information in the operation information to obtain at least one storage business class; determining the storage component set corresponding to the at least one storage business class, and clustering each storage component set according to the load information in the operation information to obtain at least one storage component class.

[0050] In actual applications, the same cloud disk may have multiple data blocks. Data blocks of the same type of cloud disk may show similar patterns in the distribution of addresses and access popularity. Therefore, we can first perform coarse-grained classification on all cloud disk services, and then cluster the data blocks of the same type of cloud disk services. This helps to analyze the access popularity of data blocks and improve the efficiency of subsequent data placement.

[0051] Among them, the business information in the operational information can be understood as the information of the cloud disk business in the user dimension. The business information can include user level, application type, cloud disk size, cloud disk creation time, etc. By classifying the cloud disk storage business based on business information, it is possible to complete the coarse-grained classification of the cloud disk from the user dimension. Since the classification is based on the business information related to the user, the classification standard is related to the business information in the user dimension. By classifying the storage business in the user dimension and associating the classification of the storage business with the user business usage, more accurate analysis results can be obtained when analyzing the same type of storage business in the future.

[0052] During specific implementation, after classifying the storage services according to the business information, at least one storage service class is obtained. The storage service class is a class composed of storage services of the same type. After classifying the storage services, multiple different types of storage service classes can be obtained. Subsequently, it is necessary to cluster the data blocks for different types of storage service classes, that is, to further cluster the storage service classes of each type, so as to cluster the data blocks of storage services with the same characteristics together. Therefore, it is necessary to determine the storage component set corresponding to each storage service class. The storage component set can be understood as a set composed of data blocks corresponding to each cloud disk storage service in the cloud disk storage service class. Since each cloud disk service will have multiple data blocks, and the data blocks of the same cloud disk service class can form a storage component set, the data blocks in the storage component set show similar patterns in the distribution of addresses and access heat, which is convenient for subsequent heat statistics. After determining the storage component set, each storage component set can be clustered according to the load information in the operation information. The load information in the operation information can be understood as the load characteristic information of the data block. The load characteristics may include BPS timing information, read-write ratio, I / OPS (I / O Operations Per Second, input / output processing rate) timing information, etc. The storage components are clustered according to the load characteristic information, and data blocks with the same load characteristics are aggregated together to facilitate the subsequent statistics of the heat information of each type of cloud disk data block. The storage component class is each type of cloud disk data block class obtained after clustering. The storage components, i.e., data blocks, in the same storage component class have the same load characteristics, such as similar processing rates, similar read-write ratios, or similar response times.

[0053] In a specific embodiment of the present specification, cloud disk storage services are classified according to business information in the operation information, such as user level and application type, to obtain multiple different types of storage business classes, including storage business classes A, B, and C. Then, the storage component sets A1, B1, and C1 corresponding to each type of storage business class are determined, and each storage component set is clustered based on the load information in the operation information to obtain multiple storage component classes in each storage component set. For example, the storage component classes in the storage component set A1 include a1 and a2, and the storage component classes in the storage component set B1 include b1, b2, and b3.

[0054] Based on this, by coarse-grained classification of cloud disk services from the user dimension and fine-grained clustering of cloud disk services in the cloud disk business category from the load dimension, data blocks under the same type of cloud disk services can be clustered together, which helps to analyze the access popularity of data blocks and improve statistical accuracy.

[0055] Furthermore, in order to be able to classify storage services, it is necessary to determine classification standards based on business information. Specifically, the storage services are classified according to the business information in the operation information to obtain at least one storage business class, including: determining at least one business classification identifier based on the business information in the operation information; classifying the storage services according to the at least one business classification identifier to obtain at least one storage business class.

[0056] Among them, the business classification identifier can be understood as a classification standard for classifying storage services. The business classification identifier can be determined based on business information. For example, if the business information includes user level and cloud disk size, the determined business classification identifier can be user level and cloud disk size. Subsequent classification can be carried out according to these two business classification identifiers. For example, user level greater than level 10 and cloud disk size greater than 100G are classified as one category.

[0057] In practical applications, after the storage services are classified, a classification label can be determined for each storage service class obtained by classification, such as a certain type of cloud disk storage service is a log data disk, and a certain type of cloud disk storage service is a system disk.

[0058] In a specific embodiment of this specification, a service classification identifier is determined based on service information in the operation information. The service classification identifiers are A and B. Multiple storage services are classified according to A and B to obtain multiple different types of storage service classes.

[0059] Based on this, by determining the business classification identifier according to the business information in the operation information, the storage business can be classified according to the business classification identifier in the future, and the coarse-grained classification of the cloud disk can be completed from the user dimension.

[0060] Furthermore, since each storage business class includes multiple storage businesses, and each storage business corresponds to multiple storage components, when determining the storage component set corresponding to each storage business class, it is necessary to first determine the storage components of each storage business. Specifically determining the storage component set corresponding to the at least one storage business class includes: determining the storage components corresponding to each storage business in the at least one storage business class; combining the storage components corresponding to each storage business class, and obtaining the storage component set corresponding to the at least one storage business class according to the combination result.

[0061] Among them, the storage components corresponding to each storage business can be understood as the storage components of each storage business in the same type of storage business class. For example, storage business class A includes storage business a and storage business b, storage business a corresponds to storage component a1 and storage component a2, storage business b corresponds to storage component b1 and storage component b2, then the storage component set corresponding to the storage business class A includes storage components a1, a2, b1 and b2.

[0062] In actual applications, in order to cluster the storage components of a storage service class, it is first necessary to determine the storage component set corresponding to the storage service class. To determine the storage component set, it is necessary to first determine the storage components of each storage service in the storage service class. After determining the storage components of each storage service, all storage components can be combined to obtain the storage component set of the storage service class. In specific implementation, after obtaining the storage component set of the storage service class, the storage services included in the storage service class can be clustered and divided into multiple storage service subclasses. The storage components in the storage component set are then divided according to each storage service subclass, thereby obtaining the storage component class corresponding to each storage service subclass. According to the above method, the purpose of clustering each storage component set according to the load information in the operation information is achieved.

[0063] Based on this, by determining the storage component set corresponding to each storage business class, storage components of the same type of storage business class can be clustered to cluster data blocks with similar access patterns, facilitating subsequent heat statistics.

[0064] Step 204: Count access information corresponding to the at least one storage component class based on the operation information, and generate analysis data corresponding to the storage component according to the access information.

[0065] In practical applications, after identifying each storage component class, we can calculate access information for each storage component class based on operational information. This access information can be understood as the access frequency and probability of each storage component class. By calculating the access frequency and probability of each cloud disk data block within a certain time period, we can use this as the primary indicator for measuring data block popularity. Once the access information for each storage component class is determined, we can generate analytical data for all storage components. This analytical data can be understood as the data obtained from performing a heat analysis on the data blocks.

[0066] Among them, since the storage component class includes multiple different data blocks, and since the storage component class is obtained based on load feature clustering, the heat access of the storage components in a storage component class is similar. Therefore, when counting access information, a storage component class can be used as an object for statistics. After the access information corresponding to each storage component class is determined through statistics, that is, after the access information of each storage component is obtained, the analysis data of the storage component can be generated based on the access information. In specific implementation, the analysis data can be an OAH (Operation-Analysis-Based Heatmap) heat table. In the data structure of the heat table, each table entry is represented by a tuple (BlockIndex, HeatDegree), where BlockIndex represents the number of the data block and HeatDegree represents the access heat of the data block. Since the heat table is obtained based on operational information analysis, it can better reflect the access status of each data block, which facilitates subsequent data migration and placement based on the heat table. In other embodiments, the analysis data can also be other types of data files that store heat information, such as document type data files, graph structure type data files, etc.

[0067] Furthermore, in order to accurately count the access information corresponding to each storage component class, it is first necessary to calculate the access frequency and probability corresponding to each storage component class. Specifically, the access information corresponding to the at least one storage component class is counted based on the operation information, including: calculating the access frequency information and access probability information corresponding to the at least one storage component class based on the operation information; and using the access frequency information and the access probability information as the access information corresponding to the at least one storage component class.

[0068] The access frequency information represents the access frequency of storage components in the storage component class within a certain time period, and the access probability information represents the access probability of storage components in the storage component class within a certain time period. By calculating the access frequency and access probability information, the access frequency and access probability information can be used as the access information of the storage component class, which can then be used to calculate the popularity information of the storage component class based on the access information.

[0069] In specific implementations, access probability and frequency information can be calculated based on the I / O trace information of each storage component in the storage component class. Specifically, access information can be calculated based on the input and output trace information of each data block within a certain time period. In practical applications, statistical operations can be performed by a pre-trained statistical model to improve statistical efficiency. Specifically, each data block and its corresponding trace information are input into the statistical model, and the statistical model outputs access information for each data block.

[0070] Based on this, by calculating the access frequency and access probability corresponding to each storage component class, the access information of each storage component class is determined, and the analysis data of the storage component class can be calculated subsequently based on the access information.

[0071] Furthermore, in order to generate analytical data containing the heat information of each data block, it is necessary to first calculate the heat value of each data block. Specifically, the analytical data corresponding to the storage component is generated based on the access information, including: calculating the heat parameter corresponding to the storage component based on the access information corresponding to the at least one storage component class; obtaining the identification parameter of the storage component, and associating the identification parameter with the heat parameter, and generating the analytical data corresponding to the storage component based on the association result.

[0072] The heat parameter can be understood as the heat value parameter of the storage component. Based on the access information of each storage component class, the heat parameter of each storage component contained in each storage component class can be calculated, thereby determining the heat parameters of all storage components. The identification parameter of the storage component can be understood as the component identification parameter of each component. By associating the component identification parameter of each component with its heat parameter, a tuple corresponding to each component is generated. Based on the tuple parameters of each component, analytical data can be generated.

[0073] In practical applications, a pre-trained model can be used to calculate the heat value based on the access probability and frequency of storage components. By inputting the access information of each storage component into the pre-trained model, the heat value of each storage component output by the pre-trained model can be obtained. After obtaining the heat value of each storage component, an OAH heat table can be generated to facilitate subsequent data migration and placement based on the OAH heat table.

[0074] Based on this, by calculating the heat parameters of each storage component according to the access information of the storage component class, the heat parameters of each storage component are associated with its corresponding component identifier, thereby generating analysis data containing the heat values ​​of all storage components, so as to facilitate the subsequent execution of data placement strategies based on the analysis data.

[0075] Step 206: Migrate the storage data in the storage component based on the analysis data according to a preset data processing strategy to obtain a migrated storage component.

[0076] Among them, the preset data processing strategy can be understood as a preset strategy for migrating data. By combining the analysis data according to the preset processing strategy, the storage data in the current storage component can be migrated, that is, the migration position of each storage component is determined, and the storage data of the storage component is migrated to obtain the migrated storage component.

[0077] In actual applications, the purpose of pre-setting data processing strategies for data migration based on analytical data is to place data blocks in appropriate locations to improve the performance and user experience of cloud services while optimizing storage costs. In typical scenarios, analysis can be performed based on two dimensions: server location and configuration.

[0078] Furthermore, the preset data processing strategy is a location processing strategy; wherein, data migration is performed on the storage data in the storage component based on the analysis data in accordance with the preset data processing strategy to obtain the migrated storage component, including: determining the target storage component in the storage component based on the analysis data in accordance with the location processing strategy; determining the storage node to be received based on the target storage component, and migrating the target storage data in the target storage component to the storage node to be received to obtain the migrated target storage component.

[0079] The location processing strategy can be understood as a server location partitioning strategy. Specifically, the location processing strategy places data blocks that need to be migrated in a storage area server that is closer to the user. The primary purpose of the location processing strategy is to reduce network transmission time and latency, thereby improving the user experience. The target storage component selected from all storage components based on the analyzed data according to the location processing strategy is the component that requires data migration. The target storage component can be a popular storage component. By migrating the target storage component to the receiving storage node, network transmission time and latency are reduced.

[0080] In practical applications, a target storage component can be used to determine a receiving storage node. The target storage component represents the data block to be migrated. The user who created or used the target storage component can be determined based on the target storage component. A storage node closest to the user is then selected as the receiving storage node. The target storage data of the target storage component is then transferred to the receiving storage node, completing the data migration process. In specific implementations, if the target storage component has multiple users, a suitable storage node can be selected as the receiving node based on the locations of the multiple users. Alternatively, the number or frequency of use of the data block by each user can be determined, and based on this number or frequency, the first user can be selected for data migration.

[0081] In a specific embodiment of the present specification, a target data block is selected from all data blocks based on analysis data according to a location processing strategy, the target data block is data block A, the user using data block A is determined to be user A, a storage node to be received is selected based on the IP location of user A, and the data content of data block A is migrated to the storage node to be received.

[0082] Based on this, by migrating the storage data of the storage components according to the location processing strategy, the purpose of reducing network transmission time and delay and improving user experience is achieved.

[0083] Furthermore, the preset data processing strategy is a configuration processing strategy; wherein, data migration is performed on the storage data in the storage component based on the analysis data according to the preset data processing strategy to obtain the migrated storage component, including: determining a first storage component and a second storage component in the storage component based on the analysis data according to the location processing strategy, wherein the access information corresponding to the first storage component and the second storage component is different; determining a first preset storage node and a second preset storage node, wherein the storage configurations corresponding to the first preset storage node and the second preset storage node are different; migrating the first storage data in the first storage component to the first preset storage node to obtain the migrated first storage component, and migrating the second storage data in the second storage component to the second preset storage node to obtain the migrated second storage component.

[0084] In practical applications, in addition to partitioning by server location, you can also partition by hardware configuration. Specifically, if the default data processing strategy is a configuration processing strategy, which can be understood as a strategy for partitioning based on hardware configuration, the configuration processing strategy will classify data blocks into two categories: cold data and hot data, and then place them in storage areas with different performance. Since cold data is typically rarely accessed, it can be placed in low-cost storage areas, while hot data is frequently accessed and should be placed in high-cost but faster-transferring storage areas. The purpose of the configuration processing strategy is to reduce costs and improve efficiency.

[0085] In the case where the preset data processing strategy is a location processing strategy, based on the analysis data according to the location processing strategy, the first storage component and the second storage component in the storage component may have different access information. That is, the first storage component may be a data block of the hot data type, and the second storage component may be a data block of the cold data type; or the first storage component may be a data block of the cold data type, and the second storage component may be a data block of the hot data type. In this specification, the data block of the first storage component being a data block of the hot data type and the second storage component being a data block of the cold data type is used as an example for description.

[0086] In actual applications, according to the configuration processing strategy, cold data needs to be placed in a low-cost storage area, and hot data needs to be placed in a high-cost storage area. Therefore, when the preset data processing strategy is a location processing strategy, two storage areas can be first determined, namely a first preset storage node and a second preset storage node. The storage configurations of these two storage nodes are different, namely, one is a high-cost storage area and the other is a low-cost storage area. After selecting the first storage node and the second storage node from the storage nodes, the first storage node can be migrated to the first preset storage node and the second storage node can be migrated to the second preset storage node, respectively, thereby migrating the data blocks of the hot data class to the high-cost storage area and the data blocks of the cold data class to the low-cost storage area.

[0087] Based on this, by migrating the storage data of the storage components according to the configuration processing strategy, the storage cost can be reduced and the storage transmission efficiency can be improved.

[0088] In specific implementation, in addition to using location processing strategies and configuration processing strategies for data migration respectively, the two dimensions of location and configuration can also be integrated to achieve multi-dimensional and multi-level optimization of the overall block storage service performance and cost. The specific comprehensive processing strategy can be to select the appropriate storage node for data migration based on the weights of the two.

[0089] Furthermore, since the cloud storage service is always in operation, the operational information is also changing during the continuous operation process, which may require data migration every once in a while. The specific method also includes: generating data migration instructions at preset time intervals; and obtaining the operational information of the storage business in response to the data migration instructions.

[0090] Among them, the preset time interval can be understood as the time interval for data migration set in advance, such as data migration once every day. The data migration instruction can be understood as the instruction to trigger data migration. The data migration instruction can be received by the analysis system. The collection module in the analysis system will collect operation information over the past period of time from the operation system, and analyze it to generate analysis data. The execution module will then execute the corresponding data processing strategy according to the analysis data, thereby completing the data migration processing.

[0091] Based on this, by regularly migrating data, we can ensure that the cloud storage service can continuously optimize the performance of the storage service and reduce storage costs without stopping operations.

[0092] Furthermore, when storing business data of a storage service, the location where the data needs to be stored can also be determined based on the analyzed data, thereby enabling reasonable data placement during the data storage phase, reducing subsequent data migration operations and lowering storage costs. Specifically, the method further includes: determining, in response to a data storage instruction, the data to be stored and the target storage service corresponding to the data to be stored, wherein the target storage service has the same business information as the storage service; selecting a target storage component from the storage components corresponding to the target storage service based on the analyzed data; and storing the data to be stored in the target storage component.

[0093] In actual applications, if a storage service has other storage services with similar or identical services, data placement can be based on the analysis data of the other storage services. For example, if a user already uses a system cloud disk and now uses the cloud disk service to create a new system cloud disk, the business data of the new system cloud disk can be stored based on the popularity analysis of the previous system cloud disk.

[0094] Among them, the data storage instruction can be understood as an instruction for data storage. After the system receives the data storage instruction, it can determine the data to be stored that needs to be stored this time, as well as the target storage business corresponding to the data to be stored. The business information of the target storage business is the same as the business information of the storage business, which can be understood as the target storage business being similar or identical to the storage business, such as both are used to store system disk data or both are used to store logs. When the business information of two storage businesses is the same, it means that the heat analysis data of the two storage businesses in the future will be similar, and the analysis data of the storage business can be used as the analysis data of the target storage business for data storage. This avoids the problem of the target storage business not having historical heat analysis data, and avoids the subsequent re-data migration of the stored data after obtaining the historical heat analysis data, thereby reducing the pressure on the storage system and reducing storage costs.

[0095] In specific implementations, when placing the business data of a target storage business, it can be determined whether there is a storage business with the same business information as the target storage business. If so, a target storage component can be selected from the storage components corresponding to the target storage business based on the storage business's analysis data. The target storage component is then the storage component suitable for storing the business data. If there is no storage business with the same business information as the target storage business, it means that the target storage business has no reference analysis data. In this case, the data can be stored according to the default storage policy, which can be determined based on actual conditions.

[0096] This specification provides a data processing method, comprising obtaining operational information of a storage service, determining at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component; collecting access information corresponding to the at least one storage component class based on the operational information, and generating analytical data corresponding to the storage component based on the access information; and migrating storage data in the storage component based on the analytical data according to a preset data processing strategy to obtain the migrated storage component. By obtaining and analyzing the operational information of the storage service, clustering the storage components of the storage service based on the operational information, and determining multiple different storage component classes, it is convenient to subsequently collect access information statistics for different types of storage components. After collecting access information for each storage component class based on the operational information, analytical data corresponding to the storage component can be generated based on the access information. Because the analytical data is generated based on the operational data combined with the characteristics of the storage service type, the accuracy of the analytical data is improved. When migrating storage data based on the analytical data according to the preset data processing strategy, data placement can be better performed based on data popularity, improving storage service performance and reducing storage costs. Furthermore, after the storage data is placed based on data popularity, better storage services can be provided to users, ensuring user data access performance and improving user experience.

[0097] The following combined Figure 3 , taking the application of the data processing method provided in this specification in the cloud disk business as an example, the data processing method is further explained. Figure 3 A flowchart of a data processing method provided in one embodiment of this specification is shown, which specifically includes the following steps.

[0098] Step 302: Collect initial operation information of the storage service, perform data preprocessing on the initial operation information, and obtain operation information.

[0099] In one achievable manner, a data migration instruction is generated at a preset time interval, and operation information of the cloud disk service is obtained in response to the data migration instruction.

[0100] Step 304: Classify the storage service according to the service information in the operation information to obtain at least one storage service class.

[0101] In one feasible method, the business classification identifier is determined based on the business information in the operation information. The business classification identifier is the user level and the cloud disk creation time. The cloud disk business with a user level greater than 10 and a creation time earlier than 3 months is divided into one category, and the rest are divided into another category, and finally cloud disk business category A and cloud disk business category B are obtained.

[0102] Step 306: Determine a storage component set corresponding to at least one storage service class, and cluster each storage component set according to the load information in the operation information to obtain at least one storage component class.

[0103] In one achievable method, data blocks corresponding to each cloud disk service in cloud disk service class A are determined and formed into a data block set a corresponding to cloud disk service class A. Data blocks corresponding to each cloud disk service in cloud disk service class B are also formed into a data block set b corresponding to cloud disk service class B. Clustering data block set a is specifically implemented by clustering each cloud disk service in cloud disk service class A based on load information to obtain cloud disk service subclasses "a1, a2, ...an." The data blocks in data block set a are then divided based on the clustering results for each cloud disk service subclass, thereby obtaining data block subsets corresponding to each cloud disk service subclass. The data block subsets corresponding to each cloud disk service subclass are then used as data block classes. Accordingly, clustering data block set b in the same manner as described above can obtain the data block classes corresponding to storage service class B after clustering.

[0104] Step 308: Collect access information corresponding to at least one storage component class based on the operation information.

[0105] In one feasible manner, the access frequency and access probability of each data block class are calculated based on the operation information, and the access probability and access frequency are used as the access information.

[0106] Step 310: Calculate a heat parameter corresponding to a storage component based on access information corresponding to at least one storage component class.

[0107] In one implementable manner, the heat value of each data block is calculated based on the access information corresponding to each data block class.

[0108] Step 312: Obtain identification parameters of the storage component, associate the identification parameters with the heat parameters, and generate analysis data corresponding to the storage component according to the association result.

[0109] In one achievable manner, the data block number identifier of each data block class is determined, the data block number identifier of each data block class is associated with its corresponding heat value, and a heat table is generated according to the association result.

[0110] Step 314: Migrate the storage data in the storage component based on the analysis data according to a preset data processing strategy to obtain a migrated storage component.

[0111] In one feasible method, the target data block is determined in the data block based on the analysis data according to the location processing strategy, and the storage node to be received corresponding to the target data block is determined. The storage node to be received is the storage node closer to the user, and the storage data corresponding to the target data block is migrated to the storage node to be received.

[0112] In one feasible manner, a first data block and a second data block are determined in a data block based on analysis data according to a location processing strategy, the first data block is a hot data block, and the second data block is a cold data block. The first data block is migrated to a first preset storage node, and the second data block is migrated to a second preset storage node. The first preset storage node is a high-cost storage area, and the second preset storage node is a low-cost storage area.

[0113] This specification provides a data processing method that obtains and analyzes operational information of cloud disk services, clusters the storage components of cloud disk services based on the operational information, and determines multiple different data block classes, facilitating subsequent access information statistics for different types of data blocks. After the access information for each data block class is statistically analyzed based on the operational information, analytical data corresponding to the data block can be generated based on the access information. Since the analytical data is generated based on the operational data combined with the characteristics of the cloud disk service type, the accuracy of the analytical data is improved. When data is migrated based on the analytical data according to a preset data processing strategy, data placement can be better performed according to data popularity, improving storage service performance and reducing storage costs. At the same time, after the storage data is placed according to data popularity, better storage services can be provided to users, ensuring user data access performance and improving user experience.

[0114] See also Figure 4 , Figure 4 FIG. 4 shows a schematic diagram of a data processing system according to an embodiment of the present specification. The system includes a storage service end 402 and an analysis service end 404. Specifically,

[0115] The analysis server 404 is configured to obtain operational information of the storage service of the storage service client, determine at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component; collect access information corresponding to the at least one storage component class based on the operational information, and generate analysis data corresponding to the storage component based on the access information;

[0116] The storage service end 402 is used to migrate the storage data in the storage component based on the analysis data according to a preset data processing strategy to obtain a migrated storage component.

[0117] In actual applications, the analysis server is deployed with a collection module, a data module, and an analysis module, and the storage service can be deployed with an execution module. The collection module is used to collect operational data from the operation system of the storage service. This module can obtain data through an API interface, log files, or other means, and send the data to the data module. The data module is used to store the operational data sent by the collection module for use by the analysis module, and can also store the analysis results of the analysis module for execution by the execution module. The analysis module is used to analyze the operational data to be analyzed in the storage module. This module can use various machine learning algorithms, statistical models, and other analysis techniques to determine user types, load categories, hot and cold data blocks, etc., and then send the analysis results to the storage module. The execution module is used to obtain the analysis results from the storage module and perform data scheduling based on the analysis results.

[0118] This specification provides a data processing system, which includes a storage service end and an analysis service end, wherein the analysis service end is used to obtain operational information of the storage service end for the storage service end, determine at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component; based on the operational information, statistics are collected on access information corresponding to the at least one storage component class, and analysis data corresponding to the storage component is generated based on the access information; the storage service end is used to migrate the storage data in the storage component based on the analysis data according to a preset data processing strategy to obtain the migrated storage component. By obtaining and analyzing the operational information of the storage service, clustering the storage components of the storage service according to the operational information, and determining multiple different storage component classes, it is convenient to perform statistics on access information for different types of storage components in the future. After the access information of each storage component class is calculated based on the operational information, analysis data corresponding to the storage component can be generated based on the access information. Since the analysis data is generated based on the operational data combined with the characteristics of the storage service type, the accuracy of the analysis data is improved. When migrating stored data based on analysis data according to preset data processing strategies, data can be placed better according to data popularity, improving storage service performance and reducing storage costs. At the same time, after placing stored data according to data popularity, storage services can be better provided to users, ensuring user data access performance and improving user experience.

[0119] Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Figure 5 FIG1 shows a schematic diagram of the structure of a data processing device provided by an embodiment of this specification. Figure 5 As shown, the device includes:

[0120] The determination module 502 is configured to obtain operation information of a storage service and determine at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component;

[0121] A statistics module 504 is configured to collect access information corresponding to the at least one storage component class based on the operation information, and generate analysis data corresponding to the storage component according to the access information;

[0122] The migration module 506 is configured to perform data migration on the storage data in the storage component based on the analysis data according to a preset data processing strategy to obtain a migrated storage component.

[0123] Optionally, the determining module 502 is further configured to collect initial operation information of the storage service; and perform data preprocessing on the initial operation information to obtain operation information.

[0124] Optionally, the determination module 502 is further configured to classify the storage business according to the business information in the operation information to obtain at least one storage business class; determine the storage component set corresponding to the at least one storage business class, and cluster each storage component set according to the load information in the operation information to obtain at least one storage component class.

[0125] Optionally, the determination module 502 is further configured to determine at least one business classification identifier based on the business information in the operation information; and classify the storage business according to the at least one business classification identifier to obtain at least one storage business class.

[0126] Optionally, the determination module 502 is further configured to determine the storage components corresponding to each storage business in the at least one storage business class; combine the storage components corresponding to each storage business class, and obtain a storage component set corresponding to the at least one storage business class based on the combination result.

[0127] Optionally, the statistical module 504 is further configured to calculate access frequency information and access probability information corresponding to the at least one storage component class based on the operation information; and use the access frequency information and the access probability information as the access information corresponding to the at least one storage component class.

[0128] Optionally, the statistical module 504 is further configured to calculate the heat parameters corresponding to the storage component based on the access information corresponding to the at least one storage component class; obtain the identification parameters of the storage component, and associate the identification parameters with the heat parameters, and generate analysis data corresponding to the storage component based on the association results.

[0129] Optionally, the migration module 506 is further configured to determine the target storage component in the storage component based on the analysis data in accordance with the location processing strategy; determine the storage node to be received based on the target storage component, and migrate the target storage data in the target storage component to the storage node to be received to obtain the migrated target storage component.

[0130] Optionally, the migration module 506 is further configured to determine a first storage component and a second storage component in the storage component based on the analysis data in accordance with the location processing strategy, wherein the access information corresponding to the first storage component and the second storage component is different; determine a first preset storage node and a second preset storage node, wherein the storage configurations corresponding to the first preset storage node and the second preset storage node are different; migrate the first storage data in the first storage component to the first preset storage node to obtain the migrated first storage component, and migrate the second storage data in the second storage component to the second preset storage node to obtain the migrated second storage component.

[0131] Optionally, the device further includes a response module configured to generate a data migration instruction at a preset time interval; and obtain operation information of the storage service in response to the data migration instruction.

[0132] Optionally, the device also includes a storage module, configured to determine the data to be stored and the target storage business corresponding to the data to be stored in response to a data storage instruction, wherein the target storage business has the same business information as the storage business; select a target storage component from the storage component corresponding to the target storage business based on the analysis data; and store the data to be stored in the target storage component.

[0133] This specification provides a data processing device, comprising: a determination module configured to obtain operational information of a storage business and determine at least one storage component class corresponding to the storage business, wherein the storage business has a corresponding storage component; a statistics module configured to collect access information corresponding to the at least one storage component class based on the operational information, and generate analysis data corresponding to the storage component based on the access information; and a migration module configured to migrate the storage data in the storage component based on the analysis data in accordance with a preset data processing strategy to obtain the migrated storage component. By obtaining and analyzing the operational information of the storage business, clustering the storage components of the storage business based on the operational information, and determining multiple different storage component classes, it is convenient to subsequently collect access information statistics for different types of storage components. After the access information of each storage component class is collected based on the operational information, analysis data corresponding to the storage component can be generated based on the access information. Since the analysis data is generated based on the operational data combined with the characteristics of the storage business type, the accuracy of the analysis data is improved. When migrating stored data based on analysis data according to preset data processing strategies, data can be placed better according to data popularity, improving storage service performance and reducing storage costs. At the same time, after placing stored data according to data popularity, storage services can be better provided to users, ensuring user data access performance and improving user experience.

[0134] The above is a schematic diagram of a data processing device according to this embodiment. It should be noted that the technical solution of the data processing device and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the data processing device, please refer to the description of the technical solution of the above-mentioned data processing method.

[0135] Figure 6 6 shows a block diagram of a computing device 600 according to one embodiment of the present disclosure. Components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.

[0136] The computing device 600 also includes an access device 640 that enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0137] In one embodiment of the present specification, the above components of the computing device 600 and Figure 6 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 6 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0138] Computing device 600 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 600 may also be a mobile or stationary server.

[0139] The processor 620 is configured to execute the following computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by the processor.

[0140] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned data processing method are of the same concept. For details not described in detail in the technical scheme of the computing device, please refer to the description of the technical scheme of the above-mentioned data processing method.

[0141] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.

[0142] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical scheme of the above-mentioned data processing method.

[0143] An embodiment of the present specification further provides a computer program product, including a computer program or instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.

[0144] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-mentioned data processing method.

[0145] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0146] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0147] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0148] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0149] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A data processing method, comprising: Acquire operation information of a storage service, and determine at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component; collecting access information corresponding to the at least one storage component class based on the operation information, and generating analysis data corresponding to the storage component according to the access information; The storage data in the storage component is migrated based on the analysis data according to a preset data processing strategy to obtain a migrated storage component.

2. The method according to claim 1, wherein obtaining storage service operation information comprises: Collect initial operational information of storage services; Screening the initial operation information according to a preset data cleaning strategy to determine target operation information; The target operation information is removed from the initial operation information to obtain operation information.

3. The method according to claim 1, determining at least one storage component class corresponding to the storage service comprises: Classifying the storage service according to the service information in the operation information to obtain at least one storage service class; A storage component set corresponding to the at least one storage service class is determined, and each storage component set is clustered according to load information in the operation information to obtain at least one storage component class.

4. The method according to claim 3, wherein the storage service is classified according to the service information in the operation information to obtain at least one storage service class, comprising: Determine at least one business classification identifier based on the business information in the operation information; The storage service is classified according to the at least one service classification identifier to obtain at least one storage service class.

5. The method according to claim 3, determining the storage component set corresponding to the at least one storage service class comprises: Determining a storage component corresponding to each storage service in the at least one storage service class; The storage components corresponding to each storage service class are combined, and a storage component set corresponding to the at least one storage service class is obtained according to the combination result.

6. The method according to claim 1, wherein collecting statistics on access information corresponding to the at least one storage component class based on the operation information comprises: Calculate access frequency information and access probability information corresponding to the at least one storage component class based on the operation information; The access frequency information and the access probability information are used as access information corresponding to the at least one storage component class.

7. The method according to claim 1, generating analysis data corresponding to the storage component according to the access information, comprising: Calculating a heat parameter corresponding to the storage component according to access information corresponding to the at least one storage component class; An identification parameter of the storage component is obtained, and the identification parameter is associated with the heat parameter, and analysis data corresponding to the storage component is generated according to the association result.

8. The method according to claim 1, wherein the preset data processing strategy is a location processing strategy; in, Migrating the stored data in the storage component based on the analysis data according to a preset data processing strategy to obtain a migrated storage component includes: determining a target storage component among the storage components based on the analysis data according to the location processing strategy; A storage node to be received is determined based on the target storage component, and target storage data in the target storage component is migrated to the storage node to be received, so as to obtain the migrated target storage component.

9. The method according to claim 1, wherein the preset data processing strategy is a configuration processing strategy; in, Migrating the stored data in the storage component based on the analysis data according to a preset data processing strategy to obtain a migrated storage component includes: determining a first storage component and a second storage component in the storage component based on the analysis data according to the location processing strategy, wherein the first storage component and the second storage component respectively correspond to different access information; Determining a first preset storage node and a second preset storage node, wherein the first preset storage node and the second preset storage node respectively correspond to different storage configurations; Migrate the first storage data in the first storage component to the first preset storage node to obtain the migrated first storage component, and migrate the second storage data in the second storage component to the second preset storage node to obtain the migrated second storage component.

10. The method according to any one of claims 1 to 9, further comprising: Generate data migration instructions according to preset time intervals; In response to the data migration instruction, operation information of the storage service is obtained.

11. The method according to any one of claims 1 to 9, further comprising: determining, in response to the data storage instruction, data to be stored and a target storage service corresponding to the data to be stored, wherein the target storage service has the same service information as the storage service; selecting a target storage component from the storage components corresponding to the target storage service according to the analysis data; The data to be stored is stored in the target storage component.

12. A data processing system, comprising a storage service end and an analysis service end, wherein: The analysis server is configured to obtain operational information of a storage service for the storage service client, determine at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component; collect access information corresponding to the at least one storage component class based on the operational information, and generate analysis data corresponding to the storage component based on the access information; The storage service end is used to migrate the storage data in the storage component based on the analysis data according to a preset data processing strategy to obtain a migrated storage component.

13. A data processing device comprising: a determination module configured to obtain operation information of a storage service and determine at least one storage component class corresponding to the storage service, wherein the storage service has a corresponding storage component; a statistics module configured to collect access information corresponding to the at least one storage component class based on the operation information, and generate analysis data corresponding to the storage component according to the access information; The migration module is configured to perform data migration on the storage data in the storage component based on the analysis data according to a preset data processing strategy to obtain a storage component after migration.

14. A computing device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.

16. A computer program product comprising a computer program or instructions, which implement the steps of the method according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • Systems, methods and devices for implementing data management in a distributed data storage system

    CA2867589A1

  • Sub-tree migration method and device for realizing metadata load balancing

    CN111338801A

  • Accelerate data access in memory systems via data stream segregation

    CN111699477A

  • Load balancing prediction method, device and system and storage medium

    CN116069594A

  • Virtual machine migration method and device of cloud data center

    CN116450282A