Data processing method and apparatus
Through an intelligent hierarchical system, the service data access log in the storage system is analyzed, the access heat is determined according to the hot and cold identification algorithm, and data migration is carried out, which solves the pressure problem of traditional data processing methods on the storage system, and improves the operating performance of the storage system and the accuracy of data hierarchy.
Patent Information
- Application Number
- PCT/CN2024/092990
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-06
- Filing Date
- 2024-05-14
- Publication Date
- 2025-05-30
AI Technical Summary
Traditional data processing methods in the storage system scan the data access time recorded by the metadatabase in a timely manner, resulting in an increase in the pressure on the metadatabase, affecting other business operations of the main IO path, and increasing the risk of failure spread, affecting the normal operation of the storage system.
The intelligent hierarchical system is adopted to obtain the access log of business data, analyze the access heat according to the user-configured hot and cold recognition algorithm, and send data migration instructions to the storage system to realize automatic data hierarchy.
It reduces the computing resource usage of the storage system, reduces the impact on the storage system, improves the operating performance of the storage system, and realizes more accurate automatic layering of hot and cold data.
Smart Images

Figure CN2024092990_30052025_PF_FP_ABST
Abstract
Description
Data processing method and device
[0001] This application claims priority to Chinese Patent Application No. 202311566468.2, filed with the State Intellectual Property Office of China on November 22, 2023, entitled “A Method, Apparatus, and Other Devices for Data Processing,” the entire contents of which are incorporated herein by reference. This application also claims priority to Chinese Patent Application No. 202410171571.5, filed with the State Intellectual Property Office of China on February 6, 2024, entitled “A Method and Apparatus for Data Processing,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of storage, and in particular to a data processing method and device. Background Art
[0003] Users use cloud storage systems to manage massive amounts of data. To further manage storage costs, they need to perform storage tiering based on the frequency of object usage. For example, frequently accessed data can be placed in the frequently accessed tier to reduce the overhead of access requests, while infrequently accessed data can be placed in the infrequent, archived, or even deep archived access tier to reduce storage costs.
[0004] Traditional data processing involves periodically scanning the metadata database's data access times in the background of the storage system, switching storage tiers based on changes in access times. This data processing approach increases the pressure on the metadata database, impacting other business operations on the primary input / output (IO) path, and even increasing the risk of fault propagation, leading to service unavailability. In other words, this data processing approach can impact the normal operation of the storage system.
[0005] Summary of the Invention
[0006] The present application provides a data processing method and device for improving the operating performance of a storage system.
[0007] In order to achieve the above objectives, this application adopts the following technical solutions.
[0008] In a first aspect, an embodiment of the present application provides a data processing method, which is applied to an intelligent tiering system, the intelligent tiering system is used to perform intelligent tiering processing on business data in a storage system, the storage system includes at least one storage path, the method includes: obtaining first business data, the first business data is stored in a target storage path in the storage system, wherein the storage type of the first business data is an intelligent tiering type; obtaining an access log of the first business data; analyzing the access log of the first business data according to a hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data; sending a data migration instruction to the storage system according to the access heat of the first business data, so that the storage node of the storage system performs data migration on the first business data.
[0009] In the above method, the storage system can be a system that can support storage tiering, and the intelligent tiering system is an out-of-band system, that is, another system deployed outside the storage system, which can support a richer range of hot and cold identification algorithms (including algorithms for hot and cold identification based on the last access time, algorithms for hot and cold identification based on access frequency, and algorithms for hot and cold identification based on life cycle access ratio). In this way, when the hot and cold identification algorithm corresponding to the target storage path is subsequently used for hot and cold identification, the log can be analyzed at a deeper level, making the final access heat more accurate, and thus being able to more intelligently achieve automatic tiering of hot and cold data. In addition, compared to the traditional data processing method of recording redundant fields in the storage system, the intelligent tiering system uses the log of the first business data for analysis, without adding redundant fields to the storage system. This not only reduces the storage space occupied, thereby saving users' storage costs, but also eliminates the need for the storage system itself to analyze whether objects need to be migrated, thereby effectively reducing the impact on the storage system and improving the operating performance of the storage system.
[0010] In one implementation, the hot and cold identification algorithms are stored in an intelligent grading system.
[0011] In the above implementation, the hot and cold identification algorithm is not stored in the storage system, but in another system independent of the storage system (i.e., the intelligent grading system). This can greatly reduce the space occupied in the storage system and reduce storage costs.
[0012] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on the access ratio of the life cycle. According to the hot and cold identification algorithm configured by the user for the target storage path, the access log of the first business data is analyzed to obtain the access heat of the first business data, including: analyzing the access log of the first business data to obtain the access ratio of the first business data in the target life cycle; comparing the access ratio with the target access threshold corresponding to the target life cycle to obtain the access heat of the first business data.
[0013] In the above implementation method, the access heat is determined based on the data indicators of the first business data, and the hot and cold identification algorithms of different storage paths are different, and the data indicators obtained are also different. This can be more in line with the business scenario of the first business data. When the hot and cold identification algorithm of the target storage path is used for hot and cold identification based on the access ratio of the life cycle, the first business data here is one or more business data whose object creation time belongs to the target life cycle. This means that the embodiment of the present application can analyze multiple business data belonging to the target life cycle segment at the same time, so as to facilitate subsequent batch data migration and improve data migration efficiency. In addition, through the data indicator of the access ratio of the target life cycle, the business scenario that fits the first business data can be obtained, so that the final access heat is more accurate.
[0014] In one implementation, analyzing the access log of the first business data to obtain the access ratio of the first business data in the target life cycle includes: obtaining the second business data, where the second business data is the business data stored in the target storage path; analyzing the access log of the second business data to obtain the access ratio of the second business data in at least one life cycle; and determining the access threshold of each life cycle according to the access ratio in at least one life cycle.
[0015] In the above implementation, since the second business data is the business data stored in the target storage path, and the access thresholds of different life cycles are related to their corresponding access ratios, this means that the embodiment of the present application can classify business data with different life cycles based on a batch of business data with similar life cycles, thereby realizing automatic migration of business data in a special business scenario (for example, a scenario where access behaviors to objects with the same life cycle are relatively consistent).
[0016] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on access frequency. According to the hot and cold identification algorithm configured by the user for the target storage path, the access log of the first business data is analyzed to obtain the access heat of the first business data, including: analyzing the access log of the first business data to obtain the target access frequency of the first business data; obtaining N historical access frequencies of the first business data within a preset time, where N is a positive integer; determining the changing trends of the N historical access frequencies and the target access frequency to obtain the access heat of the first business data.
[0017] In the above implementation, when the hot and cold identification algorithm of the target storage path is used to perform hot and cold identification based on the access frequency, this means that the access popularity of the first business data conforms to a certain change pattern. Therefore, the intelligent grading system not only needs to analyze and obtain the target access frequency of the first business data, but also needs to obtain N historical access frequencies, and then accurately estimate the access popularity of the first business data through the obtained change trends of these access frequencies, so as to realize the automatic migration of business data in another special business scenario (for example, a scenario where the access frequency has a change pattern).
[0018] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on the last access time. According to the hot and cold identification algorithm configured by the user for the target storage path, the access log of the first business data is analyzed to obtain the access popularity of the first business data, including: analyzing the access log of the first business data to obtain the last access time of the first business data; calculating the target time interval between the last access time and the current time; comparing the target time interval and the preset time interval to obtain the access popularity of the first business data.
[0019] In the above implementation, when the hot and cold identification algorithm of the target path is used for hot and cold identification based on the last access time, this algorithm that determines the access heat by comparing the target time interval and the preset time interval is relatively simple, not only has high identification efficiency, but also can be applied to most business scenarios.
[0020] In one implementation, the target storage path includes one or more of the following: a bucket of object storage, a volume of block storage, a directory of file storage, and an instance of a database.
[0021] In the above implementation, the target storage path includes one or more of the above, which means that the storage system in the embodiment of the present application does not specifically refer to a certain storage system, but can be widely applied to various types of storage systems, such as object storage systems, block storage systems, file storage systems, and database-related systems, so as to provide extremely cost-effective storage services in various types of storage systems.
[0022] In the second aspect, an embodiment of the present application provides a data processing device for implementing a method as described in the first aspect or any one of the implementation methods of the first aspect. Specifically, the device includes: an acquisition module for acquiring first business data, the first business data is stored in a target storage path in a storage system, wherein an intelligent grading system is used to perform intelligent grading processing on the business data in the storage system, the storage system includes at least one storage path, and the storage type of the first business data is an intelligent grading type; the acquisition module is also used to acquire an access log of the first business data; the processing module is used to analyze the access log of the first business data according to a hot and cold identification algorithm configured by the user for the target storage path, and obtain the access heat of the first business data; the sending module is used to send a data migration instruction to the storage system according to the access heat of the first business data, so that the storage node of the storage system performs data migration on the first business data.
[0023] In one implementation, the hot and cold identification algorithms are stored in an intelligent grading system.
[0024] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on the access ratio of the life cycle. The processing module is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module, specifically used to analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle; a processing module, specifically used to compare the access ratio with the target access threshold corresponding to the target life cycle to obtain the access heat of the first business data.
[0025] In one implementation, the processing module is specifically used to analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle, and also includes: a processing module, specifically used to obtain the second business data, the second business data is the business data stored in the target storage path; the processing module is specifically used to analyze the access log of the second business data to obtain the access ratio of the second business data in at least one life cycle; the processing module is specifically used to determine the access threshold of each life cycle according to the access ratio in at least one life cycle.
[0026] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on access frequency, and a processing module is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module, specifically used to analyze the access log of the first business data to obtain the target access frequency of the first business data; a processing module, specifically used to obtain N historical access frequencies of the first business data within a preset time, where N is a positive integer; a processing module, specifically used to determine the changing trend of the N historical access frequencies and the target access frequency to obtain the access heat of the first business data.
[0027] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on the last access time, and a processing module is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module, specifically used to analyze the access log of the first business data to obtain the last access time of the first business data; a processing module, specifically used to calculate the target time interval between the last access time and the current time; a processing module, specifically used to compare the target time interval and the preset time interval to obtain the access heat of the first business data.
[0028] In one implementation, the target storage path includes one or more of the following: a bucket of object storage, a volume of block storage, a directory of file storage, and an instance of a database.
[0029] In a third aspect, embodiments of the present application provide a computing device cluster, comprising at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the data processing method of the first aspect or any possible implementation of the first aspect.
[0030] In a fourth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, enables the computing device cluster to execute the data processing method in the first aspect or any possible implementation of the first aspect.
[0031] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the data processing method in the first aspect or any possible implementation of the first aspect.
[0032] The technical effects produced by any implementation method in the above-mentioned second to fifth aspects and each aspect can refer to the above-mentioned first aspect and the corresponding implementation method in the first aspect, and the repetitions will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] FIG1 is a system architecture diagram corresponding to a storage system provided in an embodiment of the present application;
[0034] FIG2 is a system architecture diagram corresponding to an object storage system provided in an embodiment of the present application;
[0035] FIG3 is a system schematic diagram corresponding to a data processing system provided in an embodiment of the present application;
[0036] FIG4 is a first schematic diagram of an interface for configuring a storage bucket provided in an embodiment of the present application;
[0037] FIG5 is a second schematic diagram of an interface for configuring a storage bucket provided in an embodiment of the present application;
[0038] FIG6 is a schematic diagram of a scenario of data migration in an object storage system provided by an embodiment of the present application;
[0039] FIG7 is an interaction diagram of a method for performing data processing provided in an embodiment of the present application;
[0040] FIG8 is a graph showing a trend of changes in access popularity provided by an embodiment of the present application;
[0041] FIG9 is a schematic diagram of a method for performing data processing provided in an embodiment of the present application;
[0042] FIG10 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;
[0043] FIG11 is a schematic diagram of the structure of a computing device provided by the present application;
[0044] FIG12 is a schematic diagram of the structure of a computing device cluster provided by the present application;
[0045] FIG13 is a schematic diagram of a structure of a network connection between computing devices provided by the present application. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. In order to facilitate the clear description of the technical solutions in the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design solutions. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.
[0047] To facilitate understanding of the technical solutions provided in the embodiments of the present application, the following terms are first introduced:
[0048] 1. Storage system:
[0049] A storage system is composed of various storage nodes that store programs and data, control components, and equipment (hardware) and algorithms (software) for managing information adjustments. It can be broadly divided into object storage systems, block storage systems, and file storage systems.
[0050] 2. Object Storage System:
[0051] Its primary operating object is the object. Object storage is generally represented by a universally unique identifier (UUID). Data and metadata are packaged together as a whole object and stored in a large pool. It is understandable that the object storage system is a system used to provide the Object Storage Service (OBS). The basic components of OBS are buckets and objects. Buckets are containers for storing objects in OBS. Each bucket has its own storage category, access permissions, region, and other attributes. Users locate buckets on the internet using the bucket's access domain name. Objects are the basic unit of data storage in OBS. An object is actually a collection of file data and its related attribute information. Specifically, it can include three parts: key value, metadata, and data.
[0052] The key value can be a character sequence used to uniquely represent the object name. This character sequence is unique in the current bucket, that is, each object in a bucket has a unique object key value; metadata is the object's description information, including system metadata and user metadata. These metadata are uploaded to OBS in the form of key-value pairs; data is the data content of the file.
[0053] 3. Intelligent grading system (also known as out-of-band system):
[0054] The intelligent grading system is used to provide a plug-in hot and cold identification algorithm to analyze logs and realize automatic stratification of hot and cold data. The intelligent grading system and storage system can be deployed on the same device (e.g., a server). The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The embodiments of this application will not limit the number of servers.
[0055] 4. Garbage Collection (GC):
[0056] Garbage collection is a memory management mechanism that releases memory when it is no longer needed. Garbage collection can be categorized as either memory-based or external memory-based. Memory-based garbage collection is integrated into high-level programming languages like Java, Go, and C#, while external memory-based garbage collection is widely used in solid-state drives (SSDs).
[0057] To facilitate understanding of the technical solutions provided by the embodiments of the present application, a brief introduction to the relevant technologies of the embodiments of the present application is given below:
[0058] Please refer to Figure 1, which is a system architecture diagram corresponding to a storage system provided in an embodiment of the present application. As shown in Figure 1, the storage system 10 can be a distributed storage system, which can specifically include a client cluster 110, a server 120, and a storage node cluster 130. Among them, the client cluster 110 can include one or more clients, and the number of clients is not limited here. As shown in Figure 1, the client cluster 110 can specifically include client 111, client 112, client 113, and client 114. Among them, each client in the client cluster 110 can respectively establish a network connection with the server 120, so that each client can interact with the server 120 through the network connection.
[0059] It is understandable that each client in the client cluster 110 can be a terminal device installed with a target application (for example, an application for accessing business data), and the terminal device can include a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, car terminal, smart TV and other smart terminals with data processing capabilities. The target application can be a social application, a multimedia application (for example, a video application), an entertainment application (for example, a game application), an information flow application, an educational application, a live broadcast application, etc. Among them, the target application here can be an independent application or an embedded sub-application integrated in a certain application (for example, a mini-program, etc.), which will not be limited here.
[0060] The server 120 can be a server corresponding to the target application. The server 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. Among them, the embodiment of the present application will not limit the number of servers. As shown in Figure 1, the server 120 can also be connected to each storage node in the storage node cluster 130 through a network, so that the server 120 can manage or access any storage node in the storage node cluster 130.
[0061] Among them, the network connection in the embodiment of the present application is not limited to the connection method. It can be directly or indirectly connected through wired communication, directly or indirectly connected through wireless communication, or through other methods. This application does not limit this.
[0062] The storage node cluster 130 may include one or more storage nodes. There is no limit on the number of storage nodes. Each storage node may be deployed in multiple data centers. These multiple data centers may be located in the same area or in different areas. This application does not limit this. Each data center may include one or more physical devices with storage capabilities, such as servers, mobile phones, tablet computers, read-only memory (ROM), flash memory, hard disk drive (HDD) or solid state drive (SSD), etc. This application is specific to physical devices with storage capabilities.
[0063] For ease of understanding, the storage system in the embodiments of the present application can be used as an example of an object storage system to illustrate how to analyze the logs of first business data (i.e., business data with an intelligent tiering storage type) based on the intelligent tiering system, thereby achieving automatic tiering of hot and cold data. Further, please refer to Figure 2, which is a system architecture diagram corresponding to an object storage system provided in an embodiment of the present application. As shown in Figure 2, the object storage system 20 can include a client 200, a service layer 210, and a metadata layer 220.
[0064] Among them, client 200 can be any client in the client cluster shown in Figure 1 above, and the client 200 can also deploy N applications, which can specifically include application 201, application 202, ..., application 20N. The user corresponding to client 200 can use the account on client 200 to log in to the cloud management platform, create a storage bucket, configure the bucket name, and obtain the bucket domain name through the cloud management platform. After detecting the user's operation (such as creating a bucket, configuring the bucket name, hot and cold identification algorithm, etc.), the cloud management platform can issue a creation instruction, which includes information such as the bucket name and bucket domain name, and notifies the object storage service node to create a bucket and save information such as the bucket name and bucket domain name.
[0065] After creating a bucket, the user can access the bucket domain name through client 200, locate the corresponding bucket, and upload data (such as uploading an object) or download data (such as downloading an object) from the bucket. Taking application 201 as an example, client 200 can respond to user operations on application 201 and generate an object upload request for a specific object, thereby uploading the object to a bucket created by the client.
[0066] It's understood that users can create one or more buckets. When there are many business scenarios, you can create a bucket for each business scenario and subsequently store objects in the bucket that matches the business scenario. For example, in a multimedia scenario, objects associated with entertainment can be stored in bucket 1, objects associated with sports can be stored in bucket 2, and objects associated with pets can be stored in bucket 3.
[0067] The service layer 210 can interact with the client 200 and provide external front-end services. Foreground services can be services that request the client 200 to read or write data. The service layer 210 can include lifecycle management 211 and background tasks (e.g., garbage collection 212). Lifecycle management 211 is used to scan the metadata layer 220 to calculate whether objects need to be migrated based on the last access time; garbage collection 212 is used to regularly clean up redundant data and free up storage space.
[0068] The metadata layer 220 includes bucket metadata 221 and object metadata 222. For clusters storing metadata, partitioning is often used to dynamically split the metadata to accommodate the growing number of metadata entries. Each partition manages different data entries. Bucket metadata 221 not only records partitioning rules but also bucket metadata, such as the region to which the bucket belongs. Object metadata 222 records object metadata, such as its attributes, size, upload time, and policy information. Object metadata 222 also includes a write-ahead log (WAL). This log requires that all modifications to the storage system be written to a log before being submitted, ensuring the atomicity and durability of the object storage system. Assuming the hard disk data is intact, the write-ahead log allows the storage system to recover to its pre-crash state after a crash, guided by the log, to avoid data loss. This technology is widely used to store metadata in storage systems (file, object, database, and column storage).
[0069] In current technical solutions, when a user accesses an object (e.g., object 1) in object storage system 20, the object storage service provided by object storage system 20 records relevant information in object storage system 20 based on the user's access request and the current time. For example, if an append write method is used to add a last access time field to the object metadata of object 1 (e.g., object metadata D1 shown in FIG2 ), this will result in the existence of both the original object metadata and new object metadata with the added last access time field (e.g., object metadata D2 shown in FIG2 ) in object storage system 20. This not only requires a large amount of additional storage space, but also results in significant read / write amplification.
[0070] In order to free up storage space, garbage collection 212 is required to promptly clean up redundant data (i.e., delete object metadata D1). If object 1 is frequently accessed, object metadata for object 1 will be frequently updated, which will place more burden and overhead on garbage collection. In addition, based on lifecycle management 211, object storage system 20 can periodically scan objects from object metadata 222 and perform simple analysis and processing to determine whether they need to switch access layers. For example, intelligent grading can be performed based on lifecycle rules such as the latest access timestamp (i.e., the last access time) and the last modification time recorded in the metadata database. However, this method of directly scanning and calculating within the object storage system 20 has already affected the normal business processing of the object storage system 20. If a failure occurs, it will also indirectly affect the serviceability of the object storage system 20. Therefore, it can be seen that traditional data processing methods will affect the normal operation of the object storage system 20.
[0071] In order to solve the above problems, an embodiment of the present application provides a data processing method, which is applied to an intelligent tiering system, and the intelligent tiering system is used to perform intelligent tiering processing on business data in a storage system, where the storage system includes at least one storage path. In this method, the intelligent tiering system can obtain first business data (i.e., business data whose storage type is an intelligent tiering type, for example, an intelligent tiering object in an object storage system), and the first business data is stored in a target storage path in the storage system. Furthermore, the intelligent tiering system can obtain the access log of the first business data, and analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path, so as to obtain the access heat of the first business data. Then, the intelligent tiering system can send a data migration instruction to the storage system according to the access heat of the first business data, so that the storage node of the storage system performs data migration on the first business data.
[0072] In an embodiment of the present application, the storage system may be a system that supports storage tiering, and the intelligent tiering system is an out-of-band system, that is, another system deployed outside the storage system, which can support a richer set of hot and cold identification algorithms (including algorithms for hot and cold identification based on the last access time, algorithms for hot and cold identification based on access frequency, and algorithms for hot and cold identification based on life cycle access ratio). In this way, when the hot and cold identification algorithm corresponding to the target storage path is subsequently used for hot and cold identification, a deeper analysis of the log can be performed, making the final access heat more accurate, thereby enabling a more intelligent implementation of automatic tiering of hot and cold data. In addition, compared to the traditional data processing method of recording redundant fields in the storage system, the intelligent tiering system uses the log of the first business data for analysis, without adding redundant fields to the storage system. This not only reduces the storage space occupied, thereby saving the user's storage cost, but also eliminates the need for the storage system itself to analyze whether the object needs to be migrated, thereby effectively reducing the impact on the storage system and improving the operating performance of the storage system.
[0073] The following is a system diagram of the data processing method provided by the embodiment of the present application:
[0074] Please refer to Figure 3, which is a system diagram corresponding to a data processing system provided in an embodiment of the present application. As shown in Figure 3, the data processing system can be an out-of-band intelligent hierarchical management system based on a split architecture, that is, the data processing system can include a storage system 30 and an intelligent hierarchical system 31. The storage system 30 is a system deployed on the data plane, and the intelligent hierarchical system 31 is a system deployed on the control plane.
[0075] If the data processing system is deployed in the cloud, that is, the storage system 30 and the intelligent tiering system 31 shown in Figure 3 are both systems deployed in the cloud, then this means that the data processing method provided in the embodiments of the present application can be applied to the field of cloud storage. Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as storage system 30) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage nodes (also referred to as storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.
[0076] Currently, storage systems utilize a method for creating logical volumes. When creating a logical volume, physical storage space is allocated for each logical volume. This physical storage space may consist of disks on a specific storage node or several storage nodes. When a client stores data on a logical volume, it stores the data on a file system. The file system divides the data into multiple parts, each of which is an object. An object contains not only the data but also additional information such as the data identifier (ID entity). The file system writes each object to the physical storage space of the logical volume and records the storage location information for each object. Therefore, when a client requests access to data, the file system can provide access based on the storage location information for each object.
[0077] The storage system allocates physical storage space to logical volumes by pre-dividing the physical storage space into stripes based on the estimated capacity of the objects to be stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the Redundant Array of Independent Disks (RAID) groupings. A logical volume can be understood as a stripe, thereby allocating physical storage space to the logical volume.
[0078] As shown in FIG3 , the storage system 30 may be a system that supports data tiering and may be used to store X business data, where X is a positive integer. In the embodiment of the present application, the data types of the business data may include two types: one in which the user directly specifies the storage tier (for example, directly specifying storage in storage tier L1), and the other in which data is automatically migrated based on access popularity (i.e., an intelligent tiering type).
[0079] It should be understood that the embodiment of the present application can classify the data stored in the storage system 30 into different categories (for example, hot data, warm data, and cold data) according to the access popularity of the business data (obtained after evaluation by data indicators such as access frequency) to facilitate better management and storage of data. It is understandable that the hot data here refers to data with high access frequency and is important to business and applications. These data usually require fast and efficient access and processing, and therefore need to be stored in a high-performance, low-latency storage layer L1 (for example, solid-state drives, memory, etc.); warm data refers to data with moderate access frequency and a certain importance to business and applications. These data usually do not need to be accessed and processed as quickly as hot data, but also require reliable storage and access within a certain period of time, so they can be stored in a storage layer L2 with lower cost and larger capacity (for example, disk array); cold data refers to data with lower access frequency and less important to business and applications. These data usually need to be stored for a long time, but do not require frequent access and processing, so they can be stored in a storage layer L3 with lower cost and larger capacity (for example, tape library).
[0080] As shown in Figure 3, when the storage system 30 is an object storage system, the business data stored in the storage system 30 can be objects, specifically including Object 1, Object 2, Object 3, Object 4, and Object 5. Object 1 and Object 4 are both objects in the storage system 30 whose storage type belongs to the intelligent tiering type. Object 2 can be hot data stored in storage layer L1 as specified by the user, Object 3 can be warm data stored in storage layer L2 as specified by the user, and Object 5 can be cold data stored in storage layer L3 as specified by the user. When the storage system 30 is running, an event record called a log is generated. Each log line can record a description of the relevant operation, such as the date and time of the event, the user, and the action. Therefore, when a user performs a system operation (e.g., a read operation, a write operation, or a delete operation) on any of the above five objects, the storage system 30 can record the log corresponding to the system operation.
[0081] The intelligent tiering system 31 can be a system deployed outside the storage system 30 to implement automatic data tiering, and its final analysis results can be applied to the storage system 30. The intelligent tiering system 31 can include a storage subsystem 311, a heat recognition plug-in 312, and a big data analysis platform 313.
[0082] The storage subsystem 311 can be used to store R intelligent tiered objects in storage buckets, where an intelligent tiered object refers to an object whose storage type belongs to the intelligent tiering type, and R is a positive integer. For example, the storage subsystem 311 can store object names, version numbers, and data indicators. The storage subsystem 311 can be a key-value storage subsystem, i.e., a key-value (kv) storage system. Users can configure different hot and cold identification algorithms (i.e., hot and cold identification algorithms) on the storage buckets according to different business scenarios.
[0083] The intelligent grading system 31 may store multiple hot / cold identification algorithms, specifically including a first hot / cold identification algorithm, a second hot / cold identification algorithm, and a third hot / cold identification algorithm. The first hot / cold identification algorithm may be an algorithm that identifies hot / cold based on the last access time; the second hot / cold identification algorithm may be an algorithm that identifies hot / cold based on the access frequency; and the third hot / cold identification algorithm may be an algorithm that identifies hot / cold based on the access ratio within the life cycle.
[0084] Because intelligent tiering system 31 includes n tiering services, it can be deployed in multiple primary zones, such as Zone 1, Zone 2, ..., Zone n. In this case, objects stored in storage subsystem 311 are associated with these n primary zones. The n tiering services in intelligent tiering system 31 do not interfere with or affect each other.
[0085] Among them, the intelligent grading system 31 can determine the newly uploaded intelligent grading objects in the storage system 30 by consuming the index (index) incremental pre-write log, and then store the obtained intelligent grading objects in the storage subsystem 311, for example, the object name, version number, etc. of the intelligent grading objects. Then, when the intelligent grading system 31 calls the grading service corresponding to each area and performs persistent management on the objects in the storage system 30, the object storage rate can be effectively improved. Among them, the index here is a sorted data structure in the storage system 30, which is used to uniquely indicate a certain object; WAL can be used to ensure the atomicity and persistence of data operations.
[0086] The big data analysis platform 313 can not only store massive logs into the lake, but also perform offline analysis on the logs in the storage system 30. It is understandable that if the first business data obtained by the intelligent grading system 31 is any object in the storage system 30 (for example, object 1), the big data analysis platform 313 can obtain the access log of the first business data from the storage system 30, and then analyze the access log of the first business data based on the hot and cold identification algorithm configured by the user for the target storage path (that is, the storage path where the first business data is stored in the storage system 30), so as to obtain its corresponding data indicators. Among them, the data indicators here can be any one of the last access time, access frequency or life cycle access ratio.
[0087] In the embodiment of the present application, the storage bucket where the first business data is located may be referred to as a target storage bucket. If the hot and cold identification algorithm is the first hot and cold identification algorithm or the second hot and cold identification algorithm, the first business data here may be any intelligent grading object in the target storage bucket. If the hot and cold identification algorithm is the third hot and cold identification algorithm, the first business data here may be multiple intelligent grading objects in the target storage bucket whose object creation time belongs to the same life cycle. The life cycle refers to the time period to which the object creation time belongs. For example, if Object 1 and Object 4 shown in FIG3 are both intelligent grading objects in the same target storage bucket, the object creation time of Object 1 is December 3, 2023, and the object creation time of Object 4 is December 4, 2023, then if a life cycle is the time period from December 1, 2023 to December 15, 2023, the first business data belonging to the life cycle may include Object 1 and Object 4.
[0088] The heat identification plug-in 312 can be used to perform heat identification on the first business data based on the data indicators obtained by the big data analysis platform 313 to obtain the access heat of the first business data, and then send a data migration instruction to the storage system 30 based on the access heat. For example, the intelligent grading system can compare the storage layer corresponding to the access heat (i.e., the second storage layer) with the current storage layer of the first business data (i.e., the first storage layer) to determine whether the target object needs to be migrated. When the first storage layer is different from the second storage layer, it is necessary to send a data migration instruction to the storage system 30 so that the storage node of the storage system 30 migrates the first business data from the first storage layer to the second storage layer.
[0089] It is understandable that users can configure different hot and cold identification algorithms for different storage paths. For example, in the object storage system 30, when creating a bucket, users can configure various properties of the bucket, which may include: the default storage type of the bucket, access rights, the region to which it belongs, the hot and cold identification algorithm, etc. For ease of understanding, please further refer to Figure 4, which is a schematic diagram of an interface for configuring a bucket provided in an embodiment of the present application. As shown in Figure 4, user a can be a user who logs in to the cloud management platform using an account on the client, and the interface J1 can be a guide interface provided by the object storage system based on the object storage service. The interface J1 includes business controls for creating buckets (for example, the control K1 shown in Figure 4).
[0090] When user a triggers control K1, the client corresponding to user a can respond to the trigger and display a creation interface (e.g., interface J2 shown in FIG4 ). User a can then set the name of the currently created bucket and select the default storage type in interface J2. Interface J2 can include multiple default storage types, including intelligent tiering, standard storage, low-frequency access, and archive storage. The default storage type indicates the default storage tier for objects subsequently uploaded to the bucket. Furthermore, interface J2 can include a business control (e.g., control K2 shown in FIG4 ) for displaying multiple hot and cold identification algorithms. These multiple hot and cold identification algorithms can include an algorithm for hot and cold identification based on the last access time (i.e., a first hot and cold identification algorithm), an algorithm for hot and cold identification based on access frequency (i.e., a second hot and cold identification algorithm), and an algorithm for hot and cold identification based on the access ratio within the life cycle (i.e., a third hot and cold identification algorithm).
[0091] It should be understood that when the hot and cold identification algorithm configured by user a for the storage bucket is the first hot and cold identification algorithm, user a can directly perform a trigger operation on the first hot and cold identification algorithm, thereby causing the client to display the algorithm configuration interface for the first hot and cold identification algorithm (for example, interface J3 shown in Figure 4), wherein the interface J3 can include settings for relevant thresholds of the first hot and cold identification algorithm (for example, conversion days).
[0092] When user a completes the configuration, a new storage bucket is successfully created. The triggering operation in the embodiment of the present application may include contact operations such as clicking and long pressing, or non-contact operations such as voice and gestures, which will not be limited here.
[0093] Optionally, when the storage system is an object storage system, the user can also configure different hot and cold identification algorithms for the sub-paths of each folder in the bucket. For ease of understanding, please refer to Figure 5. Figure 5 is a second schematic diagram of an interface for configuring a bucket provided in an embodiment of the present application. As shown in Figure 5, user b can be a user who logs in to the cloud management platform using an account on the client, and interface J4 can be a details display interface for displaying multiple folders of a bucket (for example, bucket 1) created by user b, and one folder corresponds to one sub-path. Among them, the interface J4 can also include controls K3 and K4. Control K3 can be used to upload objects to bucket 1, and control K4 can be used to create a new folder in bucket 1.
[0094] As shown in Figure 5, user b has created two folders in bucket 1, which may specifically include folder 1 and folder 2. User b can configure different hot and cold identification algorithms for the sub-paths of each folder. Taking folder 1 as an example, user b can perform a trigger operation on file 1 so that the client corresponding to user b displays the sub-path configuration interface (for example, interface J5 shown in Figure 5). Among them, the interface J5 may include business controls for displaying multiple hot and cold identification algorithms (for example, control K6 shown in Figure 5). The multiple hot and cold identification algorithms here may include an algorithm for hot and cold identification based on the last access time (i.e., a first hot and cold identification algorithm), an algorithm for hot and cold identification based on access frequency (i.e., a second hot and cold identification algorithm), and an algorithm for hot and cold identification based on the access ratio within the life cycle (i.e., a third hot and cold identification algorithm).
[0095] Of course, if the hot and cold recognition algorithm configured by user b for the sub-path of folder 1 is the first hot and cold recognition algorithm, then user b can also directly perform a trigger operation on the first hot and cold recognition algorithm, thereby causing user b's client to display the algorithm configuration interface for the first hot and cold recognition algorithm (for example, interface J3 shown in Figure 4).
[0096] It should be noted that the interfaces and controls shown in Figures 4 and 5 are merely some reference forms of expression. In actual business scenarios, developers can make relevant designs based on product requirements. The embodiments of this application do not limit the specific forms of the interfaces and controls involved.
[0097] Further, for ease of understanding, please refer to Figure 6, which is a schematic diagram of a scenario for data migration in an object storage system provided by an embodiment of the present application. As shown in Figure 6, R buckets can be created in the object storage system, where R is a positive integer. For ease of explanation, three buckets can be used as an example, specifically including bucket 1, bucket 2, and bucket 3. Among them, these three buckets can be created by the same user based on different business scenarios, or they can be created by different users separately, and this will not be limited here.
[0098] As shown in Figure 6, the hot and cold identification algorithm here is configured for the storage path corresponding to the bucket. For example, the hot and cold identification algorithm configured for bucket 1 can be hot and cold identification algorithm S1, which is used to perform hot and cold identification based on the last access time; the hot and cold identification algorithm configured for bucket 2 can be hot and cold identification algorithm S2, which is used to perform hot and cold identification based on the access frequency; the hot and cold identification algorithm configured for bucket 3 can be hot and cold identification algorithm S3, which is used to perform hot and cold identification based on the access ratio within the life cycle.
[0099] It should be understood that the intelligent grading system can be used to perform intelligent grading processing on the intelligent grading objects in the object storage system, that is, to analyze the access log of the first business data in the storage system and realize the data migration of the intelligent grading objects in each storage bucket. It can be understood that the storage layers in the embodiment of the present application may include M, where M is a positive integer greater than 1. For ease of understanding, only 3 are taken as an example here, which may specifically include storage layer L1, storage layer L2 and storage layer L3. Among them, the access popularity of the business data stored in the storage layer L1 is lower than the access popularity of the business data stored in the storage layer L2, and the access popularity of the business data stored in the storage layer L2 is lower than the access popularity of the business data stored in the storage layer L3.
[0100] Among them, the storage layer L1 here can be a standard storage layer (for example, a disk for storing hot data), which is used to provide high-performance, high-reliability, and high-availability object storage service storage, and is suitable for business scenarios that require frequent access to data, such as big data, mobile applications, hot videos, social pictures and other business scenarios.
[0101] Storage tier L2 can be a low-frequency access storage tier (for example, disks used to store warm data). It provides highly reliable, low-cost, real-time access storage services. This is suitable for business scenarios where data is infrequently accessed but requires fast access when needed, such as file synchronization / sharing and enterprise backup. Compared to storage tier L1, storage tier L2 offers the same data durability, throughput, and access latency, but at a lower cost.
[0102] The storage layer L3 here can be an archive storage layer (for example, a tape for storing cold data), which is suitable for business scenarios where data is rarely accessed, such as data archiving, long-term backup, and offline analysis (for example, model training in machine learning or big data analysis).
[0103] For example, when the first business data obtained by the intelligent grading system is an intelligent grading object in bucket 1, the intelligent grading object can obtain the access log of the intelligent grading object and analyze the access log of the intelligent grading object according to the hot and cold identification algorithm S1 corresponding to bucket 1 to obtain the access heat of the intelligent grading object, and then send a data migration instruction to the object storage system according to the access heat of the intelligent grading object, so that the storage node of the object storage system performs data migration on the intelligent grading object.
[0104] For example, if the first storage layer of the intelligent tiering object (i.e., the current storage layer) is storage layer L1, then when the intelligent tiering system determines that the second storage layer of the intelligent tiering object (i.e., the storage layer corresponding to the access heat) is still storage layer L1, it means that the second storage layer is the same as the first storage layer, and there is no need to migrate the intelligent tiering object at this time; when the intelligent tiering system determines that the second storage layer of the intelligent tiering object is storage layer L2 or storage layer L3, it means that the second storage layer is different from the first storage layer. At this time, the intelligent tiering system needs to send a data migration instruction to the object storage system to enable the storage node of the object storage system to migrate the intelligent tiering object from the first storage layer to the second storage layer.
[0105] Similarly, the intelligent grading system can refer to the data management method for the intelligent grading objects in the above-mentioned bucket 1, and perform data management on the intelligent grading objects in bucket 2 and bucket 3 in turn, which will not be further described here. It can be understood that different hot and cold identification algorithms need to pay attention to different indicators. For example, the object indicator corresponding to the hot and cold identification algorithm S1 is the last access time. For example, it is necessary to count the access requests of objects in the buckets under different users, that is, to record the timestamp of the last access, and then convert the timestamp according to certain rules (for example, add 30 days to the last access timestamp as the next conversion time), and then flush the converted timestamp into the storage subsystem (for example, the kv storage system), so as to facilitate the subsequent use of the global secondary index of the kv storage system and analyze data indicators more quickly; the object indicator corresponding to the hot and cold identification algorithm S2 is the access frequency, so it is necessary to count the number of access requests per day. The object indicator corresponding to the hot and cold identification algorithm S2 is the access ratio of the life cycle. Then, it is necessary to provide hot and cold identification results (i.e., access heat) for different scenarios based on the heat recognition plug-in, so that the grading service in the subsequent intelligent grading system can start distributed jobs based on the algorithm calculation results and trigger the hot and cold migration of intelligent grading objects.
[0106] This shows that, unlike traditional data processing methods, the present embodiment does not require the storage system itself to analyze whether objects need to be migrated. Therefore, there is no need to add redundant fields to the storage system. Instead, an intelligent tiering system deployed outside the storage system regularly analyzes the logs within the storage system. This can greatly save the computing resources of the storage system and effectively improve the operating performance of the storage system. In addition, this data processing method of analysis by the intelligent tiering system can conduct a deeper analysis of the logs, obtaining data indicators that are more in line with business scenarios, so as to facilitate the subsequent improvement of the rationality of object migration.
[0107] Further, please refer to Figure 7, which is an interactive diagram of a method for data processing provided by an embodiment of the present application. As shown in Figure 7, the method can be jointly executed by an intelligent tiering system and a storage system, wherein the intelligent tiering system is used to perform intelligent tiering processing on business data in the storage system, and the storage system includes at least one storage path. The storage system may include an object storage system, a block storage system, a file storage system, and a storage system related to a database. The method may include at least steps S701-S705:
[0108] Step S701: The intelligent grading system obtains first business data.
[0109] In this embodiment of the present application, the storage path where the first business data is stored in the storage system may be referred to as the target storage path. The storage type of the first business data is the intelligent tiering type. For example, in an object storage system, the first business data here may be referred to as an intelligent tiering object (object 1 or object 4 shown in FIG. 3 above), and the storage bucket where the intelligent tiering object is located may be referred to as the target storage bucket.
[0110] It is understandable that the intelligent tiering system can obtain business data of the intelligent tiering type from the storage system through indexing, and then store the obtained business data in the storage subsystem of the intelligent tiering system through the regional tiering service to which the business data belongs, so as to facilitate subsequent data management. In other words, the first business data here can be consumed by the intelligent tiering system from the storage system (for example, the storage system shown in Figure 3), or can be obtained from the storage subsystem of the intelligent tiering system (for example, the storage subsystem 311 shown in Figure 3). The method of obtaining the first business data will not be limited here.
[0111] Step S702: The intelligent grading system obtains an access log of the first business data.
[0112] Specifically, the intelligent grading system obtains the logs in the storage system based on the big data analysis platform, and then can determine the access logs of the first business data from the obtained logs. Among them, the way in which the intelligent grading system obtains the access logs in the storage system may include the following two ways. For example, the first way is: the intelligent grading system sends a log consumption instruction to the storage system, so that the storage system pushes the logs within the data grading period (for example, one day or 12 hours) to the intelligent grading system. Among them, the data grading period can be dynamically adjusted according to the actual business situation (for example, the resource computing power and network parameters of the big data analysis platform, etc.), which will not be limited here. The second way is that the storage system actively pushes the logs in the storage system to the intelligent grading system on a regular basis based on the prior configuration.
[0113] Among them, the logs in the storage system may include user behavior logs (also known as user behavior tracks, traffic logs, etc.). Simply put, it is the behavioral data generated each time a user accesses business data in the storage system (for example, access, browsing, searching, and clicking, etc.). In the specific implementation of this application, when data related to user behavior is involved, when the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data shall comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0114] For example, when the storage system is an object storage system, if a user accesses an object in a storage bucket, the object storage service will generate a user behavior log based on the user's access information (e.g., user ID, access time, terminal ID, operation, etc.). The big data analysis platform in the intelligent grading system can consume user behavior logs, and its strong resource computing power can facilitate the cumulative processing of hundreds of billions of user access behavior records within a data grading cycle (e.g., one day or 12 hours, etc.). For example, when the time difference between the current timestamp and the previous grading timestamp reaches the data grading cycle, the big data analysis platform in the intelligent grading system can consume the logs in the storage system from the storage system.
[0115] In step S703 , the intelligent grading system analyzes the access log of the first business data according to the hot-cold identification algorithm configured by the user for the target storage path to obtain the access popularity of the first business data.
[0116] The access heat here refers to one of the following: hot, warm, or cold. The hot and cold identification algorithm can be stored in an intelligent tiering system, effectively reducing storage space usage in the storage system. The target storage path can include one or more of the following: an object storage bucket, a block storage volume, a file storage directory, and a database instance.
[0117] For example, when the storage system is an object storage system, the target storage path here can be the storage path corresponding to the target storage bucket where the first business data is located, or the storage subpath of the first business data in the target storage bucket (that is, the storage subpath corresponding to the folder of the first business data in the target storage bucket).
[0118] For ease of understanding, the hot and cold recognition algorithms in the embodiments of the present application can be taken as three examples, specifically including a first hot and cold recognition algorithm (for example, the hot and cold recognition algorithm S1 shown in FIG6 above), a second hot and cold recognition algorithm (for example, the hot and cold recognition algorithm S2 shown in FIG6 above) and a third hot and cold recognition algorithm (for example, the hot and cold recognition algorithm S3 shown in FIG6 above). Based on this, the hot and cold recognition algorithm here can be any one of the first hot and cold recognition algorithm, the second hot and cold recognition algorithm or the third hot and cold recognition algorithm.
[0119] The first hot and cold identification algorithm is used to identify hot and cold data based on the last access time. This algorithm's identification rules are relatively simple and applicable to most business scenarios, particularly big data scenarios, where objects not accessed for a long time need to be migrated to cooler storage tiers to save costs. Big data refers to data sets that cannot be captured, managed, and processed within a certain timeframe using conventional software tools. These are massive, high-growth, and diverse information assets that require new processing models to enhance decision-making, insight discovery, and process optimization. With the advent of the cloud era, big data has attracted increasing attention. Big data requires specialized technologies to effectively process large amounts of data within a timeframe. Technologies suitable for big data include massively parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the internet, and scalable storage systems.
[0120] The second hot / cold identification algorithm is used to identify hot / cold objects based on access frequency. This algorithm requires a certain period of learning. Specifically, it collects statistics and analyzes the number of times a batch of objects have been accessed over that period. When the access frequency statistics conform to a certain access trend distribution (for example, a normal distribution), the batch of objects is migrated. For example, if the object is a currently showing movie, its popularity changes in line with a normal distribution, that is, it first increases from low to high, and then decreases from high to low.
[0121] The third hot and cold identification algorithm is used to perform hot and cold identification based on the access ratio within the life cycle. The third hot and cold identification algorithm can classify objects created in different life cycles and is more suitable for business scenarios where the access behavior of objects in the same life cycle is relatively consistent. For example, when the object uploaded by the user is a short video, the access popularity of the short video will gradually decrease over time. For example, when the object uploaded by the user is a reference document, the access popularity of the reference document will remain high in a special time period (for example, graduation season).
[0122] It should be understood that if the hot / cold identification algorithm is the first hot / cold identification algorithm, the intelligent grading system can analyze the access log of the first business data to obtain the last access time of the first business data, calculate the target time interval between the last access time and the current time, and then compare the target time interval with the preset time interval to obtain the access popularity of the first business data. The preset time intervals here may include one or more. For ease of illustration, the embodiment of the present application may use two as an example, specifically including a first time interval and a second time interval, where the second time interval is greater than the first time interval. Both the first time interval and the second time interval can be dynamically adjusted based on business needs. For example, the second time interval can be 90 days, and the first time interval can be 30 days.
[0123] If the target time interval is less than or equal to the first time interval, the intelligent grading system can determine that the access popularity of the first business data is hot; if the target time interval is greater than the first time interval and less than the second time interval, the intelligent grading system can determine that the access popularity of the first business data is warm; if the target time interval is greater than or equal to the second time interval, the intelligent grading system can determine that the access popularity of the first business data is cold.
[0124] For example, if the first storage layer of the first business data (e.g., any one of the intelligent grading objects in bucket 1 shown in FIG6 ) is storage layer L1 shown in FIG6 , when the intelligent grading system determines that the target time interval between the current time and the last access time falls within the time period of 30-90 days, that is, the first business data has not been accessed by the user for 30 days, the access popularity of the first business data has changed from hot to warm. Optionally, when the intelligent grading system determines that the target time interval between the current time and the last access time is greater than 90 days, that is, the first business data has not been accessed by the user for 90 days, the access popularity of the first business data has directly changed from hot to cold.
[0125] Optionally, if the hot and cold identification algorithm is the second hot and cold identification algorithm, the intelligent grading system can analyze the access log of the first business data to obtain the number of times the first business data accesses the first business data within the data grading period (i.e., the target access frequency). Then, the intelligent grading system can obtain N historical access frequencies of the first business data within a preset time, where N is a positive integer. Each of the N historical access frequencies here is an access frequency counted before the data grading period. Furthermore, the intelligent grading system can determine the changing trends of the N historical access frequencies and the target access frequency, thereby obtaining the access heat of the first business data.
[0126] For ease of explanation, in the embodiment of the present application, the access popularity of the first business data before log analysis may be determined as the first access popularity, and the access popularity of the first business data after log analysis may be determined as the second access popularity.
[0127] If the trend of change indicates that it will decline subsequently, then when the first access popularity is not the lowest access popularity, the intelligent grading system may determine the access popularity lower than the first access popularity as the second access popularity; when the first access popularity is the lowest access popularity, the intelligent grading system may continue to determine the first access popularity as the second access popularity.
[0128] If the trend of change indicates that it will rise in the future, when the first access popularity is not the highest access popularity, the intelligent grading system can determine the access popularity higher than the first access popularity as the second access popularity; when the first access popularity is the highest access popularity, the intelligent grading system can continue to determine the first access popularity as the second access popularity.
[0129] If the change trend indicates that it will remain unchanged in the future, the intelligent grading system continues to determine the first access popularity as the second access popularity.
[0130] For ease of understanding, further reference is made to Table 1, which is a statistical table associated with access frequency provided in an embodiment of the present application. The statistical table may include H intelligent grading objects in a target bucket (e.g., bucket 2 shown in FIG. 6 ), and the access frequencies corresponding to each intelligent grading object in multiple grading periods. For ease of illustration, H here can be taken as an example of 3, and the access frequency here can be taken as an example of 8, as shown in Table 1:
[0131] Table 1
[0132] As shown in Table 1, if access frequency 8 is the target access frequency counted by the intelligent grading system in the data grading period of December 8, then access frequency 1 can be the access frequency counted in the historical grading period of December 1, access frequency 2 can be the access frequency counted in the historical grading period of December 2, access frequency 3 can be the access frequency counted in the historical grading period of December 3, and so on. Access frequency 7 can be the access frequency counted in the historical grading period of December 7.
[0133] It can be understood that the change trend of the first business data can be determined according to the distribution corresponding to the business scenario where the first business data is located. For example, the distribution here can be a normal distribution.
[0134] To facilitate understanding of the changing trends of Object 1 shown in Table 1, please refer to Figure 8, which is a changing trend graph for determining access popularity, provided in an embodiment of the present application. As shown in Figure 8, this changing trend graph is generated based on the access frequencies of Object 1 in Table 1. It can be used to clearly demonstrate the relationship between date and access frequency, facilitating observation of subsequent changes in access popularity of Object 1.
[0135] As shown in Figure 8, the changing trends of these eight access frequencies conform to a certain distribution (e.g., a normal distribution). That is, between December 1st and December 6th, the changing trend of object 1 gradually increased, while between December 6th and December 8th, the changing trend of object 1 showed a decreasing trend. In other words, the intelligent grading system can determine that the changing trend of object 1 first increased and then decreased. If the first access popularity of object 1 is not the lowest access popularity, the intelligent grading system can determine the access popularity lower than the first access popularity as the second access popularity.
[0136] For object 2 shown in Table 1 above, observing the historical access frequencies corresponding to eight consecutive historical classification periods shows that from December 1 to December 8, the change trend of object 2 is an upward trend. However, between December 7 and December 8, the change trend of object 2 increases significantly but the amplitude is not high. At this time, it can be determined that the change trend of object 2 also conforms to a certain normal distribution (for example, the upward trend of the normal distribution). Based on this, when the first access popularity of object 2 is not the highest access popularity, the intelligent classification system can directly determine the access popularity that is higher than the first access popularity as the second access popularity.
[0137] For object 3 shown in Table 1 above, observing the historical access frequencies corresponding to eight consecutive historical classification periods shows that object 3 was not accessed between December 1st and December 7th, while on December 8th, object 3's access frequency was 1, indicating that it was accessed once. If object 3's first access popularity is cold, then in this case, it is necessary to determine whether object 3 needs to be migrated back to the frequently accessed tier based on object 3's business scenario. For example, if object 3's business scenario is a video scenario, it can be assumed that object 3's changing trend conforms to the upward trend of a normal distribution. In this case, the intelligent classification system can determine the access popularity higher than the first access popularity as the second access popularity. If object 3's business scenario is a security scanning scenario (for example, a fixed weekly scan), it can be assumed that object 3's changing trend conforms to the distribution of security scanning scenarios. In this case, the intelligent classification system can continue to determine the first access popularity as the second access popularity. When object 3, which is cold data, is accessed, the second cold and hot identification algorithm can determine that object 3 does not need to be migrated back to the frequently accessed tier.
[0138] Optionally, if the hot / cold identification algorithm is the third hot / cold identification algorithm, then when the storage system is an object storage system, the first business data can be an intelligently graded object determined in the target storage bucket, whose object creation time falls within the target lifecycle. Furthermore, the intelligent grading system can analyze the access logs of the first business data to obtain an access ratio of the first business data within the target lifecycle, and then compare the access ratio with a target access threshold corresponding to the target lifecycle to determine the access popularity of the first business data.
[0139] The division of life cycle segments can be dynamically adjusted according to actual needs. For example, with 15 days as a segment, the object life cycle distribution can be divided into multiple time periods, such as within 15 days, 15-30 days, 30-45 days, 45-60 days, 60-75 days, 75-90 days, 90-105 days, 105-120 days, 120-135 days, 135-150 days, 150-165 days, 165-180 days, and more than 180 days.
[0140] For ease of understanding, please further refer to Table 2, which is a schematic table provided in an embodiment of the present application for determining the access ratio within a data classification cycle. For ease of explanation, the life cycle segments of the embodiment of the present application are only taken as an example of 3 segments, which may specifically include within 15 days, 15-30 days, and 30-45 days. If the data classification cycle of the intelligent classification system is one day, and the last object classification timestamp is timestamp 1 (for example, October 31 00:00:00), then when the current timestamp is timestamp 2 (for example, November 1 00:00:00), it means that the time difference between the current timestamp and the last classification timestamp reaches the data classification cycle. When the intelligent classification system determines that the hot and cold identification algorithm of the target storage bucket is the third hot and cold identification algorithm, it can determine the access ratio within the following life cycle segments. As shown in Table 2:
[0141] Table 2
[0142] It should be noted that the items shown in Table 2 are only a form of expression for reference. In actual business scenarios, other items can be established according to needs (for example, the object name within a certain object creation time period, the current storage layer, etc.). The embodiment of this application does not limit the specific form of Table 2.
[0143] The intelligent grading system can identify any lifecycle shown in Table 2 (taking the lifecycle period of less than 15 days as an example) as the target lifecycle. In this case, the first business data refers to the intelligently graded objects in the target bucket whose object creation time falls within the lifecycle period of October 17th to October 31st. The intelligent grading system can then analyze the access logs of the first business data to determine the number of object accesses within the data grading period (e.g., 10 as shown in Table 2) and the number of objects created within the target lifecycle (e.g., 50 as shown in Table 2). Based on the number of object accesses and the number of objects, the system can determine the access ratio (e.g., 20%) for the target lifecycle period of October 17th to October 31st.
[0144] Furthermore, the intelligent grading object may obtain a target access threshold corresponding to a target lifecycle, and obtain access popularity of the first business data by comparing the access ratio with the target access threshold corresponding to the target lifecycle.
[0145] Among them, the target access threshold can be determined from the access thresholds corresponding to M storage layers, and each of the M access thresholds is determined based on the media information of the corresponding storage layer, and the media information includes storage capacity, access cost, and storage cost. The access threshold here is used to indicate the minimum access ratio of the corresponding storage layer. Exemplarily, if the M storage layers are the three storage layers shown in Figure 6 above, then the M thresholds here can include three access thresholds, specifically including the access threshold corresponding to storage layer L1 (for example, 35%), the access threshold corresponding to storage layer L2 (for example, 10%), and the access threshold corresponding to storage layer L3 (for example, 5%).
[0146] For Table 2 above, if the target lifecycle is less than 15 days, the first business data here refers to the 50 objects created between October 17 and October 31. Since the access ratio for this target lifecycle is 20%, which is greater than the access threshold corresponding to storage tier L2, the intelligent tiering object can determine that the access popularity of these 50 objects is warm. If the target lifecycle is between 15 and 30 days, the first business data here refers to the 20 objects created between October 2 and October 16. Since the access ratio for this target lifecycle is 50%, which is greater than the access threshold corresponding to storage tier L1, the access popularity of these 20 objects is hot. If the target lifecycle is between 30 and 45 days, the first business data here refers to the 250 objects created between September 17 and October 1. Since the access ratio for this target lifecycle is 4%, which is greater than the access threshold corresponding to storage tier L3, the access popularity of these 250 objects is cold.
[0147] Optionally, the target access threshold here is determined from the access thresholds of multiple life cycles. For example, the intelligent grading system can obtain the second business data, and then analyze the access log of the second business data to obtain the access ratio of the second business data in at least one life cycle. Furthermore, the intelligent grading system can determine the access threshold of each life cycle separately based on the access ratio in at least one life cycle. The second business data here can be the business data stored in the target storage path, that is, the historical business data belonging to the same business scenario as the first business data.
[0148] For example, if the first business data is related to seasonal changes, the intelligent grading system can obtain the following access thresholds after analyzing the access log of the second business data, which may include: the access threshold in spring (for example, 50%), the access threshold in summer (for example, 90%), the access threshold in autumn (for example, 50%), and the access threshold in winter (for example, 10%). If the target life cycle belongs to summer, the target access threshold is the access threshold in summer. If the access ratio of the first business data in the target life cycle is less than or equal to the target access threshold, the second access heat of the first business data is determined to be an access heat lower than or equal to the first access heat; if the access ratio of the first business data in the target life cycle is greater than the target access threshold, the second access heat of the first business data is determined to be an access heat higher than or equal to the first access heat. Among them, the situation where the second access heat is equal to the first access heat is because the first access heat is the maximum value and cannot be changed.
[0149] Step S704: The intelligent grading system sends a data migration instruction to the storage system according to the access popularity of the first business data.
[0150] Specifically, the intelligent tiering system can determine the storage layer (i.e., the second storage layer) to which the first business data needs to be migrated based on the access popularity of the first business data (i.e., the second access popularity), and then determine whether it is necessary to send a data migration instruction to the storage system by comparing whether the first storage layer is consistent with the second storage layer. If the second storage layer is the same as the first storage layer, there is no need to send a data migration instruction. If the second storage layer is different from the first storage layer, it is necessary to generate a data migration instruction for migrating the first business data from the first storage layer to the second storage layer, and send the data migration instruction to the storage system.
[0151] Step S705: The storage node of the storage system performs data migration on the first business data.
[0152] Specifically, the storage nodes of the storage system can migrate the first business data from the first storage layer to the second storage layer based on the data migration instruction, thereby achieving migration of the first business data.
[0153] It can be seen from this that the storage system can be a system that can support storage tiering, and the intelligent tiering system belongs to an out-of-band system, that is, another system deployed outside the storage system, which can support a richer set of hot and cold identification algorithms (including algorithms for hot and cold identification based on the last access time, algorithms for hot and cold identification based on access frequency, and algorithms for hot and cold identification based on life cycle access ratio). In this way, when the hot and cold identification algorithm corresponding to the target storage path is used for hot and cold identification, the log can be analyzed in a deeper level, so that the access heat obtained in the end is more accurate, and thus the automatic tiering of hot and cold data can be realized more intelligently. When the identified second storage layer is different from the current storage layer of the first business data (i.e., the first storage layer), the intelligent tiering system realizes the automatic tiering of hot and cold data by sending a data migration instruction to the storage system. In addition, the embodiment of the present application does not need to add redundant fields to the storage system, nor does it need the storage system itself to analyze whether the object needs to be migrated. Instead, an intelligent tiering system deployed outside the storage system analyzes the log, which can effectively reduce the impact on the storage system and thus improve the operating performance of the storage system.
[0154] The following describes the method provided in the embodiment of the present application from the perspective of a single system in conjunction with the accompanying drawings.
[0155] Please refer to Figure 9, which is a schematic diagram of a method for data processing provided by an embodiment of the present application. As shown in Figure 9, the method can be executed by an intelligent tiering system, which is independent of the storage system and is used to perform intelligent tiering processing on business data in the storage system, where the storage system includes at least one storage path. The method may include at least steps S901-S904:
[0156] Step S901: Acquire first business data.
[0157] The first business data may be stored in a target storage path in a storage system, and the storage type of the first business data is an intelligent tiering type. The storage system may include an object storage system, a block storage system, a file storage system, and a storage system related to a database.
[0158] Step S902: Obtain access logs of first business data.
[0159] The access log of the first business data is obtained by the intelligent grading system from the storage system, and the access log refers to the behavior data (access, browsing, searching, clicking, etc.) generated when the user accesses the first business data.
[0160] Step S903 : Analyze the access log of the first business data according to the hot / cold identification algorithm configured by the user for the target storage path to obtain the access popularity of the first business data.
[0161] The target storage path may include one or more of the following: a bucket for object storage, a volume for block storage, a directory for file storage, and an instance of a database. Specifically, the intelligent grading system may call a big data analysis platform to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the data index of the first business data, and then call a heat identification plug-in to determine the access heat of the first business data according to the data index. The access heat here refers to one of the heats among hot, warm, or cold.
[0162] The hot and cold identification algorithm may be any one of a first hot and cold identification algorithm, a second hot and cold identification algorithm, and a third hot and cold identification algorithm. The first hot and cold identification algorithm is used to perform hot and cold identification based on the last access time; the second hot and cold identification algorithm is used to perform hot and cold identification based on the access frequency; and the third hot and cold identification algorithm is used to perform hot and cold identification based on the access ratio within the life cycle.
[0163] When the storage system is an object storage system, if the hot and cold identification algorithm is the first hot and cold identification algorithm or the second hot and cold identification algorithm, the first business data here can be any intelligent grading object in the target storage bucket. If the hot and cold identification algorithm is the third hot and cold identification algorithm, the first business data can be an intelligent grading object in the target storage bucket whose determined object creation time belongs to the target life cycle. The target life cycle here can be any life cycle segment of the multiple life cycle segments divided by the embodiment of the present application (for example, any life cycle segment in Table 2 above).
[0164] In an optional implementation, if the hot and cold identification algorithm is the first hot and cold identification algorithm, the data indicator here is the last access time corresponding to the first business data. At this time, the intelligent grading system needs to determine the target time interval between the current time and the last access time, and then obtain the access popularity of the first business data based on the target time interval and the preset time interval.
[0165] In another optional implementation, if the hot / cold identification algorithm is the second hot / cold identification algorithm, the data metric here is the target access frequency, which is the number of times the first business data is accessed within the data classification period. In this case, the intelligent classification system needs to obtain N historical access frequencies of the first business data within a preset time period, where N is a positive integer. Furthermore, the intelligent classification system can determine the changing trend of these N historical access frequencies and the target access frequency to obtain the access popularity of the first business data.
[0166] In another optional implementation, if the hot / cold identification algorithm is the third hot / cold identification algorithm, the data metric here is the access ratio of the target lifecycle. The intelligent grading system needs to determine the target access threshold corresponding to the target lifecycle and then compare the target access threshold with the access ratio to determine the access popularity of the first business data.
[0167] Step S904: Send a data migration instruction to the storage system according to the access popularity of the first business data.
[0168] The data migration instruction is used to instruct the storage node of the storage system to migrate the first business data from the first storage layer to the second storage layer.
[0169] The specific implementation of steps S901-S904 can be found in the description of steps S701-S704 in the embodiment corresponding to FIG7 above, and will not be repeated here.
[0170] In this embodiment, the intelligent tiering system is an out-of-band management system that implements a plug-in hot / cold identification algorithm. This significantly reduces the impact on the original storage system and provides highly cost-effective storage services. Furthermore, this plug-in hot / cold identification algorithm management can address diverse business scenarios, broadening its application scope.
[0171] Further, referring to Figure 10, Figure 10 is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. As shown in Figure 10, the data processing device 1 may include at least one of an acquisition module 1001, a processing module 1002, and a sending module 1003. These modules can perform the response functions of the various devices in the above method embodiments.
[0172] In a possible implementation, the data processing device 1 can be used to implement the functions of the intelligent grading system in FIG. 7 .
[0173] Specifically, the acquisition module 1001 is used to obtain the first business data, and the first business data is stored in the target storage path in the storage system, wherein the intelligent grading system is used to perform intelligent grading processing on the business data in the storage system, the storage system includes at least one storage path, and the storage type of the first business data is an intelligent grading type; the acquisition module 1001 is also used to obtain the access log of the first business data; the processing module 1002 is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path, and obtain the access heat of the first business data; the sending module 1003 is used to send a data migration instruction to the storage system according to the access heat of the first business data, so that the storage node of the storage system performs data migration on the first business data.
[0174] In one implementation, the hot and cold identification algorithms are stored in an intelligent grading system.
[0175] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on the access ratio of the life cycle. The processing module 1002 is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: the processing module 1002 is specifically used to analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle; the processing module 1002 is specifically used to compare the access ratio with the target access threshold corresponding to the target life cycle to obtain the access heat of the first business data.
[0176] In one implementation, the processing module 1002 is specifically used to analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle, and also includes: the processing module 1002 is specifically used to obtain the second business data, and the second business data is the business data stored in the target storage path; the processing module 1002 is specifically used to analyze the access log of the second business data to obtain the access ratio of the second business data in at least one life cycle; the processing module 1002 is specifically used to determine the access threshold of each life cycle according to the access ratio in at least one life cycle.
[0177] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on access frequency, and a processing module 1002 is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module 1002, specifically used to analyze the access log of the first business data to obtain the target access frequency of the first business data; a processing module 1002, specifically used to obtain N historical access frequencies of the first business data within a preset time, where N is a positive integer; a processing module 1002, specifically used to determine the changing trends of the N historical access frequencies and the target access frequency to obtain the access heat of the first business data.
[0178] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on the last access time, and a processing module 1002 is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module 1002, specifically used to analyze the access log of the first business data to obtain the last access time of the first business data; a processing module 1002, specifically used to calculate the target time interval between the last access time and the current time; a processing module 1002, specifically used to compare the target time interval and the preset time interval to obtain the access heat of the first business data.
[0179] In one implementation, the target storage path includes one or more of the following: a bucket of object storage, a volume of block storage, a directory of file storage, and an instance of a database.
[0180] The specific implementation of the acquisition module 1001, the processing module 1002 and the sending module 1003 can be found in the description of steps S701 to S705 in the embodiment corresponding to FIG7 , which will not be further described here. In addition, the description of the beneficial effects of adopting the same method will not be repeated either.
[0181] The modules in the data processing device 1 can be implemented in software or hardware. For example, the implementation of the acquisition module 1001 will be described below, taking the acquisition module 1001 as an example. Similarly, the implementation of the processing module 1002 and the sending module 1003 can refer to the implementation of the acquisition module 1001.
[0182] As an example of a software functional unit, the acquisition module 1001 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the acquisition module 1001 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0183] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0184] As an example of a hardware functional unit, acquisition module 1001 may include at least one computing device, such as a server. Alternatively, acquisition module 1001 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0185] The multiple computing devices included in acquisition module 1001 can be distributed in the same region or in different regions. The multiple computing devices included in acquisition module 1001 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in acquisition module 1001 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0186] It should be noted that, in other embodiments, the acquisition module 1001 can be used to execute any step in the data processing method, the processing module 1002 can be used to execute any step in the data processing method, and the sending module 1003 can be used to execute any step in the XXX data processing method. The steps that the acquisition module 1001, the processing module 1002, and the sending module 1003 are responsible for implementing can be specified as needed. The full functions of the data processing device 1 are realized by respectively implementing different steps in the data processing method through the acquisition module 1001, the processing module 1002, and the sending module 1003.
[0187] Further, please refer to Figure 11, which is a schematic diagram of the structure of a computing device provided by this application. As shown in Figure 11, computing device 1100 includes: a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. Processor 1104, memory 1106, and communication interface 1108 communicate with each other via bus 1102. Computing device 1100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1100.
[0188] Bus 1102 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG11 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1102 may include a path for transmitting information between various components of computing device 1100 (e.g., memory 1106, processor 1104, and communication interface 1108).
[0189] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0190] The memory 1106 may include volatile memory, such as random access memory (RAM). The memory 1106 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0191] Memory 1106 stores executable program code. Processor 1104 executes the executable program code to implement the functions of acquisition module 1001, processing module 1002, and sending module 1003 in the aforementioned data processing device 1, thereby implementing the data processing method. In other words, memory 1106 stores instructions for executing the data processing method.
[0192] The communication interface 1108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.
[0193] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0194] Further, referring to Figure 12, Figure 12 is a schematic diagram of the structure of a computing device cluster provided by this application. As shown in Figure 12, the computing device cluster includes at least one computing device 1200. The memory 1206 of one or more computing devices 1200 in the computing device cluster may store the same instructions for executing the data processing method.
[0195] In some possible implementations, the memory 1206 of one or more computing devices 1200 in the computing device cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 1200 can jointly execute the instructions for executing the data processing method.
[0196] It should be noted that the memory 1206 in different computing devices 1200 in the computing device cluster can store different instructions, each used to execute part of the functions of the data processing device 1. In other words, the instructions stored in the memory 1206 in different computing devices 1200 can implement the functions of one or more of the above-mentioned acquisition module 1001, processing module 1002, and sending module 1003 in the aforementioned data processing device 1.
[0197] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. Further, please refer to Figure 13, which is a structural diagram of a network connection between computing devices provided by the present application. As shown in Figure 13, two computing devices 1300A and 1300B are connected via a network. Specifically, the connection to the network is made through a communication interface in each computing device. In this type of possible implementation, the memory 1306 in the computing device 1300A stores instructions for executing the functions of the above-mentioned acquisition module 1001 and processing module 1002 in the aforementioned data processing device 1. At the same time, the memory 1306 in the computing device 1300B stores instructions for executing the functions of the sending module 1003 in the data processing device 1.
[0198] The connection method between the computing device clusters shown in Figure 13 can be considered to be that the data processing method provided in this application needs to obtain data (for example, the first business data and its access log) and analyze the access log, so it is considered to entrust the functions implemented by the acquisition module 1001 and the processing module 1002 to the computing device 1300A for execution.
[0199] It should be understood that the functionality of the computing device 1300A shown in FIG13 may also be implemented by multiple computing devices 1300. Similarly, the functionality of the computing device 1300B may also be implemented by multiple computing devices 1300.
[0200] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the data processing method or data processing method described in the above embodiments.
[0201] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the data processing method or data processing method in the above embodiment.
[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method, characterized in that: The method is applied to an intelligent grading system, the intelligent grading system is used to perform intelligent grading processing on business data in a storage system, the storage system includes at least one storage path, and the method includes: Acquire first business data, where the first business data is stored in a target storage path in the storage system, wherein the storage type of the first business data is an intelligent grading type; Obtaining an access log of the first business data; Analyze the access log of the first business data according to the hot-cold identification algorithm configured by the user for the target storage path to obtain the access popularity of the first business data; According to the access popularity of the first business data, a data migration instruction is sent to the storage system, so that the storage node of the storage system performs data migration on the first business data.
2. The method according to claim 1, characterized in that The hot and cold identification algorithm is stored in the intelligent grading system.
3. The method according to claim 1 or 2, characterized in that: The hot-cold identification algorithm is used to perform hot-cold identification based on the access ratio of the life cycle. The hot-cold identification algorithm configured by the user for the target storage path analyzes the access log of the first business data to obtain the access heat of the first business data, including: Analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle; The access ratio is compared with a target access threshold corresponding to the target life cycle to obtain access popularity of the first business data.
4. The method according to claim 3, characterized in that The analyzing the access log of the first business data to obtain the access ratio of the first business data in the target life cycle also includes: Acquire second business data, where the second business data is business data stored in the target storage path; Analyze the access log of the second business data to obtain an access ratio of the second business data in at least one life cycle; An access threshold for each life cycle is determined according to the access ratio in the at least one life cycle.
5. The method according to claim 1 or 2, characterized in that: The hot and cold identification algorithm is used to perform hot and cold identification based on access frequency, and the hot and cold identification algorithm configured by the user for the target storage path analyzes the access log of the first business data to obtain the access heat of the first business data, including: Analyze the access log of the first business data to obtain the target access frequency of the first business data; Obtain N historical access frequencies of the first business data within a preset time, where N is a positive integer; Determine the change trends of the N historical access frequencies and the target access frequency to obtain the access popularity of the first business data.
6. The method according to claim 1 or 2, characterized in that: The hot and cold identification algorithm is used to perform hot and cold identification based on the last access time, and the hot and cold identification algorithm configured by the user for the target storage path analyzes the access log of the first business data to obtain the access heat of the first business data, including: Analyze the access log of the first business data to obtain the last access time of the first business data; Calculate the target time interval between the last access time and the current time; The target time interval and the preset time interval are compared to obtain the access popularity of the first business data.
7. The method according to any one of claims 1 to 6, characterized in that: The target storage path includes one or more of the following: Object storage buckets, block storage volumes, file storage directories, and database instances.
8. A data processing device, characterized in that: include: an acquisition module, configured to acquire first business data, wherein the first business data is stored in a target storage path in a storage system, wherein the intelligent grading system is configured to perform intelligent grading processing on the business data in the storage system, wherein the storage system includes at least one storage path, and the storage type of the first business data is an intelligent grading type; The acquisition module is further used to acquire an access log of the first business data; a processing module, configured to analyze the access log of the first business data according to a hot-cold identification algorithm configured by a user for the target storage path, and obtain the access popularity of the first business data; A sending module is used to send a data migration instruction to the storage system according to the access popularity of the first business data, so that the storage node of the storage system performs data migration on the first business data.
9. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.
10. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7.
11. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A librados-based distributed NFS system and a construction method thereof
CN109783438A
Data migration deployment method based on access heat
CN110008199A
Data processing method and device, data access method and device and computer equipment
CN112100293A
Hierarchical storage method and device for file data
CN116303280A
Data storage method and device, equipment and storage medium
CN116860177A
Cited By
Data management method based on distributed storage system
CN120255824A
Computer data storage method and system based on Internet of Things
CN120353800A
Data caching method, system and device and storage medium
CN120583063A
Cold and hot data storage optimization method based on large model
CN120596490A
Data dynamic migration method and device, electronic equipment, storage medium and program
CN120723166A