Data processing method and device
Through an intelligent hierarchical system, analyzing the business data access logs in the storage system, using the hot and cold recognition algorithm to determine the access heat, and perform data migration, solving the pressure problem of traditional data processing methods on the storage system, and achieving efficient data hierarchy and storage performance improvements.
Patent Information
- Application Number
- CN202410171571.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-02-06
- Publication Date
- 2025-05-23
AI Technical Summary
Traditional data processing methods record data access time by scanning the metadatabase regularly in the storage system, resulting in an increase in the pressure of the metadatabase, affecting other business operations of the main IO path, and increasing the risk of failure spread, affecting the normal operation of the storage system.
The intelligent hierarchical system is adopted to obtain the access log of business data, use the user-configured hot and cold recognition algorithm to analyze the access heat, and send data migration instructions to the storage system to realize automatic data hierarchy.
It reduces the impact on the storage system, reduces storage space usage, saves storage costs, improves the operating performance of the storage system, and realizes smarter automatic layering of hot and cold data.
Smart Images

Figure CN120029529A_ABST
Abstract
Description
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 22, 2023, with application number 202311566468.2 and application name “A method, device and other equipment for data processing”, all contents of which are incorporated by reference in this application. Technical Field
[0002] The present application relates to the field of storage, and in particular to a data processing method and device. Background Art
[0003] Users use cloud storage systems to manage massive amounts of data. To further manage storage costs, they need to perform storage tiering based on the frequency of object use. For example, frequently accessed data can be placed in a frequently accessed tier to reduce the overhead of access requests, and infrequently accessed data can be placed in a low-frequency or archived or even deep archived access tier to reduce storage costs.
[0004] The traditional data processing method is to periodically scan the data access time recorded in the metadata database in the background of the storage system to switch the storage layer according to the change of access time. This data processing method will increase the pressure on the metadata database, affect other business operations of the main input / output (IO) path, and even increase the risk of fault propagation, resulting in service unavailability. In other words, this data processing method will affect the normal operation of the storage system. Summary of the invention
[0005] The present application provides a data processing method and device for improving the operating performance of a storage system.
[0006] In order to achieve the above-mentioned purpose, this application adopts the following technical solution.
[0007] In a first aspect, an embodiment of the present application provides a data processing method, which is applied to an intelligent grading system, the intelligent grading system is used to perform intelligent grading processing on business data in a storage system, the storage system includes at least one storage path, the method includes: obtaining first business data, the first business data is stored in a target storage path in the storage system, wherein the storage type of the first business data is an intelligent grading type; obtaining an access log of the first business data; analyzing the access log of the first business data according to a hot and cold identification algorithm configured by a user for the target storage path to obtain access heat of the first business data; and sending a data migration instruction to the storage system according to the access heat of the first business data, so that the storage node of the storage system migrates data for the first business data.
[0008] In the above method, the storage system can be a system that can support storage tiering, and the intelligent tiering system belongs to an out-of-band system, that is, another system deployed outside the storage system, which can support richer hot and cold identification algorithms (including algorithms for hot and cold identification based on the last access time, algorithms for hot and cold identification based on the access frequency, and algorithms for hot and cold identification based on the access ratio of the life cycle). In this way, when the hot and cold identification algorithm corresponding to the target storage path is used for hot and cold identification, the log can be analyzed more deeply, so that the final access heat is more accurate, and then the automatic tiering of hot and cold data can be realized more intelligently. In addition, compared with the traditional data processing method of recording redundant fields in the storage system, the intelligent tiering system uses the log of the first business data for analysis, without adding redundant fields in the storage system, which not only reduces the storage space occupied, thereby saving the user's storage cost, but also does not need the storage system itself to analyze whether the object needs to be migrated, thereby effectively reducing the impact on the storage system, thereby improving the operating performance of the storage system.
[0009] In one implementation, the hot and cold identification algorithm is stored in an intelligent grading system.
[0010] In the above implementation, the hot and cold identification algorithm is not stored in the storage system, but in another system (i.e., the intelligent grading system) independent of the storage system. This can greatly reduce the space occupied in the storage system and reduce storage costs.
[0011] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on an access ratio of a life cycle. According to the hot and cold identification algorithm configured by a user for a target storage path, an access log of the first business data is analyzed to obtain the access heat of the first business data, including: analyzing the access log of the first business data to obtain the access ratio of the first business data in the target life cycle; comparing the access ratio with a target access threshold corresponding to the target life cycle to obtain the access heat of the first business data.
[0012] In the above implementation, the access heat is determined based on the data indicator of the first business data, and the hot and cold identification algorithms of different storage paths are different, and the obtained data indicators are also different, which can be more in line with the business scenario of the first business data. When the hot and cold identification algorithm of the target storage path is used for hot and cold identification based on the access ratio of the life cycle, the first business data here is one or more business data whose object creation time belongs to the target life cycle. This means that the embodiment of the present application can analyze multiple business data belonging to the target life cycle segment together, so as to facilitate subsequent batch data migration and improve data migration efficiency. In addition, through the data indicator of the access ratio of the target life cycle, the business scenario that fits the first business data can be obtained, so that the final access heat is more accurate.
[0013] In one implementation, analyzing the access log of the first business data to obtain the access ratio of the first business data in the target life cycle includes: acquiring the second business data, where the second business data is the business data stored in the target storage path; analyzing the access log of the second business data to obtain the access ratio of the second business data in at least one life cycle; and determining the access threshold of each life cycle according to the access ratio in at least one life cycle.
[0014] In the above implementation, since the second business data is the business data stored in the target storage path, and the access thresholds of different life cycles are related to their corresponding access ratios, this means that the embodiment of the present application can classify business data with different life cycles according to a batch of business data with similar life cycles, thereby realizing automatic migration of business data in a special business scenario (for example, a scenario where access behaviors of objects with the same life cycle are relatively consistent).
[0015] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on access frequency. According to the hot and cold identification algorithm configured by the user for the target storage path, the access log of the first business data is analyzed to obtain the access heat of the first business data, including: analyzing the access log of the first business data to obtain the target access frequency of the first business data; obtaining N historical access frequencies of the first business data within a preset time, where N is a positive integer; determining the changing trends of the N historical access frequencies and the target access frequency to obtain the access heat of the first business data.
[0016] In the above implementation, when the hot and cold identification algorithm of the target storage path is used to perform hot and cold identification based on the access frequency, this means that the access popularity of the first business data conforms to a certain changing pattern. Therefore, the intelligent grading system not only needs to analyze and obtain the target access frequency of the first business data, but also needs to obtain N historical access frequencies, and then accurately estimate the access popularity of the first business data through the changing trends of these access frequencies, so as to realize the automatic migration of business data in another special business scenario (for example, a scenario where the access frequency has a changing pattern).
[0017] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on the last access time. According to the hot and cold identification algorithm configured by the user for the target storage path, the access log of the first business data is analyzed to obtain the access heat of the first business data, including: analyzing the access log of the first business data to obtain the last access time of the first business data; calculating the target time interval between the last access time and the current time; comparing the target time interval and the preset time interval to obtain the access heat of the first business data.
[0018] In the above implementation, when the hot and cold identification algorithm of the target path is used for hot and cold identification based on the last access time, this algorithm for determining the access heat by comparing the target time interval and the preset time interval is relatively simple, not only has high identification efficiency, but also can be applied to most business scenarios.
[0019] In one implementation, the target storage path includes one or more of the following: a bucket of object storage, a volume of block storage, a directory of file storage, and an instance of a database.
[0020] In the above implementation, the target storage path includes one or more of the above, which means that the storage system in the embodiment of the present application does not specifically refer to a certain storage system, but can be widely applied to various storage systems, such as object storage systems, block storage systems, file storage systems, and database-related systems, so as to provide extremely cost-effective storage services in various types of storage systems.
[0021] In the second aspect, the embodiment of the present application provides a data processing device for implementing a method as described in the first aspect or any one of the implementation methods of the first aspect. Specifically, the device includes: an acquisition module for acquiring first business data, the first business data is stored in a target storage path in a storage system, wherein an intelligent grading system is used to perform intelligent grading processing on business data in the storage system, the storage system includes at least one storage path, and the storage type of the first business data is an intelligent grading type; the acquisition module is also used to acquire an access log of the first business data; a processing module is used to analyze the access log of the first business data according to a hot and cold identification algorithm configured by a user for the target storage path, and obtain the access heat of the first business data; a sending module is used to send a data migration instruction to the storage system according to the access heat of the first business data, so that the storage node of the storage system performs data migration on the first business data.
[0022] In one implementation, the hot and cold identification algorithm is stored in an intelligent grading system.
[0023] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on an access ratio of a life cycle, and a processing module is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module, specifically used to analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle; a processing module, specifically used to compare the access ratio with the target access threshold corresponding to the target life cycle to obtain the access heat of the first business data.
[0024] In one implementation, a processing module is specifically used to analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle, and also includes: a processing module, specifically used to obtain the second business data, the second business data is the business data stored in the target storage path; a processing module, specifically used to analyze the access log of the second business data to obtain the access ratio of the second business data in at least one life cycle; a processing module, specifically used to determine the access threshold of each life cycle according to the access ratio in at least one life cycle.
[0025] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on access frequency, and a processing module is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module, specifically used to analyze the access log of the first business data to obtain the target access frequency of the first business data; a processing module, specifically used to obtain N historical access frequencies of the first business data within a preset time, where N is a positive integer; a processing module, specifically used to determine the changing trends of the N historical access frequencies and the target access frequency to obtain the access heat of the first business data.
[0026] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on the last access time, and a processing module is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module, specifically used to analyze the access log of the first business data to obtain the last access time of the first business data; a processing module, specifically used to calculate the target time interval between the last access time and the current time; a processing module, specifically used to compare the target time interval and the preset time interval to obtain the access heat of the first business data.
[0027] In one implementation, the target storage path includes one or more of the following: a bucket of object storage, a volume of block storage, a directory of file storage, and an instance of a database.
[0028] In a third aspect, an embodiment of the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory. The processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the data processing method in the first aspect or any possible implementation of the first aspect.
[0029] In a fourth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, enables the computing device cluster to execute the data processing method in the first aspect or any possible implementation of the first aspect.
[0030] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, including computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the data processing method in the first aspect or any possible implementation of the first aspect.
[0031] The technical effects produced by any implementation method in the above-mentioned second to fifth aspects and each aspect can refer to the above-mentioned first aspect and the corresponding implementation method in the first aspect, and the repetitive parts will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a system architecture diagram corresponding to a storage system provided in an embodiment of the present application;
[0033] Figure 2 is a system architecture diagram corresponding to an object storage system provided in an embodiment of the present application;
[0034] Figure 3 A system schematic diagram corresponding to a data processing system provided in an embodiment of the present application;
[0035] Figure 4 This is a schematic diagram of an interface for configuring a storage bucket provided in an embodiment of the present application. Figure 1 ;
[0036] Figure 5 This is a schematic diagram of an interface for configuring a storage bucket provided in an embodiment of the present application. Figure 2 ;
[0037] Figure 6 This is a schematic diagram of a scenario of performing data migration in an object storage system provided by an embodiment of the present application;
[0038] Figure 7 It is an interactive diagram of a method for performing data processing provided in an embodiment of the present application;
[0039] Figure 8 It is a change trend graph for determining access popularity provided by an embodiment of the present application;
[0040] Fig. 9 A schematic diagram of a method for performing data processing provided in an embodiment of the present application;
[0041] Fig.10 is a structural schematic diagram of a data processing device provided in an embodiment of the present application;
[0042] Fig.11 It is a structural schematic diagram of a computing device provided by the present application;
[0043] Fig.12 It is a structural diagram of a computing device cluster provided by this application;
[0044] Fig.13 It is a structural diagram of a network connection between computing devices provided by the present application. DETAILED DESCRIPTION
[0045] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. Among them, in order to facilitate the clear description of the technical solutions in the embodiments of the present application, in the embodiments of the present application, the words "first", "second" and the like are used to distinguish the same items or similar items with substantially the same functions and effects. It can be understood by those skilled in the art that the words "first", "second" and the like do not limit the quantity and execution order, and the words "first", "second" and the like do not limit necessarily different. At the same time, in the embodiments of the present application, the words "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design solutions. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.
[0046] To facilitate understanding of the technical solutions provided in the embodiments of the present application, the relevant terms involved in the embodiments of the present application are first introduced:
[0047] 1. Storage system:
[0048] A storage system is a system composed of various storage nodes for storing programs and data, control components, and equipment (hardware) and algorithms (software) for managing information adjustments. It can be broadly divided into object storage systems, block storage systems, and file storage systems.
[0049] 2. Object storage system:
[0050] Its main operation object is the object. Object storage is generally manifested in the form of a universally unique identifier (UUID). Data and metadata are packaged together as a whole object in a large pool. It can be understood that the object storage system is a system for providing object storage service (OBS). The basic components of OBS are buckets (storage buckets) and objects. Buckets are containers for storing objects in OBS. Each bucket has its own storage category, access rights, and region. Users locate buckets on the Internet through the bucket's access domain name. Objects are the basic unit of data storage in OBS. An object is actually a collection of a file's data and its related attribute information, which can specifically include three parts: key value, metadata, and data.
[0051] Among them, the key value can be a character sequence used to uniquely represent the object name. The character sequence is unique in the current bucket, that is, each object in a bucket has a unique object key value; metadata, that is, the description information of the object, including system metadata and user metadata, which are uploaded to OBS in the form of key-value pairs; data, that is, the data content of the file.
[0052] 3. Intelligent grading system (also known as out-of-band system)
[0053] The intelligent grading system is used to provide a plug-in hot and cold identification algorithm to analyze the logs and realize the automatic stratification of hot and cold data. The intelligent grading system and the storage system here can be deployed on the same device (for example, a server), and the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. The embodiment of the present application will not limit the number of servers.
[0054] 4. Garbage Collection (GC):
[0055] Garbage collection is a memory management mechanism. When the memory is no longer needed, it needs to be released to make room for storage. Garbage collection can be divided into two categories: memory-oriented and external memory-oriented. The memory-oriented garbage collection mechanism is integrated into high-level programming languages such as Java, Go, and C#, while the external memory-oriented garbage collection mechanism is widely used in solid-state drives (SSDs).
[0056] To facilitate understanding of the technical solutions provided by the embodiments of the present application, the relevant technologies of the embodiments of the present application are briefly introduced:
[0057] See also Figure 1 , Figure 1 1 is a system architecture diagram corresponding to a storage system provided in an embodiment of the present application. Figure 1 As shown, the storage system 10 may be a distributed storage system, and may specifically include a client cluster 110, a server 120, and a storage node cluster 130. The client cluster 110 may include one or more clients, and the number of clients is not limited. Figure 1 As shown, the client cluster 110 may specifically include client 111, client 112, client 113 and client 114. Each client in the client cluster 110 may respectively establish a network connection with the server 120, so that each client may exchange data with the server 120 through the network connection.
[0058] It is understandable that each client in the client cluster 110 can be a terminal device with a target application (e.g., an application for accessing business data) installed, and the terminal device can include a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a car terminal, a smart TV, and other smart terminals with data processing functions. The target application can be a social application, a multimedia application (e.g., a video application), an entertainment application (e.g., a game application), an information flow application, an educational application, a live broadcast application, etc. Among them, the target application here can be an independent application or an embedded sub-application (e.g., a small program, etc.) integrated in a certain application, which will not be limited here.
[0059] The server 120 may be a server corresponding to the target application. The server 120 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The embodiment of the present application does not limit the number of servers. Figure 1 As shown, the server 120 may also establish a network connection with each storage node in the storage node cluster 130 , so that the server 120 can manage or access any storage node in the storage node cluster 130 .
[0060] Among them, the network connection in the embodiments of the present application does not limit the connection method. It can be directly or indirectly connected through wired communication, directly or indirectly connected through wireless communication, or through other methods, which are not limited in the present application.
[0061] The storage node cluster 130 may include one or more storage nodes. The number of storage nodes is not limited here. Each storage node may be deployed in multiple data centers. The multiple data centers may be located in the same area or in different areas. This application does not limit this. Each data center may include one or more physical devices with storage functions, such as servers, mobile phones, tablet computers, read-only memory (ROM), flash memory, hard disk drive (HDD) or solid state drive (SSD), etc. This application is for physical devices with storage functions.
[0062] For ease of understanding, the storage system in the embodiment of the present application can be taken as an example of an object storage system to illustrate how to analyze the log of the first business data (i.e., business data of the storage type of the intelligent grading type) according to the intelligent grading system, thereby realizing automatic stratification of hot and cold data. Figure 2 , Figure 2 1 is a system architecture diagram corresponding to an object storage system provided in an embodiment of the present application. Figure 2 As shown, the object storage system 20 may include a client 200 , a service layer 210 , and a metadata layer 220 .
[0063] The client 200 may be the above Figure 1 Any client in the client cluster shown in the figure may also have N applications deployed on the client 200, which may specifically include application 201, application 202, ..., application 20N. The user corresponding to the client 200 may use an account on the client 200 to log in to the cloud management platform, create a storage bucket, configure a bucket name, and obtain a bucket domain name through the cloud management platform. After detecting the user's operation (such as creating a bucket, configuring a bucket name, hot and cold identification algorithms, etc.), the cloud management platform may issue a creation instruction, which includes information such as the bucket name and bucket domain name, and notify the object storage service node to create a bucket and save information such as the bucket name and bucket domain name.
[0064] After creating a storage bucket, the user can also access the bucket domain name by operating the client 200, locate the corresponding bucket, upload data (such as uploading an object) and download data (such as downloading an object) in the bucket. Taking the application 201 as an example, the client 200 can respond to the user's operation on the application 201 and generate an object upload request for a certain object to upload the object to a storage bucket created by it.
[0065] It is understandable that users can create one or more buckets. When there are many business scenarios, a bucket can be created for a business scenario, and objects can be stored in buckets that match their business scenarios in subsequent use. For example, in a multimedia scenario, objects associated with the entertainment type can be stored in bucket 1, objects associated with the sports type can be stored in bucket 2, and objects associated with the pet type can be stored in bucket 3.
[0066] The service layer 210 can be used to interact with the client 200 and provide foreground services to the outside. The foreground services can be services for the client 200 to request to read or write data. The service layer 210 may include lifecycle management 211 and background tasks (e.g., garbage collection 212). The lifecycle management 211 is used to scan the metadata layer 220 to calculate whether an object needs to be migrated based on the last access time; the garbage collection 212 is used to regularly clean up redundant data and release storage space.
[0067] The metadata layer 220 includes bucket metadata 221 and object metadata 222. For clusters storing metadata, in order to adapt to the growing demand for the number of metadata entries, partitioning technology is usually used to dynamically split the metadata, and each partition (also called partition) after the split manages different data entries. Bucket metadata 221 can not only be used to record partitioning rules, but also record metadata information of buckets, such as the region to which the bucket belongs; object metadata 222 can record metadata of objects, such as object attributes, size, upload time, policy information, etc. Object metadata 222 also includes a write ahead log (WAL), which requires that the modification operations of the storage system must be written to the log before submission, which can ensure the atomicity and persistence of the object storage system. In the case that the hard disk data is not damaged, the write ahead log allows the storage system to recover to the state before the crash under the guidance of the log after the crash, avoiding data loss. This technology is widely used to store metadata information in (file, object, database, column) storage systems.
[0068] In the current technical solution, when a user accesses an object (e.g., object 1) in the object storage system 20, the object storage service provided by the object storage system 20 will record relevant information in the object storage system 20 according to the user's access request and the current time. For example, in the append write method, in the object metadata of object 1 (e.g., Figure 2 The object metadata D shown 1 ), which will result in the existence of both the original object metadata and the new object metadata with the last access time field added in the object storage system 20 (for example, Figure 2 The object metadata D shown 2 ), which not only requires more additional storage space, but also brings more read-write amplification.
[0069] In order to free up storage space, garbage collection 212 is required to clean up redundant data in a timely manner (i.e., delete object metadata D 1 ), if the object 1 is frequently accessed, the object metadata of the object 1 will be frequently updated, which will bring more burden and overhead to garbage collection. In addition, the object storage system 20 can periodically scan the object from the object metadata 222 based on the life cycle management 211, and perform simple analysis and processing on it to determine whether it needs to switch the access layer. For example, intelligent classification is performed according to the life cycle rules of the latest access timestamp (i.e., the last access time) and the last modification time recorded in the metadata database. However, this method of directly scanning and calculating and processing inside the object storage system 20 has affected the normal business processing of the object storage system 20. If a failure occurs, it will also indirectly affect the serviceability of the object storage system 20. It can be seen that the traditional data processing method will affect the normal operation of the object storage system 20.
[0070] In order to solve the above problems, an embodiment of the present application provides a data processing method, which is applied to an intelligent grading system, and the intelligent grading system is used to perform intelligent grading processing on business data in a storage system, where the storage system includes at least one storage path. In this method, the intelligent grading system can obtain first business data (i.e., business data whose storage type is an intelligent grading type, for example, an intelligent grading object in an object storage system), and the first business data is stored in a target storage path in the storage system. Furthermore, the intelligent grading system can obtain an access log of the first business data, and analyze the access log of the first business data according to a hot and cold identification algorithm configured by a user for the target storage path, so as to obtain the access heat of the first business data. Then, the intelligent grading system can send a data migration instruction to the storage system according to the access heat of the first business data, so that the storage node of the storage system performs data migration on the first business data.
[0071] In the embodiment of the present application, the storage system may be a system that can support storage tiering, and the intelligent tiering system belongs to an out-of-band system, that is, another system deployed outside the storage system, which can support richer hot and cold identification algorithms (including algorithms for hot and cold identification based on the last access time, algorithms for hot and cold identification based on the access frequency, and algorithms for hot and cold identification based on the access ratio of the life cycle). In this way, when the hot and cold identification algorithm corresponding to the target storage path is used for hot and cold identification, the log can be analyzed at a deeper level, so that the final access heat is more accurate, and the automatic tiering of hot and cold data can be realized more intelligently. In addition, compared with the traditional data processing method of recording redundant fields in the storage system, the intelligent tiering system uses the log of the first business data for analysis, without adding redundant fields in the storage system, which not only reduces the storage space occupied, thereby saving the user's storage cost, but also does not need the storage system itself to analyze whether the object needs to be migrated, thereby effectively reducing the impact on the storage system, thereby improving the operating performance of the storage system.
[0072] The following is a system diagram of the data processing method provided by the embodiment of the present application:
[0073] See also Figure 3 , Figure 3 1 is a system diagram corresponding to a data processing system provided in an embodiment of the present application. Figure 3 As shown, the data processing system can be an out-of-band intelligent hierarchical management system based on a separated architecture, that is, the data processing system can include a storage system 30 and an intelligent hierarchical system 31, the storage system 30 is a system deployed on the data plane, and the intelligent hierarchical system 31 is a system deployed on the control plane.
[0074] If the data processing system is deployed in the cloud, Figure 3 The storage system 30 and the intelligent grading system 31 shown are both systems deployed on the cloud, which means that the data processing method provided in the embodiment of the present application can be applied to the field of cloud storage. Cloud storage is a new concept extended and developed from the concept of cloud computing. A distributed cloud storage system (hereinafter referred to as storage system 30) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of storage nodes of various types (storage nodes are also referred to as storage nodes) in the network through application software or application interfaces to work together and provide external data storage and business access functions.
[0075] At present, the storage method of the storage system is: create a logical volume, and when creating a logical volume, allocate physical storage space for each logical volume. The physical storage space may be composed of disks of a storage node or several storage nodes. The client stores data on a logical volume, that is, stores the data on the file system. The file system divides the data into many parts, each of which is an object. The object contains not only data but also additional information such as data identification (ID entity, ID). The file system writes each object into the physical storage space of the logical volume, and the file system records the storage location information of each object, so that when the client requests to access the data, the file system can allow the client to access the data according to the storage location information of each object.
[0076] The process of the storage system allocating physical storage space to a logical volume is as follows: based on the estimated capacity of the objects stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the grouping of independent redundant disk arrays (RAID, Redundant Array of Independent Disks), the physical storage space is pre-divided into stripes. A logical volume can be understood as a stripe, thereby allocating physical storage space to the logical volume.
[0077] like Figure 3 As shown, the storage system 30 may be a system supporting data tiering, and the storage system 30 may be used to store X business data, where X is a positive integer. In the embodiment of the present application, the data type of the business data may include two types, one of which is the type of storage layer directly specified by the user (for example, directly specifying storage in storage layer L 1 ), and the other is the type that automatically migrates based on the popularity of access (i.e., the intelligent grading type).
[0078] It should be understood that the embodiment of the present application can classify the data stored in the storage system 30 into different categories (e.g., hot data, warm data, and cold data) according to the access popularity of the business data (obtained after evaluation by data indicators such as access frequency) to facilitate better management and storage of data. It is understandable that the hot data here refers to data with high access frequency and for business and applications. These data usually require fast and efficient access and processing, and therefore need to be stored in a high-performance, low-latency storage layer L 1 (For example, solid-state drives, memory, etc.); warm data refers to data with moderate access frequency and a certain degree of importance to business and applications. This data usually does not need to be accessed and processed as quickly as hot data, but also requires reliable storage and access within a certain period of time. Therefore, it can be stored in a storage layer L with lower cost and larger capacity. 2(For example, disk arrays); cold data refers to data that is less frequently accessed and less important to business and applications. This data usually needs to be stored for a long time, but does not need to be frequently accessed and processed. Therefore, it can be stored in a storage layer L with lower cost and larger capacity. 3 (For example, a tape library).
[0079] like Figure 3 As shown, when the storage system 30 is an object storage system, the business data stored in the storage system 30 may be objects, specifically including object 1, object 2, object 3, object 4, and object 5. Among them, object 1 and object 4 are objects of the storage type of intelligent grading type in the storage system 30, and object 2 may be stored in the storage layer L specified by the user. 1 The hot data in the storage layer L can be stored in the storage layer L specified by the user. 2 The object 5 can be the warm data stored in the storage layer L specified by the user. 3 When the storage system 30 is running, an event record called a log is generated, and each line of the log can record the date, time, user, action and other related operation descriptions of the event. Therefore, when a user performs a system operation (for example, a read operation, a write operation or a delete operation) on any of the above five objects, the storage system 30 can record the log corresponding to the system operation.
[0080] The intelligent grading system 31 may be a system deployed outside the storage system 30 for realizing automatic data grading, and the analysis results finally obtained may act on the storage system 30. The intelligent grading system 31 may include a storage subsystem 311, a heat recognition plug-in 312, and a big data analysis platform 313.
[0081] The storage subsystem 311 here can be used to store the intelligent grading objects in R storage buckets, wherein the intelligent grading objects here refer to objects whose storage type belongs to the intelligent grading type, and R is a positive integer. For example, the storage subsystem 311 can store object names, version numbers, and data indicators. The storage subsystem 311 can be a key-value storage subsystem, i.e., a key-value (kv) storage system. Users can configure different hot and cold identification algorithms (i.e., hot and cold identification algorithms) on the storage buckets according to different business scenarios.
[0082] The intelligent grading system 31 may store a variety of hot and cold identification algorithms, which may specifically include a first hot and cold identification algorithm, a second hot and cold identification algorithm, and a third hot and cold identification algorithm. The first hot and cold identification algorithm may be an algorithm for performing hot and cold identification based on the last access time; the second hot and cold identification algorithm may be an algorithm for performing hot and cold identification based on the access frequency; and the third hot and cold identification algorithm may be an algorithm for performing hot and cold identification based on the access ratio within the life cycle.
[0083] Since the intelligent grading system 31 includes n grading services, it means that the intelligent grading system 31 can be deployed in multiple main areas such as area 1, area 2, ..., area n, etc. At this time, the objects stored in the storage subsystem 311 are related to the n main areas. Among them, the n grading services in the intelligent grading system 31 do not interfere with each other and do not affect each other.
[0084] Among them, the intelligent grading system 31 can determine the newly uploaded intelligent grading objects in the storage system 30 by consuming the index incremental pre-write log, and then store the acquired intelligent grading objects in the storage subsystem 311, for example, the object name, version number, etc. of the intelligent grading objects are stored. Then, when the intelligent grading system 31 calls the grading services corresponding to each area and performs persistent management on the objects in the storage system 30, the object storage rate can be effectively improved. Among them, the index here is a sorted data structure in the storage system 30, which is used to uniquely indicate a certain object; WAL can be used to ensure the atomicity and persistence of data operations.
[0085] The big data analysis platform 313 can not only store massive logs into the lake, but also perform offline analysis on the logs in the storage system 30. It is understandable that if the first business data obtained by the intelligent grading system 31 is any object in the storage system 30 (for example, object 1), the big data analysis platform 313 can obtain the access log of the first business data from the storage system 30, and then analyze the access log of the first business data based on the hot and cold identification algorithm configured by the user for the target storage path (that is, the storage path where the first business data is stored in the storage system 30), so as to obtain its corresponding data indicators. Among them, the data indicator here can be any one of the last access time, access frequency or life cycle access ratio.
[0086] In this embodiment of the present application, the storage bucket where the first business data is located may be referred to as the target storage bucket. If the hot and cold identification algorithm is the first hot and cold identification algorithm or the second hot and cold identification algorithm, the first business data here may be any intelligent grading object in the target storage bucket. If the hot and cold identification algorithm is the third hot and cold identification algorithm, the first business data here may be multiple intelligent grading objects in the target storage bucket whose object creation time belongs to the same life cycle. The life cycle refers to the time period to which the object creation time belongs. For example, if Figure 3 The objects 1 and 4 shown are both intelligent hierarchical objects in the same target storage bucket. The object creation time of object 1 is December 3, 2023, and the object creation time of object 4 is December 4, 2023. Therefore, if a life cycle is from December 1, 2023 to December 15, 2023, the first business data belonging to the life cycle may include object 1 and object 4.
[0087] The heat identification plug-in 312 can be used to perform heat identification on the first business data according to the data indicators obtained by the big data analysis platform 313 to obtain the access heat of the first business data, and then send a data migration instruction to the storage system 30 based on the access heat. For example, the intelligent grading system can compare the storage layer corresponding to the access heat (i.e., the second storage layer) with the current storage layer of the first business data (i.e., the first storage layer) to determine whether the target object needs to be migrated. When the first storage layer is different from the second storage layer, it is necessary to send a data migration instruction to the storage system 30 so that the storage node of the storage system 30 migrates the first business data from the first storage layer to the second storage layer.
[0088] It is understandable that users can configure different hot and cold identification algorithms for different storage paths. For example, in the object storage system 30, when creating a bucket, users can configure various properties of the bucket, including: the default storage type, access rights, region, hot and cold identification algorithm, etc. of the bucket. For easier understanding, please refer to Figure 4 , Figure 4 This is a schematic diagram of an interface for configuring a storage bucket provided in an embodiment of the present application. Figure 1 .like Figure 4 As shown, user a can be a user who logs in to the cloud management platform using an account on the client. 1 It can be a guide interface provided by the object storage system based on the object storage service. 1 The business controls for creating buckets are included in Figure 4 The control K shown 1 ).
[0089] When user a targets control K1 When the trigger operation is executed, the client corresponding to user a can respond to the trigger operation and display the creation interface (for example, Figure 4 The interface J shown 2 ). Then, user a can 2 In the interface, set the name of the currently created bucket and select the default storage type. 2 Multiple default storage types may be included, including intelligent tiering type, standard storage type, low-frequency access storage type, and archive storage type. The default storage type here is used to indicate the default storage tier corresponding to the objects subsequently uploaded to the bucket. 2 A business control for displaying multiple hot and cold identification algorithms may also be included (e.g., Figure 4 The control K shown 2 ), the multiple hot and cold identification algorithms here may include an algorithm for performing hot and cold identification based on the last access time (i.e., a first hot and cold identification algorithm), an algorithm for performing hot and cold identification based on the access frequency (i.e., a second hot and cold identification algorithm), and an algorithm for performing hot and cold identification based on the access ratio within the life cycle (i.e., a third hot and cold identification algorithm).
[0090] It should be understood that when the hot and cold identification algorithm configured by user a for the storage bucket is the first hot and cold identification algorithm, user a can directly perform a trigger operation on the first hot and cold identification algorithm, thereby causing the client to display an algorithm configuration interface for the first hot and cold identification algorithm (for example, Figure 4 The interface J shown 3 ), where the interface J 3 The setting of relevant thresholds (eg, conversion days) of the first hot and cold identification algorithm may be included.
[0091] When user a completes the configuration, a new storage bucket can be successfully created. The triggering operation in the embodiment of the present application may include contact operations such as clicking and long pressing, and may also include non-contact operations such as voice and gestures, which will not be limited here.
[0092] Optionally, when the storage system is an object storage system, users can also configure different hot and cold identification algorithms for the sub-paths of each folder in the bucket. For easier understanding, please refer to Figure 5 , Figure 5 This is a schematic diagram of an interface for configuring a storage bucket provided in an embodiment of the present application. Figure 2 .like Figure 5 As shown, user b can be a user who logs in to the cloud management platform using an account on the client. 4It can be a details display interface, used to display multiple folders of a bucket (for example, bucket 1) created by user b, where one folder corresponds to one subpath. 4 You can also include controls K 3 and control K 4 . Control K 3 Can be used to upload objects to bucket 1, control K 4 Can be used to create a new folder in bucket 1.
[0093] like Figure 5 As shown, user B has created two folders in bucket 1, specifically folder 1 and folder 2. User B can configure different hot and cold identification algorithms for the sub-paths of each folder. Taking folder 1 as an example, user B can perform a trigger operation on file 1 to make the client corresponding to user B display the sub-path configuration interface (for example, Figure 5 The interface J shown 5 ). Among them, the interface J 5 The business controls for displaying multiple hot and cold identification algorithms may be included (for example, Figure 5 The control K shown 6 ), the multiple hot and cold identification algorithms here may include an algorithm for performing hot and cold identification based on the last access time (i.e., a first hot and cold identification algorithm), an algorithm for performing hot and cold identification based on the access frequency (i.e., a second hot and cold identification algorithm), and an algorithm for performing hot and cold identification based on the access ratio within the life cycle (i.e., a third hot and cold identification algorithm).
[0094] Of course, if the hot and cold identification algorithm configured by user b for the sub-path of folder 1 is the first hot and cold identification algorithm, user b can also directly perform a trigger operation on the first hot and cold identification algorithm, thereby causing the client of user b to display an algorithm configuration interface for the first hot and cold identification algorithm (for example, Figure 4 The interface J shown 3 ).
[0095] It should be noted that Figure 4 and Figure 5 The interfaces and controls shown are only some forms of expression for reference. In actual business scenarios, developers can make relevant designs according to product requirements. The embodiments of the present application do not limit the specific forms of the interfaces and controls involved.
[0096] For further understanding, please see Figure 6 , Figure 6 Schematic diagram of a scenario of data migration in an object storage system provided by an embodiment of the present application. Figure 6As shown, R storage buckets can be created in the object storage system, where R is a positive integer. For ease of explanation, 3 storage buckets can be used as an example, specifically including bucket 1, bucket 2, and bucket 3. Among them, these three storage buckets can be created by the same user according to different business scenarios, or they can be created by different users respectively, which will not be limited here.
[0097] like Figure 6 As shown, the hot and cold identification algorithm here is configured for the storage path corresponding to the storage bucket. For example, the hot and cold identification algorithm configured for storage bucket 1 can be the hot and cold identification algorithm S 1 , the hot and cold identification algorithm S 1 Used to identify hot and cold based on the last access time; the hot and cold identification algorithm configured for bucket 2 can be the hot and cold identification algorithm S 2 , the hot and cold identification algorithm S 2 Used to identify hot and cold based on access frequency; the hot and cold identification algorithm configured for bucket 3 can be the hot and cold identification algorithm S 3 , the hot and cold identification algorithm S 3 Used for hot and cold identification based on the access ratio within the life cycle.
[0098] It should be understood that the intelligent grading system can be used to perform intelligent grading processing on the intelligent grading objects in the object storage system, that is, to analyze the access log of the first business data in the storage system to achieve data migration of the intelligent grading objects in each storage bucket. It can be understood that the storage layer in the embodiment of the present application can include M, where M is a positive integer greater than 1. For ease of understanding, only 3 are taken as an example here, which can specifically include storage layer L 1 , storage layer L 2 And the storage layer L 3 Among them, the storage layer L 1 The access popularity of business data stored in the storage layer is lower than that of the storage layer L 2 The access popularity of the business data stored in the storage layer L 2 The access popularity of business data stored in the storage layer is lower than that of the storage layer L 3 The access popularity of business data stored in .
[0099] Among them, the storage layer L 1 It can be a standard storage layer (for example, a disk for storing hot data), which is used to provide high-performance, high-reliability, and high-availability object storage service storage. It is suitable for business scenarios that require frequent access to data, such as big data, mobile applications, hot videos, social pictures, and other business scenarios.
[0100] The storage layer L here 2It can be a low-frequency access storage layer (for example, a disk for storing warm data) to provide highly reliable, low-cost real-time access storage services. It is suitable for business scenarios that require infrequent access but also require fast access to data when needed, such as file synchronization / sharing, enterprise backup, etc. 1 In contrast, storage layer L 2 Same data persistence, throughput, and access latency, but at a lower cost.
[0101] The storage layer L here 3 It can be an archive storage layer (for example, a tape for storing cold data), which is suitable for business scenarios where data is rarely accessed, such as data archiving, long-term backup, and offline analysis (for example, model training in machine learning or big data analysis).
[0102] Exemplarily, when the first business data acquired by the intelligent grading system is an intelligent grading object in bucket 1, the intelligent grading object can obtain the access log of the intelligent grading object and identify the hot and cold data according to the cold and hot identification algorithm S corresponding to bucket 1. 1 , analyze the access log of the intelligent grading object to obtain the access popularity of the intelligent grading object, and then send a data migration instruction to the object storage system according to the access popularity of the intelligent grading object, so that the storage node of the object storage system performs data migration on the intelligent grading object.
[0103] For example, if the first storage layer (i.e., the current storage layer) of the intelligent grading object is storage layer L 1 , then, when the intelligent grading system determines that the second storage layer of the intelligent grading object (i.e., the storage layer corresponding to the access heat) is still the storage layer L 1 , it means that the second storage layer is the same as the first storage layer, and there is no need to migrate the intelligent tiering object; when the intelligent tiering system determines that the second storage layer of the intelligent tiering object is storage layer L 2 or storage layer L 3 , it means that the second storage layer is different from the first storage layer. At this time, the intelligent tiering system needs to send a data migration instruction to the object storage system so that the storage node of the object storage system migrates the intelligent tiering object from the first storage layer to the second storage layer.
[0104] Similarly, the intelligent grading system can refer to the data management method of the intelligent grading object in the above-mentioned bucket 1, and manage the data of the intelligent grading objects in buckets 2 and 3 in turn, which will not be further described here. It can be understood that different hot and cold identification algorithms need to pay attention to different indicators. For example, the hot and cold identification algorithm S 1The corresponding object index is the last access time. For example, it is necessary to count the access requests of objects in the storage buckets of different users, that is, to record the timestamp of the last access, and then convert the timestamp according to certain rules (for example, add 30 days to the last access timestamp as the next conversion time), and then flush the converted timestamp into the storage subsystem (for example, the kv storage system), so as to facilitate the subsequent use of the global secondary index of the kv storage system and analyze data indicators more quickly; hot and cold identification algorithm S 2 The corresponding object indicator is access frequency, so it is necessary to count the number of access requests per day. The hot and cold identification algorithm S 2 The corresponding object indicator is the access ratio of the life cycle. Then, according to the heat recognition plug-in, it is necessary to provide hot and cold recognition results (i.e., access heat) of different scenarios, so that the grading service in the subsequent intelligent grading system can start distributed jobs based on the algorithm calculation results and trigger the hot and cold migration of intelligent grading objects.
[0105] It can be seen that the embodiment of the present application does not need to analyze whether the object needs to be migrated by the storage system itself like the traditional data processing method, so there is no need to add redundant fields in the storage system. Instead, an intelligent grading system deployed outside the storage system regularly analyzes the logs in the storage system, which can greatly save the computing resources of the storage system and effectively improve the operating performance of the storage system. In addition, this data processing method analyzed by the intelligent grading system can perform a deeper analysis of the logs and obtain data indicators that are more in line with the business scenario, so as to facilitate the subsequent improvement of the rationality of object migration.
[0106] For further information, see Figure 7 , Figure 7 This is an interactive diagram of a method for data processing provided by an embodiment of the present application. Figure 7 As shown, the method can be performed by an intelligent grading system and a storage system, wherein the intelligent grading system is used to perform intelligent grading processing on business data in the storage system, and the storage system includes at least one storage path. The storage system may include an object storage system, a block storage system, a file storage system, and a storage system related to a database. The method may at least include steps S701-S705:
[0107] Step S701: the intelligent grading system obtains first business data.
[0108] In this embodiment of the present application, the storage path where the first business data is stored in the storage system may be referred to as the target storage path. The storage type of the first business data is the intelligent grading type. For example, in an object storage system, the first business data here may be referred to as an intelligent grading object (the above Figure 3The storage bucket where the intelligent grading object is located is called the target storage bucket.
[0109] It is understandable that the intelligent grading system can obtain business data of the intelligent grading type from the storage system through the index, and then store the obtained business data in the storage subsystem of the intelligent grading system through the regional grading service to which the business data belongs, so as to facilitate subsequent data management. In other words, the first business data here can be the business data obtained by the intelligent grading system from the storage system (for example, Figure 3 The storage system shown in FIG. 1 may also be consumed by a storage subsystem in the intelligent tiering system (e.g., Figure 3 The first business data is obtained from the storage subsystem 311 as shown, and the method for obtaining the first business data is not limited here.
[0110] Step S702: the intelligent grading system obtains an access log of the first business data.
[0111] Specifically, the intelligent grading system obtains the logs in the storage system based on the big data analysis platform, and then can determine the access logs of the first business data from the obtained logs. Among them, the intelligent grading system can obtain the access logs in the storage system in the following two ways. For example, the first way is: the intelligent grading system sends a log consumption instruction to the storage system, so that the storage system pushes the logs within the data grading period (for example, one day or 12 hours) to the intelligent grading system. Among them, the data grading period can be dynamically adjusted according to the actual business situation (for example, the resource computing power and network parameters of the big data analysis platform), which will not be limited here. The second way is that the storage system actively pushes the logs in the storage system to the intelligent grading system on a regular basis based on the prior configuration.
[0112] Among them, the logs in the storage system may include user behavior logs (also known as user behavior tracks, traffic logs, etc.). In simple terms, it is the behavior data generated each time a user accesses business data in the storage system (for example, access, browsing, searching, and clicking, etc.). In the specific implementation of this application, when the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data shall comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0113] For example, when the storage system is an object storage system, if a user accesses an object in a storage bucket, the object storage service will generate a user behavior log based on the user's access information (e.g., user ID, access time, terminal ID, operation, etc.). The big data analysis platform in the intelligent grading system can consume user behavior logs, and its strong resource computing power can facilitate the cumulative processing of hundreds of billions of user access behavior records within the data grading cycle (e.g., one day or 12 hours, etc.). For example, when the time difference between the current timestamp and the previous grading timestamp reaches the data grading cycle, the big data analysis platform in the intelligent grading system can consume the logs in the storage system from the storage system.
[0114] Step S703: the intelligent grading system analyzes the access log of the first business data according to the hot-cold identification algorithm configured by the user for the target storage path to obtain the access popularity of the first business data.
[0115] The access heat here refers to one of the heats of hot, warm or cold. The hot and cold identification algorithm can be stored in an intelligent grading system, thereby effectively reducing the storage space occupied in the storage system. The target storage path can include one or more of the following: a bucket of object storage, a volume of block storage, a directory of file storage, and an instance of a database.
[0116] For example, when the storage system is an object storage system, the target storage path here can be the storage path corresponding to the target storage bucket where the first business data is located, or the storage subpath of the first business data in the target storage bucket (that is, the storage subpath corresponding to the folder of the first business data in the target storage bucket).
[0117] For ease of understanding, the hot and cold identification algorithms in the embodiment of the present application can be taken as three examples, which can specifically include a first hot and cold identification algorithm (for example, the above Figure 6 The hot and cold identification algorithm S 1 ), the second hot and cold identification algorithm (for example, the above Figure 6 The hot and cold identification algorithm S 2 ) and the third hot and cold identification algorithm (for example, the above Figure 6 The hot and cold identification algorithm S 3 ), based on this, the hot and cold recognition algorithm here can be any one of the first hot and cold recognition algorithm, the second hot and cold recognition algorithm or the third hot and cold recognition algorithm.
[0118] Among them, the first hot and cold identification algorithm is used to identify hot and cold based on the last access time. The identification rules of the first hot and cold identification algorithm are relatively simple and are applicable to most business scenarios, especially big data scenarios, which require objects that have not been accessed for a long time to be migrated to a colder storage layer and save costs. Among them, big data refers to a collection of data that cannot be captured, managed and processed by conventional software tools within a certain time range. It is a massive, high-growth and diversified information asset that requires a new processing model to have stronger decision-making power, insight discovery and process optimization capabilities. With the advent of the cloud era, big data has also attracted more and more attention. Big data requires special technologies to effectively process a large amount of data within the tolerance time. Technologies suitable for big data include large-scale parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet and scalable storage systems.
[0119] The second hot and cold identification algorithm is used to identify hot and cold objects based on access frequency. The second hot and cold identification algorithm needs to be learned through a certain historical period, that is, it mainly counts and analyzes the number of times a batch of objects are accessed in the historical period. When the access frequency statistics are consistent with a certain access trend distribution (for example, normal distribution), the batch of objects will be migrated. For example, when the object is a movie currently being shown, its popularity changes more in line with the normal distribution, that is, it changes from low to high first, and then from high to low.
[0120] The third hot and cold identification algorithm is used to perform hot and cold identification based on the access ratio within the life cycle. The third hot and cold identification algorithm can classify objects created in different life cycles and is more suitable for business scenarios where the access behavior of objects in the same life cycle is relatively consistent. For example, when the object uploaded by the user is a short video, the access popularity of the short video will gradually decrease over time. For another example, when the object uploaded by the user is a reference document, the access popularity of the reference document will remain high in a special time period (for example, graduation season).
[0121] It should be understood that if the hot and cold identification algorithm is the first hot and cold identification algorithm, the intelligent grading system can analyze the access log of the first business data, obtain the last access time of the first business data, and calculate the target time interval between the last access time and the current time, and then compare the target time interval and the preset time interval to obtain the access heat of the first business data. Among them, the preset time interval here can include one or more. For the sake of convenience, the embodiment of the present application can take 2 as an example, which can specifically include a first time interval and a second time interval, and the second time interval is greater than the first time interval. The first time interval and the second time interval here can be dynamically adjusted according to business needs. For example, the second time interval can be 90 days, and the first time interval can be 30 days.
[0122] If the target time interval is less than or equal to the first time interval, the intelligent grading system can determine that the access popularity of the first business data is hot; if the target time interval is greater than the first time interval and less than the second time interval, the intelligent grading system can determine that the access popularity of the first business data is warm; if the target time interval is greater than or equal to the second time interval, the intelligent grading system can determine that the access popularity of the first business data is cold.
[0123] For example, if the first service data (for example, the above Figure 6 The first storage layer of any intelligent grading object in the storage bucket 1 shown in FIG. Figure 6 The storage layer L shown 1 When the intelligent grading system determines that the target time interval between the current time and the last access time belongs to the time period of 30-90 days, that is, the first business data has not been accessed by the user for 30 days, and the access heat of the first business data has changed from hot to warm. Optionally, when the intelligent grading system determines that the target time interval between the current time and the last access time is greater than 90 days, that is, the first business data has not been accessed by the user for 90 days, and the access heat of the first business data directly changes from hot to cold.
[0124] Optionally, if the hot and cold identification algorithm is the second hot and cold identification algorithm, the intelligent grading system can analyze the access log of the first business data to obtain the number of times the first business data accesses the first business data within the data grading period (i.e., the target access frequency), and then the intelligent grading system can obtain N historical access frequencies of the first business data within a preset time, where N is a positive integer. Each of the N historical access frequencies here is an access frequency counted before the data grading period. Furthermore, the intelligent grading system can determine the changing trends of the N historical access frequencies and the target access frequency, thereby obtaining the access heat of the first business data.
[0125] For ease of explanation, the embodiment of the present application may determine the access popularity of the first business data before log analysis as the first access popularity, and determine the access popularity of the first business data after log analysis as the second access popularity.
[0126] If the trend of change indicates that it will decline in the future, then when the first access heat is not the lowest access heat, the intelligent grading system can determine the access heat lower than the first access heat as the second access heat; when the first access heat is the lowest access heat, the intelligent grading system can continue to determine the first access heat as the second access heat.
[0127] If the trend of change indicates that it will rise in the future, when the first access popularity is not the highest access popularity, the intelligent grading system can determine the access popularity higher than the first access popularity as the second access popularity; when the first access popularity is the highest access popularity, the intelligent grading system can continue to determine the first access popularity as the second access popularity.
[0128] If the change trend indicates that it will remain unchanged in the future, the intelligent grading system continues to determine the first access popularity as the second access popularity.
[0129] For ease of understanding, further, please refer to Table 1, which is a statistical table associated with access frequency provided by an embodiment of the present application. The statistical table may include target storage buckets (for example, the above Figure 6 The H intelligent grading objects in the storage bucket 2) shown in the figure, and the access frequencies corresponding to each intelligent grading object in multiple grading periods. For the convenience of explanation, H can be taken as 3 as an example, and the access frequency can be taken as 8 as an example, as shown in Table 1:
[0130] Table 1
[0131]
[0132] As shown in Table 1, if access frequency 8 is the target access frequency counted by the intelligent grading system in the data grading period of December 8, then access frequency 1 can be the access frequency counted in the historical grading period of December 1, access frequency 2 can be the access frequency counted in the historical grading period of December 2, access frequency 3 can be the access frequency counted in the historical grading period of December 3, and so on. Access frequency 7 can be the access frequency counted in the historical grading period of December 7.
[0133] It can be understood that the change trend of the first business data can be determined according to the distribution corresponding to the business scenario where the first business data is located. For example, the distribution here can be a normal distribution.
[0134] To understand the changing trend of object 1 shown in Table 1, please refer to Figure 8 , Figure 8 This is a trend chart for determining access popularity provided by an embodiment of the present application. Figure 8 As shown, the change trend graph is generated based on the fitting of the access frequencies of object 1 in Table 1, and can be used to clearly display the relationship between the date and the access frequency, so as to facilitate the observation of the changes in the subsequent access popularity of object 1.
[0135] like Figure 8As shown, the changing trends of these 8 access frequencies conform to a certain distribution (e.g., normal distribution), that is, between December 1 and December 6, the changing trend of object 1 is gradually increasing, and between December 6 and December 8, the changing trend of object 1 is decreasing, that is, the intelligent grading system can determine that the changing trend of object 1 is increasing first and then decreasing. If the first access heat of object 1 is not the lowest access heat, the intelligent grading system can determine the access heat lower than the first access heat as the second access heat.
[0136] For object 2 shown in Table 1 above, by observing the historical access frequencies corresponding to 8 consecutive historical classification periods, it can be seen that from December 1 to December 8, the change trend of object 2 is rising, but between December 7 and December 8, the change trend of object 2 is obviously rising but the amplitude is not high. At this time, it can be determined that the change trend of object 2 also conforms to a certain normal distribution (for example, the rising trend of the normal distribution). Based on this, when the first access heat of object 2 does not belong to the highest access heat, the intelligent classification system can directly determine the access heat higher than the first access heat as the second access heat.
[0137] For object 3 shown in Table 1 above, it can be seen from the historical access frequencies corresponding to the eight consecutive historical classification cycles that object 3 has not been accessed between December 1 and December 7, and the access frequency of object 3 on December 8 is 1, that is, it has been accessed once. If the first access heat of object 3 is cold, then in this case, it is necessary to determine whether object 3 needs to be migrated back to the frequent access layer according to the business scenario of object 3. For example, if the business scenario of object 3 is a video scenario, it can be considered that the change trend of object 3 conforms to the rising trend of normal distribution, and the intelligent classification system can determine the access heat higher than the first access heat as the second access heat. If the business scenario of object 3 is a security scanning scenario (for example, a fixed scan once a week), it can be considered that the change trend of object 3 conforms to the distribution under the security scanning scenario. At this time, the intelligent classification system can continue to determine the first access heat as the second access heat. When object 3 belonging to cold data is accessed, according to the second cold and hot identification algorithm, it can be determined that object 3 does not need to be migrated back to the frequent access layer.
[0138] Optionally, if the hot and cold identification algorithm is the third hot and cold identification algorithm, then when the storage system is an object storage system, the first business data here can be an intelligently graded object whose object creation time is determined in the target storage bucket and belongs to the target life cycle. Further, the intelligent grading system can analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle, and then compare the access ratio with the target access threshold corresponding to the target life cycle to obtain the access heat of the first business data.
[0139] Among them, the division of life cycle segments can be dynamically adjusted according to actual needs. For example, taking 15 days as a division segment, the object life cycle distribution is divided into multiple time periods, such as within 15 days, 15-30 days, 30-45 days, 45-60 days, 60-75 days, 75-90 days, 90-105 days, 105-120 days, 120-135 days, 135-150 days, 150-165 days, 165-180 days, more than 180 days, etc.
[0140] For ease of understanding, further, please refer to Table 2, which is a schematic table provided in an embodiment of the present application for determining the access ratio within a data classification cycle. For ease of explanation, the life cycle segments of the embodiment of the present application only take 3 segments as an example, which may specifically include within 15 days, 15-30 days, and 30-45 days. If the data classification cycle of the intelligent classification system is one day, and the last object classification timestamp is timestamp 1 (for example, October 31st 00:00:00), then when the current timestamp is timestamp 2 (for example, November 1st 00:00:00), it means that the time difference between the current timestamp and the last classification timestamp reaches the data classification cycle. When the intelligent classification system determines that the hot and cold identification algorithm of the target storage bucket is the third hot and cold identification algorithm, it can determine the access ratios within the following life cycle segments. As shown in Table 2:
[0141] Table 2
[0142]
[0143] It should be noted that the items shown in Table 2 are merely a form of expression for reference. In actual business scenarios, other items (for example, the object name within a certain object creation time period, the current storage layer, etc.) can also be established according to needs. The embodiments of the present application do not limit the specific form of Table 2.
[0144] The intelligent grading system can determine any life cycle shown in Table 2 (taking the life cycle period within 15 days as an example) as the target life cycle. In this case, the first business data here refers to the intelligent grading object in the target storage bucket whose object creation time belongs to the life cycle from October 17 to October 31. Then, the intelligent grading system can analyze the access log of the first business data, determine the number of objects accessed within the data grading period (for example, 10 as shown in Table 2) and the number of objects created within the target life cycle (for example, 50 as shown in Table 2), and determine the access ratio of the target life cycle from October 17 to October 31 (for example, 20%) based on the number of objects and the number of objects.
[0145] Further, the intelligent classification object can obtain the target access threshold corresponding to the target life cycle, and obtain the access heat of the first service data by comparing the access ratio with the target access threshold corresponding to the target life cycle.
[0146] Among them, the target access threshold can be determined from the access thresholds corresponding to M storage layers. Each of the M access thresholds is determined based on the medium information of the corresponding storage layer. The medium information includes storage capacity, access cost, and storage cost. Here, the access threshold is used to indicate the minimum access ratio of the corresponding storage layer. Exemplarily, if the M storage layers are the 3 storage layers shown above Figure 6 shown, then the M thresholds here can include 3 access thresholds, specifically including the access threshold corresponding to storage layer L 1 (for example, 35%), the access threshold corresponding to storage layer L 2 (for example, 10%), the access threshold corresponding to storage layer L 3 (for example, 5%).
[0147] For Table 2 above, if the target life cycle is within 15 days, then the first service data here refers to 50 objects created from October 17th to October 31st. Since the access ratio of this target life cycle is 20%, which is greater than the access threshold corresponding to storage layer L 2 Therefore, the intelligent classification object can confirm that the access heat of these 50 objects is all warm. If the target life cycle is in the time period of 15 - 30 days, then the first service data here refers to 20 objects created from October 2nd to October 16th. Since the access ratio of this target life cycle is 50%, which is greater than the access threshold corresponding to storage layer L 1 Therefore, the access heat of these 20 objects is all hot. If the target life cycle is in the time period of 30 - 45 days, then the first service data here refers to 250 objects created from September 17th to October 1st. Since the access ratio of this target life cycle is 4%, which is greater than the access threshold corresponding to storage layer L 3 Therefore, the access heat of these 250 objects is all cold.
[0148] Optionally, the target access threshold here is determined from the access thresholds of multiple life cycles. For example, the intelligent classification system can obtain the second service data, and then analyze the access log of the second service data to obtain the access ratio of the second service data within at least one life cycle. Further, the intelligent classification system can respectively determine the access threshold of each life cycle according to the access ratio within at least one life cycle. Among them, the second service data here can be the service data already stored in the target storage path, that is, the historical service data belonging to the same service scenario as the first service data.
[0149] For example, if the first business data is related to seasonal changes, the intelligent grading system can obtain the following access thresholds after analyzing the access log of the second business data, which may include: the access threshold in spring (for example, 50%), the access threshold in summer (for example, 90%), the access threshold in autumn (for example, 50%), and the access threshold in winter (for example, 10%). If the target life cycle belongs to summer, the target access threshold is the access threshold in summer. If the access ratio of the first business data in the target life cycle is less than or equal to the target access threshold, the second access heat of the first business data is determined to be an access heat lower than or equal to the first access heat; if the access ratio of the first business data in the target life cycle is greater than the target access threshold, the second access heat of the first business data is determined to be an access heat higher than or equal to the first access heat. Among them, the situation where the second access heat is equal to the first access heat is because the first access heat is the maximum value and cannot be changed.
[0150] Step S704: the intelligent classification system sends a data migration instruction to the storage system according to the access popularity of the first business data.
[0151] Specifically, the intelligent grading system can determine the storage layer (i.e., the second storage layer) to which the first business data needs to be migrated based on the access popularity of the first business data (i.e., the second access popularity), and then determine whether it is necessary to send a data migration instruction to the storage system by comparing whether the first storage layer is consistent with the second storage layer. If the second storage layer is the same as the first storage layer, there is no need to send a data migration instruction. If the second storage layer is different from the first storage layer, it is necessary to generate a data migration instruction for migrating the first business data from the first storage layer to the second storage layer, and send the data migration instruction to the storage system.
[0152] Step S705: The storage node of the storage system performs data migration on the first business data.
[0153] Specifically, the storage node of the storage system can migrate the first business data from the first storage layer to the second storage layer based on the data migration instruction, thereby achieving the migration of the first business data.
[0154] It can be seen that the storage system can be a system that can support storage tiering, and the intelligent tiering system belongs to an out-of-band system, that is, another system deployed outside the storage system, which can support richer hot and cold identification algorithms (including algorithms for hot and cold identification based on the last access time, algorithms for hot and cold identification based on the access frequency, and algorithms for hot and cold identification based on the access ratio of the life cycle). In this way, when the hot and cold identification algorithm corresponding to the target storage path is used for hot and cold identification, the log can be analyzed at a deeper level, so that the access heat obtained in the end is more accurate, and then the automatic tiering of hot and cold data can be realized more intelligently. When the identified second storage layer is different from the current storage layer (i.e., the first storage layer) of the first business data, the intelligent tiering system realizes the automatic tiering of hot and cold data by sending a data migration instruction to the storage system. In addition, the embodiment of the present application does not need to add redundant fields in the storage system, nor does it need to be analyzed by the storage system itself whether the object needs to be migrated, but an intelligent tiering system deployed outside the storage system analyzes the log, which can effectively reduce the impact on the storage system, thereby improving the operating performance of the storage system.
[0155] The following describes the method provided in the embodiment of the present application from the perspective of a single system in conjunction with the accompanying drawings.
[0156] See also Fig. 9 , Fig. 9 Schematic diagram of a method for data processing provided by an embodiment of the present application. Fig. 9 As shown, the method can be executed by an intelligent grading system, which is independent of the storage system. The intelligent grading system is used to perform intelligent grading processing on business data in the storage system, and the storage system includes at least one storage path. The method can at least include steps S901-S904:
[0157] Step S901, obtaining first business data.
[0158] The first business data may be stored in a target storage path in the storage system, and the storage type of the first business data is an intelligent tiering type. The storage system may include an object storage system, a block storage system, a file storage system, and a storage system related to a database.
[0159] Step S902: Obtain access logs of first business data.
[0160] The access log of the first business data is obtained by the intelligent classification system from the storage system, and the access log refers to the behavior data (access, browsing, searching, clicking, etc.) generated when the user accesses the first business data.
[0161] Step S903: Analyze the access log of the first business data according to the hot-cold identification algorithm configured by the user for the target storage path to obtain the access popularity of the first business data.
[0162] The target storage path may include one or more of the following: a bucket of object storage, a volume of block storage, a directory of file storage, and an instance of a database. Specifically, the intelligent grading system may call the big data analysis platform, and analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the data index of the first business data, and then call the heat identification plug-in to determine the access heat of the first business data according to the data index. The access heat here refers to a heat of hot, warm or cold.
[0163] The hot and cold identification algorithm may be any one of the first hot and cold identification algorithm, the second hot and cold identification algorithm, and the third hot and cold identification algorithm. The first hot and cold identification algorithm is used to perform hot and cold identification based on the last access time; the second hot and cold identification algorithm is used to perform hot and cold identification based on the access frequency; and the third hot and cold identification algorithm is used to perform hot and cold identification based on the access ratio within the life cycle.
[0164] When the storage system is an object storage system, if the hot and cold identification algorithm is the first hot and cold identification algorithm or the second hot and cold identification algorithm, the first business data here can be any intelligent grading object in the target storage bucket. If the hot and cold identification algorithm is the third hot and cold identification algorithm, the first business data can be an intelligent grading object in the target storage bucket whose determined object creation time belongs to the target life cycle. The target life cycle here can be any life cycle segment of the multiple life cycle segments divided by the embodiment of the present application (for example, any life cycle segment in Table 2 above).
[0165] In an optional implementation, if the hot and cold identification algorithm is the first hot and cold identification algorithm, the data indicator here is the last access time corresponding to the first business data. At this time, the intelligent grading system needs to determine the target time interval between the current time and the last access time, and then can obtain the access heat of the first business data based on the target time interval and the preset time interval.
[0166] In another optional implementation, if the hot and cold identification algorithm is the second hot and cold identification algorithm, the data indicator here is the target access frequency, which is the number of times the first business data is accessed within the data classification cycle. At this time, the intelligent classification system needs to obtain N historical access frequencies of the first business data within a preset time, where N is a positive integer. Further, the intelligent classification system can determine the change trend of these N historical access frequencies and the target access frequency to obtain the access heat of the first business data.
[0167] In another optional implementation, if the hot and cold identification algorithm is the third hot and cold identification algorithm, the data indicator here is the access ratio of the target life cycle. The intelligent grading system needs to determine the target access threshold corresponding to the target life cycle, and then obtain the access heat of the first business data by comparing the target access threshold and the access ratio.
[0168] Step S904: Send a data migration instruction to the storage system according to the access popularity of the first business data.
[0169] The data migration instruction is used to instruct the storage node of the storage system to migrate the first business data from the first storage layer to the second storage layer.
[0170] The specific implementation of steps S901-S904 can be found in the above Figure 7 The description of steps S701-S704 in the corresponding embodiment will not be repeated here.
[0171] In the embodiment of the present application, since the intelligent grading system is an out-of-band management system that builds a plug-in hot and cold identification algorithm, it can greatly reduce the impact on the original storage system and provide an extremely cost-effective storage service. In addition, this plug-in hot and cold identification algorithm management can cope with different business scenarios, making its application range wider.
[0172] For further information, see Fig.10 , Fig.10 Schematic diagram of a data processing device provided in an embodiment of the present application. Fig.10 As shown, the data processing device 1 may include at least one of an acquisition module 1001, a processing module 1002 and a sending module 1003. These modules may execute the response functions of each device in the above method embodiment.
[0173] In a possible implementation, the data processing device 1 can be used to implement Figure 7 The function of the intelligent grading system in .
[0174] Specifically, the acquisition module 1001 is used to acquire the first business data, and the first business data is stored in the target storage path in the storage system, wherein the intelligent grading system is used to perform intelligent grading processing on the business data in the storage system, and the storage system includes at least one storage path, and the storage type of the first business data is an intelligent grading type; the acquisition module 1001 is also used to acquire the access log of the first business data; the processing module 1002 is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path, and obtain the access heat of the first business data; the sending module 1003 is used to send a data migration instruction to the storage system according to the access heat of the first business data, so that the storage node of the storage system migrates the first business data.
[0175] In one implementation, the hot and cold identification algorithm is stored in an intelligent grading system.
[0176] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on an access ratio of a life cycle, and a processing module 1002 is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module 1002, which is specifically used to analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle; a processing module 1002, which is specifically used to compare the access ratio with the target access threshold corresponding to the target life cycle to obtain the access heat of the first business data.
[0177] In one implementation, the processing module 1002 is specifically used to analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle, and also includes: the processing module 1002 is specifically used to obtain the second business data, the second business data is the business data stored in the target storage path; the processing module 1002 is specifically used to analyze the access log of the second business data to obtain the access ratio of the second business data in at least one life cycle; the processing module 1002 is specifically used to determine the access threshold of each life cycle according to the access ratio in at least one life cycle.
[0178] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on access frequency, and a processing module 1002 is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module 1002, which is specifically used to analyze the access log of the first business data to obtain the target access frequency of the first business data; a processing module 1002, which is specifically used to obtain N historical access frequencies of the first business data within a preset time, where N is a positive integer; a processing module 1002, which is specifically used to determine the changing trends of the N historical access frequencies and the target access frequency to obtain the access heat of the first business data.
[0179] In one implementation, a hot and cold identification algorithm is used to perform hot and cold identification based on the last access time, and a processing module 1002 is used to analyze the access log of the first business data according to the hot and cold identification algorithm configured by the user for the target storage path to obtain the access heat of the first business data, including: a processing module 1002, which is specifically used to analyze the access log of the first business data to obtain the last access time of the first business data; a processing module 1002, which is specifically used to calculate the target time interval between the last access time and the current time; a processing module 1002, which is specifically used to compare the target time interval with the preset time interval to obtain the access heat of the first business data.
[0180] In one implementation, the target storage path includes one or more of the following: a bucket of object storage, a volume of block storage, a directory of file storage, and an instance of a database.
[0181] The specific implementation of the acquisition module 1001, the processing module 1002 and the sending module 1003 can be found in the above Figure 7 The description of steps S701 to S705 in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated.
[0182] The above modules in the above data processing device 1 can be implemented by software or hardware. Exemplarily, the implementation of the acquisition module 1001 is described below by taking the acquisition module 1001 as an example. Similarly, the implementation of the processing module 1002 and the sending module 1003 can refer to the implementation of the acquisition module 1001.
[0183] As an example of a software functional unit, the acquisition module 1001 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the acquisition module 1001 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with similar geographical locations. Among them, usually a region may include multiple AZs.
[0184] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0185] As an example of a hardware functional unit, the acquisition module 1001 may include at least one computing device, such as a server, etc. Alternatively, the acquisition module 1001 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0186] The multiple computing devices included in the acquisition module 1001 can be distributed in the same region or in different regions. The multiple computing devices included in the acquisition module 1001 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the acquisition module 1001 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0187] It should be noted that, in other embodiments, the acquisition module 1001 can be used to execute any step in the data processing method, the processing module 1002 can be used to execute any step in the data processing method, and the sending module 1003 can be used to execute any step in the XXX data processing method. The steps that the acquisition module 1001, the processing module 1002, and the sending module 1003 are responsible for implementing can be specified as needed. The full functions of the data processing device 1 are realized by respectively implementing different steps in the data processing method through the acquisition module 1001, the processing module 1002, and the sending module 1003.
[0188] For further information, see Fig.11 , Fig.11 is a schematic diagram of the structure of a computing device provided by this application. Fig.11 As shown, the computing device 1100 includes: a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate through the bus 1102. The computing device 1100 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1100.
[0189] The bus 1102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.11 The bus 1102 may include a path for transmitting information between various components of the computing device 1100 (eg, the memory 1106, the processor 1104, and the communication interface 1108).
[0190] The processor 1104 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0191] The memory 1106 may include a volatile memory, such as a random access memory (RAM). The memory 1106 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0192] The memory 1106 stores executable program codes, and the processor 1104 executes the executable program codes to respectively implement the functions of the acquisition module 1001, the processing module 1002, and the sending module 1003 in the aforementioned data processing device 1, thereby implementing the data processing method. That is, the memory 1106 stores instructions for executing the data processing method.
[0193] The communication interface 1108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1100 and other devices or a communication network.
[0194] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0195] For further information, see Fig.12 , Fig.12 This is a schematic diagram of the structure of a computing device cluster provided by this application. Fig.12 As shown, the computing device cluster includes at least one computing device 1200. The memory 1206 in one or more computing devices 1200 in the computing device cluster may store the same instructions for executing the data processing method.
[0196] In some possible implementations, the memory 1206 of one or more computing devices 1200 in the computing device cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 1200 may jointly execute instructions for executing the data processing method.
[0197] It should be noted that the memory 1206 in different computing devices 1200 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the data processing device 1. That is, the instructions stored in the memory 1206 in different computing devices 1200 can implement the functions of one or more of the acquisition module 1001, processing module 1002, and sending module 1003 in the aforementioned data processing device 1.
[0198] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. For further information, see Fig.13 , Fig.13 This is a schematic diagram of a structure in which computing devices are connected via a network. Fig.13 As shown, two computing devices 1300A and 1300B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1306 in the computing device 1300A stores instructions for executing the functions of the acquisition module 1001 and the processing module 1002 in the aforementioned data processing device 1. At the same time, the memory 1306 in the computing device 1300B stores instructions for executing the functions of the sending module 1003 in the data processing device 1.
[0199] Fig.13 The connection method between the computing device clusters shown can be considered to be that the data processing method provided in this application needs to obtain data (for example, the first business data and its access log) and analyze the access log, so it is considered to entrust the functions implemented by the acquisition module 1001 and the processing module 1002 to the computing device 1300A for execution.
[0200] It should be understood that Fig.13 The functions of the computing device 1300A shown in FIG. 1300A may also be completed by multiple computing devices 1300. Similarly, the functions of the computing device 1300B may also be completed by multiple computing devices 1300.
[0201] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the data processing method or data processing method in the above embodiment.
[0202] The present application embodiment also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to execute the data processing method or data processing method in the above embodiment.
[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method, characterized in that: The method is applied to an intelligent grading system, the intelligent grading system is used to perform intelligent grading processing on business data in a storage system, the storage system includes at least one storage path, and the method includes: Acquire first business data, where the first business data is stored in a target storage path in the storage system, wherein the storage type of the first business data is an intelligent grading type; Obtaining an access log of the first business data; Analyze the access log of the first business data according to the hot-cold identification algorithm configured by the user for the target storage path to obtain the access popularity of the first business data; According to the access popularity of the first business data, a data migration instruction is sent to the storage system, so that the storage node of the storage system performs data migration on the first business data.
2. The method according to claim 1, characterized in that The hot and cold identification algorithm is stored in the intelligent grading system.
3. The method according to claim 1 or 2, characterized in that: The hot-cold identification algorithm is used to perform hot-cold identification based on the access ratio of the life cycle. The hot-cold identification algorithm configured by the user for the target storage path analyzes the access log of the first business data to obtain the access heat of the first business data, including: Analyze the access log of the first business data to obtain the access ratio of the first business data in the target life cycle; The access ratio is compared with a target access threshold corresponding to the target life cycle to obtain access popularity of the first business data.
4. The method according to claim 3, characterized in that The analyzing the access log of the first business data to obtain the access ratio of the first business data in the target life cycle also includes: Acquire second business data, where the second business data is business data stored in the target storage path; Analyze the access log of the second business data to obtain an access ratio of the second business data in at least one life cycle; An access threshold for each life cycle is determined according to the access ratio in the at least one life cycle.
5. The method according to claim 1 or 2, characterized in that: The hot and cold identification algorithm is used to perform hot and cold identification based on access frequency, and the hot and cold identification algorithm configured by the user for the target storage path analyzes the access log of the first business data to obtain the access heat of the first business data, including: Analyze the access log of the first business data to obtain the target access frequency of the first business data; Obtain N historical access frequencies of the first business data within a preset time, where N is a positive integer; Determine the change trends of the N historical access frequencies and the target access frequency to obtain the access popularity of the first business data.
6. The method according to claim 1 or 2, characterized in that: The hot and cold identification algorithm is used to perform hot and cold identification based on the last access time, and the hot and cold identification algorithm configured by the user for the target storage path analyzes the access log of the first business data to obtain the access heat of the first business data, including: Analyze the access log of the first business data to obtain the last access time of the first business data; Calculate the target time interval between the last access time and the current time; The target time interval and the preset time interval are compared to obtain the access popularity of the first business data.
7. The method according to any one of claims 1 to 6, characterized in that: The target storage path includes one or more of the following: Object storage buckets, block storage volumes, file storage directories, and database instances.
8. A data processing device, characterized in that: include: an acquisition module, configured to acquire first business data, wherein the first business data is stored in a target storage path in a storage system, wherein the intelligent grading system is configured to perform intelligent grading processing on the business data in the storage system, wherein the storage system includes at least one storage path, and the storage type of the first business data is an intelligent grading type; The acquisition module is further used to acquire an access log of the first business data; a processing module, configured to analyze the access log of the first business data according to a hot-cold identification algorithm configured by a user for the target storage path, and obtain the access popularity of the first business data; A sending module is used to send a data migration instruction to the storage system according to the access popularity of the first business data, so that the storage node of the storage system performs data migration on the first business data.
9. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.
10. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 7.
11. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 7.