Data compression method and related equipment
By judging based on the attribute information of logic blocks, compressing only when the total data scale of cold data of multiple data pages reaches the threshold, the problem of resource waste in the prior art is solved and more efficient data compression is achieved.
Patent Information
- Application Number
- CN202411751593.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-30
AI Technical Summary
The existing data compression method still consumes system resources when there is only a small amount of cold data or no cold data in the data page, resulting in a large amount of I/O resources being wasted.
By obtaining the attribute information of the logic block, it is determined whether the total data scale of the multiple data pages in the target period is greater than or equal to the first threshold. If it is greater than or equal to, the data included in the logic block is compressed; otherwise, compression is not performed.
Reduces I/O overhead, saves system resources, and avoids unnecessary compression operations.
Smart Images

Figure CN119937904A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a data compression method and related equipment. Background Art
[0002] With the continuous development of database technology, data compression has become a key means to improve storage efficiency, optimize performance, and reduce storage and operating costs. In practical applications, how to choose a compression method has become an urgent problem to be solved.
[0003] In one data compression method, a row-level scanning granularity is used. The compression task traverses all data pages in the data table and scans each row of data in each data page one by one. By recording the timestamp and compression status of each row of data, uncompressed cold data is determined from the data page. In the case where there is compression benefit for cold data, compression is performed on the data page even if the cold data is only part of the data in the data page where it is located.
[0004] In this compression method, the compression task needs to traverse all data page by page and read the timestamp and compression status of the data line by line to determine whether to perform compression. For scenarios where there is only a small amount of cold data or no cold data in the data page, this method still consumes system resources, resulting in a large amount of input / output (I / O) resource waste. Summary of the invention
[0005] The present application provides a data compression method and related devices for reducing I / O overhead and saving system resources.
[0006] In a first aspect, the present application provides a data compression method, comprising:
[0007] Acquire attribute information of a logic block, the logic block includes multiple data pages, and the attribute information indicates the data scale of the data included in the multiple data pages in multiple time periods. In a scheme where the attribute information indicates that the total data scale of the data included in the multiple data pages in the target time period is greater than or equal to a first threshold, compress the data included in the logic block and store the compressed result. The target time period includes at least one time period among the multiple time periods, and the time interval between the target time period and the current moment is greater than or equal to the time threshold. If the first attribute information indicates that the total data scale of the data included in the multiple data pages in the target time period is less than the first threshold, determine not to compress the data included in the logic block. That is, do not process the data included in the logic block.
[0008] In the present application, multiple data pages are mapped to logical blocks, and whether to compress the data included in the logical block is determined based on the attribute information of the logical block. Specifically, the attribute information indicates the data scale of the multiple data pages included in the logical block in different time periods. If the total data scale of the cold data in these multiple data pages is greater than or equal to the first threshold, the data included in the logical block is compressed. Otherwise, the data included in the logical block is not compressed. Among them, the timestamp of the cold data is in the target time period. In other words, for the logical block whose total data scale of the cold data is less than the first threshold, the compression operation will not be performed, which reduces the I / O overhead and saves system resources.
[0009] In some optional implementations of the first aspect, the attribute information includes time period information and data scale information corresponding to the time period information. The time period information indicates the time period in which the timestamp of the data is located. The data scale information indicates the amount of data, and the first threshold includes a data amount threshold. Alternatively, the data scale information indicates the number of data rows, and the first threshold includes a data row number threshold.
[0010] In the present application, the attribute information includes the time period information and the data scale corresponding to the time period information. Based on the attribute information, the relationship between the total data scale of the target time period and the first threshold value can be determined, thereby determining whether the data of the corresponding logical block needs to be compressed, which provides technical support for the implementation of the technical solution of the present application and improves the feasibility of the technical solution. In addition, the data scale includes the amount of data or the number of data rows, which enriches the implementation method and application scenarios of the technical solution of the present application.
[0011] In some optional implementations of the first aspect, before obtaining the attribute information of the logic block, a first correspondence set is also obtained, the first correspondence set indicating that the identifiers of multiple logic blocks correspond to the multiple attribute information one by one. The identifier of the logic block is obtained, and the identifier of the logic block uniquely indicates the logic block. Then, the attribute information of the logic block is obtained, specifically, according to the identifier of the logic block, the attribute information corresponding to the logic block is obtained from the first correspondence set.
[0012] In the present application, an attribute information is maintained for each logic block, and a one-to-one correspondence between the identifiers of multiple logic blocks and multiple attribute information is stored in the first correspondence set. Then, based on the identifier of the logic block, the attribute information can be determined from the first correspondence set. This provides technical support for the implementation of the technical solution of the present application and improves the feasibility of the technical solution.
[0013] In some optional implementations of the first aspect, there are multiple possibilities for obtaining the identification of the logic block. Optionally, a second correspondence set is obtained, the second correspondence set including the correspondence between the identification of the logic block and the number of the data page, wherein the identification of one logic block corresponds to the numbers of multiple data pages, and the number of one data page uniquely corresponds to the identification of one logic block. The number of the data page currently being read is obtained. According to the number of the data page currently being read, the identification of the logic block is determined from the second correspondence set.
[0014] In the present application, based on the second correspondence set, the identifier of the logical block where the currently read data page is located is directly determined, which is simple to operate and saves computing resources.
[0015] In some optional implementations of the first aspect, the identification of the logical block may also be obtained based on the following method. The number of the data page currently being read, the data amount of the data page, and the data amount of the logical block are obtained. According to the number of the data page currently being read, the data amount of the data page, and the data amount of the logical block, the number of the target data page in the logical block is determined, and the target data page is the starting data page or the ending data page of the logical block. According to the number of the target data page and the target data amount, the identification of the logical block is determined, and the target data amount is the data amount of the logical block, or the data amount of any data page in the logical block.
[0016] In the present application, the target data page in the logic block where the data page currently being read is located is determined by the number of the data page currently being read, and then combined with the target data volume, the identification of the logic block is determined. The target data page can be either the starting data page of the logic block or the ending data page of the logic block; the target data volume can be either the data volume of the data page or the data volume of the logic block, which provides multiple possibilities for the implementation of the technical solution of the present application and enriches the application scenarios and implementation methods.
[0017] In some optional implementations of the first aspect, the identification of the logical block can also be obtained based on the following method. The number of the data page currently being read and the target data volume are obtained, where the target data volume is the data volume of the logical block, or the data volume of any one of the multiple data pages. The target identification is determined based on the number of the data page currently being read and the target data volume. The identification of the logical block is determined based on the identification corresponding to the logical page currently being read and the first correspondence set, where the target identification is the same as the identification of the logical block, or the target identification is smaller than or larger than the identification of the logical block, and the identification of the logical block is closest to the target identification in the correspondence set.
[0018] In the present application, based on the number of the data page currently being read and the target data volume, the target identifier is determined, and then the target identifier is compared with the identifier of the logic block included in the first correspondence set to determine the identifier of the logic block where the data page currently being read is located. Since there are multiple possibilities for the relationship between the identifier of the logic block and the target data volume defined by the system configuration, there are also multiple implementation methods for the server to determine the identifier of the logic block, which enriches the application scenarios of the technical solution of the present application and improves flexibility.
[0019] In some optional implementations of the first aspect, compressing the data included in the logical block refers to obtaining a data page where cold data is located from multiple data pages, and the timestamp of the cold data is located in the target time period. The compression operation is performed on the data page where the cold data is located. In other words, the data actually compressed is the cold data in the logical block that is located in the target time period.
[0020] In the present application, the compressed data is cold data in the target time period. The probability of cold data being accessed again is low, and compressing it also complies with the law of data access.
[0021] In a second aspect, the present application provides a data compression device, comprising:
[0022] An acquisition unit, configured to acquire attribute information of a logic block, wherein the logic block includes a plurality of data pages, and the attribute information indicates data scales of data included in the plurality of data pages in a plurality of time periods;
[0023] a processing unit, configured to compress the data included in the logic block and store the compressed result if the attribute information indicates that the total data size of the data included in the multiple data pages in a target time period is greater than or equal to a first threshold; wherein the target time period includes at least one time period among the multiple time periods, and the time interval between the target time period and the current time is greater than or equal to a time threshold;
[0024] The processing unit is further configured to determine not to compress the data included in the logic block if the attribute information indicates that a total data size of the data included in the plurality of data pages in a target time period is smaller than the first threshold.
[0025] The data compression device is used to implement the method shown in the aforementioned first aspect, or any possible implementation of the first aspect, which will not be described in detail here.
[0026] In a third aspect, the present application provides a computing device, the computing device comprising a processor and a memory. The processor of the computing device is used to execute instructions stored in the memory, so that the computing device implements the method shown in the first aspect or any possible implementation of the first aspect.
[0027] In a fourth aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster implements the method disclosed in the first aspect, or any possible implementation manner of the first aspect.
[0028] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on a processor, implement the method shown in the aforementioned first aspect or any possible implementation of the first aspect; or, when the instruction is executed by a computer device cluster, the computer device cluster implements the method disclosed in the first aspect or any possible implementation of the first aspect.
[0029] In a sixth aspect, the present application provides a computer-readable storage medium, in which computer program instructions are stored. When the computer program instructions are executed on a processor, the method shown in the aforementioned first aspect or any possible implementation of the first aspect is implemented; or, when the computer program instructions are executed by a computer device cluster, the computer device cluster implements the method disclosed in the first aspect or any possible implementation of the first aspect.
[0030] The beneficial effects shown in any one of the second to sixth aspects are similar to those of the first aspect, or any possible implementation method of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 A schematic diagram of a system architecture provided for an embodiment of the present application;
[0032] Figure 2 Another schematic diagram of a system architecture provided for an embodiment of the present application;
[0033] Figure 3 A schematic diagram of a flow chart of a data compression method provided in an embodiment of the present application;
[0034] Figure 4 A schematic diagram of a logic block provided in an embodiment of the present application;
[0035] Figure 5 A schematic diagram of a histogram is provided for an embodiment of the present application;
[0036] Figure 6 Another flowchart of the data compression method provided in the embodiment of the present application;
[0037] Figure 7 A schematic diagram of the structure of a data compression device provided in an embodiment of the present application;
[0038] Figure 8 A schematic diagram of a structure of a computing device provided in an embodiment of the present application;
[0039] Fig. 9 A schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0040] Fig.10 Another structural diagram of a computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] The embodiments of the present application provide a data compression method and related devices for reducing I / O overhead and saving system resources.
[0042] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0043] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged in appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attribute in the embodiments of the present application when describing. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment containing a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment. In addition, "at least one" refers to one or more, and "multiple" refers to two or more. "And / or", describes the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B, can represent: A exists alone, A and B exist simultaneously, and B exists alone, wherein A, B can be singular or plural. The character " / " generally represents that the associated objects before and after are a kind of "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0044] First, see Figure 1 and Figure 2 , Figure 1 and Figure 2 All of them are schematic diagrams of system architecture provided in embodiments of the present application.
[0045] like Figure 1 As shown, the terminal device establishes a communication connection with the server, and the terminal device accesses the data stored on the server based on the communication connection, or uses the service provided by the server. In the data compression method provided in the embodiment of the present application, the user can trigger the server to perform a data compression operation through the terminal device. The server responds to the instruction from the terminal device, and based on the instruction, performs data compression on the data stored on the server, or the data managed by the server. The specific implementation process of data compression is described in detail later, and will not be expanded here for the time being.
[0046] The technical solution provided in the embodiment of the present application can also be applied in cloud computing scenarios. Figure 2 As shown, tenants log in to the cloud service platform through the Internet using their registered account and password on the cloud service platform through their terminal devices. The cloud service platform manages the infrastructure, which includes multiple data centers located in different regions, such as Figure 2 The region 1 shown includes cloud data center 1 and cloud data center 2, and region 2 includes cloud data center 3 and cloud data center 4. Each cloud data center is provided with multiple servers, and service instances (including at least one of virtual machines, containers, and dedicated hosts) are run on the servers.
[0047] In the embodiment of the present application, a data compression service is deployed in the business instance. The tenant purchases the cloud service through the client on the cloud service platform. The tenant sends a call request to the cloud service platform, which is used to request the cloud service from the cloud service platform. The specific content of the cloud service includes providing the tenant with the data compression method described later.
[0048] In summary, the data compression method provided in the embodiment of the present application can be applied to servers, virtual machines, cloud service platforms, etc. In the following description, the server is taken as an example of the execution subject. The implementation process of other types of execution subjects is similar.
[0049] It should be noted that the terminal device mentioned in the embodiments of the present application may be a device with wireless transceiver functions, specifically referring to user equipment (UE), access terminal, subscriber unit, user station, mobile station, remote station, remote terminal, mobile device, user terminal, wireless communication equipment, user agent or user device. The terminal device may also be a satellite phone, a cellular phone, a smart phone, a wireless data card, a wireless modem, a machine type communication device, a cordless phone, a session initiation protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, an on-board device, a communication device carried on a high-altitude aircraft, a wearable device, a drone, a robot, a terminal in device-to-device (D2D) communication, a terminal in vehicle to everything (V2X), a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a mixed reality (MR), a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city ... This application does not limit the wireless terminal in the city, the wireless terminal in the smart home, or the terminal equipment in the future communication network.
[0050] In addition, in the embodiments of the present application, there may be more or fewer terminal devices in this communication system. The number and type of terminal devices are determined according to actual needs and are not specifically limited here.
[0051] See below. Figure 3 , Figure 3 A schematic diagram of a data compression method provided in an embodiment of the present application includes:
[0052] 301. Obtain attribute information of a logic block, where the logic block includes multiple data pages, and the attribute information indicates the data scale of the data included in the multiple data pages in multiple time periods.
[0053] A database is running on the server, or the server manages the database. The data in the database is stored in the form of data pages, and the amount of data in each data page is the same. In the embodiment of the present application, multiple continuous data pages are mapped to a logical block, and this operation can also be understood as logically dividing the logical blocks, and each logical block includes multiple continuous data pages. Moreover, the amount of data in each logical block is the same. In other words, the number of data pages included in each logical block is the same.
[0054] For example, see Figure 4 , Figure 4 A schematic diagram of a logic block provided in an embodiment of the present application. Figure 4 In the illustrated embodiment, N logic blocks are provided, each logic block includes n data pages. Logic block 1 includes data pages 0 to n-1, and logic block 2 includes data pages n to 2n-1.
[0055] In the embodiment of the present application, the granularity of the logic block is controllable, that is, the amount of data included in the logic block or the number of logic blocks can be set based on the needs of the actual application. Optionally, the amount of data included in the logic block or the number of logic blocks can be set by the user or defined by system parameters.
[0056] For example, assume that the data size of each data page in the database is fixed to 30MB. If the data size of the logical block set by the user is 500MB, then each logical block includes 16 data pages. If the data size of the logical block set by the user is 600MB, then each logical block includes 20 data pages.
[0057] For example, assuming that the storage capacity of the database is 100 GB, and the data volume of each data page is fixed at 64 MB. If the number of logical blocks set by the user is 200, then the data volume of each logical block is 512 MB, including 8 data pages. If the number of logical blocks set by the user is 100, then the data volume of each logical block is 1024 MB, including 16 data pages.
[0058] During the operation of the database, the database responds to data read / write operations, and the data in the database is updated accordingly, including being read, writing new data, being modified, being deleted, etc. The server obtains the attribute information of the logical block, which indicates the data size of the data in the logical block at different time periods. In other words, the attribute information reflects the distribution status of the data in the logical block.
[0059] Specifically, the attribute information includes time period information and data scale information corresponding to the time period information. The time period information indicates the time period of the data timestamp. The data timestamp indicates the time when the data was last updated or the last time the data was updated. The data scale information indicates the amount of data or the number of data rows, which is used to reflect the data scale of the corresponding time period.
[0060] In practical applications, attribute information can be expressed in many forms, including histograms, tables, key-value pairs, etc. The following diagram is used to provide a more intuitive explanation of attribute information. Figure 5 , Figure 5 A schematic diagram of a histogram is provided for an embodiment of the present application.
[0061] like Figure 5 As shown, the histogram can intuitively display the amount of data in each time period, reflecting the data scale of the data included in the logical block in each time period. For example, the timestamp of the data with a volume of m1 in the logical block is in the time period from t0 to t1, the timestamp of the data with a volume of m2 is in the time period from t1 to t2, the timestamp of the data with a volume of m3 is in the time period from t2 to t3, and the timestamp of the data with a volume of m4 is in the time period from t3 to t4.
[0062] It should be noted that Figure 5 The histogram shown is an example of an equal-width histogram, that is, Figure 5 In practical applications, the histogram may also be in the form of an equal-height histogram, a compressed histogram, etc., which is not specifically limited here.
[0063] Optionally, the attribute information may also be represented in a table format. For example, the following Table 1 is an example of the attribute information:
[0064] Table 1
[0065]
[0066] Table 1 is a reflection of the data size of each logical block in each time period in the data table t. As shown in Table 1, the data volume of the logical block chunk_0 in the T1 period is M1, the data volume in the T2 period is M2, the data volume in the T3 period is M3, and the data volume in the T4 period is M4.
[0067] Similarly, the duration of each time period in Table 1 may be the same or different, and is not specifically limited here. In addition, the time period may also be represented by a start time and an end time, such as 00:00-12:00.
[0068] It should be noted that the aforementioned Figure 5Table 1 and Table 2 all take the data volume as an example. In actual applications, the data volume can also be the number of data rows, which is not limited here. In addition, the attribute information of the logic block can be stored in the system table, fork file, or other devices accessible by the server, which is not limited here.
[0069] In an embodiment of the present application, the attribute information includes time period information and the data size corresponding to the time period information. Based on the attribute information, the relationship between the total data size of the target time period and the first threshold value can be determined, thereby determining whether the data of the corresponding logical block needs to be compressed, which provides technical support for the implementation of the technical solution of the present application and improves the feasibility of the technical solution.
[0070] 302. The attribute information indicates whether the total data size of the data included in the multiple data pages in the target time period is greater than or equal to the first threshold. If so, execute step 303; if not, execute step 304.
[0071] The server obtains the attribute information of the logic block, analyzes the data scale of each time period indicated by the attribute information, and determines whether the total data scale of the cold data included in the logic block reaches the first threshold. If the total data scale of the cold data reaches the first threshold, that is, the total data scale of the cold data is greater than or equal to the first threshold, it means that compression of the cold data is necessary. At this time, step 303 is executed, that is, the data included in the first logic block is compressed. Otherwise, step 304 is executed, that is, it is determined not to compress the data included in the logic block.
[0072] The so-called cold data refers to data whose timestamp is in the target time period, and the time interval between the target time period and the current time is greater than or equal to the time threshold. The specific value of the time threshold is set based on the needs of the actual application, and can be defined by the user or specified by the system parameters, and is not limited here.
[0073] In addition, the target period includes at least one of the multiple periods indicated by the attribute information. If the target period is one of the multiple periods, then the time interval between the target period and the current moment refers to the time interval between the end moment of the one period and the current moment. If the target period is multiple of the multiple periods, then the multiple periods included in the target period are continuous. And the time interval between the target period and the current moment refers to the time interval between the end moment of the multiple continuous periods and the current moment.
[0074] For example, Figure 5 For example, assuming that the current time is t4, the period of each histogram bucket in the figure is 1 day. The histogram bucket refers to each bar box in the figure, for example, the bar box representing the data volume m1 corresponding to t0 to t1 is a histogram bucket.
[0075] If the time threshold is 3 days, then the target period is the period from t0 to t1, and the total data size of the target period is m1.
[0076] If the time threshold is 1 day, then the target period is t0 to t3, including the three periods t0 to t1, t1 to t2, and t2 to t3. The total data size of the target period is m1+m2+m3.
[0077] In some optional implementations, if the time periods indicated by the attribute information are equal time periods, the time threshold may be an integer multiple of the time period indicated by the time period information. Figure 5 For example, assuming the time threshold is 4 days, the duration of each period in the figure can be 0.5 days, 1 day, 2 days, etc. By setting the time threshold to an integer multiple of the period information, it can be directly determined whether all the data in each histogram bucket is cold data or not, and there is no need to further determine which data in each histogram bucket is cold data and which data is not cold data. This makes the determination of cold data easier.
[0078] For example, assuming the time threshold is 2 days, Figure 5 The duration of each histogram bucket is 1 day, so the server can directly determine whether all the data in a histogram bucket is cold data. If the time threshold is 3 days, the duration of each histogram bucket is 2 days, and if the corresponding period of a histogram bucket is 2-4 days before the current time, then the server needs to further analyze which data in the histogram bucket has a time difference of more than or equal to 2 days from the current time, so as to determine the cold data included in the histogram bucket.
[0079] In addition, the aforementioned step 301 introduces that the data scale includes the data volume or the number of data rows, and correspondingly, the first threshold includes the data volume threshold or the data row number scale. The specific value of the first threshold is also set based on the needs of the actual application, and can be defined by the user or specified by the system parameter, which is not specifically limited here.
[0080] For example, it is assumed that the time set by the compression scheduling strategy is 2 days, that is, the time threshold is 2 days, and the duration of each time period defined by the attribute information is 1 day. Figure 5 For example, suppose Figure 5 The times t1 to t4 are 22:00 on November 9, 22:00 on November 10, 22:00 on November 11, 22:00 on November 12, and 22:00 on November 13, respectively.
[0081] If the current time is t4, then the time periods from t0 to t1 and from t1 to t2 are target time periods, and the total amount of data corresponding to these two time periods is m1+m2. The server compares the total amount of data m1+m2 with the data amount threshold to determine whether the scale of cold data reaches the first threshold.
[0082] 303. Compress the data included in the logic block and store the compressed result.
[0083] The server compresses the data included in the logical block only when the total data size of the cold data in the logical block is greater than or equal to the first threshold. That is to say, in the embodiment of the present application, compression is not performed as soon as cold data is present, but the relationship between the total data size of the cold data and the first threshold is considered. Then, by reasonably setting the value of the first threshold, not only can the waste of resources caused by frequent compression be avoided, but also the data compression of cold data can be determined to be profitable. In other words, the first threshold can be understood as a parameter used to balance computing resources and compression benefits.
[0084] The server compresses the data included in the logic block by determining the data page where the cold data is located from the multiple data pages included in the logic block, and performing a compression operation on the data page where the cold data is located. The timestamp of the cold data is in the target time period. In other words, the data compressed by the server is the data whose timestamp is in the target time period.
[0085] Specifically, the server starts from the starting data page of the multiple data pages included in the logical block, and sequentially obtains the cold data included in each data page until the data size of the cold data reaches the total data size of the multiple data pages in the target time period. If there is a compression gain in the compression of the cold data, the compression operation is performed on the data page where the cold data is located. In other words, the data page compressed by the server includes cold data and has compression gain. Then, in the process of performing the compression operation, the server needs to determine the starting data page of the logical block and determine whether there is a compression gain. This process is described in detail below. When compressing the data included in the logical block, the server knows the identifier of the logical block. Based on the identifier of the logical block, the server can determine the starting data page among the multiple data pages included in the logical block. There are many possible ways for the server to determine the starting data page, which are described below.
[0086] In some optional implementations, the server may obtain a mapping relationship between an identifier indicating a logic block and a number of a data page, and determine the starting data page included in the logic block based on the mapping relationship. Specifically, the server obtains a second correspondence set, and the second correspondence set includes a correspondence between the identifier of the logic block and the number of the data page, wherein the identifier of one logic block corresponds to the numbers of multiple data pages, and the number of a data page uniquely corresponds to the identifier of one logic block. The server determines the starting data page included in the logic block based on the second correspondence set and the identifier of the logic block.
[0087] Exemplarily, the second correspondence set may be as shown in Table 2 below:
[0088] Table 2
[0089] chunk_id page_number chunk_0 page_0, page_1...page_n-1 …… …… chunk_N page_Nn, page_Nn+1……page_Nn+n-1
[0090] As shown in Table 2, the multiple data pages included in the logic block chunk_0 are numbered page_0, page_1, ... page_n-1. In the database, the data page numbers are continuous. This means that the starting data page included in the logic block chunk_0 is data page page_0. Based on the identification of the logic block, the server can determine the numbers of multiple data pages that match the identification of the logic block from the second corresponding relationship, thereby determining the starting data page.
[0091] It should be noted that Table 2 is only an example of the second correspondence set. In practical applications, the second correspondence set is not limited to the form of a table, but can also be a key-value pair, or other forms that can reflect the correspondence between the identifier of the logic block and the number of the data page, which is not specifically limited here. In addition, the second correspondence set can be stored locally on the server, remotely, or on other devices accessible to the server, which is not specifically limited here.
[0092] In some optional implementations, the server may obtain a correspondence between a logic block and a starting data page included therein. Based on the correspondence and an identifier of the logic block, the server may determine the starting data page included in the logic block.
[0093] Exemplarily, it is assumed that the relationship between the identifier of the logic block and the starting data page satisfies the following formula 1:
[0094] chunk_id=start-page_number×page_size……Formula 1
[0095] Among them, chunk_id represents the logical block identifier, start-page_number represents the number of the starting data page, and page_size represents the amount of data in the data page. In the database, the number of data pages is fixed and cannot be modified manually. The amount of data in each data page is the same, and the server can obtain this information from the system parameters. Then, based on the identifier of the logical block and the amount of data on the data page, combined with formula 1, the server can determine the number of the starting data page of the logical block.
[0096] Exemplarily, it is assumed that the relationship between the identifier of the logic block and the starting data page satisfies the following formula 2:
[0097] chunk_id=start-page_number×chunk_size...Formula 2
[0098] Wherein, chunk_size indicates the data size of the data page, which is defined by the user or specified by the system parameter, and the server can obtain the information. Combining the identification of the logic block with formula 2, the number of the starting data page of the logic block can be determined.
[0099] In the embodiment of the present application, there are multiple possible ways for the server to determine the starting data page included in the logic block, which enriches the implementation methods and application scenarios of the technical solution of the present application and can be selected based on the needs of actual applications, thereby improving the flexibility of the technical solution.
[0100] The server obtains cold data included in the data page, which means that the server reads the timestamp of the data in the data page and determines the time interval between the timestamp and the current time. If the time interval is greater than or equal to the time threshold, the data is determined to be cold data.
[0101] In addition, the server may not read all the data pages included in the logical block. This is because the server has already obtained the total data size of the cold data in the logical block before reading the logical block. Then, in the process of sequentially reading the data pages of the logical block, when the data size of the cold data read by the server reaches the total data size indicated by the attribute information, it means that the cold data included in the logical block has been read, and the data that has not been read is not cold data, and there is no need to compress it. The server does not need to read the remaining data, which can also save I / O overhead.
[0102] After obtaining the cold data, the server pre-compresses the cold data to determine whether there is any compression benefit from the compression of the cold data. The so-called pre-compression means that compression is performed at the memory level without changing the physical form of data storage. The compression benefit is determined by comparing the storage space occupied by the cold data before compression with the storage space occupied after compression. It can be understood that after data compression, in addition to storing the compressed data itself, the parameters used for decompression must also be stored. Therefore, the compressed data includes compressed data and related parameters. If the storage space occupied by the compressed data is less than the storage space occupied by the data before compression, it is considered that data compression is profitable and can be compressed. Otherwise, there is no compression benefit and the data is not compressed.
[0103] In the embodiment of the present application, in the scheme of compressing the data included in the logical block, the operation of obtaining cold data may not traverse all data pages included in the logical block until the scale of the cold data indicated by the first attribute information is reached, thereby saving I / O overhead. In addition, the server will compress the cold data only when it is determined that there is a compression benefit for the cold data, thereby reducing the storage space occupied.
[0104] 304. Determine whether to uncompress the data included in the logical block.
[0105] In the scenario where the total data size of the cold data indicated by the attribute information is less than the first threshold, it means that the cold data included in the logical block has not yet met the compression condition, and the server will not compress the data included in the logical block. It can also be understood that in the current compression task, the data in the logical block is kept unchanged.
[0106] In an embodiment of the present application, multiple data pages are mapped to a logical block, and whether to compress the data included in the logical block is determined based on the attribute information of the logical block. Specifically, the attribute information indicates the data scale of the multiple data pages included in the logical block in different time periods. If the total data scale of the cold data in the multiple data pages is greater than or equal to the first threshold, the data included in the logical block is compressed. Otherwise, the data included in the logical block is not compressed. Among them, the timestamp of the cold data is in the target time period. In other words, for the logical block whose total data scale of the cold data is less than the first threshold, the compression operation will not be performed, which reduces the I / O overhead and saves system resources.
[0107] In the aforementioned step 301, it is introduced that the server obtains the attribute information of the logic block. Before step 301, the server also obtains the first correspondence set and the identifier of the logic block. The first correspondence set indicates that the identifiers of multiple logic blocks correspond to multiple attribute information one by one, and the identifier of the logic block uniquely indicates the logic block. Then, the server obtains the attribute information of the logic block, specifically, obtains the attribute information from the first correspondence set according to the identifier of the logic block.
[0108] Further, the first correspondence set includes correspondences between identifiers of multiple logic blocks and multiple attribute information, each identifier of a logic block uniquely corresponds to one attribute information, and each attribute information uniquely corresponds to one identifier of a logic block. Alternatively, it can be understood that the server maintains one attribute information for each logic block, and the attribute information can be in the form of the aforementioned histogram, system table, etc.
[0109] Optionally, the first correspondence set may be stored locally on the server, such as in a system table or a fork file, or in the cloud, or on a device accessible to other servers.
[0110] The server also needs to obtain the identification of the logical block, which is actually the identification of the logical block where the data page currently read is located when the server performs the compression scheduling task. In the embodiment of the present application, there are many possible ways for the server to obtain the identification of the logical block, which are described below.
[0111] In some optional implementations, the server may also obtain a second correspondence set, the second correspondence set including a correspondence between the identifier of the logic block and the number of the data page, wherein the identifier of one logic block corresponds to the numbers of multiple data pages, and the number of one data page uniquely corresponds to the identifier of one logic block. The number of the data page currently being read is obtained. According to the number of the data page currently being read, the identifier of the logic block is determined from the second correspondence set.
[0112] Specifically, the server obtains the number of the data page currently being read, queries the second correspondence set, and can obtain the identification of the logic block. The specific implementation method and storage location of the second correspondence set have been introduced in the previous description of Table 2, and will not be repeated here.
[0113] In an embodiment of the present application, the server can also directly determine the identifier of the logical block where the currently read data page is located based on the second correspondence set, which is simple to operate and saves computing resources.
[0114] In some optional implementations, the server obtains the number of the data page currently being read, the data volume of the data page, and the data volume of the logical block. According to the number of the data page currently being read, the data volume of the data page, and the data volume of the logical block, the number of the target data page in the logical block is determined, and the target data page is the starting data page or the ending data page of the logical block. According to the number of the target data page and the target data volume, the identifier of the logical block is determined, and the target data volume is the data volume of the logical block, or the data volume of any data page in the logical block.
[0115] Specifically, the data page is divided into multiple logical blocks, starting from the first data page of the database, and mapping multiple consecutive data pages to one logical block. In addition, the server can determine the number of the first data page of the database based on the system parameters. First, the server can determine the number of data pages included in each logical block based on the data volume of the data page and the data volume of the logical block. Then, according to the number of the data page currently read, the number of data pages included in each logical block and the number of the first data page of the database, the number of the starting data page of the logical block where the data page currently read is located, or the number of the ending data page, can be determined. The relationship between the identification of the logical block, the number of the target data page, and the target data volume is also defined in the system configuration. Combined with this relationship, the server can determine the identification of the first logical block.
[0116] Optionally, the target data page is the starting data page of the logical block, and the target data volume is the data volume of the data page. That is to say, the relationship between the identifier of the logical block and the starting data page satisfies the aforementioned formula 1. For example, based on the data volume of the logical block and the data volume of the data page, it is determined that each logical block includes 5 data pages. The first data page of the database is numbered 0. The number of the data page currently read by the server is 12, and the server determines that the data page currently read is a data page in the third logical block, and the data pages included in the third logical block are numbered 9 to 14. The third logical block is the logical block where the data page currently being read is located. Then, the starting data page of the third logical block is numbered 9. If the data volume of each data page is 30MB, then the identifier of the logical block where the data page currently being read is located is 9×30=270.
[0117] Optionally, the target data page is the starting data page of the logical block, and the target data volume is the data volume of the logical block. That is to say, the relationship between the identifier of the logical block and the starting data page satisfies the aforementioned formula 2. For example, based on the data volume of the logical block and the data volume of the data page, it is determined that each logical block includes 5 data pages. The first data page of the database is numbered 0. The number of the data page currently read by the server is 12, and the server determines that the data page currently read is a data page in the third logical block, and the data pages included in the third logical block are numbered 9 to 14. The third logical block is the logical block where the data page currently being read is located. Then, the starting data page of the third logical block is numbered 9. If the data volume of the logical block is 90MB, then the identifier of the logical block where the data page currently being read is located is 9×90=810.
[0118] Optionally, the target data page is the end data page of the logic block, and the target data volume is the data volume of the data page. In other words, the relationship between the identifier of the logic block and the end data page satisfies Formula 3:
[0119] chunk_id=end-page_number×page_size...Formula 3
[0120] Wherein, end-page_number indicates the number of the ending data page. For example, based on the data volume of the logical block and the data volume of the data page, it is determined that each logical block includes 5 data pages. The first data page of the database is numbered 0. The data page currently read by the server is numbered 12, and the server determines that the data page currently read is a data page in the third logical block, and the data pages included in the third logical block are numbered 9 to 14. The third logical block is the logical block where the data page currently read is located. Then, the ending data page of the third logical block is numbered 14. If the data volume of each data page is 30MB, then the identifier of the logical block where the data page currently read is located is 14×30=420.
[0121] Optionally, the target data page is the end data page of the logic block, and the target data volume is the data volume of the logic block. In other words, the relationship between the identifier of the logic block and the end data page satisfies Formula 4:
[0122] chunk_id=end-page_number×chunk_size……Formula 4
[0123] For example, based on the data volume of the logical block and the data volume of the data page, it is determined that each logical block includes 5 data pages. The first data page of the database is numbered 0. The data page currently read by the server is numbered 12, and the server determines that the data page currently read is a data page in the third logical block, and the data pages included in the third logical block are numbered 9 to 14. The third logical block is the logical block where the data page currently read is located. Then, the number of the end data page of the third logical block is 14. If the data volume of each data page is 90MB, then the identifier of the logical block where the data page currently read is located is 14×90=1260.
[0124] In the embodiment of the present application, the target data page in the logic block where the data page currently read is located is determined by the number of the data page currently read by the server, and then combined with the target data volume, the identifier of the logic block where the data page currently read is located is determined. The target data page can be the starting data page of the logic block or the ending data page of the logic block; the target data volume can be the data volume of the data page or the data volume of the logic block, which provides multiple possibilities for the implementation of the technical solution of the present application and enriches the application scenarios and implementation methods.
[0125] In some optional implementations, the server obtains the number of the data page currently being read and the target data volume, where the target data volume is the data volume of the logical block, or the data volume of any one of the multiple data pages. The target identifier is determined based on the number of the data page currently being read and the target data volume. Then, based on the identifier corresponding to the logical page currently being read and the first correspondence set, the identifier of the logical block where the data page currently being read is located is determined, and the target identifier is the same as the identifier of the logical block where the data page currently being read is located, or the target identifier is smaller than or larger than the identifier of the logical block where the data page currently being read is located, and the identifier of the logical block where the data page currently being read is located in the first correspondence set is closest to the target identifier.
[0126] Specifically, the system configuration defines the relationship between the identifier of the logical block and the target data volume, which can be specifically the relationship shown in any one of the aforementioned formulas 1 to 4. In addition, the first correspondence set includes identifiers of multiple logical blocks. If the target identifier determined by the server based on the number of the data page currently being read and the target data volume is included in the first correspondence set, then the target identifier is the identifier of the logical block where the data page currently being read is located. If the target identifier is not in the first correspondence set, the server can determine the identifier of the logical block where the data page currently being read is located based on the target identifier and the first correspondence set. The following is a detailed explanation with examples.
[0127] Exemplarily, it is assumed that the target data volume is the data volume of the data page, and the identifier of the logical block and the data volume of the data page satisfy the relationship defined by the aforementioned formula one. For example, the first data page in the database is numbered 1, each logical block includes 3 data pages, and the identifiers of the logical blocks included in the first correspondence set are 10, 40, and 60. The data page currently read by the server is numbered 5, and the data volume of the data page is 10. The server determines the target identifier to be 50. Since the identifier of the logical block and the data volume of the data page satisfy the aforementioned formula one, that is, the system configuration defines the identifier of the logical block based on the number of the starting data page. Then, the server determines from the first correspondence set that the identifier 40 of the logical block that is smaller than the target identifier and closest to the target identifier is the identifier of the logical block where the data page currently being read is located.
[0128] Exemplarily, it is assumed that the target data volume is the data volume of the data page, and the identifier of the logical block and the data volume of the data page satisfy the relationship defined by the aforementioned formula three. For example, the first data page in the database is numbered 1, each logical block includes 3 data pages, and the identifiers of the logical blocks included in the first correspondence set are 10, 40, and 60. The data page currently read by the server is numbered 5, and the data volume of the data page is 10. The server determines the target identifier to be 50. Since the identifier of the logical block and the data volume of the data page satisfy the aforementioned formula three, that is, the system configuration defines the identifier of the logical block based on the number of the end data page. Then, the server determines from the first correspondence set that the identifier 60 of the logical block that is larger than the target identifier and closest to the target identifier is the identifier of the logical block where the data page currently being read is located.
[0129] Similarly, the target data volume is also the data volume of the logical block, and the identification of the logical block and the number of logical blocks satisfy the relationship defined by the aforementioned formula 2 or the aforementioned formula 4. Similar to the principles of the previous two examples, the server determines the identification of the logical block where the currently read data page is located, which will not be repeated here.
[0130] In an embodiment of the present application, the server determines the target identifier based on the number of the data page currently being read and the target data volume, and then compares the target identifier with the identifier of the logical block included in the first correspondence set to determine the identifier of the logical block where the data page currently being read is located. Since there are multiple possibilities for the relationship between the identifier of the logical block and the target data volume defined by the system configuration, there are also multiple implementation methods for the server to determine the identifier of the logical block where the data page currently being read is located, which enriches the application scenarios of the technical solution of the present application and improves flexibility. In addition, an embodiment of the present application provides multiple solutions to determine the identifier of the logical block where the data page currently being read is located, which can be flexibly selected based on the needs of actual applications, enriching the implementation methods and application scenarios of the technical solution of the present application.
[0131] In addition, the embodiment of the present application maintains an attribute information for each logic block, and stores a one-to-one correspondence between the identifiers of multiple logic blocks and multiple attribute information in the first correspondence set, so that the server determines the attribute information corresponding to the logic block from the first correspondence set based on the identifier of the logic block where the currently read data page is located. This provides technical support for the implementation of the technical solution of the present application and improves the feasibility of the technical solution.
[0132] In the previous description, the compression of a logical block is taken as an example. In actual applications, data pages in the database may be mapped to multiple logical blocks. The following is a diagram to illustrate this scenario.
[0133] In some optional implementations, the server completes the compression scheduling task in a serial manner, that is, determines one by one whether the data included in each logical block needs to be compressed. Figure 6 , Figure 6 A flowchart of a data compression method provided in an embodiment of the present application.
[0134] like Figure 6 As shown, the compression scheduling task starts, the server obtains the attribute information of the logic block where the currently read data page is located, and determines whether the cold data scale of the logic block indicated by the attribute information has reached a first threshold. If so, the data included in the logic block is compressed. If not, it is determined not to compress the data included in the logic block, and the attribute information of the next logic block is obtained to determine whether to compress the data included in the next logic block based on the attribute information. Until each logic block in the database is traversed, the compression of the database is completed.
[0135] In some optional implementations, the server can also complete the compression scheduling task in a parallel manner, that is, the server simultaneously determines the attribute information of multiple logical blocks corresponding to the database to determine whether to compress the data included in the multiple logical blocks. In other words, the server synchronously performs the compression scheduling task for the data of each logical block.
[0136] It is understandable that no matter which method is used, the principle of the server compressing the data included in a logical block is the same as mentioned above. Figure 3 The principles described in the related embodiments are similar and will not be repeated here.
[0137] Next, the related equipment provided in the embodiments of the present application is described.
[0138] See also Figure 7 , Figure 7 A schematic diagram of the structure of a data compression device provided in an embodiment of the present application. Figure 7 As shown, the data compression device 700 includes an acquisition unit 701 and a processing unit 702 .
[0139] In some optional implementations, the acquisition unit 701 is used to acquire attribute information of a logic block, where the logic block includes multiple data pages, and the attribute information indicates the data scale of the data included in the multiple data pages in multiple time periods.
[0140] The processing unit 702 is configured to compress the data included in the logic block and store the compressed result if the attribute information indicates that the total data size of the data included in the multiple data pages in the target time period is greater than or equal to the first threshold value. The target time period includes at least one time period among the multiple time periods, and the time interval between the target time period and the current time is greater than or equal to the time threshold value.
[0141] The processing unit 702 is further configured to determine not to compress the data included in the logical block if the attribute information indicates that the total data size of the data included in the multiple data pages in the target time period is smaller than a first threshold.
[0142] In some optional implementations, the attribute information includes time period information and data size information corresponding to the time period information. The time period information indicates the time period in which the timestamp of the data is located.
[0143] The data scale information indicates the data volume, and the first threshold value includes a data volume threshold value. Alternatively, the data scale information indicates the number of data rows, and the first threshold value includes a data row number threshold value.
[0144] In some optional implementations, the acquisition unit 701 is further used to: acquire a first correspondence set, the first correspondence set indicating that the identifiers of multiple logic blocks correspond to the multiple attribute information one by one. Acquire the identifier of the logic block, the identifier of the logic block indicates the logic block. The acquisition unit 701 is specifically used to acquire the attribute information corresponding to the logic block from the first correspondence set according to the identifier of the logic block.
[0145] In some optional implementations, the acquisition unit 701 is specifically used to: acquire a second correspondence set, the second correspondence set including a correspondence between the identifier of a logic block and the number of a data page, wherein the identifier of one logic block corresponds to the numbers of multiple data pages, and the number of one data page uniquely corresponds to the identifier of one logic block. Acquire the number of the data page currently being read. Determine the identifier of the logic block from the second correspondence set according to the number of the data page currently being read.
[0146] In some optional implementations, the acquisition unit 701 is specifically used to: acquire the number of the data page currently being read, the data volume of the data page, and the data volume of the logic block, where the data page currently being read is included in the first logic block. According to the number of the data page currently being read, the data volume of the data page, and the data volume of the logic block, the number of the target data page in the first logic block is determined, where the target data page is the starting data page or the ending data page of the first logic block. According to the number of the target data page and the target data volume, the identifier of the first logic block is determined, where the target data volume is the data volume of the first logic block, or the data volume of any one data page in the first logic block.
[0147] In some optional implementations, the acquisition unit 701 is specifically used to: acquire the number of the data page currently being read and the target data volume, the target data volume being the data volume of the first logical block, or the data volume of any one of the multiple data pages. Determine the target identifier according to the number of the data page currently being read and the target data volume. Determine the identifier of the first logical block according to the identifier corresponding to the logical page currently being read and the first correspondence set, the target identifier is the same as the identifier of the first logical block, or the target identifier is smaller than or greater than the identifier of the first logical block, and the identifier of the first logical block is closest to the target identifier in the first correspondence set.
[0148] In some optional implementations, the processing unit 702 is specifically configured to: obtain a data page where cold data is located from the multiple data pages, where the timestamp of the cold data is in the target time period, and perform a compression operation on the data page where the cold data is located.
[0149] The acquisition unit 701 and the processing unit 702 may be implemented by software or hardware. For example, the implementation of the processing unit 702 is described below by taking the processing unit 702 as an example. Similarly, the implementation of the acquisition unit 701 may refer to the implementation of the processing unit 702.
[0150] As an example of a software functional unit, the processing unit 702 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the processing unit 702 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with similar geographical locations. Typically, a region may include multiple AZs.
[0151] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0152] As an example of a hardware functional unit, the processing unit 702 may include at least one computing device, such as a server, etc. Alternatively, the processing unit 702 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0153] The multiple computing devices included in the processing unit 702 can be distributed in the same region or in different regions. The multiple computing devices included in the processing unit 702 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the processing unit 702 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0154] It should be noted that the acquisition unit 701 and the processing unit 702 respectively implement different steps in the data compression method to implement all functions of the data compression device 700. The data compression device 700 is used to implement the data compression method provided in the embodiment of the present application, which will not be described in detail here.
[0155] See also Figure 8 , Figure 8 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. The computing device 800 includes a processor 801, a communication interface 802, a bus 803, and a memory 804. The processor 801, the communication interface 802, and the memory 804 communicate with each other via the bus 803. In practical applications, communication can also be achieved through other means such as wireless transmission, which is not limited here.
[0156] It should be understood that the present application does not limit the number of processors and memories in the computing device 800 .
[0157] The processor 801 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP) or a digital signal processor (DSP).
[0158] The communication interface 802 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 800 and other devices or a communication network.
[0159] The bus 803 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 The bus 803 is represented by only one line, but it does not mean that there is only one bus or one type of bus. The bus 803 may include a path for transmitting information between various components of the computing device 800 (for example, the memory 804, the processor 801, and the communication interface 802).
[0160] The memory 804 may include a volatile memory, such as a random access memory (RAM). The memory 804 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0161] The memory 804 stores executable program codes, and the processor 801 executes the executable program codes to respectively implement the functions of the acquisition unit 701 and the processing unit 702, thereby implementing the data compression method. That is, the memory 804 stores instructions for executing the data compression method.
[0162] The embodiment of the present application also provides a computing device cluster, which includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some optional implementations, the computing device can also be a terminal device such as a desktop computer or a laptop computer.
[0163] See also Fig. 9 and Fig.10 , Fig. 9 and Fig.10 All of them are structural schematic diagrams of the computing device cluster provided in embodiments of the present application.
[0164] like Fig. 9 As shown, the computing device cluster includes at least one computing device 800. The memory 804 in one or more computing devices 800 in the computing device cluster may store the same instructions for executing the data compression method provided in the embodiment of the present application.
[0165] In some possible implementations, the memory 804 of one or more computing devices 800 in the computing device cluster may also store partial instructions for executing the data compression method. In other words, the combination of one or more computing devices 804 may jointly execute instructions for executing the data compression method.
[0166] It should be noted that the memory 804 in different computing devices 800 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the data compression device. That is, the instructions stored in the memory 804 in different computing devices 800 can implement the functions of one or more units in the acquisition unit 701 and the processing unit 702.
[0167] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig.10 A possible implementation is shown. Fig.10 As shown, two computing devices 800A and 800B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 804 in the computing device 800A stores instructions for executing the functions of the acquisition unit 701. At the same time, the memory 804 in the computing device 800B stores instructions for executing the functions of the processing unit 702.
[0168] Fig.10 The connection method between the computing device clusters shown can be based on the consideration that in the data compression method provided in the present application, processing operations and operations other than processing operations are performed separately, that is, the function of the acquisition unit 701 is considered to be executed by the computing device 800A, and the function of the processing unit 702 is considered to be executed by the computing device 800B.
[0169] It should be understood that Fig.10The functions of the computing device 800A shown in FIG. 8 may also be completed by multiple computing devices 800. Similarly, the functions of the computing device 800B may also be completed by multiple computing devices 800.
[0170] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Fig. 9 and Fig.10 The connection method of the computing device cluster will not be described in detail here.
[0171] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computer device, the at least one computer device executes the above data compression method.
[0172] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to perform the above-mentioned data compression method.
[0173] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data compression method, characterized in that: include: Acquire attribute information of a logic block, the logic block including a plurality of data pages, the attribute information indicating data scales of data included in the plurality of data pages in a plurality of time periods; If the attribute information indicates that the total data size of the data included in the multiple data pages in the target time period is greater than or equal to a first threshold, compressing the data included in the logic block and storing the compressed result; wherein the target time period includes at least one time period among the multiple time periods, and the time interval between the target time period and the current time is greater than or equal to the time threshold; If the first attribute information indicates that the total data size of the data included in the plurality of data pages in the target time period is smaller than the first threshold, it is determined not to compress the data included in the logical block.
2. The method according to claim 1, characterized in that The attribute information includes time period information and data scale information corresponding to the time period information; The time period information indicates the time period in which the timestamp of the data is located; The data scale information indicates the data volume, and the first threshold includes a data volume threshold; or, The data scale information indicates the number of data rows, and the first threshold includes a data row number threshold.
3. The method according to claim 1 or 2, characterized in that: Before acquiring the attribute information of the logic block, the method further includes: Acquire a first correspondence set, wherein the first correspondence set indicates a one-to-one correspondence between identifiers of a plurality of logic blocks and a plurality of attribute information; Acquire an identifier of the logic block, wherein the identifier of the logic block indicates the logic block; The obtaining of attribute information of the logic block includes: According to the identifier of the logic block, the attribute information corresponding to the logic block is acquired from the first correspondence set.
4. The method according to claim 3, characterized in that The obtaining the identification of the logic block includes: Acquire a second correspondence set, the second correspondence set including correspondences between identifiers of logic blocks and numbers of data pages, wherein an identifier of one logic block corresponds to numbers of multiple data pages, and a number of a data page uniquely corresponds to an identifier of one logic block; Get the number of the data page currently being read; According to the number of the data page currently being read, the identifier of the logic block is determined from the second correspondence set.
5. The method according to any one of claims 1 to 4, characterized in that The compressing the data included in the logic block includes: Acquire a data page where cold data is located from the multiple data pages, where a timestamp of the cold data is in the target time period; A compression operation is performed on the data page where the cold data is located.
6. A data compression device, characterized in that: include: An acquisition unit, configured to acquire attribute information of a logic block, wherein the logic block includes a plurality of data pages, and the attribute information indicates data scales of data included in the plurality of data pages in a plurality of time periods; a processing unit, configured to compress the data included in the logic block and store the compressed result if the attribute information indicates that the total data size of the data included in the multiple data pages in a target time period is greater than or equal to a first threshold; wherein the target time period includes at least one time period among the multiple time periods, and the time interval between the target time period and the current time is greater than or equal to a time threshold; The processing unit is further configured to determine not to compress the data included in the logic block if the attribute information indicates that a total data size of the data included in the plurality of data pages in a target time period is smaller than the first threshold.
7. The device according to claim 6, characterized in that The attribute information includes time period information and data scale information corresponding to the time period information; The time period information indicates the time period in which the timestamp of the data is located; The data scale information indicates the data volume, and the first threshold includes a data volume threshold; or, The data scale information indicates the number of data rows, and the first threshold includes a data row number threshold.
8. The device according to claim 6 or 7, characterized in that The acquisition unit is further used for: Acquire a first correspondence set, wherein the first correspondence set indicates a one-to-one correspondence between identifiers of a plurality of logic blocks and a plurality of attribute information; Acquire an identifier of the logic block, wherein the identifier of the logic block indicates the logic block; The acquisition unit is specifically configured to acquire the attribute information corresponding to the logic block from the first correspondence set according to the identifier of the logic block.
9. The device according to claim 8, characterized in that The acquisition unit is specifically used for: Acquire a second correspondence set, the second correspondence set including correspondences between identifiers of logic blocks and numbers of data pages, wherein an identifier of one logic block corresponds to numbers of multiple data pages, and a number of a data page uniquely corresponds to an identifier of one logic block; Get the number of the data page currently being read; According to the number of the data page currently being read, the identifier of the logic block is determined from the second correspondence set.
10. The device according to any one of claims 6 to 9, characterized in that The processing unit is specifically used for: Acquire a data page where cold data is located from the multiple data pages, where a timestamp of the cold data is in the target time period; A compression operation is performed on the data page where the cold data is located.
11. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to implement the method according to any one of claims 1 to 5.
12. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the method according to any one of claims 1 to 5 is implemented.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes computer program instructions, which, when executed by a computing device cluster, implement the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Data processing method, device thereof and flash-memory storage system
CN101526923A
Data compression method and device
CN109802684A
Data compression method and device and computer readable storage medium
CN111984610A
Cold data identification method and flash memory device
CN116149549A
Block Compression in a Key / Value Store
US20140215170A1
Cited By
Data compression method and related device
WO2026113307A1