Data storage method and device, equipment, medium and program product

By dynamically determining partition weights and compression methods based on data characteristics, the problems of low data storage efficiency and waste of resources in the prior art are solved, and flexible and efficient data storage and query are realized.

CN120491906AActive Publication Date: 2025-08-15INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510948966.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-08-15
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

When the existing data storage methods face the increase in data volume, there are problems of space waste and inefficiency. Especially in the dynamic data access mode, static partitioning and compression strategies cannot adapt to changes in the data access mode, resulting in waste of storage resources and degradation of computing performance.

Method used

By determining the partition weight based on the characteristics of the data to be stored, dynamically match the target data partition, and selecting a suitable compression method for storage, a layered compression decision tree is built to adapt to different data characteristics, and a secondary query index is built to improve query efficiency.

Benefits of technology

It realizes efficient utilization of data storage, improves storage efficiency and query speed, reduces manual intervention, dynamically adapts to changes in data access mode, and avoids the shortcomings of a single compression algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491906A_ABST
    Figure CN120491906A_ABST
Patent Text Reader

Abstract

The invention provides a data storage method and device, equipment, a medium and a program product, which can be applied to the technical fields of computers and data storage. The method comprises the steps of determining partition weights of to-be-stored data according to data features of the to-be-stored data; determining a target data partition corresponding to the to-be-stored data from the plurality of data partitions according to the partition weight; determining a target compression mode of the to-be-stored data from a plurality of compression modes according to the target data partition; and compressing the to-be-stored data according to the target compression mode, and storing the compressed to-be-stored data to the target data partition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computer technology and data storage, and in particular to a data storage method, apparatus, device, medium and program product. Background Art

[0002] During system operation, a large amount of data is generated. This data can reflect the system's operation process and provide a better understanding of the system's working status. Data storage ensures data persistence, so it is usually necessary to store this data. However, with the rapid development of information technology, the data generated during system operation is also increasing. However, data storage methods often waste space and are inefficient. Summary of the Invention

[0003] In view of the above problems, the present application provides a data storage method, apparatus, device, medium and program product.

[0004] According to the first aspect of the present application, a data storage method is provided, including: determining a partition weight of the data to be stored based on data characteristics of the data to be stored; determining a target data partition corresponding to the data to be stored from multiple data partitions based on the partition weight; determining a target compression method for the data to be stored from multiple compression methods based on the target data partition; and compressing the data to be stored and storing it in the target data partition based on the target compression method.

[0005] The second aspect of the present application provides a data storage device, including: a first determination module, used to determine the partition weight of the data to be stored according to the data characteristics of the data to be stored; a second determination module, used to determine the target data partition corresponding to the data to be stored from multiple data partitions according to the partition weight; a third determination module, used to determine the target compression method of the data to be stored from multiple compression methods according to the target data partition; and a first storage module, used to compress the data to be stored according to the target compression method and then store it in the target data partition.

[0006] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0007] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0008] The fifth aspect of the present application further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.

[0009] According to the embodiments of the present application, the partition weights of the data to be stored are determined based on the data characteristics of the data to be stored. The corresponding target data partitions can be matched according to the characteristics of the data to be stored, thereby achieving classified storage of the data to be stored and rationally utilizing storage space. Different data partitions correspond to different compression methods. Compressing the data to be stored according to the target compression method corresponding to the target data partition allows the data to be compressed using an appropriate compression method, increasing the diversification of data compression, avoiding the problem of a single compression algorithm being unable to balance speed and compression rate, and improving the efficiency of data storage. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:

[0011] Figure 1 An application scenario diagram of the data storage method, apparatus, device, medium, and program product according to an embodiment of the present application is shown.

[0012] Figure 2 A flow chart of a data storage method according to an embodiment of the present application is shown.

[0013] Figure 3 A flow chart of a data storage method according to another embodiment of the present application is shown.

[0014] Figure 4 A system block diagram of the data storage method according to an embodiment of the present application is shown.

[0015] Figure 5 A structural block diagram of a data storage device according to an embodiment of the present application is shown.

[0016] Figure 6 A block diagram of an electronic device suitable for implementing a data storage method according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0017] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.

[0018] The terms used herein are only for describing specific embodiments and are not intended to limit this application. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0020] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0021] In the technical solution of this application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0022] In the scenario of using personal information for automated decision-making, the methods, devices, and systems provided in the embodiments of the present application all provide users with corresponding operation portals for users to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge and skills, and have reached a certain level of professionalism.

[0023] Data storage usually relies on static threshold divisions (such as fixed time windows or single access frequency) and is bound to a single compression algorithm. It cannot dynamically adapt to fluctuations in data access patterns. As a result, old data with high frequency access is misclassified as cold data due to outdated rules, inefficient compression algorithms exacerbate query latency, and manual parameter adjustments are required on a regular basis. The monitoring system and storage engine lack real-time linkage, ultimately resulting in a vicious cycle of wasted storage resources and degraded computing performance. It is difficult to balance the storage efficiency and response speed requirements in large-scale real-time scenarios.

[0024] In one example, row-based data storage resulted in high redundancy when storing time-series data, with cold data occupying space for long periods. Static compression strategies, which employed fixed compression algorithms and did not dynamically adjust based on data access patterns, resulted in insufficient compression rates for cold data and high decompression latency for hot data. As data volumes grew, the maintenance cost of time-range-based B+ tree indexes (a data structure) skyrocketed, and query response times increased exponentially. Cross-partition query efficiency was low: Static partitioning strategies (such as monthly partitioning) required cross-partition queries to merge multiple tables, resulting in significant I / O overhead and inefficient cross-partition queries. Data storage methods that relied on preset time thresholds (such as "archiving data for more than three years") lacked dynamic adaptability and could not adjust storage strategies based on real-time query hotspots. Storage parameters required manual intervention, making it difficult to adapt to dynamic changes in business access patterns.

[0025] In view of this, an embodiment of the present application provides a data storage method, including: determining the partition weight of the data to be stored based on the data characteristics of the data to be stored; determining the target data partition corresponding to the data to be stored from multiple data partitions based on the partition weight; determining the target compression method of the data to be stored from multiple compression methods based on the target data partition; and compressing the data to be stored and storing it in the target data partition based on the target compression method.

[0026] Figure 1 An application scenario diagram of the data storage method, apparatus, device, medium, and program product according to an embodiment of the present application is shown.

[0027] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a network 101 and a server 102. The network 101 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0028] Server 102 may be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal devices (for example only). The backend management server may analyze and process received data such as user requests, and feed back the processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.

[0029] It should be noted that the data storage method provided in the embodiment of the present application can generally be executed by the server 102. Accordingly, the data storage device provided in the embodiment of the present application can generally be set in the server 102. The data storage method provided in the embodiment of the present application can also be executed by a server or server cluster that is different from the server 102 and can communicate with the server 102. Accordingly, the data storage device provided in the embodiment of the present application can also be set in a server or server cluster that is different from the server 102 and can communicate with the server 102.

[0030] It should be understood that Figure 1 The number of networks and servers in the embodiment is only for illustration. Any number of networks and servers may be provided as required.

[0031] Figure 2 A flow chart of a data storage method according to an embodiment of the present application is shown.

[0032] like Figure 2 As shown, the data storage method of this embodiment includes operations S210 to S240.

[0033] In operation S210 , partition weights of the data to be stored are determined according to data characteristics of the data to be stored.

[0034] In operation S220 , a target data partition corresponding to the data to be stored is determined from the plurality of data partitions according to the partition weights.

[0035] In operation S230 , a target compression method for the data to be stored is determined from a plurality of compression methods according to the target data partition.

[0036] In operation S240 , the data to be stored is compressed according to the target compression method and then stored in the target data partition.

[0037] The data characteristics of the data to be stored may be characteristics indicating the importance of the data to be stored. For example, the importance of the data to be stored may be determined based on the data memory of the data to be stored. The partition weight of the data to be stored may be determined based on the importance of the data to be stored. If the importance of the data to be stored is high, the partition weight may be large; if the importance of the data to be stored is low, the partition weight may be small.

[0038] Different partition weights can correspond to different data partitions. A table of correspondence between partition weights and data partitions can be established based on the data partitions corresponding to different partition weights. Based on the partition weight of the data to be stored, the target data partition corresponding to the data to be stored is determined from the table of correspondence between partition weights and data partitions.

[0039] Different data partitions may correspond to different compression methods. A correspondence table between data partitions and compression methods may be established based on the compression methods corresponding to the different data partitions. Based on the target data partition, a target compression method for the data to be stored is determined from the correspondence table between data partitions and compression methods. The data to be stored is compressed according to the target compression method and then stored in the target data partition.

[0040] According to the embodiments of the present application, the partition weights of the data to be stored are determined based on the data characteristics of the data to be stored. The corresponding target data partitions can be matched according to the characteristics of the data to be stored, thereby achieving classified storage of the data to be stored and rationally utilizing storage space. Different data partitions correspond to different compression methods. Compressing the data to be stored according to the target compression method corresponding to the target data partition allows the data to be compressed using an appropriate compression method, increasing the diversification of data compression, avoiding the problem of a single compression algorithm being unable to balance speed and compression rate, and improving the efficiency of data storage.

[0041] According to an embodiment of the present application, the data features include a time feature and a frequency feature of the data to be stored. The time feature indicates the generation time of the data to be stored, and the frequency feature indicates the number of times the data to be stored is accessed within a preset time period.

[0042] According to an embodiment of the present application, the partition weight of the data to be stored is determined according to the data characteristics of the data to be stored, including: determining the first partition time sub-weight of the data to be stored according to the time characteristics, the business cycle constant of the data to be stored and the weight adjustment factor; determining the partition frequency sub-weight of the data to be stored according to the frequency characteristics, the access frequency normalization coefficient of the data to be stored and the weight adjustment factor; determining the partition weight of the data to be stored according to the sum of the first partition time sub-weight and the partition frequency sub-weight.

[0043] The time feature can be the timestamp of the data to be stored. The frequency feature can be the number of times the data to be stored is accessed within a preset time period (e.g., 7 days).

[0044] The partition weight of the data to be stored It can be calculated by the following formula:

[0045] (1)

[0046] in, Indicates the time sub-weight of the first partition. Represents the weight adjustment factor, which can also be regarded as the time decay factor (Time Decay Factor). It can be 0.6. Represents the base of natural logarithms. The data storage duration can be calculated from the time when the data to be stored is generated to the present time. is the business cycle constant of the data to be stored. For example, the business cycle constant in the financial system may be 90. Represents the partition frequency sub-weight. Represents frequency characteristics. The normalized coefficient representing the access frequency of the data to be stored may be, for example, the peak number of accesses to the data to be stored in history.

[0047] According to an embodiment of the present application, the first partition time sub-weight is calculated by the weight adjustment factor and the business cycle constant, and is added to the partition frequency sub-weight to obtain the partition weight, which is more suitable for data with business cycles, such as data from financial systems.

[0048] According to an embodiment of the present application, data characteristics include time characteristics, frequency characteristics and data volume of the data to be stored. The time characteristics represent the usage time range of the data to be stored, and the frequency characteristics represent the number of times the data to be stored is accessed within a preset time period.

[0049] According to an embodiment of the present application, determining the partition weight of the data to be stored based on the data characteristics of the data to be stored may also include: determining the second partition time sub-weight of the data to be stored based on the ratio of the frequency characteristic and the time characteristic and the partition time weight coefficient; determining the partition data quantum weight of the data to be stored based on the data volume of the data to be stored and the partition data volume weight coefficient, wherein the sum of the partition time weight coefficient and the partition data volume weight coefficient is a fixed value; determining the partition weight of the data to be stored based on the sum of the second partition time sub-weight and the partition data quantum weight.

[0050] The amount of data to be stored can be the data size of the data to be stored. It can also be calculated using the following formula:

[0051] (2)

[0052] in, Indicates the time sub-weight of the second partition. Indicates the partition time weight coefficient. Indicates the usage time range of the data to be stored. Represents the partition data quantum weight. Indicates the partition data volume weight coefficient. Indicates the amount of data to be stored. Indicates the maximum amount of data that may be generated by the data to be stored.

[0053] According to the embodiments of the present application, the present application provides a variety of calculation formulas for partition weights, which can be flexibly selected according to needs, thereby improving the flexibility of data storage.

[0054] According to an embodiment of the present application, the multiple data partitions include hot data partitions and cold data partitions.

[0055] According to an embodiment of the present application, determining a target data partition corresponding to the data to be stored from multiple data partitions based on the partition weight may include: when the partition weight is greater than or equal to a preset partition weight threshold, determining the hot data partition as the target data partition corresponding to the data to be stored; when the partition weight is less than the preset partition weight threshold, determining the cold data partition as the target data partition corresponding to the data to be stored.

[0056] The data stored in the hot data partition may be data with a relatively high number of accesses, and the data stored in the cold data partition may be data with a relatively low number of accesses.

[0057] According to the calculation formula of the partition weight, the larger the partition weight, the more times the data to be stored is accessed or the shorter the storage time. Therefore, when the partition weight is greater than or equal to the preset partition weight threshold, it means that the number of accesses to the data to be stored is relatively high or the storage time is relatively short, so the hot data partition can be determined as the target data partition corresponding to the data to be stored. When the partition weight is less than the preset partition weight threshold, it means that the number of accesses to the data to be stored is relatively low or the storage time is relatively long, so the cold data partition can be determined as the target data partition corresponding to the data to be stored. The partition weight threshold can be, for example, 0.6, 0.65, 0.7, 0.75, etc.

[0058] Partition weights can be used to flexibly match the data to be stored to the corresponding target data partition.

[0059] According to an embodiment of the present application, the multiple compression methods include a hot data compression method and a cold data compression method, and the compression rate of the hot data compression method is lower than the compression rate of the cold data compression method.

[0060] Among them, according to the target data partition, the target compression method of the data to be stored is determined from multiple compression methods, including: when the target data partition is a hot data partition, the hot data compression method is determined from multiple compression methods as the target compression method of the data to be stored; when the target data partition is a cold data partition, the cold data compression method is determined from multiple compression methods as the target compression method of the data to be stored.

[0061] Since the data in the hot data partition is accessed frequently, a low-latency compression method is required in the hot data partition so that it can be decompressed quickly during the query process. For example, the hot data compression method can use a dictionary-based lossless data compression algorithm, which can achieve a good compression rate while ensuring high speed. The compression level of the hot data compression method can be set to 1~3, and the compressed block size can be set to 1~3. 4MB. The compression ratio of cold data compression can be The compression level of cold data compression can be set to 10~22, and the compressed block size is ≥16MB. For example, the original 1TB of equipment replacement logs from 3 years ago only takes up 200-250GB after being compressed by cold data compression. The compression rate of hot data compression can be At the same time, the decompression speed reaches 500MB / s.

[0062] Since the data in the cold data partition is accessed less frequently, a compression method with a higher compression ratio can be used. For example, the cold data compression method can use an algorithm with a high compression ratio and a medium compression speed.

[0063] The problem of a single compression algorithm failing to balance speed and compression ratio can also be addressed by constructing a hierarchical compression decision tree. For example, a three-level decision tree based on data type, access pattern, and hardware characteristics can be constructed to dynamically select the optimal compression algorithm. When the time and frequency characteristics of the data to be stored exceed a threshold and are stored on a solid-state drive, hot data compression can be selected. When the time and frequency characteristics of the data to be stored are less than a threshold and are stored on a solid-state drive, cold data compression can be selected. In other cases, hybrid compression can be adaptively selected.

[0064] According to the embodiments of the present application, different compression methods are matched for cold data partitions and hot data partitions, so that data with different data characteristics can be compressed using the corresponding compression methods, thereby improving the flexibility of data storage and avoiding the problems of insufficient compression rate of cold data or high decompression delay of hot data caused by a single compression method.

[0065] According to an embodiment of the present application, the cold data partition includes multiple cold data sub-partitions, and the method further includes: constructing a secondary query index based on the multiple cold data sub-partitions and the cold data compression method, so as to query data in the multiple cold data sub-partitions based on the secondary query index.

[0066] Since the data stored in the hot data partition is processed using a hot data compression method, its decompression speed is relatively fast. Therefore, when the data in the hot data partition is queried, directly decompressing it will not have much impact on the query speed. The data stored in the cold data partition is processed using a cold data compression method, which has a higher compression ratio, and its decompression speed is not as fast as the hot data compression method. Therefore, a secondary query index can be constructed for multiple cold data sub-partitions in the cold data partition. Specifically, the cold data partition can include multiple cold data sub-partitions, and the cold data compression method corresponding to each cold data sub-partition can be different, for example, the compressed data block size can be different. A secondary query index can be constructed based on the partition identifiers of the multiple cold data sub-partitions and the corresponding cold data compression methods.

[0067] According to an embodiment of the present application, by constructing a secondary query index for a cold data partition, the data query efficiency in the cold data partition can be improved.

[0068] Figure 3 A flow chart of a data storage method according to another embodiment of the present application is shown.

[0069] like Figure 3 As shown, the data storage method of this embodiment includes operations S310 to S330.

[0070] In operation S310 , when target data is queried based on data stored in a data partition, the data stored in the hot data partition is first decompressed according to a hot data compression method to obtain hot data partition decompressed data.

[0071] In operation S320 , when the decompressed data of the hot data partition does not include the target data, a first target cold data sub-partition where the target data is located is determined from the plurality of cold data sub-partitions according to the secondary query index.

[0072] In operation S330 , a second decompression is performed on the data stored in the first target cold data sub-partition according to the cold data compression method to obtain target data, and a decompression speed of the second decompression is lower than a decompression speed of the first decompression.

[0073] After the data is stored in multiple data partitions, the target data can be queried in the multiple data partitions. When querying the target data, since the decompression speed of the hot data partition is faster, the data stored in the hot data partition can be first decompressed according to the hot data compression method to obtain the hot data partition decompressed data. Determine whether the target data is included in the hot data partition decompressed data. If the target data is included in the hot data partition decompressed data, the target data has been queried in the hot data partition and there is no need to query in the cold data partition. If the target data is not included in the hot data partition decompressed data, the target data is not queried in the hot data partition. The first target cold data sub-partition where the target data is located can be determined from the cold data sub-partition based on the secondary query index. If the first target cold data sub-partition is not determined in the secondary query index, all cold data sub-partitions can be directly decompressed.

[0074] The data stored in the first target cold data sub-partition may be subjected to a second decompression according to the cold data compression method, and the target data may be obtained from the decompressed data.

[0075] The decompression speed of the second decompression is lower than the compression speed of the first decompression. Therefore, the data in the cold data sub-partition can be decompressed in parallel while the target data is queried in the hot data partition data, thereby improving the efficiency of data query.

[0076] Data in hot data partitions can be responded to extremely quickly, and the hot data compression method can have the functions of real-time decompression and memory caching, thereby reducing latency. For example, the latency of a single query can be less than 50ms, which is less than the more than 200ms in the common example.

[0077] The secondary query index allows for rapid location of the first target cold data subpartition. Queries on these subpartitions can be decompressed using multi-threaded parallelization. For example, for a log query spanning three years, the response time can be reduced from the typical 120 seconds to just 8 seconds.

[0078] According to an embodiment of the present application, when querying target data, the response speed of querying the target data is improved by preferentially accessing the hot data partition and then accessing the cold data partition.

[0079] According to an embodiment of the present application, the data storage method may further include: generating a hot data partition query log when data stored in the hot data partition is queried; generating a cold data partition query log when data stored in the cold data partition is queried.

[0080] According to an embodiment of the present application, corresponding query logs may be generated when data in different data partitions are queried, so as to analyze the data query situations in different data partitions and perform adaptive adjustments.

[0081] According to an embodiment of the present disclosure, a hot data partition includes multiple hot data sub-partitions, each of the multiple hot data sub-partitions corresponds to a different partition weight interval, and each of the multiple cold data sub-partitions corresponds to a different partition weight interval.

[0082] According to an embodiment of the present application, the data storage method may further include: when the hot data partition query log indicates that the first actual partition weight of the data stored in the first target hot data sub-partition among multiple hot data sub-partitions does not match the partition weight interval corresponding to the first target hot data sub-partition, determining the first actual partition weight interval corresponding to the first target hot data sub-partition based on the first actual partition weight; when the first data sub-partition corresponding to the first actual partition weight interval is one of the multiple hot data sub-partitions, storing the data stored in the first target hot data sub-partition to the first data sub-partition; when the first data sub-partition is one of the multiple cold data sub-partitions, decompressing the data stored in the first target hot data sub-partition according to the hot data compression method, and compressing it according to the cold data compression method and then storing it in the first data sub-partition.

[0083] Hot data partition query logs can be analyzed to obtain query status for data stored in the hot data partitions. Actual partition weights for data stored in multiple hot data sub-partitions can be calculated based on the hot data partition query logs. If the first actual partition weight of data stored in a first target hot data sub-partition among the multiple hot data sub-partitions does not match the partition weight interval corresponding to the first target hot data sub-partition, a first actual partition weight interval corresponding to the first target hot data sub-partition can be determined based on the first actual partition weight. A determination is then made as to whether the first data sub-partition corresponding to the first actual partition weight interval belongs to a hot data partition or a cold data partition. If the first data sub-partition is one of the multiple hot data sub-partitions, the data stored in the first target hot data sub-partition can be directly stored in the first data sub-partition. If the first data sub-partition is one of the multiple cold data sub-partitions, the data stored in the first target hot data sub-partition can be decompressed according to a hot data compression method, compressed according to a cold data compression method, and then stored in the first data sub-partition.

[0084] According to an embodiment of the present application, when the first actual partition weight interval corresponding to the first actual partition weight does not exist, a new first data sub-partition can be created based on the first actual partition weight interval, and the data stored in the first target hot data sub-partition can be processed based on the first actual partition weight interval and then stored in the first data sub-partition.

[0085] According to an embodiment of the present application, the data storage method may further include: when the cold data partition query log indicates that the second actual partition weight of the data stored in the second target cold data sub-partition among multiple cold data sub-partitions does not match the partition weight interval corresponding to the second target cold data sub-partition, determining the second actual partition weight interval corresponding to the second target cold data sub-partition based on the second actual partition weight; when the second data sub-partition corresponding to the second actual partition weight interval is one of multiple cold data sub-partitions, storing the data stored in the second target cold data sub-partition to the second data sub-partition; when the second data sub-partition is one of multiple hot data sub-partitions, decompressing the data stored in the second target cold data sub-partition according to the cold data compression method, and compressing it according to the hot data compression method and then storing it in the second data sub-partition.

[0086] Cold data partition query logs can be analyzed to obtain query status for data stored in the cold data partitions. Actual partition weights for data stored in multiple cold data sub-partitions can be calculated based on the cold data partition query logs. If the second actual partition weight of data stored in a second target hot data sub-partition among the multiple cold data sub-partitions does not match the partition weight interval corresponding to the second target cold data sub-partition, a second actual partition weight interval corresponding to the second target cold data sub-partition can be determined based on the second actual partition weight. A determination is then made as to whether the second data sub-partition corresponding to the second actual partition weight interval belongs to a cold data partition or a hot data partition. If the second data sub-partition is one of the multiple cold data sub-partitions, the data stored in the second target cold data sub-partition can be directly stored in the second data sub-partition. If the second data sub-partition is one of the multiple hot data sub-partitions, the data stored in the second target cold data sub-partition can be decompressed according to a cold data compression method, compressed according to a hot data compression method, and then stored in the second data sub-partition.

[0087] When the second actual partition weight interval corresponding to the second actual partition weight does not exist, a new second data sub-partition can be created based on the second actual partition weight interval, and the data stored in the second target cold data sub-partition can be processed based on the second actual partition weight interval and then stored in the second data sub-partition.

[0088] In some cases, once data is marked as cold, it cannot be reactivated. However, this application analyzes cold data partition query logs and hot data partition query logs, recalculates partition weights, and automatically migrates sudden hot data (such as logs from three years ago triggered by an audit) from the cold data partition to the hot data partition. For example, in a bank's asset change system, a batch of equipment operation logs from six months ago is frequently queried (daily visits increase from 5 to 200 times). The system increases its partition weight from 0.32 to 0.81 through the partition query log, triggering the following actions: migrate the data from the cold data partition to the hot data partition, compress the data using hot data compression; update the global index to point to the new location; and preheat the associated cache to memory.

[0089] By analyzing the query logs of cold data partitions and hot data partitions, the data stored in the cold data partitions and the data stored in the hot data partitions are dynamically adjusted, and the granularity and compression strategy of data storage are optimized, which makes it easier to manage the stored data.

[0090] According to an embodiment of the present application, the data storage method may further include: when the hot data partition query log indicates that the query frequency of data stored in the second target hot data sub-partition among multiple hot data sub-partitions is greater than a first preset query threshold, splitting the data stored in the second target hot data sub-partition into multiple and storing them in multiple hot data partition sub-units; when the hot data partition query log indicates that the query frequency of data stored in multiple third target hot data sub-partitions among multiple hot data sub-partitions is less than or equal to the first preset query threshold, merging the data stored in multiple third target hot data sub-partitions into one hot data sub-partition.

[0091] In the case where the hot data partition query log indicates that the query frequency of the data stored in the second target hot data sub-partition among the multiple hot data sub-partitions is greater than the first preset query threshold, it means that the data query activity in the second target hot data sub-partition is relatively high. Therefore, in order to improve the data query efficiency, the data stored in the second target hot data sub-partition can be split into multiple parts and stored in multiple hot data partition sub-units, each of which can serve as a new hot data sub-partition. In the case where the hot data partition query log indicates that the query frequency of the data stored in multiple third target hot data sub-partitions among the multiple hot data sub-partitions is less than or equal to the first preset query threshold, it means that the data query activity in the multiple third target hot data sub-partitions is relatively low. Therefore, in order to avoid wasting space for low-activity data, the data stored in the multiple third target hot data sub-partitions can be merged into one hot data sub-partition.

[0092] According to an embodiment of the present application, when the cold data partition query log indicates that the query frequency of data stored in the third target cold data sub-partition among multiple cold data sub-partitions is greater than the second preset query threshold, the data stored in the third target cold data sub-partition is split into multiple and stored in multiple cold data partition sub-units; when the cold data partition query log indicates that the query frequency of data stored in multiple fourth target cold data sub-partitions among multiple cold data sub-partitions is less than or equal to the second preset query threshold, the data stored in multiple fourth target cold data sub-partitions are merged into one cold data sub-partition.

[0093] In the case where the cold data partition query log indicates that the query frequency of the data stored in the third target cold data sub-partition among the multiple cold data sub-partitions is greater than the second preset query threshold, it means that the data query activity in the third target cold data sub-partition is relatively high. Therefore, in order to improve the data query efficiency, the data stored in the third target cold data sub-partition can be split into multiple parts and stored in multiple cold data partition sub-units, and each cold data partition sub-unit can serve as a new cold data sub-partition. In the case where the cold data partition query log indicates that the query frequency of the data stored in multiple fourth target cold data sub-partitions among the multiple cold data sub-partitions is less than or equal to the second preset query threshold, it means that the data query activity in the multiple fourth target cold data sub-partitions is relatively low. Therefore, in order to avoid wasting space for low-activity data, the data stored in the multiple fourth target cold data sub-partitions can be merged into one cold data sub-partition.

[0094] The hot data partition query log and the cold data partition query log can be analyzed periodically, for example, once every 24 hours, to timely adjust the data stored in the hot data partition and the data stored in the cold data partition.

[0095] According to the embodiments of the present application, by adjusting the partitions storing data in real time, the efficiency of reasonable utilization of data storage space can be improved.

[0096] Figure 4 A system block diagram of the data storage method according to an embodiment of the present application is shown.

[0097] like Figure 4 As shown, the system 400 includes a data acquisition module 410 , a dynamic partition controller 420 , a layered compression engine 430 , and an adaptive optimization module 440 .

[0098] The data acquisition module 410 extracts data to be stored (e.g., time series data) from the system and extracts data features (e.g., timestamp, device ID, operation type, and other fields). The dynamic partition controller 420 is used to divide and adjust data partitions. The layered compression engine 430 compresses the data in each data partition according to the compression method corresponding to each partition. The adaptive optimization module 440 dynamically updates the storage strategy based on the system's workload using a machine learning model.

[0099] The layered compression engine 430 may also include a compression algorithm performance monitoring unit and an exception handling unit. The compression algorithm performance monitoring unit can generate real-time statistics on the compression ratio, compression time, and decompression throughput of each partition. The exception handling unit can automatically switch to the hot data compression mode and log an error when it detects a compression failure in the cold data compression mode.

[0100] Adaptive optimization module 440 can include a prediction model and a resource scheduler. The prediction model can be trained based on a long short-term memory (LSTM) network. By inputting historical query time, device ID, operation type, and other information, it outputs a prediction of the hot data distribution over the next 24 hours. Based on the prediction results, the resource scheduler can pre-allocate solid-state drive (SSD) cache space and pre-heat high-frequency data.

[0101] If a batch of historical data suddenly experiences high-frequency access due to audit requirements (partition weight increases from 0.3 to 0.8), the system automatically migrates it to the hot data partition in the next cycle and switches to hot data compression. In scenarios with fluctuating traffic, throughput fluctuations are reduced by 70% (compared to a fixed policy). Furthermore, by using machine learning to predict business trends (such as data pre-warming before quarterly reports), partitioning strategies can be adjusted in advance, reducing operation and maintenance labor costs by 90%.

[0102] An adaptive optimization mechanism with closed-loop feedback can also be designed to address the lag caused by manual parameter adjustment. This adaptive optimization mechanism comprises a closed-loop system from the monitoring layer to the analysis layer to the execution layer. The monitoring layer collects over 20 metrics in real time, including query latency, CPU / disk load, and cache hit rate. The analysis layer uses a predictive model to predict the distribution of access hotspots for the next three days. The execution layer can trigger data migration and compression policy changes 72 hours in advance based on these predictions. This solution enables minute-by-minute policy implementation. For example, it can automatically expand hot data partitions and preload historical order data before major e-commerce promotions, or temporarily switch to high compression mode when detecting nighttime batch jobs.

[0103] Based on the above data storage method, the present application also provides a data storage device. Figure 5 The device is described in detail.

[0104] Figure 5 A structural block diagram of a data storage device according to an embodiment of the present application is shown.

[0105] like Figure 5 As shown, the data storage device 500 of this embodiment includes a first determination module 510 , a second determination module 520 , a third determination module 530 and a first storage module 540 .

[0106] The first determining module 510 is configured to determine the partition weight of the data to be stored according to the data characteristics of the data to be stored. In one embodiment, the first determining module 510 may be configured to execute the operation S210 described above, which will not be described in detail here.

[0107] The second determining module 520 is configured to determine a target data partition corresponding to the data to be stored from the plurality of data partitions according to the partition weights. In one embodiment, the second determining module 520 may be configured to execute the operation S220 described above, which will not be described in detail herein.

[0108] The third determination module 530 is configured to determine a target compression method for the data to be stored from a plurality of compression methods according to the target data partition. In one embodiment, the third determination module 530 may be configured to execute the operation S230 described above, which will not be described in detail here.

[0109] The first storage module 540 is configured to compress the data to be stored according to the target compression method and then store it in the target data partition. In one embodiment, the first storage module 540 can be configured to perform the operation S240 described above, which will not be described in detail here.

[0110] According to an embodiment of the present application, the data features include a time feature and a frequency feature of the data to be stored, wherein the time feature indicates the generation time of the data to be stored, and the frequency feature indicates the number of times the data to be stored is accessed within a preset time period.

[0111] According to an embodiment of the present application, the first determination module 510 for determining the partition weight of the data to be stored according to the data characteristics of the data to be stored includes:

[0112] A first determining unit, configured to determine a first partition time sub-weight of the data to be stored according to a time feature, a business cycle constant of the data to be stored, and a weight adjustment factor;

[0113] A second determining unit is configured to determine a partition frequency sub-weight of the data to be stored based on the frequency feature, the access frequency normalization coefficient of the data to be stored, and the weight adjustment factor;

[0114] The third determining unit is configured to determine the partition weight of the data to be stored according to the sum of the first partition time sub-weight and the partition frequency sub-weight.

[0115] According to an embodiment of the present application, data characteristics include time characteristics, frequency characteristics and data volume of the data to be stored. The time characteristics represent the usage time range of the data to be stored, and the frequency characteristics represent the number of times the data to be stored is accessed within a preset time period.

[0116] According to an embodiment of the present application, the first determination module 510 for determining the partition weight of the data to be stored according to the data characteristics of the data to be stored includes:

[0117] a fourth determining unit, configured to determine a second partition time sub-weight of the data to be stored according to a ratio of the frequency feature to the time feature and the partition time weight coefficient;

[0118] a fifth determining unit, configured to determine a partition data quantum weight of the data to be stored according to the data volume of the data to be stored and the partition data volume weight coefficient, wherein the sum of the partition time weight coefficient and the partition data volume weight coefficient is a fixed value;

[0119] The sixth determining unit is configured to determine the partition weight of the data to be stored according to the sum of the second partition time sub-weight and the partition data sub-weight.

[0120] According to an embodiment of the present application, the multiple data partitions include hot data partitions and cold data partitions.

[0121] According to an embodiment of the present application, the second determining module 520 for determining a target data partition corresponding to the data to be stored from a plurality of data partitions according to the partition weights includes:

[0122] a seventh determining unit, configured to determine the hot data partition as a target data partition corresponding to the data to be stored when the partition weight is greater than or equal to a preset partition weight threshold;

[0123] An eighth determining unit is configured to determine the cold data partition as a target data partition corresponding to the data to be stored when the partition weight is less than a preset partition weight threshold.

[0124] According to an embodiment of the present application, the multiple compression methods include a hot data compression method and a cold data compression method, and the compression rate of the hot data compression method is lower than the compression rate of the cold data compression method.

[0125] According to an embodiment of the present application, the third determination module 530 for determining a target compression method for the data to be stored from multiple compression methods according to the target data partition includes:

[0126] a ninth determining unit, configured to, when the target data partition is a hot data partition, determine the hot data compression mode as a target compression mode for the data to be stored from among the plurality of compression modes;

[0127] The tenth determining unit is configured to determine, when the target data partition is a cold data partition, a cold data compression method from a plurality of compression methods as a target compression method for the data to be stored.

[0128] According to an embodiment of the present application, the cold data partition includes multiple cold data sub-partitions.

[0129] According to an embodiment of the present application, the data storage device 500 further includes:

[0130] The building module is used to build a secondary query index according to the multiple cold data sub-partitions and the cold data compression method, so as to query data in the multiple cold data sub-partitions according to the secondary query index.

[0131] According to an embodiment of the present application, the data storage device 500 further includes:

[0132] a first obtaining module, configured to, when querying target data based on the data stored in the data partition, first decompress the data stored in the hot data partition according to the hot data compression method to obtain hot data partition decompressed data;

[0133] a fourth determining module, configured to determine, when the decompressed data of the hot data partition does not include the target data, a first target cold data sub-partition where the target data is located from the plurality of cold data sub-partitions according to the secondary query index;

[0134] The second obtaining module is configured to perform a second decompression on the data stored in the first target cold data sub-partition according to the cold data compression method to obtain target data, wherein the decompression speed of the second decompression is lower than the decompression speed of the first decompression.

[0135] According to an embodiment of the present application, the data storage device 500 further includes:

[0136] A first generating module is configured to generate a hot data partition query log when data stored in the hot data partition is queried;

[0137] The second generating module is configured to generate a cold data partition query log when data stored in the cold data partition is queried.

[0138] According to an embodiment of the present application, the hot data partition includes multiple hot data sub-partitions, each of the multiple hot data sub-partitions corresponds to a different partition weight interval, and each of the multiple cold data sub-partitions corresponds to a different partition weight interval.

[0139] According to an embodiment of the present application, the data storage device 500 further includes:

[0140] a fifth determining module, configured to determine, based on the first actual partition weight, a first actual partition weight interval corresponding to the first target hot data subpartition when the hot data partition query log indicates that the first actual partition weight of the data stored in the first target hot data subpartition among the multiple hot data subpartitions does not match the partition weight interval corresponding to the first target hot data subpartition;

[0141] a second storage module, configured to store the data stored in the first target hot data subpartition into the first data subpartition when the first data subpartition corresponding to the first actual partition weight interval is one of the plurality of hot data subpartitions;

[0142] The third storage module is used to decompress the data stored in the first target hot data sub-partition according to the hot data compression method when the first data sub-partition is one of multiple cold data sub-partitions, and compress the data according to the cold data compression method and then store it in the first data sub-partition.

[0143] According to an embodiment of the present application, the data storage device 500 further includes:

[0144] a sixth determining module, configured to determine, based on the second actual partition weight, a second actual partition weight interval corresponding to the second target cold data subpartition when the cold data partition query log indicates that the second actual partition weight of the data stored in the second target cold data subpartition among the multiple cold data subpartitions does not match the partition weight interval corresponding to the second target cold data subpartition;

[0145] a fourth storage module, configured to store the data stored in the second target cold data subpartition into the second data subpartition when the second data subpartition corresponding to the second actual partition weight interval is one of the plurality of cold data subpartitions;

[0146] The fifth storage module is used to decompress the data stored in the second target cold data subpartition according to the cold data compression method when the second data subpartition is one of multiple hot data subpartitions, and compress the data according to the hot data compression method and then store it in the second data subpartition.

[0147] According to an embodiment of the present application, the data storage device 500 further includes:

[0148] a sixth storage module, configured to split the data stored in the second target hot data subpartition into multiple parts and store the split data into the multiple hot data partition sub-units when the hot data partition query log indicates that the query frequency of the data stored in the second target hot data subpartition among the multiple hot data subpartitions is greater than a first preset query threshold;

[0149] a first merging module, configured to merge the data stored in the plurality of third target hot data subpartitions into one hot data subpartition when the hot data partition query log indicates that the query frequency of the data stored in the plurality of third target hot data subpartitions in the plurality of hot data subpartitions is less than or equal to a first preset query threshold;

[0150] a sixth storage module, configured to split the data stored in the third target cold data subpartition into multiple parts and store the split data into the multiple cold data partition sub-units if the cold data partition query log indicates that the query frequency of the data stored in the third target cold data subpartition among the multiple cold data subpartitions is greater than a second preset query threshold;

[0151] The second merging module is used to merge the data stored in multiple fourth target cold data sub-partitions into one cold data sub-partition when the cold data partition query log indicates that the query frequency of the data stored in multiple fourth target cold data sub-partitions in the multiple cold data sub-partitions is less than or equal to the second preset query threshold.

[0152] According to embodiments of the present application, any multiple modules among the first determination module 510, the second determination module 520, the third determination module 530, and the first storage module 540 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the first determination module 510, the second determination module 520, the third determination module 530, and the first storage module 540 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the first determination module 510 , the second determination module 520 , the third determination module 530 and the first storage module 540 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0153] Figure 6 A block diagram of an electronic device suitable for implementing a data storage method according to an embodiment of the present application is shown.

[0154] like Figure 6As shown, an electronic device 600 according to an embodiment of the present application includes a processor 601, which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 602 or a program loaded from a storage unit 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.

[0155] Various programs and data required for the operation of the electronic device 600 are stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in the one or more memories.

[0156] According to an embodiment of the present application, electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to bus 604. Electronic device 600 may also include one or more of the following components connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or modem. Communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 610 as needed, so that computer programs read from the removable media can be installed into storage section 608 as needed.

[0157] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.

[0158] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 602 and / or RAM 603 described above and / or one or more memories other than ROM 602 and RAM 603.

[0159] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided in the embodiments of the present application.

[0160] The computer program executes the above functions defined in the system / device of the embodiment of the present application when the computer program is executed by the processor 601. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0161] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0162] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from a removable medium 611. When the computer program is executed by the processor 601, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0163] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0165] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.

[0166] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.

Claims

1. A data storage method, characterized in that: The method comprises: Determining the partition weight of the data to be stored according to the data characteristics of the data to be stored; determining, according to the partition weights, a target data partition corresponding to the data to be stored from a plurality of data partitions; Determining a target compression method for the data to be stored from a plurality of compression methods according to the target data partition; According to the target compression method, the data to be stored is compressed and then stored in the target data partition.

2. The method according to claim 1, characterized in that The data characteristics include a time characteristic and a frequency characteristic of the data to be stored, wherein the time characteristic represents the time when the data to be stored is generated, and the frequency characteristic represents the number of times the data to be stored is accessed within a preset time period; The step of determining the partition weight of the data to be stored according to the data characteristics of the data to be stored includes: Determining a first partition time sub-weight of the data to be stored according to the time feature, a business cycle constant of the data to be stored, and a weight adjustment factor; Determining a partition frequency sub-weight of the data to be stored according to the frequency feature, the access frequency normalization coefficient of the data to be stored, and the weight adjustment factor; The partition weight of the to-be-stored data is determined according to the sum of the first partition time sub-weight and the partition frequency sub-weight.

3. The method according to claim 1, characterized in that The data characteristics include a time characteristic, a frequency characteristic, and a data volume of the data to be stored, wherein the time characteristic represents a usage time range of the data to be stored, and the frequency characteristic represents a number of times the data to be stored is accessed within a preset time period; The determining, based on the data characteristics of the data to be stored, the partition weights of the data to be stored includes: Determining a second partition time sub-weight of the data to be stored according to a ratio of the frequency feature to the time feature and a partition time weight coefficient; Determining a partition data quantum weight of the data to be stored according to the data volume of the data to be stored and a partition data volume weight coefficient, wherein the sum of the partition time weight coefficient and the partition data volume weight coefficient is a fixed value; The partition weight of the data to be stored is determined according to the sum of the second partition time sub-weight and the partition data sub-weight.

4. The method according to claim 1, wherein The multiple data partitions include hot data partitions and cold data partitions; The step of determining a target data partition corresponding to the data to be stored from a plurality of data partitions according to the partition weights includes: When the partition weight is greater than or equal to a preset partition weight threshold, determining the hot data partition as the target data partition corresponding to the data to be stored; When the partition weight is less than the preset partition weight threshold, the cold data partition is determined as the target data partition corresponding to the data to be stored.

5. The method according to claim 4, characterized in that The multiple compression modes include a hot data compression mode and a cold data compression mode, and the compression rate of the hot data compression mode is lower than the compression rate of the cold data compression mode; Wherein, determining the target compression mode of the data to be stored from a plurality of compression modes according to the target data partition includes: In a case where the target data partition is the hot data partition, determining the hot data compression mode as the target compression mode for the data to be stored from the multiple compression modes; In a case where the target data partition is the cold data partition, the cold data compression method is determined as the target compression method for the data to be stored from the multiple compression methods.

6. The method according to claim 5, characterized in that The cold data partition includes a plurality of cold data sub-partitions, and the method further includes: A secondary query index is constructed according to the multiple cold data sub-partitions and the cold data compression method, so as to query data in the multiple cold data sub-partitions according to the secondary query index.

7. The method according to claim 6, characterized in that The method further comprises: In a case where target data is queried based on the data stored in the data partition, first decompressing the data stored in the hot data partition according to the hot data compression method to obtain hot data partition decompressed data; In a case where the target data is not included in the decompressed data of the hot data partition, determining a first target cold data sub-partition where the target data is located from the multiple cold data sub-partitions according to the secondary query index; The data stored in the first target cold data sub-partition is subjected to a second decompression according to the cold data compression method to obtain the target data, wherein a decompression speed of the second decompression is lower than a decompression speed of the first decompression.

8. The method according to claim 7, characterized in that The method further comprises: When the data stored in the hot data partition is queried, generating a hot data partition query log; When the data stored in the cold data partition is queried, a cold data partition query log is generated.

9. The method according to claim 8, characterized in that The hot data partition includes a plurality of hot data sub-partitions, each of the plurality of hot data sub-partitions corresponds to a different partition weight interval, and each of the plurality of cold data sub-partitions corresponds to a different partition weight interval; The method further comprises: When the hot data partition query log indicates that a first actual partition weight of data stored in a first target hot data subpartition among the multiple hot data subpartitions does not match the partition weight interval corresponding to the first target hot data subpartition, determining a first actual partition weight interval corresponding to the first target hot data subpartition based on the first actual partition weight; When the first data subpartition corresponding to the first actual partition weight interval is one of the multiple hot data subpartitions, storing the data stored in the first target hot data subpartition into the first data subpartition; When the first data subpartition is one of the multiple cold data subpartitions, the data stored in the first target hot data subpartition is decompressed according to the hot data compression method, and compressed according to the cold data compression method before being stored in the first data subpartition.

10. The method according to claim 9, characterized in that The method further comprises: If the cold data partition query log indicates that a second actual partition weight of data stored in a second target cold data subpartition among the multiple cold data subpartitions does not match the partition weight interval corresponding to the second target cold data subpartition, determining a second actual partition weight interval corresponding to the second target cold data subpartition based on the second actual partition weight; When the second data subpartition corresponding to the second actual partition weight interval is one of the multiple cold data subpartitions, storing the data stored in the second target cold data subpartition into the second data subpartition; When the second data subpartition is one of the multiple hot data subpartitions, the data stored in the second target cold data subpartition is decompressed according to the cold data compression method, and compressed according to the hot data compression method before being stored in the second data subpartition.

11. The method according to claim 9, characterized in that The method further comprises: When the hot data partition query log indicates that the query frequency of data stored in a second target hot data subpartition among the multiple hot data subpartitions is greater than a first preset query threshold, splitting the data stored in the second target hot data subpartition into multiple parts and storing the data in the multiple hot data partition sub-units; When the hot data partition query log indicates that the query frequency of data stored in multiple third target hot data sub-partitions among the multiple hot data sub-partitions is less than or equal to the first preset query threshold, merging the data stored in the multiple third target hot data sub-partitions into one hot data sub-partition; If the cold data partition query log indicates that the query frequency of data stored in a third target cold data sub-partition among the multiple cold data sub-partitions is greater than a second preset query threshold, splitting the data stored in the third target cold data sub-partition into multiple pieces and storing the pieces into multiple cold data partition sub-units; When the cold data partition query log indicates that the query frequency of data stored in multiple fourth target cold data sub-partitions among the multiple cold data sub-partitions is less than or equal to the second preset query threshold, the data stored in the multiple fourth target cold data sub-partitions are merged into one cold data sub-partition.

12. A data storage device, characterized in that: The device comprises: A first determining module, configured to determine a partition weight of the data to be stored according to data characteristics of the data to be stored; a second determining module, configured to determine, from a plurality of data partitions, a target data partition corresponding to the data to be stored according to the partition weight; a third determining module, configured to determine a target compression method for the data to be stored from a plurality of compression methods according to the target data partition; The first storage module is used to compress the data to be stored according to the target compression method and then store it in the target data partition.

13. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Working condition data management method, device, equipment and medium

    CN116842223A

  • Data processing method and device, storage medium and computer equipment

    CN119166079A

  • Data processing method and device, computer equipment and storage medium

    CN119322592A

  • Method for managing computing devices, electronic device and computer storage medium

    US20210216376A1

  • Automatically tuning logical partition weights

    US20240330071A1