Data storage method, device, computer equipment, storage medium

By aggregating and grouping massive Key-Value data and generating key-value data using the method of aggregating and grouping and information summary algorithm, the problem of slow storage speed of massive data is solved, and the effect of fast storage and query is achieved.

CN116595033BActive Publication Date: 2025-07-22IND BANK CO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310568850.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-19
Publication Date
2025-07-22
Estimated Expiration
2043-05-19

AI Technical Summary

Technical Problem

The prior art is slow to store massive Key-Value data, and cannot complete storage within 5-10 hours, which is beyond the acceptable range of daily data processing.

Method used

The data to be stored is aggregated and grouped according to the preset indicators and time dimensions, and multiple groups of aggregated data are generated, and information summary data is generated using the information summary algorithm to combine it with the time dimension to generate key-value data, and stored in the key-value database.

Benefits of technology

By compressing the generation of grouping and key-value data, the data storage speed is improved, and the query request can be responded to the query request within milliseconds, which significantly improves the efficiency of data query.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116595033B_ABST
    Figure CN116595033B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a data storage method, apparatus, computer device, and storage medium. The method includes: obtaining data to be stored, where the data to be stored is data whose data volume exceeds a preset quantity threshold; aggregating and grouping the data to be stored according to preset metrics and time dimensions to obtain multiple groups of aggregated data, where each group of aggregated data in the multiple groups of aggregated data is a row, and the multiple groups of aggregated data form multi-row and multi-column data, and there are corresponding delimiters for each row and each column of data; generating a key according to the preset metrics and time dimensions, and combining the key with the multiple groups of aggregated data according to the delimiter to obtain key-value data; and storing the key-value data in a key-value database. By using this method, the speed of storing a large amount of data in Key-Value type data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and in particular, to a data storage method, apparatus, computer device, and storage medium. Background Art

[0002] With the development of Internet technology, various different business applications generate a large amount of data. When querying after storing a large amount of data, in order to improve the query speed, Key-Value type data is usually used for storage.

[0003] However, for Key-Value type data, it takes 5 - 10 hours to store tens of billions of massive data at one time, which exceeds the acceptable range of daily data processing.

[0004] Therefore, there is an urgent need for a method that can improve the storage speed of massive Key-Value type data. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, apparatus, computer device, and storage medium that can improve the storage speed of massive Key-Value type data.

[0006] In a first aspect, the present disclosure provides a data storage method, characterized in that the method includes:

[0007] Obtain data to be stored, where the data to be stored is data whose data volume exceeds a preset quantity threshold;

[0008] Aggregate and group the data to be stored according to a preset index and time dimension to obtain multiple groups of aggregated data, where each group of aggregated data in the multiple groups of aggregated data is a row, and the multiple groups of aggregated data form multi-row and multi-column data, and there are corresponding delimiters for each row and each column of data;

[0009] Generate a key according to the preset index and time dimension, and combine the key with the multiple groups of aggregated data according to the delimiter to obtain key-value data;

[0010] Store the key-value data in a key-value database.

[0011] In one embodiment, the generating a key according to the preset index and time dimension, and combining the key with the multiple groups of aggregated data according to the delimiter to obtain key-value data includes:

[0012] Generate information summary data for the preset index by using an information digest algorithm;

[0013] Combine the information summary data and the time dimension to obtain a key corresponding to the time dimension;

[0014] According to the time dimension, combine the key and the multiple sets of aggregated data according to the row delimiter corresponding to each row and the column delimiter corresponding to each column to obtain key-value data, where the length of the key is less than a preset byte threshold.

[0015] In one embodiment, after storing the key-value data in the key-value database, the method further includes:

[0016] In response to receiving a query request for the key-value data, obtain a query condition corresponding to the query request, where the query condition at least includes: a query time range;

[0017] Match a key corresponding to the query time range from the key-value data to obtain a query key;

[0018] Use the query key to obtain first aggregated data in the key-value data.

[0019] In one embodiment, after matching a key corresponding to the query time range from the key-value data to obtain a query key, the method further includes:

[0020] Obtain the information summary data corresponding to the query key;

[0021] Use the query summary data in the query condition to filter the information summary data corresponding to the query key to determine the target key in the query key;

[0022] Correspondingly, the step of using the query key to obtain first aggregated data in the key-value data includes:

[0023] Use the target key to obtain second aggregated data in the key-value data.

[0024] In one embodiment, after using the query key to obtain first aggregated data in the key-value data, the method further includes:

[0025] Use a preset data query statement in the query condition to filter in the first aggregated data to obtain target aggregated data;

[0026] The step of using the target key to obtain second aggregated data in the key-value data includes: using a preset data query statement in the query condition to filter in the second aggregated data to obtain target aggregated data.

[0027] In one embodiment, after using the query key to obtain first aggregated data in the key-value data, the method further includes:

[0028] Obtain the display data volume threshold and cache quantity threshold of the front-end display interface, divide the first aggregated data into multiple pages of display data according to the display quantity threshold and the cache quantity threshold, and cache the multiple pages of display data into the front-end display interface;

[0029] In response to the display of the first aggregated data, sequentially display each page of the display data on the front-end display interface, and delete the displayed display data from the cache of the front-end display interface;

[0030] In response to the number of pages of the display data cached in the front-end display interface being less than a preset target page threshold, obtain the first aggregated data that has not been cached into the front-end interface;

[0031] Divide the first aggregated data that has not been cached into multiple pages of display data according to the display quantity threshold and the cache quantity threshold until all the first aggregated data queried is displayed through the front-end display interface.

[0032] In a second aspect, the present disclosure also provides a data storage device. The device includes:

[0033] A data acquisition module, configured to acquire data to be stored, where the data to be stored is data whose data volume exceeds a preset quantity threshold;

[0034] An aggregation grouping module, configured to aggregate and group the data to be stored according to a preset metric and time dimension to obtain multiple groups of aggregated data, where each group of aggregated data in the multiple groups of aggregated data is a row, and the multiple groups of aggregated data form multiple rows and multiple columns of data, and there are corresponding delimiters for each row and each column of data;

[0035] A key-value combination module, configured to generate a key according to the preset metric and time dimension, and combine the key with the multiple groups of aggregated data according to the delimiter to obtain key-value data;

[0036] A data storage module, configured to store the key-value data into a key-value database.

[0037] In a third aspect, the present disclosure also provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps in any of the above method embodiments are implemented.

[0038] In a fourth aspect, the present disclosure also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0039] In a fifth aspect, the present disclosure also provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the steps in any of the above method embodiments.

[0040] In the above embodiments, the data to be stored is aggregated and grouped according to preset metrics and time dimensions to obtain multiple groups of aggregated data. Each group of aggregated data in the multiple groups of aggregated data is a row, and the multiple groups of aggregated data form multi-row and multi-column data, and there are corresponding delimiters for each row and each column of data. It is possible to compress the data to be stored according to metrics and time dimensions. Furthermore, keys are generated according to the preset metrics and time dimensions, and the keys and the multiple groups of aggregated data are combined according to the delimiters to obtain key-value data. Different from the traditional method of generating a key for each data and combining it with the value, processing using metrics and time dimensions can compress and group a large amount of data to be stored to obtain multiple groups of aggregated data. And the keys are also generated according to the preset metrics and time dimensions. Therefore, after the keys and the aggregated data are combined, each key and the aggregated data can be quickly matched to obtain key-value data. This improves the speed of data storage. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 It is an application environment diagram of the data storage method in an embodiment;

[0043] Figure 2 It is a flowchart of the data storage method in an embodiment;

[0044] Figure 3 It is a flowchart of step S206 in an embodiment;

[0045] Figure 4 It is a flowchart after step S208 in an embodiment;

[0046] Figure 5 It is a flowchart after step S404 in an embodiment;

[0047] Figure 6 It is a flowchart after step S406 in an embodiment;

[0048] Figure 7Schematic block diagram of a data storage device in an embodiment;

[0049] Figure 8 Internal structure schematic diagram of a computer device in an embodiment. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present disclosure clearer and more understandable, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure, but not to limit the present disclosure.

[0051] It should be noted that the terms "first", "second", etc. in the description and claims of this article and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0052] In this article, the term "and / or" is only a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0053] The embodiments of the present disclosure provide a data storage method, which can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the data server 104 and the storage server 106 through the network respectively. The terminal 102 can obtain the data to be stored in the data server 104. The data to be stored is usually data whose data volume exceeds a preset quantity threshold. The terminal 102 aggregates the data to be stored according to preset metrics and time dimensions to obtain multiple groups of aggregated data. Among them, each group of aggregated data in the multiple groups of aggregated data is a row, and the multiple groups of aggregated data can form multi-row and multi-column data, and there are corresponding delimiters for each row and each column of data. The terminal 102 can generate a key according to the preset metrics and time dimensions, and combine the key with the multiple groups of aggregated data according to the delimiter to obtain key-value data. The terminal 102 can store the key-value data in the database of the storage server 106. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc. The data server 104 and the storage server 106 can be implemented by an independent server or a server cluster composed of multiple servers.

[0054] In one embodiment, as Figure 2 shown, a data storage method is provided. Taking the terminal 102 in Figure 1 as an example, the method includes the following steps:

[0055] S202, obtain the data to be stored, where the data to be stored is data whose data volume exceeds a preset quantity threshold.

[0056] Among them, the preset quantity threshold can usually be set according to the actual application scenario. Usually, the application scenario of this case is a scenario of storing a large amount of data. Therefore, the preset quantity threshold can be more than ten thousand pieces of data, or it can be the data volume of the data to be stored, such as 1Tb of data. In some embodiments of the present disclosure, the quantity threshold is not absolutely limited.

[0057] Specifically, when a large amount of data to be stored needs to be stored, the data to be stored can be obtained from the business system or the data server.

[0058] In some exemplary embodiments, the data computing framework MapReduce (MR2) of the big data cluster can be used to store the data to be stored in HDFS (distributed file system).

[0059] S204, aggregate and group the data to be stored according to preset metrics and time dimensions to obtain multiple groups of aggregated data.

[0060] Among them, each set of aggregated data in the multiple sets of aggregated data is one line, and the multiple sets of aggregated data form multi-line and multi-column data. There are corresponding delimiters for each row and each column of data. The delimiter can be, for example, " / ", "*", "#", etc. In some embodiments of the present disclosure, the delimiter is not limited. Aggregation grouping usually groups the data to be stored according to certain metrics and time dimensions. Data with the same metrics and the same time dimension can be a set of aggregated data. The preset metrics usually vary according to the type of data to be stored. For example, in the financial field, the metrics can usually be the account opening age corresponding to the data to be stored and the deposit and withdrawal amount corresponding to the data to be stored. In the cloud computing field, the metrics can usually be the storage area of the data to be stored. The time dimension can usually be a time metric, such as the data to be stored generated within ten days. The data to be stored generated within ten to twenty days.

[0061] Specifically, the data to be stored can be aggregated and grouped according to the preset metrics and time dimensions. Data with the same metrics and time dimensions can be a set of aggregated data. Multiple sets of aggregated data are obtained according to different combinations of metrics and time dimensions.

[0062] In some exemplary embodiments, for example, the preset metrics can be Metric A and Metric B. The time dimension can be one week. Aggregation grouping can be performed through a user-defined data aggregation function (UDAF). Then, the multiple sets of aggregated data finally obtained can include: 1. The data to be stored with Metric A generated within one week. 2. The data to be stored with Metric B generated within one week. 3. The data to be stored with Metric A generated outside one week. 4. The data to be stored with Metric B generated outside one week. A total of four sets of aggregated data are obtained. Each set of aggregated data can be compressed into one line. The multiple sets of aggregated data can form multi-line and multi-column data. For example, the above four sets of aggregated data can form two lines and two columns of data. There are corresponding delimiters for each row and each column of data. For example, the row delimiter can be used as / u0001, and the column delimiter can be used as / u0002. Using the row delimiter and the column delimiter can accurately determine each row and each column of data, and thus accurately find the corresponding aggregated data.

[0063] S206, generate a key according to the preset metrics and time dimensions, and combine the key with the multiple sets of aggregated data according to the delimiter to obtain key-value data.

[0064] Specifically, a preset metric and a time dimension can be combined to generate a key. Each metric and each time dimension can be combined to generate a type of key. Also, since multiple sets of aggregated data are obtained by aggregating and grouping according to the preset metric and time dimension, the key and the multiple sets of aggregated data can be combined according to the separator according to the preset metric and time dimension to obtain key-value data.

[0065] In some exemplary embodiments, for example, the above obtains four sets of aggregated data, and the preset metrics can be metric A and metric B. The time dimension can be one week, and then four types of keys (such as key A, key B, key C, and key D) can be obtained. Each type of key is combined with the corresponding set of aggregated data according to the preset metric and time dimension. Additionally, since the above multiple sets of aggregated data form multi-row and multi-column data, there are corresponding separators for each row and each column of data. Therefore, the key and the multiple sets of aggregated data can be combined according to the separator to obtain key-value data. Key A - U0001 aggregated data. Key B - U0002 aggregated data. Key C - U0003 aggregated data. Key D - U0004 aggregated data. Where U0001 and U0003 can be row separators. U0002 and U0004 can be column separators. It can be understood that the above is only for illustrative purposes.

[0066] S208, store the key-value data in a key-value database.

[0067] Specifically, use the MR2 framework to write a key-value data batch import program and directly store the compressed data in a key-value database.

[0068] In the above data storage method, the data to be stored is aggregated and grouped according to the preset metric and time dimension to obtain multiple sets of aggregated data. Among them, each set of aggregated data in the multiple sets of aggregated data is a row, and the multiple sets of aggregated data form multi-row and multi-column data, and there are corresponding separators for each row and each column of data. It is possible to compress the data to be stored according to the metric and time dimension. Furthermore, a key is generated according to the preset metric and time dimension, and the key and the multiple sets of aggregated data are combined according to the separator to obtain key-value data. Different from the traditional method of generating a key for each data and combining it with the value, using metrics and time dimensions for processing can compress and group a large amount of data to be stored to obtain multiple sets of aggregated data. And the key is also generated according to the preset metric and time dimension. Therefore, after the key and the aggregated data are combined, each key and the aggregated data can be quickly matched to obtain key-value data. This improves the speed of data storage.

[0069] In one embodiment, such as Figure 3As shown, generating a key according to the preset metrics and time dimension, and combining the key with the multiple sets of aggregated data according to the delimiter to obtain key-value data includes:

[0070] S302, generating message digest data from the preset metrics using a message digest algorithm.

[0071] S304, combining the message digest data and the time dimension to obtain a key corresponding to the time dimension.

[0072] S306, combining the key and the multiple sets of aggregated data according to the row delimiter corresponding to each row and the column delimiter corresponding to each column according to the time dimension to obtain key-value data, where the length of the key is less than a preset byte threshold.

[0073] Among them, the preset byte threshold can usually be 100 bytes or less than 100 bytes. The message digest algorithm can usually be the MD5 algorithm.

[0074] Specifically, if the key in the persistent file of key-value data is too long, such as exceeding 100 bytes, for 10 million rows of data, just the key will occupy 100×10 million = 1 billion bytes, nearly 1G of data, which will greatly affect the storage efficiency of the persistent file. Therefore, if the length of the preset metric is long, the byte length of the key generated from this metric will also be relatively long. Therefore, the preset metric can be processed using the MD5 algorithm to generate message digest data with a length of 32 bytes. Combining the digest information data generated for each metric and the time dimension to obtain a key corresponding to the time dimension. Since the information of the preset metric has been changed after being processed using the MD5 algorithm. Therefore, after combining the key and the aggregated data, the time dimension can be used as the combination criterion. Combining the key and multiple sets of aggregated data according to the row delimiter and column delimiter according to the time dimension to obtain key-value data.

[0075] In some exemplary embodiments, a new data output format can be defined through a user-defined table function (UDTF), and then the key and the aggregated data are combined.

[0076] In this embodiment, by using the MD5 algorithm to generate message digest data, it is possible to generate message digest data with a fixed length (32 bytes) from metrics with a relatively long length. After combining with the time dimension, it can effectively shorten the length of the generated keys and improve the query speed of subsequent retrievals. Additionally, when using the MD5 algorithm for processing, if there is only a slight difference in the preset metrics (taking 0005 and 0006 as examples), the difference in the corresponding calculated message digest data will also be significant, which can further improve the retrieval speed after retrieval. Moreover, MemStore also caches some data in memory. Since there is no need to load overly long keys, it improves the effective utilization rate of memory, can cache more data, and thus improves the retrieval efficiency.

[0077] In one embodiment, as Figure 4 shown, after storing the key-value data into the key-value database, the method further includes:

[0078] S402, in response to receiving a query request for the key-value data, obtaining a query condition corresponding to the query request, where the query condition at least includes: a query time range.

[0079] S404, matching keys corresponding to the query time range from the key-value data to obtain query keys.

[0080] S406, using the query keys to obtain first aggregated data from the key-value data.

[0081] Among them, the query time range can usually be used to retrieve data generated within a certain time period.

[0082] Specifically, when it is necessary to query the key-value data, a query request can be input into the terminal. When the terminal receives a query request for the key-value data, it can obtain a query condition corresponding to the query request. This query condition can usually be obtained from the query request. Since the key is generated using the time dimension and preset metrics. Therefore, keys corresponding to the query time range can be matched in the key-value data to obtain query keys. Then, the corresponding first aggregated data can be found using these query keys.

[0083] In some exemplary embodiments, taking the financial scenario as an example, internal auditors usually need to further confirm abnormal financial behaviors based on suspicious information (abnormal accounts) through financial event logs. The financial event log data (massive data) has characteristics such as multiple information dimensions, high update frequency, and large time span. When using traditional database query engines, it takes at least about ten minutes to retrieve one year's worth of log data. After storing the data using the solution of the embodiments of the present disclosure, the result can be returned within milliseconds under the same query conditions. Specifically as follows: The user needs to specify the account number of the account to be queried, the opening institution, and the transaction date range (required fields, no fuzzy query) and enter relevant information (query request) on the query interface. Use this transaction date range to filter the key-value data to obtain a query key, and then use the query key to obtain the data to be queried.

[0084] In this embodiment, since the key is generated according to preset metrics and time dimensions, and the data to be stored is also aggregated and grouped according to preset metrics and time dimensions to form key-value data, the dimension to be queried can be narrowed down by the time range, that is, specifically locate the data corresponding to a certain day or a preset time, thereby improving the speed of data query.

[0085] In one embodiment, as Figure 5 shown, after matching the key corresponding to the query time range from the key-value data to obtain a query key, the method further includes:

[0086] S502, obtaining the information summary data corresponding to the query key.

[0087] S504, using the query summary data in the query conditions to filter the information summary data corresponding to the query key to determine the target key in the query key.

[0088] Correspondingly, the obtaining the first aggregated data from the key-value data using the query key includes:

[0089] Obtaining the second aggregated data from the key-value data using the target key.

[0090] Specifically, after the query key is determined as above. As mentioned in the above embodiments, the query key is usually composed of preset metrics and time dimensions. Therefore, the information summary data corresponding to the query key can be obtained. The information summary data corresponding to the query key is filtered by using the query summary data carried in the query condition, and the information summary data corresponding to the query keys different from those in the query summary data is screened out, thereby determining the target key in the query key. After the target key is determined, the target key can be matched in the key-value data to obtain the second aggregated data. After the second aggregated data is obtained, a byte buffer can be established for this part of the data, and according to the line separator and column separator, the byte data is deserialized in batches into the original data.

[0091] In this embodiment, using the query summary data to further filter the query key can avoid the metrics with the same prefix from being added to the scan, quickly locate the query data, and further improve the query speed.

[0092] In one embodiment, after obtaining the first aggregated data by using the query key in the key-value data, the method further includes:

[0093] Filtering the first aggregated data by using the preset data query statement in the query condition to obtain the target aggregated data;

[0094] Obtaining the second aggregated data by using the target key in the key-value data includes: filtering the second aggregated data by using the preset data query statement in the query condition to obtain the target aggregated data.

[0095] Specifically, in order to further filter out more data that meets the requirements. The preset data query statement DSL (Domain Specified Language) can be used to perform more refined data filtering. The preset data query statement in the query condition can be used to filter the first aggregated data to obtain the target aggregated data. The preset query statement can simply and flexibly configure the data processing chain to complete a series of data processing functions such as data filtering, fuzzy matching, data transformation, and sensitive data desensitization. The preset query statement can also be used to filter the second aggregated data to obtain the target aggregated data.

[0096] In some exemplary embodiments, taking the query in the financial field as an example, for example: filter(jyje,">=",10000) can represent filtering data with transaction amounts (jyje) greater than 10,000. like(khxm,"^张.+月$") can fuzzily match customer name (khxm) data starting with 张 and ending with 月; sm3(zjhm,2) means retaining the first two bytes and then desensitizing the ID number (zjhm) according to the national standard SM3. In addition, based on actual query needs, users can set query statements, such as the range of transaction amounts, currency types, specific uses, etc. Some filter items can support fuzzy queries and truncated queries (what conditions are met for data from the Xth to the Xth position of a field). Different query statements can be set accordingly according to different query needs, and there is no absolute restriction on the query statement type in the present disclosure.

[0097] In this embodiment, by using a preset data query statement to filter the first aggregated data or the second aggregated data, target aggregated data that better meets the requirements can be obtained.

[0098] In one embodiment, Figure 6 As shown, after obtaining the first aggregated data from the key-value data using the query key, the method further includes:

[0099] S602, obtaining a display data volume threshold and a cache quantity threshold of a front-end display interface, dividing the first aggregated data into multiple pages of display data according to the display quantity threshold and the cache quantity threshold, and caching the multiple pages of display data in the front-end display interface.

[0100] S604, in response to the display of the first aggregated data, displaying each page of the display data in sequence on the front-end display interface, and deleting the displayed display data from the cache of the front-end display interface.

[0101] S606: In response to the number of pages of display data cached in the front-end display interface being less than a preset target page number threshold, the first aggregated data that has not been cached in the front-end interface is obtained.

[0102] S608: Divide the first aggregated data that has not been cached into multiple pages of display data according to the display quantity threshold and the cache quantity threshold, until all the queried first aggregated data are displayed through the front-end display interface.

[0103] The display data volume threshold may be the maximum amount of data that can be displayed on the front-end display interface at a time, and the cache volume threshold may be the maximum amount of data that can be cached on the front-end display interface.

[0104] Specifically, when the first aggregated data is obtained, the first aggregated data usually needs to be displayed on the front-end display interface. Usually, there is a limit on the amount of data per page on the front-end display interface. Therefore, when displaying the first aggregated data on the front-end display interface, it is necessary to obtain the display data volume threshold and the cache quantity threshold of the front-end interface. The first aggregated data is split according to the display quantity threshold and the cache quantity threshold to obtain multi-page display data. Then, the multi-page display data is cached in the front-end display interface. When the front-end display interface needs to display the first aggregated data, since there is a display quantity threshold on the front-end interface, the display data of each page can be sequentially displayed on the front-end interface. After the display is completed, the displayed display data can be deleted from the cache of the front-end display interface to release the cache of the front-end display interface. When the number of pages of the display data cached in the front-end display interface is relatively small, less than the preset target page threshold, the first aggregated data that has not been cached in the front-end display interface can be obtained, and then processed according to the steps such as S602 to S604 above until all the first aggregated data is displayed through the front-end interface. It can be understood that the second aggregated data or the target aggregated data can also be displayed in the above manner, which will not be repeated here.

[0105] In some exemplary embodiments, for example, the display quantity threshold of the front-end display interface is 5 pieces of data. The cache quantity threshold is 100 pieces of data. The first aggregated data includes 200 pieces of data. The first aggregated data can be split into two parts according to the above method. Each part has 100 pieces of data, and one part is cached in the front-end display interface, and the 100 pieces of data are split into 20 pages of display data. Each page of display data includes 5 pieces of data. Then the front-end display interface sequentially displays 20 pages of display data. After each page of display data is displayed, the display data can be deleted from the cache of the front-end display interface. The target business threshold can be set to 5. When there are 5 pages or less than 5 pages of the display data cached in the front-end display interface, the remaining 100 pieces of aggregated data can be split and stored in the front-end display interface in the same way, and the front-end display interface will display it.

[0106] In some specific implementation manners, it is possible to obtain a data volume that is 10 times the number of items limited on the front-end display page. When the user clicks the next page, the data in the query result set (cache) will be consumed first. When the data in the result set is almost exhausted (just exhausted / insufficient for the next query consumption), 10 more pages of data will be cached into the local result set. And through the lazy loading strategy, MemStore will cache 10 more pages of data to prepare for the next data query. Each query will perform the above loop until all data queries are completed. The user only needs to wait for 1 second to see the display result (the number of items limited on each page). At this time, the background query engine will continue to query the remaining data (when clicking the query, the data first queried will be returned first for the business personnel to confirm). When the user needs the next page of data, the system backend will return the previously queried data, and so on, finally forming an application scenario of "viewing while querying". Since the data is distributed and stored on different servers, if all data is obtained at once, there will be a relatively large network transmission burden. In addition, there will also be a memory pressure when all data is aggregated to the client node. Therefore, through this method, the memory pressure can be reduced.

[0107] In this embodiment, when the front-end display interface needs to display data, some data is returned first. When the user needs to continue browsing the data and has browsed most of the above-mentioned part of the data, some more data is returned, which can effectively reduce the system load during peak hours and improve the concurrent query performance.

[0108] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are displayed in sequence according to the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0109] Based on the same inventive concept, the embodiments of the present disclosure also provide a data storage device for implementing the data storage method described above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the following data storage devices can refer to the limitations on the data storage method in the above text, and will not be repeated here.

[0110] In one embodiment, as Figure 7As shown, a data storage device 700 is provided, including: a data acquisition module 702, an aggregation grouping module 704, a key-value combination module 706, and a data storage module 708, where:

[0111] The data acquisition module 702 is configured to acquire data to be stored, where the data to be stored is data whose data volume exceeds a preset quantity threshold;

[0112] The aggregation grouping module 704 is configured to aggregate and group the data to be stored according to a preset metric and time dimension to obtain multiple groups of aggregated data. Among them, each group of aggregated data in the multiple groups of aggregated data is one row, and the multiple groups of aggregated data form multi-row and multi-column data, and there are corresponding delimiters for each row and each column of data;

[0113] The key-value combination module 706 is configured to generate a key according to the preset metric and time dimension, and combine the key with the multiple groups of aggregated data according to the delimiter to obtain key-value data;

[0114] The data storage module 708 is configured to store the key-value data in a key-value database.

[0115] In an embodiment of the device, the key-value combination module 706 includes:

[0116] A digest algorithm processing module, configured to generate digest data by using a digest algorithm for the preset metric.

[0117] A first combination module, configured to combine the digest data and the time dimension to obtain a key corresponding to the time dimension.

[0118] A second combination module, configured to combine the key and the multiple groups of aggregated data according to the row delimiter corresponding to each row and the column delimiter corresponding to each column according to the time dimension to obtain key-value data, where the length of the key is less than a preset byte threshold.

[0119] In an embodiment of the device, the device further includes: a data query module, configured to, in response to receiving a query request for the key-value data, obtain a query condition corresponding to the query request, where the query condition at least includes: a query time range; match a key corresponding to the query time range from the key-value data to obtain a query key; and use the query key to obtain first aggregated data from the key-value data

[0120] In an embodiment of the device, the data query module is further configured to obtain digest data corresponding to the query key;

[0121] Filter the information summary data corresponding to the query key by using the query summary data in the query condition to determine the target key in the query key;

[0122] Correspondingly, the obtaining the first aggregated data by using the query key from the key-value data includes:

[0123] Obtain the second aggregated data by using the target key from the key-value data.

[0124] In one embodiment of the device, the data query module is further configured to filter in the first aggregated data by using a preset data query statement in the query condition to obtain target aggregated data;

[0125] Alternatively, filter in the second aggregated data by using a preset data query statement in the query condition to obtain target aggregated data.

[0126] In one embodiment of the device, the device further includes: a data display module, configured to obtain a display data volume threshold and a cache quantity threshold of a front-end display interface, divide the first aggregated data into multiple pages of display data according to the display quantity threshold and the cache quantity threshold, and cache the multiple pages of display data into the front-end display interface; in response to displaying the first aggregated data, sequentially display each page of the display data in the front-end display interface, and delete the displayed display data from the cache of the front-end display interface; in response to the number of pages of the display data cached in the front-end display interface being less than a preset target page threshold, obtain the first aggregated data that has not been cached into the front-end interface; divide the first aggregated data that has not been cached into multiple pages of display data according to the display quantity threshold and the cache quantity threshold until all the first aggregated data queried is displayed through the front-end display interface.

[0127] Each module in the above data storage device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to the above respective modules.

[0128] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store aggregated data, preset metrics, and time dimensions. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a data storage method.

[0129] Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of the present disclosure, and does not constitute a limitation on the computer device to which the solution of the present disclosure is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0130] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps in any of the above method embodiments are implemented.

[0131] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in any of the above method embodiments are implemented.

[0132] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in any of the above method embodiments are implemented.

[0133] It should be noted that the data to be stored, preset metrics, etc. involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0134] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided by the present disclosure can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided by the present disclosure can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided by the present disclosure can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0135] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0136] The above-described embodiments only represent several implementation manners of the present disclosure. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the patents of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present disclosure, several modifications and improvements can still be made, and these all belong to the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the appended claims.

Claims

1. A data storage method, characterized in that, The method includes: Obtain data to be stored, where the data to be stored is data with a data volume exceeding a preset quantity threshold; Aggregate and group the data to be stored according to preset metrics and time dimensions to obtain multiple groups of aggregated data. Among them, each group of aggregated data in the multiple groups of aggregated data is one row, and the multiple groups of aggregated data form multi-row and multi-column data, and there are corresponding delimiters for each row and each column of data; Generate a key according to the preset metrics and time dimensions, and combine the key with the multiple groups of aggregated data according to the delimiter to obtain key-value data; Store the key-value data in a key-value database; The step of generating a key according to the preset metrics and time dimensions, and combining the key with the multiple groups of aggregated data according to the delimiter to obtain key-value data includes: Generate information digest data for the preset metrics using an information digest algorithm; Combine the information digest data and the time dimension to obtain a key corresponding to the time dimension; According to the time dimension, combine the key and the multiple groups of aggregated data according to the row delimiter corresponding to each row and the column delimiter corresponding to each column to obtain key-value data, where the length of the key is less than a preset byte threshold.

2. The method according to claim 1, characterized in that, After storing the key-value data in the key-value database, the method further includes: In response to receiving a query request for the key-value data, obtain a query condition corresponding to the query request, where the query condition at least includes: a query time range; Match a key corresponding to the query time range from the key-value data to obtain a query key; Use the query key to obtain first aggregated data from the key-value data.

3. The method according to claim 2, wherein After matching a key corresponding to the query time range from the key-value data to obtain a query key, the method further includes: Obtain the information digest data corresponding to the query key; Use the query digest data in the query condition to filter the information digest data corresponding to the query key to determine the target key in the query key; Correspondingly, the step of using the query key to obtain first aggregated data from the key-value data includes: Use the target key to obtain second aggregated data from the key-value data.

4. The method according to claim 3, characterized in that, After using the query key to obtain first aggregated data from the key-value data, the method further includes: Use a preset data query statement in the query condition to filter in the first aggregated data to obtain target aggregated data; The step of using the target key to obtain second aggregated data from the key-value data includes: using a preset data query statement in the query condition to filter in the second aggregated data to obtain target aggregated data.

5. The method according to claim 2, characterized in that, After using the query key to obtain first aggregated data from the key-value data, the method further includes: Obtain a display quantity threshold and a cache quantity threshold of the front-end display interface, divide the first aggregated data into multiple pages of display data according to the display quantity threshold and the cache quantity threshold, and cache the multiple pages of display data in the front-end display interface; In response to the display of the first aggregated data, each page of the display data is sequentially displayed on the front-end display interface, and the displayed display data is deleted from the cache of the front-end display interface; In response to the number of pages of the display data cached in the front-end display interface being less than a preset target page threshold, the first aggregated data that has not been cached to the front-end display interface is obtained; The first aggregated data that has not been cached is divided into multiple pages of display data according to the display quantity threshold and the cache quantity threshold until all the first aggregated data retrieved is displayed through the front-end display interface.

6. A data storage device, characterized in that, The apparatus includes: A data acquisition module, configured to acquire data to be stored, where the data to be stored is data whose data volume exceeds a preset quantity threshold; An aggregation grouping module, configured to aggregate and group the data to be stored according to preset metrics and time dimensions to obtain multiple groups of aggregated data, where each group of aggregated data in the multiple groups of aggregated data is one row, and the multiple groups of aggregated data form multi-row and multi-column data, and there are corresponding delimiters for each row and each column of data; A key-value combination module, configured to generate a key according to the preset metrics and time dimensions, and combine the key with the multiple groups of aggregated data according to the delimiter to obtain key-value data; A data storage module, configured to store the key-value data in a key-value database; The key-value combination module is further configured to generate message digest data for the preset metrics by using a message digest algorithm; is further configured to combine the message digest data and the time dimension to obtain a key corresponding to the time dimension; is further configured to, according to the time dimension, combine the key and the multiple groups of aggregated data according to the row delimiter corresponding to each row and the column delimiter corresponding to each column to obtain key-value data, where the length of the key is less than a preset byte threshold.

7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Data storage method and device, computer equipment and storage medium

    CN115203159A

  • Multi-database data processing method, apparatus, computer device, and storage medium

    WO2020155760A1