A User Data Storage Method and System Based on a Cloud Platform

By adaptively adjusting the length of the sliding window, the problem that the LZ77 algorithm cannot capture periodic duplicate data is solved, which improves data compression efficiency and reduces storage space requirements.

CN120128639BActive Publication Date: 2025-07-25GUANGZHOU FUNMI NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510622774.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-07-25
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In the prior art, the LZ77 algorithm is a fixed-length sliding window, and cannot capture periodic duplicate data in user data in time, resulting in poor compression effect in some data segments.

Method used

By adaptively adjusting the length of the sliding window, the LZ77 algorithm is used to compress data according to the periodicity and repetition of the data in the buffer, including obtaining the degree of repetition and periodicity, and iteratively adjusting the length of the sliding window to achieve better compression effect.

Benefits of technology

It realizes timely capture of periodic repeated data, improves data compression effect, and reduces storage space requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128639B_ABST
    Figure CN120128639B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data compression, and specifically discloses a method and system for storing user data based on a cloud platform. The method includes the following steps: obtaining the periodicity of the data in the buffer and the degree of repetition between the data in the buffer and the data in the sliding window; when the degree of repetition is greater than a threshold, keeping the length of the sliding window unchanged and compressing the data in the buffer by the LZ77 algorithm; otherwise, obtaining the current periodicity and the current degree of repetition again according to the corrected sliding window and the buffer, and continuously iterating until the current degree of repetition is greater than the threshold; compressing the data based on the sliding window at the last iteration; moving the current sliding window, and repeating the above steps until all the data in the current buffer is compressed. The method and system for storing user data based on a cloud platform provided by the present invention can capture periodically repeated data in a timely manner by adaptively adjusting the length of the sliding window, and thus can achieve a better data compression effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data compression, and particularly to a user data storage method and system based on a cloud platform. Background Art

[0002] A cloud platform is an online service platform based on cloud computing technology, which allows users to access and use various computing resources and services through the Internet. Cloud platform providers mainly provide services such as storage, network, computing, automation, and application management for multiple industries or departments such as finance, e-commerce, and government.

[0003] With the accelerating digital transformation of financial industries such as banks, the amount of user data has increased sharply. For banks, user data has extremely high value, and this data is crucial for the daily business operations, risk assessment, etc. of banks, so it needs to be stored for a long time. However, the sharp growth of the data volume has brought huge pressure to the storage of cloud platform providers, resulting in tight storage space. Cloud platform providers have to face problems such as increased hardware procurement, energy consumption, and maintenance costs.

[0004] To solve the problem of tight storage space, data is often compressed before storage. For example, in the Chinese patent application document with the publication number CN115361026A, a self-adaptive optimization method for the LZ series compression algorithm is disclosed. By obtaining the data to be compressed, the data to be compressed is divided into multiple partitions using the length of the LZ77 sliding window. The statements in the initial partition are compressed using the LZ77 sliding window dictionary. A tag value is established for the compressed statements and the tag value of the statement is incremented each time compression is performed. Among the statements with the same data in the compressed initial partition, the statements with shorter lengths are excluded, and the statements with longer lengths are retained, and the tags of the retained statements are updated. Statements that meet the criteria are screened out to obtain an auxiliary memory dictionary. The LZ77 sliding window dictionary is used to perform parallel auxiliary compression on each other partition. Each time a statement is compressed, the auxiliary memory dictionary is updated, and the attenuation value of each statement in the auxiliary memory dictionary is obtained, and the statements with attenuation values less than the attenuation threshold are deleted.

[0005] In the above solution, the LZ77 algorithm first pre-reads data into the buffer, and then moves the data into the fixed-length sliding window. It continuously searches for the longest phrase in the buffer that can match the phrase in the fixed-length sliding window, and then marks it with a marker to achieve compression. However, actual user data will have obvious periodicity, such as regular repayment, etc. Since the above LZ77 algorithm has a fixed-length sliding window, it cannot capture this periodically repeated data in time, resulting in poor compression effect in some data segments, thus affecting the overall data compression effect. Summary of the Invention

[0006] The present invention provides a user data storage method and system based on a cloud platform, aiming to solve the technical problem in the prior art that due to the fixed-length sliding window of the LZ77 algorithm, periodic repeated data cannot be captured in time, resulting in poor compression effects in some data segments.

[0007] A user data storage method based on a cloud platform according to the present invention includes the following steps:

[0008] Obtain user data;

[0009] Process the user data using the LZ77 algorithm to obtain a buffer and a sliding window;

[0010] Obtain the periodicity of the data in the buffer and the repetition degree of the data in the buffer and the data in the sliding window; wherein, the repetition degree is the sum of the products of the frequencies of all subsequences formed by the data in the buffer appearing in the sliding window and the corresponding subsequence weights; the weight is the ratio of the number of data in the corresponding subsequence to the sum of the number of data in all subsequences;

[0011] When the repetition degree is greater than the threshold, keep the sliding window length unchanged, and compress the data in the buffer through the LZ77 algorithm; when the repetition degree is not greater than the threshold, obtain the current periodicity and the current repetition degree again according to the modified sliding window and buffer, and continuously iterate until the current repetition degree is greater than the threshold; based on the sliding window at the last iteration, compress the data in the buffer through the LZ77 algorithm; wherein, the modified sliding window length is the ceiling of the product of the sliding window length at the previous iteration and the corresponding correction coefficient; the correction coefficient is positively correlated with the current periodicity and negatively correlated with the current repetition degree;

[0012] Move the current sliding window, and repeat the above steps until all the data in the current buffer is compressed.

[0013] In the above solution, when compressing the data in the buffer, first, according to the periodicity of the data in the buffer and the repetition degree of the data in the buffer and the data in the sliding window, through multiple iterations, the sliding window length is adaptively adjusted with the goal of the repetition degree being greater than the threshold, and the data in the buffer is compressed using the LZ77 algorithm based on the modified sliding window. Compared with the fixed-length sliding window in the prior art, the present invention can capture periodic repeated data in time by adaptively adjusting the sliding window length, and thus can achieve better data compression effects.

[0014] Preferably, the periodicity is:

[0015] ;

[0016] In the formula, is thek The frequency of occurrence of a data in historical data, is the time interval between the k th occurrence and the h ( h +1)th occurrence of the th data in historical data, k is the average of the time intervals between all adjacent occurrences of the th data in historical data, k is the standard deviation of the time intervals between all adjacent occurrences of the N th data in historical data.

[0017] In the above solution, by calculating the average of the ratios of the frequencies of occurrence of all data in the buffer in historical data to the degree of dispersion of the time intervals between all adjacent occurrences of the corresponding data in historical data, the periodicity of the data in the buffer is obtained, so that the data with a higher frequency of occurrence in historical data and the data with a smaller degree of dispersion of the corresponding time intervals are emphasized, which is conducive to improving the data compression effect.

[0018] Preferably, the correction coefficient r is:

[0019] ;

[0020] In the formula, c is the periodicity, d is the degree of repetition, norm ( ) is the standard normalization function.

[0021] In the above solution, the above formula ensures that the value range of the correction coefficient r is from 0.5 to 1.5, meeting the requirement of adjusting the size of the sliding window through the correction coefficient.

[0022] Preferably, the user data includes different types of data of the user.

[0023] In the above solution, by dividing the user data into different types of data, the difference of the same type of data is small, which is convenient for data compression and improves the data compression effect.

[0024] Preferably, the subsequence includes multiple data sets, and each data set is composed of a single data or adjacent multiple data in the buffer.

[0025] Preferably, the duration of the historical data is one week, one month or one quarter.

[0026] In the above solution, historical data with different time lengths can be selected according to different data types, so as to make the calculation of data more reasonable and efficient.

[0027] Preferably, the value range of the threshold is from 0.5 to 0.7.

[0028] The present invention also provides a user data storage system based on a cloud platform, including a memory and a processor. The processor executes the computer program stored in the memory to implement the user data storage method based on the cloud platform as described in any one of the above.

[0029] The beneficial effects are:

[0030] When compressing the data in the buffer according to the solution of the present invention, first, according to the periodicity of the data in the buffer and the degree of repetition of the data in the buffer and the data in the sliding window, through multiple iterations, the length of the sliding window is adaptively adjusted with the target that the degree of repetition is greater than the threshold. Based on the modified sliding window, the LZ77 algorithm is used to compress the data in the buffer. Compared with the fixed-length sliding window in the prior art, by adaptively adjusting the length of the sliding window, the present invention can capture the periodically repeated data in time, and thus can achieve a better data compression effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0032] Figure 1 is a flowchart of the steps of the user data storage method based on the cloud platform according to the embodiment of the present invention;

[0033] Figure 2 is a structural block diagram of the user data storage system based on the cloud platform according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as a limitation to the present invention.

[0035] As Figure 1 shown, according to the first aspect of the present invention, a user data storage method based on a cloud platform is provided, including the following steps:

[0036] S1. Obtain user data.

[0037] Any user data can be extracted from the database according to the user ID through the internal service interface of the cloud platform provider. User data includes different types of data for that user. Taking bank user data as an example, it includes different types of data such as income data, consumption data, loan repayment data, and balance data.

[0038] Since there are significant differences between different types of data, in order to facilitate data compression and improve the compression effect, data of the same type is placed in the same data set. For example, all income data is placed in the income data set, and all consumption data is placed in the consumption data set.

[0039] When compressing the data, each data set of the user is compressed separately. After the compression is completed, each data set of other users is compressed until the data compression of all users is completed.

[0040] S2. Process the user data using the LZ77 algorithm to obtain a buffer and a sliding window.

[0041] It should be noted that the LZ77 algorithm itself belongs to the prior art and is based on the following principle: The LZ77 algorithm includes a buffer and a sliding window. First, the data is pre-read through the buffer, and then the data is moved into the sliding window. The longest string that is the same as the string in the sliding window is continuously searched in the buffer, and then these strings are replaced with shorter markers to achieve data compression.

[0042] Specifically, the LZ77 algorithm includes the following core steps:

[0043] 1. Pre-read the data through the buffer, move the sliding window towards the buffer direction to fill the sliding window with data;

[0044] 2. Search for the longest string that is the same as the string in the sliding window in the buffer;

[0045] 3. If the string is found, output (offset value, match length, next data), and move the sliding window towards the buffer by the match length plus one unit length; where the offset value is the distance between the string in the buffer and the same string in the sliding window, the match length is the length of the string, and the next data is the first data after the string in the buffer;

[0046] 4. If the string is not found, output (0, 0, the first data in the buffer), and move the sliding window towards the buffer by one unit length.

[0047] 5. Repeat steps 2 to 4 until all the data in the buffer is compressed.

[0048] Therefore, through this step, a buffer and a sliding window can be obtained, providing a basis for the calculations in subsequent steps.

[0049] S3. Obtain the periodicity of the data in the buffer and the degree of repetition between the data in the buffer and the data in the sliding window. Step S3 specifically includes the following steps:

[0050] S31. Obtain the degree of repetition between the data in the buffer and the data in the sliding window.

[0051] When compressing user data, if the degree of repetition between the data in the buffer and the data in the sliding window is relatively high, the LZ77 algorithm can be used to achieve a relatively high compression ratio for the data in the buffer. Therefore, it is necessary to determine the degree of repetition between the data in the buffer and the data in the sliding window.

[0052] Adjacent data in the buffer generally has a certain correlation. Multiple data sets are formed by individual data or adjacent multiple data in the buffer, and the multiple data sets are multiple subsequences. When compressing the data in the buffer, if the frequency of occurrence of the subsequences in the buffer in the sliding window is higher and the length of the subsequence is longer, it indicates that the data in the buffer can be replaced by shorter markers for these subsequences, thereby achieving timely and efficient compression of the data.

[0053] Therefore, the degree of repetition is the sum of the products of the frequencies of occurrence of all subsequences formed by the data in the buffer in the sliding window and the corresponding subsequence weights; the weight is the ratio of the number of data in the corresponding subsequence to the sum of the number of data in all subsequences.

[0054] Specifically, the calculation process of the degree of repetition is as follows:

[0055] S311. Obtain the frequencies of occurrence of all subsequences in the sliding window; it includes the following steps:

[0056] First, obtain the subsequences in the buffer.

[0057] Since the subsequences include data sets composed of individual data and data sets composed of adjacent multiple data, it is necessary to traverse the data in the buffer to obtain all subsequences. The number of all subsequences in the buffer is:

[0058] ;

[0059] In the formula, is the total number of data in the buffer.

[0060] For example, if the number of data in the buffer is 3 and they are 1, 2, and 3 respectively, then the number of all subsequences in this buffer is: = 6, and all the subsequences in the buffer are respectively: [1], [2], [3], [1, 2], [2, 3], [1, 2, 3].

[0061] Secondly, obtain the subsequences within the sliding window.

[0062] Similarly, the subsequences within the sliding window can be obtained by using the method of the above steps.

[0063] Finally, obtain the frequencies of all the subsequences in the buffer appearing within the sliding window.

[0064] For each subsequence in the buffer, traverse the sliding window, and the frequencies of all the subsequences in the buffer appearing within the sliding window can be obtained. For example, it can be used m i to represent the frequency of the i -th subsequence in the buffer appearing within the sliding window.

[0065] S312. Obtain the weights of the corresponding subsequences in the buffer.

[0066] The longer the length of the subsequence in the buffer, that is, the more data in the subsequence, the more the storage space of the data can be reduced after compressing the subsequence, and the compression effect can be greatly improved. Therefore, the longer subsequences should be given more attention, and their influence on the whole in data compression is greater, and larger weights should be given. Therefore, the weight is the ratio of the number of data in the corresponding subsequence to the sum of the number of data in all subsequences.

[0067] Specifically, the weight i of the w i -th subsequence in the buffer is:

[0068]

[0069] In the formula, is the number of data of the i -th subsequence in the buffer, is the number of all subsequences in the buffer.

[0070] S313. Obtain the degree of repetition.

[0071] The degree of repetition of the data in the buffer and the data in the sliding window D is:

[0072]

[0073] In the formula, m i is the frequency of the i -th subsequence in the buffer appearing within the sliding window, wi is the weight of the i th subsequence in the buffer, and is the number of all subsequences in the buffer.

[0074] In this step, by calculating the sum of the products of the frequencies of all subsequences in the buffer appearing in the sliding window and the corresponding subsequence weights, the degree of repetition is obtained, so that subsequences with higher frequencies and longer subsequences are emphasized, which is conducive to improving the data compression effect.

[0075] S32. Obtain the periodicity of the data in the buffer.

[0076] Actual user data often has certain regularity. Taking bank user data as an example, operations such as fixed-term deposits, loan repayments, and payment fees necessary for going to work and traveling will occur repeatedly at specific time intervals and most of the data are equal. This regularity makes some data appear frequently within a fixed time length. Then, when compressing this data, by analyzing the periodicity of this data in the historical data in the buffer, for data with stronger periodicity, it can be discovered in time and given more attention during data compression.

[0077] It can be seen from the above analysis that the higher the frequency of the data in the buffer appearing in the historical data, and the smaller the dispersion degree of the time interval between two adjacent appearances of this data, the stronger the periodicity of the data in the buffer, and more attention should be paid when compressing the data.

[0078] The specific process of step S32 is as follows:

[0079] S321. Obtain the frequency of the data in the buffer appearing in the historical data.

[0080] In this step, the duration of the historical data is one week, one month or one quarter. For example, for the payment fees necessary for going to work and traveling, which are transactions that occur every day, the payment fees within one week are used as the historical data. For loan repayments, which are transactions that occur every month, the loan repayments within one quarter are used as the historical data. Of course, the duration of the historical data can be selected according to the actual situation. The historical data is of the same type as the data in the corresponding buffer.

[0081] By traversing each data in the buffer in the historical data, the frequency of this data appearing in the historical data can be obtained. For example, can be used to represent the frequency of the k th data in the buffer appearing in the historical data.

[0082] S322. Obtain the dispersion degree of the time interval between two adjacent appearances of the data in the buffer in the historical data.

[0083] Among them, the degree of dispersion can be represented by the standard deviation of the time intervals between all adjacent occurrences of the data in the buffer in the historical data. Specifically, the degree of dispersion of the time intervals between all adjacent occurrences of the th data in the buffer in the historical data is:

[0084] ;

[0085] In the formula, is the frequency of occurrence of the k th data in the buffer in the historical data, is the time interval between the k th occurrence and the h st occurrence of the h th data in the buffer in the historical data, is the mean value of the time intervals between all adjacent occurrences of the k th data in the buffer in the historical data.

[0086] S323. Obtain the periodicity of the data in the buffer.

[0087] The said periodicity C is:

[0088] ;

[0089] In the formula, is the frequency of occurrence of the k th data in the buffer in the historical data, is the degree of dispersion, i.e., the standard deviation, of the time intervals between all adjacent occurrences of the k th data in the buffer in the historical data, N is the total number of data in the buffer.

[0090] In this step, by calculating the mean value of the ratios of the frequencies of occurrence of all data in the buffer in the historical data to the degrees of dispersion of the time intervals between all adjacent occurrences of the corresponding data in the historical data, the periodicity of the data in the buffer is obtained, so that the data with higher frequencies of occurrence in the historical data and the data with smaller degrees of dispersion of the corresponding time intervals are emphasized, which is conducive to improving the data compression effect.

[0091] In some alternative embodiments, the periodicity C can also be expressed as:

[0092] ;

[0093] In the formula, is the k th occurrence of the hThe time interval between the h nth and the (n + 1)th occurrences, N where N is the total number of data in the buffer, norm and φ( ) is the standard normalization function.

[0094] S4. When the degree of repetition is greater than the threshold, keep the length of the sliding window unchanged and compress the data in the buffer by the LZ77 algorithm. When the degree of repetition is not greater than the threshold, re-obtain the current periodicity and the current degree of repetition according to the modified sliding window and the buffer, and iterate continuously until the current degree of repetition is greater than the threshold; based on the sliding window at the last iteration, compress the data in the buffer by the LZ77 algorithm.

[0095] Here, the current periodicity and the current degree of repetition are the periodicity and the degree of repetition corresponding to each iteration.

[0096] In this step, the value range of the threshold is from 0.5 to 0.7. Preferably, the threshold is 0.6.

[0097] This step is divided into two cases, namely the case where the degree of repetition is greater than the threshold and the case where it is not greater than the threshold, which will be described separately below.

[0098] 1. When the degree of repetition is greater than the threshold:

[0099] This case indicates that the degree of repetition of the data in the buffer and the sliding window is relatively high, which means that by searching for the data in the sliding window, a relatively high compression of the data in the buffer can be achieved. Therefore, there is no need to adjust the size of the sliding window, that is, keep the length of the sliding window unchanged. The subsequent step of compressing the data in the buffer by the LZ77 algorithm belongs to the prior art and will not be elaborated here.

[0100] 2. When the degree of repetition is not greater than the threshold:

[0101] This case indicates that the degree of repetition of the data in the buffer and the sliding window is relatively low, which means that it is difficult to achieve compression of the data in the buffer by searching for the data in the sliding window. Therefore, it is necessary to adjust the size of the sliding window.

[0102] Specifically, if the data in the buffer has strong periodicity, it means that the data in the buffer can achieve relatively large compression. However, due to the limited length of the sliding window, the degree of repetition of the data in the buffer and the sliding window is relatively low. Therefore, the length of the sliding window should be increased to expand the search range in the sliding window, so as to find more identical data and thus improve the compression effect of the data in the buffer.

[0103] On the contrary, if the data in the buffer has weak periodicity, it means that the possibility of achieving relatively large compression of the data in the buffer is small. The length of the sliding window should be reduced to narrow the search range in the sliding window, so as to save the search time.

[0104] The adjusted sliding window is called the corrected sliding window. Based on the corrected sliding window and the buffer, the current periodicity and the current degree of repetition are obtained again. By continuous iteration, the size of the sliding window is continuously adjusted until the current degree of repetition obtained based on the current sliding window and the current buffer is greater than the threshold. At this time, it indicates that by adjusting the length of the sliding window, the degree of repetition between the data in the buffer and the data in the sliding window is relatively high, and a large compression can be achieved. Therefore, the data in the buffer can be compressed by the LZ77 algorithm based on the sliding window at the last iteration.

[0105] Among them, the length of the corrected sliding window is the ceiling of the product of the length of the sliding window at the previous iteration and the corresponding correction coefficient; the correction coefficient is positively correlated with the corresponding periodicity and negatively correlated with the corresponding degree of repetition.

[0106] Specifically, the length of the corrected sliding window is:

[0107]

[0108] The correction coefficient is:

[0109] ;

[0110] In the formula, is the length of the sliding window at the previous iteration, c is the periodicity, d is the degree of repetition, norm ( ) is the standard normalization function, is the ceiling function.

[0111] The above formula ensures that the value range of the correction coefficient r is from 0.5 to 1.5, meeting the requirement of adjusting the size of the sliding window through the correction coefficient.

[0112] In this step, aiming at the degree of repetition of the data in the buffer being greater than the threshold, the length of the sliding window is continuously adjusted through iteration, and then the optimal sliding window length is obtained, so that the degree of repetition between the data in the buffer and the data in the sliding window is relatively high, and thus the compression effect of the data in the buffer can be improved.

[0113] S5. Move the current sliding window and repeat the above steps until all the data in the current buffer is compressed.

[0114] Steps S1 to S4 only achieve one-time compression of the data in the buffer. Therefore, it is necessary to move the current sliding window by the number of bits of the compressed data in the direction of the buffer and repeat steps S1 to S4 until all the data in the current buffer is compressed.

[0115] In the present invention, when compressing the data in the buffer, first, according to the periodicity of the data in the buffer and the degree of repetition between the data in the buffer and the data in the sliding window, through multiple iterations, the length of the sliding window is adaptively adjusted with the goal that the degree of repetition is greater than the threshold. Based on the modified sliding window, the LZ77 algorithm is used to compress the data in the buffer. Compared with the fixed-length sliding window in the prior art, by adaptively adjusting the length of the sliding window, the present invention can timely capture the periodically repeated data, and thus can achieve a better data compression effect.

[0116] As Figure 2 shown, according to the second aspect of the present invention, there is also provided a user data storage system based on a cloud platform. The system includes a memory and a processor, and the processor executes the computer program stored in the memory to implement the user data storage method based on the cloud platform described in the first aspect of the present invention.

[0117] The system also includes other components well known to those skilled in the art such as a communication bus and a communication interface, and their settings and functions are known in the art, so they will not be described in detail here.

[0118] In the present invention, the aforementioned memory can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. For example, a computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high-bandwidth memory (HBM), a hybrid memory cube (HMC), etc., or any other medium that can be used to store the required information and can be accessed by an application program, module, or both. Any such computer storage medium can be part of the device or accessible or connectable to the device. Any application or module described in the present invention can be implemented by computer-readable / executable instructions stored or otherwise held by such a computer-readable medium.

[0119] In the description of this specification, the meaning of "a plurality" is at least two, such as two, three, or more, etc., unless otherwise specifically defined.

[0120] Although this specification has shown and described several embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, modifications and alternative forms will occur to those skilled in the art without departing from the spirit and scope of the present invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention.

Claims

1. A user data storage method based on a cloud platform, characterized in that, It includes the following steps: Obtain user data; Process the user data using the LZ77 algorithm to obtain a buffer and a sliding window; Obtain the periodicity of the data in the buffer and the degree of repetition between the data in the buffer and the data in the sliding window; wherein, the degree of repetition is the sum of the products of the frequencies of all subsequences formed by the data in the buffer appearing in the sliding window and the corresponding subsequence weights; the weight is the ratio of the number of data in the corresponding subsequence to the sum of the number of data in all subsequences; When the degree of repetition is greater than the threshold, keep the length of the sliding window unchanged and compress the data in the buffer through the LZ77 algorithm; when the degree of repetition is not greater than the threshold, obtain the current periodicity and the current degree of repetition again according to the modified sliding window and buffer, and iterate continuously until the current degree of repetition is greater than the threshold; based on the sliding window at the last iteration, compress the data in the buffer through the LZ77 algorithm; wherein, the length of the modified sliding window is the ceiling of the product of the length of the sliding window at the previous iteration and the corresponding correction coefficient; the correction coefficient is positively correlated with the current periodicity and negatively correlated with the current degree of repetition; Move the current sliding window and repeat the above steps until all the data in the current buffer is compressed.

2. The method for storing user data based on a cloud platform according to claim 1, wherein The periodicity is as follows: ; Wherein, is the frequency of occurrence of the k -th data in the buffer zone in the historical data, is the time interval between the k -th data in the buffer zone appearing for the h -th time and the h +1-th time in the historical data, is the average value of the time intervals between all adjacent occurrences of the k -th data in the buffer zone in the historical data, is the standard deviation of the time intervals between all adjacent occurrences of the k -th data in the buffer zone in the historical data, N is the total number of data in the buffer zone.

3. The user data storage method based on a cloud platform according to claim 1, wherein The correction coefficient r is as follows: ; In the formula, c is the periodicity, d is the degree of repetition, norm ( ) is the standard normalization function.

4. The user data storage method based on a cloud platform according to claim 1, characterized in that, The user data includes different types of data of the user.

5. The user data storage method based on a cloud platform according to claim 1, wherein The subsequence includes multiple data sets, and each data set is composed of a single data or adjacent multiple data in the buffer.

6. The method for storing user data based on a cloud platform according to claim 2, wherein The duration of the historical data is one week, one month or one quarter.

7. The user data storage method based on a cloud platform according to claim 1, characterized in that The value range of the threshold is from 0.5 to 0.

7.

8. A user data storage system based on a cloud platform, comprising a memory and a processor, characterized in that, The processor executes the computer program stored in the memory to implement the method for storing user data based on a cloud platform as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Efficient storage method for applet data

    CN115173866A

  • LZ series compression algorithm adaptive optimization method

    CN115361026A