User data storage method and system based on cloud platform

By adaptively adjusting the sliding window length and using the LZ77 algorithm to process user data, the problem of inability to capture periodic duplicate data in the prior art is solved, and better data compression effect is achieved.

CN120128639AActive Publication Date: 2025-06-10GUANGZHOU FUNMI NETWORK TECH CO LTD

Patent Information

Application Number
CN202510622774.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-10
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In the prior art, the LZ77 algorithm is a fixed-length sliding window, and cannot capture periodic duplicate data in user data in time, resulting in poor compression effect in some data segments.

Method used

By obtaining the periodicity of user data and the repetition of data in the buffer and data in the sliding window, the sliding window length is adaptively adjusted, and the LZ77 algorithm is used to realize the compression of data in the buffer.

Benefits of technology

It achieves a good data compression effect, can capture periodic repeated data in a timely manner, and improves the overall data compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128639A_ABST
    Figure CN120128639A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data compression, and particularly discloses a user data storage method and system based on a cloud platform, and the method comprises the following steps: obtaining the periodicity of data in a buffer area and the repetition degree of the data in the buffer area and the data in a sliding window; when the repetition degree is greater than a threshold value, keeping the length of a sliding window unchanged, and compressing data in a buffer area through an LZ77 algorithm; otherwise, the current periodicity and the current repetition degree are obtained again according to the corrected sliding window and the buffer area, and iteration is conducted continuously till the current repetition degree is larger than the threshold value; compressing the data based on the sliding window in the last iteration; and moving the current sliding window, and repeating the steps until all data in the current buffer area are compressed. According to the user data storage method and system based on the cloud platform provided by the invention, the periodically repeated data can be captured in time by adaptively adjusting the length of the sliding window, so that a better data compression effect can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data compression, and particularly relates to a user data storage method and system based on a cloud platform. Background Art

[0002] A cloud platform is an online service platform based on cloud computing technology, which allows users to access and use various computing resources and services through the Internet. Cloud platform providers mainly provide services such as storage, network, computing, automation, and application management for multiple industries or departments such as finance, e-commerce, and government.

[0003] With the accelerating digital transformation of financial industries such as banks, the amount of user data has increased sharply. For banks, user data has extremely high value, and this data is crucial for the daily business operations, risk assessment, etc. of banks, so it needs to be stored for a long time. However, the sharp growth in the amount of data has brought huge pressure to the storage of cloud platform providers, resulting in tight storage space. Cloud platform providers have to face problems such as increased hardware procurement, energy consumption, and maintenance costs.

[0004] To solve the problem of tight storage space, data is often compressed before storage. For example, in the Chinese patent application document with the publication number CN115361026A, a self-adaptive optimization method for the LZ series compression algorithm is disclosed. By obtaining the data to be compressed, the data to be compressed is divided into multiple partitions using the length of the LZ77 sliding window. The statements in the initial partition are compressed using the LZ77 sliding window dictionary. A tag value is established for the compressed statements and the tag value of the statement is incremented each time compression is performed. Among the statements with the same data in the initially compressed partition, the statements with shorter lengths are excluded and the statements with longer lengths are retained, and the tags of the retained statements are updated. Statements that meet the criteria are selected to obtain an auxiliary memory dictionary. The LZ77 sliding window dictionary is used to perform parallel auxiliary compression on each other partition. Each time a statement is compressed, the auxiliary memory dictionary is updated, and the attenuation value of each statement in the auxiliary memory dictionary is obtained, and the statements with attenuation values less than the attenuation threshold are deleted.

[0005] In the above solution, the LZ77 algorithm first pre-reads data into the buffer, and then moves the data into the fixed-length sliding window. It continuously searches for the longest phrase in the buffer that can match the phrase in the fixed-length sliding window, and then marks it with a marker to achieve compression. However, actual user data will have obvious periodicity, such as regular repayment, etc. Since the above LZ77 algorithm has a fixed-length sliding window, it cannot capture this periodically repeated data in time, resulting in poor compression effect in some data segments, thus affecting the overall data compression effect. Summary of the Invention

[0006] The present invention provides a method and system for storing user data based on a cloud platform, aiming to solve the technical problem in the prior art that due to the fixed-length sliding window of the LZ77 algorithm, periodic repeated data cannot be captured in time, resulting in poor compression effects in some data segments.

[0007] A method for storing user data based on a cloud platform according to the present invention includes the following steps: Obtain user data; Process the user data using the LZ77 algorithm to obtain a buffer and a sliding window; Obtain the periodicity of the data in the buffer and the degree of repetition between the data in the buffer and the data in the sliding window; wherein, the degree of repetition is the sum of the products of the frequencies of all subsequences formed by the data in the buffer appearing in the sliding window and the corresponding subsequence weights; the weight is the ratio of the number of data in the corresponding subsequence to the sum of the number of data in all subsequences; When the degree of repetition is greater than the threshold, keep the length of the sliding window unchanged, and compress the data in the buffer through the LZ77 algorithm; when the degree of repetition is not greater than the threshold, obtain the current periodicity and the current degree of repetition again according to the modified sliding window and buffer, and iterate continuously until the current degree of repetition is greater than the threshold; based on the sliding window at the last iteration, compress the data in the buffer through the LZ77 algorithm; wherein, the length of the modified sliding window is the ceiling of the product of the length of the sliding window at the previous iteration and the corresponding correction coefficient; the correction coefficient is positively correlated with the current periodicity and negatively correlated with the current degree of repetition; Move the current sliding window, and repeat the above steps until all the data in the current buffer is compressed.

[0008] In the above solution, when compressing the data in the buffer, first, according to the periodicity of the data in the buffer and the degree of repetition between the data in the buffer and the data in the sliding window, through multiple iterations, adaptively adjust the length of the sliding window with the goal that the degree of repetition is greater than the threshold, and use the LZ77 algorithm based on the modified sliding window to compress the data in the buffer. Compared with the fixed-length sliding window in the prior art, the present invention can capture periodic repeated data in time by adaptively adjusting the length of the sliding window, and thus can achieve better data compression effects.

[0009] Preferably, the periodicity is: ; In the formula, is the frequency of the k th data in the buffer appearing in the historical data, is the k th time the h th data in the buffer appears in the historical data and theh The time interval of the +1st occurrence, is the average of the time intervals between all adjacent occurrences of the k th data in the buffer within the historical data, is the k th data in the buffer, and is the standard deviation of the time intervals between all adjacent occurrences of the data in the historical data, N is the total number of data in the buffer.

[0010] In the above solution, by calculating the mean of the ratio of the frequency of occurrence of all data in the buffer within the historical data to the degree of dispersion of the time intervals between all adjacent occurrences of the corresponding data in the historical data, the periodicity of the data in the buffer is obtained, so that the data with a higher frequency of occurrence in the historical data and the data with a smaller degree of dispersion of the corresponding time intervals are emphasized, which is beneficial to improving the data compression effect.

[0011] Preferably, the correction coefficient r is: ; In the formula, c is the periodicity, d is the degree of repetition, norm ( ) is the standard normalization function.

[0012] In the above solution, the above formula ensures that the value range of the correction coefficient r is from 0.5 to 1.5, meeting the requirement of adjusting the size of the sliding window through the correction coefficient.

[0013] Preferably, the user data includes different types of data of the user.

[0014] In the above solution, by dividing the user data into different types of data, the difference of the same type of data is small, which is convenient for data compression and improves the data compression effect.

[0015] Preferably, the subsequence includes multiple data sets, and each data set is composed of a single data or adjacent multiple data in the buffer.

[0016] Preferably, the duration of the historical data is one week, one month or one quarter.

[0017] In the above solution, different durations of historical data can be selected according to different data types, making the calculation of data more reasonable and efficient.

[0018] Preferably, the value range of the threshold is from 0.5 to 0.7.

[0019] The present invention also provides a user data storage system based on a cloud platform, including a memory and a processor. The processor executes the computer program stored in the memory to implement the user data storage method based on the cloud platform as described in any one of the above.

[0020] Beneficial effects are as follows: When compressing the data in the buffer according to the solution of the present invention, first, according to the periodicity of the data in the buffer and the repetition degree of the data in the buffer and the data in the sliding window, through multiple iterations, the length of the sliding window is adaptively adjusted with the repetition degree greater than the threshold as the target. Based on the corrected sliding window, the LZ77 algorithm is used to compress the data in the buffer. Compared with the fixed-length sliding window in the prior art, by adaptively adjusting the length of the sliding window, the present invention can capture the periodically repeated data in time, and thus can achieve a better data compression effect. Description of the Drawings

[0021] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become easily understandable. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts, where: Figure 1 is a flowchart of the steps of the user data storage method based on the cloud platform according to the embodiment of the present invention; Figure 2 is a block diagram of the structure of the user data storage system based on the cloud platform according to the embodiment of the present invention. Detailed Embodiments

[0022] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0023] As Figure 1 shown, according to the first aspect of the present invention, a user data storage method based on a cloud platform is provided, including the following steps: S1. Obtain user data.

[0024] Any user data can be extracted from the database according to the user ID through the internal service interface of the cloud platform provider. The user data includes different types of data of the user. Taking bank user data as an example, for example, it includes different types of data such as income data, consumption data, loan repayment data, and balance data.

[0025] Since different types of data are quite different, in order to facilitate data compression and improve the compression effect, the same type of data is placed in the same data set. For example, all income data is placed in the income data set, and all consumption data is placed in the consumption data set.

[0026] When compressing data, each data set of the user is compressed separately. After the compression is completed, each data set of other users is compressed again until the data compression work of all users is completed.

[0027] S2. Use the LZ77 algorithm to process user data to obtain a buffer and a sliding window.

[0028] It should be noted that the LZ77 algorithm itself belongs to the prior art, and is based on the following principle: The LZ77 algorithm includes a buffer and a sliding window. The data is first pre-read through the buffer, and then the data is moved into the sliding window. The longest string that is the same as the string in the sliding window is continuously searched in the buffer, and then these strings are replaced with shorter tags, thereby achieving data compression.

[0029] Specifically, the LZ77 algorithm includes the following core steps: 1. Pre-read data through the buffer and move the sliding window toward the buffer to fill the sliding window with data; 2. Search the buffer for the longest string that is the same as the longest string in the sliding window. 3. If the string is found, output (offset value, match length, next data), and move the sliding window to the buffer by the match length plus one unit length; the offset value is the distance between the string in the buffer and the same string in the sliding window, the match length is the length of the string, and the next data is the first data after the string in the buffer; 4. If the string is not found, output (0,0, the first data in the buffer) and move the sliding window one unit length into the buffer.

[0030] 5. Repeat steps 2 to 4 until all the data in the buffer is compressed.

[0031] Therefore, through this step, the buffer zone and sliding window can be obtained, which provide a basis for the calculation of subsequent steps.

[0032] S3, obtaining the periodicity of the data in the buffer and the degree of repetition between the data in the buffer and the data in the sliding window. Step S3 specifically includes the following steps: S31. Obtain the degree of duplication between the data in the buffer and the data in the sliding window.

[0033] When compressing user data, if the repetition degree of the data in the buffer and the data in the sliding window is high, the LZ77 algorithm can achieve a higher compression rate for the data in the buffer. Therefore, it is necessary to determine the repetition degree of the data in the buffer and the data in the sliding window.

[0034] Adjacent data in the buffer generally have a certain correlation, and a single data or multiple adjacent data in the buffer constitute multiple data sets, and multiple data sets are multiple subsequences. When compressing the data in the buffer, if the subsequence in the buffer appears more frequently in the sliding window and the length of the subsequence is longer, it means that the data in the buffer can be replaced by shorter tags, thereby achieving timely and efficient data compression.

[0035] Therefore, the repetition degree is the sum of the product of the frequency of occurrence of all subsequences composed of data in the buffer in the sliding window and the corresponding subsequence weight; the weight is the ratio of the number of data in the corresponding subsequence to the sum of the number of data in all subsequences.

[0036] Specifically, the calculation process of the repetition degree is as follows: S311, obtaining the frequencies of occurrence of all subsequences in the sliding window; which includes the following steps: First, get the subsequence in the buffer.

[0037] Since subsequences include data sets consisting of single data and data sets consisting of multiple adjacent data, it is necessary to traverse the data in the buffer to obtain all subsequences. The number of all subsequences in the buffer for: ; In the formula, The total number of data in the buffer.

[0038] For example, if the number of data in the buffer is 3, and they are 1, 2, and 3 respectively, then the number of all subsequences in the buffer is for: =6, and all subsequences in the buffer are: [1], [2], [3], [1,2], [2,3], [1,2,3].

[0039] Next, obtain the subsequence within the sliding window.

[0040] Similarly, the subsequence within the sliding window can be obtained by using the method in the above steps.

[0041] Finally, obtain the frequency of all subsequences in the buffer appearing in the sliding window.

[0042] For each subsequence in the buffer, traverse the sliding window to obtain the frequency of occurrence of all subsequences in the buffer within the sliding window. For example, m i can be used to represent the i frequency of occurrence of the

[0043] subsequence in the buffer within the sliding window.

[0044] The longer the length of the subsequence in the buffer, that is, the more data in the subsequence, the more the storage space of the data can be reduced after compressing the subsequence, greatly improving the compression effect. Therefore, longer subsequences should receive more attention, as they have a greater impact on the overall data compression and should be given a greater weight. Thus, the weight is the ratio of the number of data in the corresponding subsequence to the sum of the number of data in all subsequences.

[0045] Specifically, the weight i of the w i subsequence in the buffer is:

[0046] In the formula, is the number of data in the i subsequence in the buffer, is the number of all subsequences in the buffer.

[0047] S313. Obtain the degree of repetition.

[0048] The degree of repetition of the data in the buffer and the data in the sliding window D is:

[0049] In the formula, m i is the frequency of occurrence of the i subsequence in the buffer within the sliding window, w i is the weight of the i subsequence in the buffer, is the number of all subsequences in the buffer.

[0050] In this step, by calculating the sum of the products of the frequencies of occurrence of all subsequences in the buffer within the sliding window and the corresponding subsequence weights, the degree of repetition is obtained, enabling subsequences with higher frequencies of occurrence and longer subsequences to be emphasized, thereby facilitating the improvement of the data compression effect.

[0051] S32. Obtain the periodicity of the data in the buffer.

[0052] Actual user data often has certain regularity. Taking bank user data as an example, operations such as fixed-term deposits, loan repayments, and payment fees necessary for going to work and traveling occur repeatedly at specific time intervals and most of the data are equal. This regularity makes certain data appear frequently within a fixed duration. Then, when compressing this data, by analyzing the periodicity of this data in the historical data within the buffer, for data with strong periodicity, it can be discovered in a timely manner and given more attention during data compression.

[0053] As can be seen from the above analysis, the higher the frequency of the data in the buffer appears in the historical data, and the smaller the degree of dispersion of the time interval between two adjacent occurrences of this data, the stronger the periodicity of the data in the buffer, and more attention should be paid when compressing the data.

[0054] The specific process of step S32 is as follows: S321. Obtain the frequency of the data in the buffer appearing in the historical data.

[0055] In this step, the duration of the historical data is one week, one month, or one quarter. For example, for the payment fees necessary for going to work and traveling, which are transactions that occur every day, the payment fees within one week are used as the historical data. For loan repayments, which are transactions that occur every month, the loan repayments within one quarter are used as the historical data. Of course, the duration of the historical data can be selected according to the actual situation. The historical data is of the same type as the data in the corresponding buffer.

[0056] By traversing each data in the buffer in the historical data, the frequency of this data appearing in the historical data can be obtained. For example, it can be represented by indicating the frequency of the k th data in the buffer appearing in the historical data.

[0057] S322. Obtain the degree of dispersion of the time interval between two adjacent occurrences of the data in the buffer in the historical data.

[0058] Among them, the degree of dispersion can be represented by the standard deviation of the time intervals between all adjacent two occurrences of the data in the buffer in the historical data. Specifically, the degree of dispersion of the time intervals between all adjacent two occurrences of the th data in the buffer in the historical data is: ; In the formula, is the frequency of the k th data in the buffer appearing in the historical data, is the k th time the h th data in the buffer appears in the historical data and theh The time interval of the +1st occurrence, is the average value of the time intervals between all adjacent occurrences of the k th data in the buffer within the historical data.

[0059] S323. Obtain the periodicity of the data in the buffer.

[0060] The said periodicity C is: ; In the formula, is the frequency of occurrence of the k th data in the buffer within the historical data, is the degree of dispersion, i.e., the standard deviation, of the time intervals between all adjacent occurrences of the k th data in the buffer within the historical data, N is the total number of data in the buffer.

[0061] In this step, by calculating the average value of the ratio of the frequency of occurrence of all data in the buffer within the historical data to the degree of dispersion of the time intervals between all adjacent occurrences of the corresponding data within the historical data, the periodicity of the data in the buffer is obtained, so that the data with a higher frequency of occurrence in the historical data and the data with a smaller degree of dispersion of the corresponding time intervals are emphasized, which is conducive to improving the data compression effect.

[0062] In some alternative embodiments, the periodicity C can also be expressed as: ; In the formula, is the time interval between the k th occurrence and the h th +1 occurrence of the h th data in the buffer within the historical data, N is the total number of data in the buffer, norm ( ) is the standard normalization function.

[0063] S4. When the degree of repetition is greater than the threshold, keep the sliding window length unchanged and compress the data in the buffer by the LZ77 algorithm. When the degree of repetition is not greater than the threshold, obtain the current periodicity and the current degree of repetition again according to the modified sliding window and the buffer, and iterate continuously until the current degree of repetition is greater than the threshold; based on the sliding window at the last iteration, compress the data in the buffer by the LZ77 algorithm.

[0064] Among them, the current periodicity and the current degree of repetition are the corresponding periodicity and degree of repetition at each iteration.

[0065] In this step, the value range of the threshold is from 0.5 to 0.7. Preferably, the threshold is 0.6.

[0066] This step is divided into two cases, namely the case where the degree of repetition is greater than the threshold and the case where it is not greater than the threshold, which will be described separately below.

[0067] 1. When the degree of repetition is greater than the threshold: This case indicates that the degree of repetition between the buffer and the data in the sliding window is relatively high, meaning that by searching for the data in the sliding window, a relatively high compression of the data in the buffer can already be achieved. Therefore, there is no need to adjust the size of the sliding window, that is, the length of the sliding window remains unchanged. The subsequent step of compressing the data in the buffer using the LZ77 algorithm belongs to the prior art and will not be elaborated here.

[0068] 2. When the degree of repetition is not greater than the threshold: This case indicates that the degree of repetition between the buffer and the data in the sliding window is relatively low, meaning that it is difficult to compress the data in the buffer by searching for the data in the sliding window. Therefore, it is necessary to adjust the size of the sliding window.

[0069] Specifically, if the data in the buffer has strong periodicity, it means that the data in the buffer can achieve a large compression. However, due to the limited length of the sliding window, the degree of repetition between the buffer and the data in the sliding window is relatively low. Therefore, the length of the sliding window should be increased to expand the search range in the sliding window, facilitate finding more identical data, and thus improve the compression effect of the data in the buffer.

[0070] On the contrary, if the data in the buffer has weak periodicity, it means that the possibility of achieving a large compression of the data in the buffer is relatively small, and the length of the sliding window should be reduced to narrow the search range in the sliding window and facilitate saving search time.

[0071] The adjusted sliding window is called the corrected sliding window. Based on the corrected sliding window and the buffer, the current periodicity and the current degree of repetition are obtained again. Through continuous iteration, the size of the sliding window is continuously adjusted until the current degree of repetition obtained based on the current sliding window and the current buffer is greater than the threshold. At this time, it indicates that through adjusting the length of the sliding window, the degree of repetition between the data in the buffer and the data in the sliding window is relatively high, and a large compression can be achieved. Therefore, the data in the buffer can be compressed using the LZ77 algorithm based on the sliding window at the last iteration.

[0072] Among them, the length of the corrected sliding window is the ceiling of the product of the length of the sliding window at the previous iteration and the corresponding correction coefficient; the correction coefficient is positively correlated with the corresponding periodicity and negatively correlated with the corresponding degree of repetition.

[0073] Specifically, the length of the corrected sliding window is:

[0074] Correction coefficient is: ; In the formula, is the sliding window length in the previous iteration, c is the periodicity, d is the degree of repetition, norm ( ) is the standard normalization function, is the ceiling function.

[0075] The above formula ensures that the value range of the correction coefficient r is from 0.5 to 1.5, meeting the requirement of adjusting the size of the sliding window through the correction coefficient.

[0076] In this step, aiming at the degree of repetition of the data in the buffer being greater than the threshold, the sliding window length is adjusted by continuous iteration, so as to obtain the optimal sliding window length, making the degree of repetition of the data in the buffer and the data in the sliding window relatively high, thereby improving the compression effect of the data in the buffer.

[0077] S5. Move the current sliding window and repeat the above steps until all the data in the current buffer is compressed.

[0078] Steps S1 to S4 only achieve one-time compression of the data in the buffer. Therefore, it is necessary to move the current sliding window by the number of bits of the compressed data towards the buffer direction and repeat steps S1 to S4 until all the data in the current buffer is compressed.

[0079] In the present invention, when compressing the data in the buffer, first, according to the periodicity of the data in the buffer and the degree of repetition of the data in the buffer and the data in the sliding window, through multiple iterations, the sliding window length is adaptively adjusted with the degree of repetition being greater than the threshold as the goal, and the LZ77 algorithm is used to compress the data in the buffer based on the corrected sliding window. Compared with the fixed-length sliding window in the prior art, by adaptively adjusting the sliding window length, the present invention can capture periodically repeated data in a timely manner, and thus can achieve a better data compression effect.

[0080] As Figure 2 shown, according to the second aspect of the present invention, a user data storage system based on a cloud platform is further provided. The system includes a memory and a processor, and the processor executes the computer program stored in the memory to implement the user data storage method based on the cloud platform described in the first aspect of the present invention.

[0081] The system further includes other components well-known to those skilled in the art, such as a communication bus and a communication interface. Their settings and functions are known in the art and thus will not be elaborated herein.

[0082] In the present invention, the aforementioned memory can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory (RRAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), an enhanced dynamic random access memory (EDRAM), a high-bandwidth memory (HBM), a hybrid memory cube (HMC), etc., or any other medium that can be used to store the required information and can be accessed by an application, a module, or both. Any such computer storage medium can be part of the device or accessible or connectable to the device. Any application or module described in the present invention can be implemented by computer-readable / executable instructions stored or otherwise held by such a computer-readable medium.

[0083] In the description of this specification, "a plurality of" means at least two, such as two, three, or more, etc., unless otherwise specifically and clearly defined.

[0084] Although this specification has shown and described multiple embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will think of many changes, alterations, and alternative ways without departing from the spirit and idea of the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein can be adopted in the practice of the present invention.

Claims

1. A user data storage method based on a cloud platform, characterized in that: The steps include: Get user data; Use LZ77 algorithm to process user data and obtain buffer and sliding window; Obtaining the periodicity of the data in the buffer and the degree of repetition between the data in the buffer and the data in the sliding window; wherein the degree of repetition is the sum of the product of the frequency of occurrence of all subsequences composed of the data in the buffer in the sliding window and the weight of the corresponding subsequence; and the weight is the ratio of the number of data in the corresponding subsequence to the sum of the number of data in all subsequences; When the repetition degree is greater than the threshold, the sliding window length is kept unchanged, and the data in the buffer is compressed by the LZ77 algorithm; when the repetition degree is not greater than the threshold, the current periodicity and the current repetition degree are obtained again according to the modified sliding window and the buffer, and the iteration is continued until the current repetition degree is greater than the threshold; based on the sliding window at the last iteration, the data in the buffer is compressed by the LZ77 algorithm; wherein, the modified sliding window length is the product of the sliding window length at the last iteration and the corresponding correction coefficient rounded up; the correction coefficient is positively correlated with the current periodicity and negatively correlated with the current repetition degree; Move the current sliding window and repeat the above steps until all the data in the current buffer is compressed.

2. The method for storing user data based on a cloud platform according to claim 1, characterized in that: The periodicity for: ; In the formula, The buffer zone k The frequency of occurrence of a data in historical data, The buffer zone k The data is the first in the historical data h Second and h +1 time interval between occurrences, The buffer zone k The mean of the time intervals between all two adjacent occurrences of a data in the historical data. The buffer zone k The standard deviation of the time intervals between all two adjacent occurrences of a data in the historical data, N The total number of data in the buffer.

3. The method for storing user data based on a cloud platform according to claim 1, characterized in that: The correction factor r for: ; In the formula, c For the periodicity, d For the degree of repetition, norm ( ) is the standard normalization function.

4. The method for storing user data based on a cloud platform according to claim 1, characterized in that: The user data includes different types of data of the user.

5. The method for storing user data based on a cloud platform according to claim 1, characterized in that: The subsequence includes multiple data sets, and each data set is composed of a single data or multiple adjacent data in the buffer.

6. The method for storing user data based on a cloud platform according to claim 2, characterized in that: The duration of the historical data is one week, one month or one quarter.

7. The method for storing user data based on a cloud platform according to claim 1, characterized in that: The threshold value ranges from 0.5 to 0.

7.

8. A user data storage system based on a cloud platform, comprising a memory and a processor, characterized in that: The processor executes the computer program stored in the memory to implement the cloud platform-based user data storage method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Efficient storage method for applet data

    CN115173866A

  • LZ series compression algorithm adaptive optimization method

    CN115361026A

  • Software-based self-test and diagnosis using on-chip memory

    US20150316605A1

Cited By

  • Food material order data storage method and system

    CN121070890A