A method, system and apparatus for data storage
Patent Information
- Application Number
- CN202311715963.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-14
AI Technical Summary
相关技术中在解决上述问题时采用的是将近期新增数据作为热数据存储在读写速度较快的存储介质内,将存量历史数据作为冷数据存储在读写速度较慢的存储介质内,由于数据的冷热程度是随机动态的,采用固定的数据热度归类方法可能会造成数据读取性能不足的风险
[0051]本发明公开了一种数据存储的方法、系统及装置,涉及存储领域,包括确定客户的基本信息及在当前周期内基本信息被访问的次数;确定客户的业务数据及在当前周期内业务数据被访问的次数,业务数据被访问的次数与基本信息被访问的次数相关;由于基本信息和业务数据被访问的次数相关,所以在计算权重热度考虑到基本信息和业务数据被访问的次数,计算得到当前周期的权重热度后,根据当前周期的权重热度和上一周期的权重热度对下一周期的权重热度进行预测,最后根据下一周期的权重热度将基本信息保存到下一周期的权重热度对应的存储位置。综合考虑到数据被访问之间的关系,同时提前预测下一周期的权重热度进而调整数据的存储位置,提高了读取性能。
Smart Images

Figure CN117708248B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of storage, and in particular to a method, system and apparatus for data storage. Background Technology
[0002] The banking system stores a large amount of business data. This data is characterized by its large volume, interrelationships between different types of business data, and significant differences in the frequency of data retrieval. As the data volume increases, query efficiency decreases when querying individual business data. Furthermore, due to the interrelationships between different types of business data, the frequency of retrieval for the specific business data typically increases or decreases as the frequency of retrieval for related business data increases or decreases. Current technologies address these issues by storing recently added data as "hot" data in high-speed storage media and existing historical data as "cold" data in slower-speed storage media. However, since the hotness or coldness of data is random and dynamic, using a fixed data hotness classification method may lead to insufficient data retrieval performance. Summary of the Invention
[0003] The purpose of this invention is to provide a data storage method, system, and apparatus that comprehensively considers the relationships between data accesses and predicts the weight of the next cycle in advance to adjust the data storage location, thereby improving read performance.
[0004] To address the aforementioned technical problems, the present invention provides a data storage method, comprising:
[0005] Determine the customer's basic information and the number of times the basic information is accessed within the current period. The basic information includes one or more combinations of customer number, customer type, customer name, ID number, ID type, and customer status.
[0006] The business data of the customer and the number of times the business data is accessed in the current period are determined. The business data includes one or more combinations of customer capital information, family member information, financial information, risk rating information, credit information, collateral information, loan information, contract information and permission information. The number of times the business data is accessed is related to the number of times the basic information is accessed.
[0007] The popularity weight of the basic information in the current period is determined based on the basic information, the number of times the basic information is accessed in the current period, the business data, and the number of times the business data is accessed in the current period. The popularity weight is positively correlated with the number of times the basic information is accessed.
[0008] Based on the weighted popularity of the basic information in the current period and the popularity weight of the basic information in the previous period, the weighted popularity of the basic information in the next period is predicted.
[0009] The basic information is saved to the storage location corresponding to the weighted heat of the next period based on the weighted heat of the next period.
[0010] On the other hand, determining the customer's basic information and the number of times that basic information is accessed within the current period includes:
[0011] Determine the basic information set A = (A1, A2, A3, ..., A...) consisting of the basic information of n customers. i ,...,A n ), 1≤i≤n, A i This contains the basic information of the i-th user;
[0012] The number of times the basic information is accessed within the current period is determined to be a1, a2, a3, ..., a i ,...,a n a i For A i The number of times it has been accessed within the current period;
[0013] Determining the customer's business data and the number of times the business data was accessed within the current period includes:
[0014] Determine the set H1, H2, H3..., H of m types of business data for each customer. j ,...,H m , 1≤j≤m, H j It is a set of the j-th type of business data for n customers;
[0015] Determine the number of times h that the business data is accessed within the current period. 1i ,h 2i ,h 3i ,...,h ji ,...,h mi h ji This represents the total number of times the j-th type of business data for the i-th user is accessed.
[0016] On the other hand, the popularity weight of the basic information in the current period is determined based on the basic information, the number of times the basic information is accessed in the current period, the business data, and the number of times the business data is accessed in the current period, including:
[0017] Based on the basic information, the number of times the basic information was accessed in the current period, the business data, the number of times the business data was accessed in the current period, and the relationship between the popularity weight formula. Determine the popularity weight of basic information in the current period;
[0018] Among them, Q i Let be the integer value of the popularity weight of the basic information of the i-th user in the current period. Main influencing factor to For related influencing factors, to The sum of is 1.
[0019] On the other hand, after determining the popularity weight of the basic information in the current period based on the basic information, the number of times the basic information is accessed in the current period, the business data, and the number of times the business data is accessed in the current period, the method further includes:
[0020] The data statistics are performed using the popularity weight as the horizontal axis and the probability density as the vertical axis. The probability density is the ratio of the number of basic information corresponding to each popularity weight to the total number of basic information.
[0021] If the statistical results of the data are left-skewed, then it is determined that there are more hot data than cold data in the basic information set.
[0022] If the statistical results of the data are right-skewed, then it is determined that there is more cold data than hot data in the basic information set.
[0023] On the other hand, after determining the popularity weight of the basic information in the current period based on the basic information, the number of times the basic information is accessed in the current period, the business data, and the number of times the business data is accessed in the current period, the method further includes:
[0024] Based on the relationship between heat weight and distribution skewness value Determine the skewness value Skew(Q) of the heat weight distribution, which represents the distribution of the rounded values of the heat weights of the basic information of all users;
[0025] in, Q is the mathematical expectation of the popularity weights of the n users in the current period. i p(Q) is the integer value of the popularity weight of the basic information of the i-th user in the current period. i ) for Q i The probability of occurrence The standard deviation of the popularity weights of the n users in the current period.
[0026] If Skew(Q) is less than 0, the distribution of the rounded values of the popularity weights of all users' basic information is left-skewed; if Skew(Q) is equal to 0, the distribution of the rounded values of the popularity weights of all users' basic information is normal; if Skew(Q) is greater than 0, the distribution of the rounded values of the popularity weights of all users' basic information is right-skewed.
[0027] Determine the popularity value of the basic information in the current period based on the distribution;
[0028] Based on the weighted popularity of the basic information in the current period and the popularity weight of the basic information in the previous period, the weighted popularity of the basic information in the next period is predicted, including:
[0029] Based on the popularity value of the basic information in the current period and the popularity value of the basic information in the previous period, the popularity value of the basic information in the next period is predicted.
[0030] On the other hand, determining the popularity value of the basic information in the current period based on the distribution includes:
[0031] Three popularity values, R1, R2, and R3, are set, with R1 having the highest popularity value and R3 having the lowest. The popularity value is positively correlated with the number of visits.
[0032] If the distribution of the integer values of the popularity weights of all users' basic information is left-skewed, then according to the first mapping relationship... Determine the popularity value of the aforementioned basic information in the current period;
[0033] If the distribution of the integer values of the popularity weights of all users' basic information is right-skewed, then according to the second mapping relationship... Determine the popularity value of the aforementioned basic information in the current period;
[0034] If the distribution of the integer values of the popularity weights of all users' basic information follows a normal distribution, then according to the third mapping relationship... Determine the popularity value of the aforementioned basic information in the current period;
[0035] Among them, Q i Let be the integer value of the popularity weight of the i-th user's basic information in the current period, M be the value that appears most frequently among the popularity weights of the n users, Qmax be the largest value among the popularity weights of the n users, Qmin be the smallest value among the popularity weights of the n users, and σ be the standard deviation of the popularity weights of the n users in the current period.
[0036] On the other hand, predicting the popularity value of the basic information in the next period based on the popularity value of the basic information in the current period and the popularity value of the basic information in the previous period includes:
[0037] According to the prediction relation r i(t+1) =R i(t) +b i(t) Predict the popularity value of the aforementioned basic information in the next period;
[0038] Where, r i(t+1) R is the predicted popularity value of the basic information of the i-th user in period t+1. i(t) Let b be the actual popularity value of the basic information of the i-th user in period t. i(t) Let b be the predicted change value of the basic information of the i-th user in period t. i(t) =β[α(R i(t) -R i(t-1) )+(1-α)b i(t-1) ]+(1-β)b i(t-1) b i(t-1) Let α be the predicted change value of the basic information of the i-th user in period t-1, where α is the discount coefficient (0 < α < 1) and β is the learning efficiency (0 < β < 1).
[0039] On the other hand, saving the basic information to the storage location corresponding to the weighted heat of the next period according to the weighted heat of the next period includes:
[0040] Determine the predicted popularity value of the basic information in the next period as R1, R2, or R3;
[0041] The basic information of the predicted popularity value R1 for the next period is saved to memory, the basic information of the predicted popularity value R2 for the next period is saved to solid-state drive, and the basic information of the predicted popularity value R3 for the next period is saved to hard disk drive.
[0042] To address the aforementioned technical problems, the present invention also provides a data storage system, comprising:
[0043] The basic information access count determination unit is used to determine the customer's basic information and the number of times the basic information is accessed in the current period. The basic information includes one or more combinations of customer number, customer type, customer name, ID number, ID type, and customer status.
[0044] The business data access frequency determination unit is used to determine the customer's business data and the number of times the business data is accessed in the current period. The business data includes one or more combinations of customer capital information, family member information, financial information, risk rating information, credit information, collateral information, loan information, contract information, and permission information. The number of times the business data is accessed is related to the number of times the basic information is accessed.
[0045] A popularity weight determination unit is used to determine the popularity weight of the basic information in the current period based on the basic information, the number of times the basic information is accessed in the current period, the business data, and the number of times the business data is accessed in the current period. The popularity weight is positively correlated with the number of times the basic information is accessed.
[0046] The prediction unit is used to predict the weight of the basic information in the next period based on the weight of the basic information in the current period and the weight of the basic information in the previous period.
[0047] A storage unit is used to save the basic information to the storage location corresponding to the weight heat of the next period according to the weight heat of the next period.
[0048] To address the aforementioned technical problems, the present invention also provides a data storage device, comprising:
[0049] Memory, used to store computer programs;
[0050] A processor, used to implement the above-described method for data storage when executing the computer program.
[0051] This invention discloses a data storage method, system, and apparatus, relating to the storage field. The method includes determining a customer's basic information and the number of times this basic information is accessed within the current period; determining the customer's business data and the number of times this business data is accessed within the current period, wherein the number of times business data is accessed is related to the number of times basic information is accessed; since the access frequency of basic information and business data is related, the weighted popularity is calculated by considering both access frequencies. After calculating the weighted popularity for the current period, the weighted popularity for the next period is predicted based on the weighted popularity of the current period and the weighted popularity of the previous period. Finally, the basic information is saved to the storage location corresponding to the weighted popularity of the next period. By comprehensively considering the relationship between data access frequencies and predicting the weighted popularity of the next period in advance to adjust the data storage location, read performance is improved. Attached Figure Description
[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the prior art and embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 A flowchart of a data storage method provided by the present invention;
[0054] Figure 2 A schematic diagram of a heat weighted normal distribution provided by the present invention;
[0055] Figure 3 A schematic diagram of a right-skewed distribution of heat weights provided by the present invention;
[0056] Figure 4 A schematic diagram of a left-skewed distribution of heat weights provided by the present invention;
[0057] Figure 5 A schematic diagram of the structure of a data storage system provided by the present invention;
[0058] Figure 6 A schematic diagram of a data storage device provided by the present invention. Detailed Implementation
[0059] The core of this invention is to provide a data storage method, system, and apparatus that comprehensively considers the relationships between data accesses and predicts the weight and popularity of the next cycle in advance to adjust the data storage location, thereby improving read performance.
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0061] Figure 1 A flowchart of a data storage method provided by the present invention, the data storage method comprising:
[0062] S11: Determine the customer's basic information and the number of times the basic information is accessed in the current period. The basic information includes one or more combinations of customer number, customer type, customer name, ID number, ID type, and customer status.
[0063] S12: Determine the customer's business data and the number of times the business data is accessed in the current period. The business data includes one or more combinations of customer capital information, family member information, financial information, risk rating information, credit information, collateral information, loan information, contract information, and authorization information. The number of times the business data is accessed is related to the number of times the basic information is accessed.
[0064] Within a banking system, the number of customers is typically at least in the millions. Therefore, the set of basic information contains at least millions of elements, each representing a customer's basic information. Simultaneously, the number of times each element in the set of basic information is accessed is collected and recorded at regular intervals.
[0065] Typically, a single business data object will have multiple associated business data objects. For example, in a banking system, basic customer information will be associated with customer capital information, family member information, financial information, risk rating information, credit information, collateral information, loan information, contract information, and authorization information. Each associated business data object is itself a collection, and any element within any of these collections must have a business relationship with an element in collection A. For instance, if collection A contains basic customer information and includes a customer named 'c', then the associated customer capital information collection within collection A must contain one or more elements representing customer 'c's' capital information data, typically linked by a customer ID or identification number.
[0066] S13: Determine the popularity weight of basic information in the current period based on basic information, the number of times basic information is accessed in the current period, business data, and the number of times business data is accessed in the current period. The popularity weight is positively correlated with the number of times basic information is accessed.
[0067] Because there are relationships between various types of business data, the access frequency of basic information usually increases or decreases when the access frequency of related business data increases or decreases. Therefore, when calculating the popularity weight, it is necessary to consider not only the number of times basic information is accessed, but also the number of times business data is accessed. Only by comprehensively considering the relationships can the popularity weight be more accurate.
[0068] S14: Based on the weight and popularity of basic information in the current period and the popularity weight of basic information in the previous period, predict the weight and popularity of basic information in the next period.
[0069] Data access may be related across different periods. Based on the weighted popularity of the previous period and the weighted popularity of the current period, the weighted popularity of the next period can be predicted.
[0070] Specifically, if the weight and popularity of basic information are both high in the previous and current periods, it indicates that the basic information is frequently accessed, so the weight of the basic information in predicting the next period should also be high. Conversely, if the weight and popularity of basic information are both low in the previous and current periods, it indicates that the basic information is not frequently accessed, so the weight of the basic information in predicting the next period should also be low.
[0071] S15: Save the basic information to the storage location corresponding to the weight heat of the next period according to the weight heat of the next period.
[0072] If the data has a low popularity weight in the next period, it means the data is less likely to be accessed in that period. Therefore, the data can be stored in a storage location with lower read / write speed to avoid occupying storage space. Conversely, if the data has a high popularity weight in the next period, it means the data is more likely to be accessed in that period. Therefore, the data can be stored in a storage location with higher read / write speed to improve access speed. The read / write speed of the storage location is positively correlated with the popularity weight in the next period.
[0073] This invention discloses a data storage method, system, and apparatus, relating to the storage field. The method includes determining a customer's basic information and the number of times this basic information is accessed within the current period; determining the customer's business data and the number of times this business data is accessed within the current period, wherein the number of times business data is accessed is related to the number of times basic information is accessed; since the access frequency of basic information and business data is related, the weighted popularity is calculated by considering both access frequencies. After calculating the weighted popularity for the current period, the weighted popularity for the next period is predicted based on the weighted popularity of the current period and the weighted popularity of the previous period. Finally, the basic information is saved to the storage location corresponding to the weighted popularity of the next period. By comprehensively considering the relationship between data access frequencies and predicting the weighted popularity of the next period in advance to adjust the data storage location, read performance is improved.
[0074] Based on the above embodiments:
[0075] In some embodiments, determining a customer's basic information and the number of times that basic information is accessed during the current period includes:
[0076] Determine the basic information set A = (A1, A2, A3, ..., A4) consisting of the basic information of n customers. i ,...,A n ), 1≤i≤n, A i This contains the basic information of the i-th user;
[0077] The number of times the basic information is accessed within the current period are determined to be a1, a2, a3, ..., ai ,...,a n a i For A i The number of times it has been accessed within the current period;
[0078] Determine the customer's business data and the number of times the business data was accessed during the current period, including:
[0079] Determine the set H1, H2, H3..., H4 for each customer's m types of business data. j ,...,H m , 1≤j≤m, H j The set of j-th type of business data for n customers;
[0080] Determine the number of times h that business data is accessed within the current period. 1i ,h 2i ,h 3i ,...,h ji ,...,h mi h ji This represents the total number of times the j-th type of business data for the i-th user is accessed.
[0081] Let n be the number of elements in the basic information set A, and let A = (A1, A2, A3, ..., A4) be the number of elements in the basic information set A. i ,...,A n That is, where 1≤i≤n. For example, in a banking system, basic customer information can be defined as a set of objects A. The number of customers is usually at least in the millions. In other words, set A contains at least millions of elements, and each element is a table of basic customer information.
[0082] Meanwhile, the system is specified to collect and record the number of accesses to existing elements in the set at regular intervals, which can be represented as a1, a2, a3, ..., a i ,...,a n The unit time for the data collection and recording interval is defined as t.
[0083] Therefore, the set of associated objects of the business data object set A can be defined as m, which can be represented as H1, H2, H3..., H... j ,...,H m Where 1 ≤ j ≤ m. Since each set of associated objects has one or more elements that are associated with an element a in set A. i Since there is a relationship between them, the number of times each associated element in each set of associated objects is accessed per unit time can be represented as h. 1i ,h 2i ,h 3i ,...,h ji,...,h mi If the set of associated objects H j There are multiple elements related to a i If a relationship exists, then h ji This is the sum of the number of visits to multiple elements. For example, if user i has two contracts in the contract information, h... ji It is the sum of the number of visits for the two contracts.
[0084] In some embodiments, the popularity weight of basic information in the current period is determined based on basic information, the number of times the basic information is accessed in the current period, business data, and the number of times the business data is accessed in the current period, including:
[0085] Based on the relationship between basic information, the number of times basic information is accessed in the current period, business data, the number of times business data is accessed in the current period, and the popularity weight formula. Determine the popularity weight of basic information in the current period;
[0086] Among them, Q i Let be the integer value of the popularity weight of the basic information of the i-th user in the current period. Main influencing factor to For related influencing factors, to The sum of is 1.
[0087] To characterize the hot / cold status of specific elements in set A, and the influence of related object elements on the hot / cold status of that element, an element a in set A is defined. i Popularity weight Q i , Where 1≤i≤n, The main influencing factor is typically configured in the range of 0.5 to 0.6. to As a related influencing factor, and to The sum of all influencing factors is 1. The influencing factors can be reasonably configured and adjusted according to specific business circumstances and the degree of influence of related relationships. It should be noted that Q after rounding is... i Always an integer.
[0088] By examining the popularity weight formula, it's easy to see that the popularity weight of each element in set A is affected not only by the number of times the element itself is accessed, but also by the number of times related elements are accessed. For example, in a banking system, if a customer's various ancillary business data, such as capital, family members, and finances, are queried multiple times in the business process, then the customer's basic information itself may also be queried multiple times.
[0089] In some embodiments, after determining the popularity weight of basic information in the current period based on basic information, the number of times basic information is accessed in the current period, business data, and the number of times business data is accessed in the current period, the method further includes:
[0090] Data statistics are performed using popularity weight as the horizontal axis and probability density as the vertical axis. Probability density is the ratio of the number of basic information items corresponding to each popularity weight to the total number of basic information items.
[0091] If the statistical results show a left-skewed distribution, then it is determined that there is more hot data than cold data in the basic information set.
[0092] If the statistical results show a right-skewed distribution, then it is determined that there is more cold data than hot data in the basic information set.
[0093] Since the elements in set A are independent of each other, the sets of related objects are independent of each other, the elements within each set of related objects are independent of each other, and the unit time interval t is fixed, all elements have the same probability of being queried. Therefore, the calculated popularity weight Q is... i It belongs to a discrete random distribution.
[0094] Furthermore, the popularity weight Q will be adjusted. i -Q n The statistical data is distributed, with the horizontal axis representing the popularity weight and the vertical axis representing the probability density. Since the value of n is relatively large, the statistical distribution can be approximated as one of a normal distribution, a right-skewed distribution, or a left-skewed distribution. Qualitative analysis shows that when the statistical distribution is left-skewed, most elements in set A have high popularity weights, representing mostly "hot" data, while a small portion represents "cold" data, and their popularity weights are much lower than the average. When the statistical distribution is right-skewed, most elements in set A have low popularity weights, representing mostly "cold" data, while a small portion represents "hot" data, and their popularity weights are much higher than the average. When the statistical distribution is normal, the popularity weights of most elements in set A are between Q... max and Q min In the middle, where Q max and Q min These represent the maximum and minimum values of the popularity weight, respectively, and only a small number of elements have maximum and minimum popularity weights.
[0095] In some embodiments, after determining the popularity weight of basic information in the current period based on basic information, the number of times basic information is accessed in the current period, business data, and the number of times business data is accessed in the current period, the method further includes:
[0096] Based on the relationship between heat weight and distribution skewness value Determine the skewness value Skew(Q) of the popularity weight distribution. The skewness value of the popularity weight distribution represents the distribution of the rounded values of the popularity weights of the basic information of all users.
[0097] in, Q is the mathematical expectation of the popularity weights of n users in the current period. i p(Q) is the integer value of the popularity weight of the basic information of the i-th user in the current period. i ) for Q i The probability of occurrence Let the standard deviation of the popularity weights of n users in the current period be denoted as .
[0098] If Skew(Q) is less than 0, the distribution of the rounded values of the popularity weights of all users' basic information is left-skewed; if Skew(Q) is equal to 0, the distribution of the rounded values of the popularity weights of all users' basic information is normal; if Skew(Q) is greater than 0, the distribution of the rounded values of the popularity weights of all users' basic information is right-skewed.
[0099] Determine the popularity value of basic information in the current period based on the distribution;
[0100] Based on the weighted popularity of basic information in the current period and the popularity weight of basic information in the previous period, the weighted popularity of basic information in the next period is predicted, including:
[0101] Based on the popularity value of basic information in the current period and the popularity value of basic information in the previous period, the popularity value of basic information in the next period is predicted.
[0102] Figure 2 A schematic diagram of a heat weighted normal distribution provided by the present invention; Figure 3 A schematic diagram of a right-skewed distribution of heat weights provided by the present invention; Figure 4 This is a schematic diagram of a left-skewed distribution of heat weights provided by the present invention.
[0103] To further quantitatively analyze and determine which of the three distributions the heat weight distribution of elements in A specifically belongs to, the skewness value of the heat weight distribution of all elements in A is calculated.
[0104] In some embodiments, determining the popularity value of basic information in the current period based on distribution includes:
[0105] Three popularity values, R1, R2, and R3, are set, with R1 having the highest popularity value and R3 having the lowest. The popularity value is positively correlated with the number of visits.
[0106] If the distribution of the integer values of the popularity weights of all users' basic information is left-skewed, then according to the first mapping relationship... Determine the popularity value of basic information in the current period;
[0107] If the distribution of the integer values of the popularity weights of all users' basic information is right-skewed, then according to the second mapping relationship... Determine the popularity value of basic information in the current period;
[0108] If the distribution of the integer values of the popularity weights of all users' basic information follows a normal distribution, then according to the third mapping relationship... Determine the popularity value of basic information in the current period;
[0109] Among them, Q i Let be the integer value of the popularity weight of the i-th user's basic information in the current period, M be the value that appears most frequently among the popularity weights of the n users, Qmax be the largest value among the popularity weights of the n users, Qmin be the smallest value among the popularity weights of the n users, and σ be the standard deviation of the popularity weights of the n users in the current period.
[0110] The popularity weights of each element in set A are categorized, specifically defined as three types of data: hot data, warm data, and cold data, corresponding to three popularity values R1, R2, and R3, respectively. Different popularity value mapping methods are used based on the different statistical distribution results of the popularity weights. The popularity of each element in set A is categorized according to the popularity weight distribution and popularity values. i ∈R, where Ri is the popularity value of the i-th basic information, and R={R1,R2,R3}.
[0111] In some embodiments, predicting the popularity value of basic information in the next period based on the popularity value of basic information in the current period and the popularity value of basic information in the previous period includes:
[0112] According to the prediction relation r i(t+1) =R i(t) +b i(t) Predict the popularity of basic information in the next period;
[0113] Where, r i(t+1) R is the predicted popularity value of the basic information of the i-th user in period t+1. i(t) Let b be the actual popularity value of the basic information of the i-th user in period t. i(t) Let b be the predicted change value of the basic information of the i-th user in period t. i(t) =β[α(R i(t) -R i(t-1) )+(1-α)b i(t-1)]+(1-β)b i(t-1) b i(t-1) Let α be the predicted change value of the basic information of the i-th user in period t-1, where α is the discount coefficient (0 < α < 1) and β is the learning efficiency (0 < β < 1).
[0114] Since the popularity values of each element in set A are calculated based on information such as the number of queries for each element in period t, they only reflect the popularity of each element in the current period t. However, the popularity of each element is dynamic. Therefore, directly adjusting the storage size of the storage medium and the storage location of business data objects based on the popularity value of period t will inevitably lead to problems such as insufficient data read performance and wasted storage performance. Therefore, it is necessary to predict the popularity value for the next period t+1 based on the current and historical popularity values, and then adjust the storage size and location at the end of period t based on the predicted popularity value for period t+1.
[0115] Where β is the learning efficiency, and 0 < β < 1, and α is the discount factor, and 0 < α < 1, with a suggested value of 0.6 to 0.85. It is easy to see that the predicted change in the current heat value over period t is correlated with the predicted change in period t-1 and the actual heat value over period t-1. Furthermore, changes in the actual heat value (for example, the heat value of element i continuously increases over multiple collection periods, but decreases in the current period) can also adjust the predicted change in the heat value in a timely manner, thus better and more promptly responding to changes in the element's heat value.
[0116] In some embodiments, basic information is saved to the storage location corresponding to the weighted popularity of the next period based on the weighted popularity of the next period, including:
[0117] Determine the predicted popularity value of basic information in the next period as R1, R2, or R3;
[0118] Save the basic information of the predicted popularity value R1 for the next period to memory, save the basic information of the predicted popularity value R2 for the next period to solid-state drive, and save the basic information of the predicted popularity value R3 for the next period to hard disk drive.
[0119] Based on the predicted heat values for each element in period t+1, the data volume of hot, warm, and cold data for period t+1 is calculated. It should be noted that the system typically stores hot data in memory for fast data read and write, and this memory can be dynamically expanded. The system memory size is adjusted accordingly based on the volume of the statistically analyzed hot data.
[0120] Furthermore, considering that warm data and cold data are stored in solid-state drives and hard disk drives respectively, since these two storage media cannot be dynamically expanded in a timely manner, but if the data query collection process has a long unit time (e.g., in "days" or "weeks"), this application can predict the changes in the amount of warm data and cold data over a long time dimension, and periodically expand or compress the storage size.
[0121] The storage location of elements with a predicted popularity value of R1 is moved to memory, the storage location of elements with a predicted popularity value of R2 is moved to the solid-state drive (SSD), and the storage location of elements with a predicted popularity value of R3 is moved to the hard disk drive (HDD). It's understandable that memory has a higher read / write speed than the SSD, and the SSD has a higher read / write speed than the HDD.
[0122] Figure 5 This invention provides a schematic diagram of a data storage system, which includes:
[0123] The basic information access count determination unit 51 is used to determine the customer's basic information and the number of times the basic information is accessed in the current period. The basic information includes one or more combinations of customer number, customer type, customer name, ID number, ID type and customer status.
[0124] The business data access frequency determination unit 52 is used to determine the customer's business data and the number of times the business data is accessed in the current period. The business data includes one or more combinations of customer capital information, family member information, financial information, risk rating information, credit information, collateral information, loan information, contract information and permission information. The number of times the business data is accessed is related to the number of times the basic information is accessed.
[0125] The popularity weight determination unit 53 is used to determine the popularity weight of basic information in the current period based on basic information, the number of times basic information is accessed in the current period, business data, and the number of times business data is accessed in the current period. The popularity weight is positively correlated with the number of times basic information is accessed.
[0126] Prediction unit 54 is used to predict the weight of basic information in the next period based on the weight of basic information in the current period and the weight of basic information in the previous period.
[0127] The storage unit 55 is used to save basic information to the storage location corresponding to the weight heat of the next period according to the weight heat of the next period.
[0128] Based on the above embodiments:
[0129] Basic information access count determination unit 51 is specifically used to determine the basic information set A = (A1, A2, A3..., A...) consisting of basic information of n customers. i ,...,A n ), 1≤i≤n, A i This contains the basic information of the i-th user;
[0130] The number of times the basic information is accessed within the current period are determined to be a1, a2, a3, ..., a i ,...,a n a i For A i The number of times it has been accessed within the current period;
[0131] Business data access frequency determination unit 52 is specifically used to determine the set H1, H2, H3..., H of m types of business data for each customer. j ,...,H m , 1≤j≤m, H j The set of j-th type of business data for n customers;
[0132] Determine the number of times h that business data is accessed within the current period. 1i ,h 2i ,h 3i ,...,h ji ,...,h mi h ji This represents the total number of times the j-th type of business data for the i-th user is accessed.
[0133] The popularity weight determination unit 53 is specifically used to determine popularity weight based on basic information, the number of times basic information is accessed in the current period, business data, the number of times business data is accessed in the current period, and the popularity weight relationship formula. Determine the popularity weight of basic information in the current period;
[0134] Among them, Q i Let be the integer value of the popularity weight of the basic information of the i-th user in the current period. Main influencing factor to For related influencing factors, to The sum of is 1.
[0135] The data statistics unit is used to perform data statistics with the popularity weight as the horizontal axis and the probability density as the vertical axis. The probability density is the ratio of the number of basic information corresponding to each popularity weight to the total number of basic information.
[0136] If the statistical results show a left-skewed distribution, then it is determined that there is more hot data than cold data in the basic information set.
[0137] If the statistical results show a right-skewed distribution, then it is determined that there is more cold data than hot data in the basic information set.
[0138] The heat weight distribution skewness value determination unit is used to determine the heat weight and distribution skewness value according to the relationship formula. Determine the skewness value Skew(Q) of the popularity weight distribution. The skewness value of the popularity weight distribution represents the distribution of the rounded values of the popularity weights of the basic information of all users.
[0139] in, Q is the mathematical expectation of the popularity weights of n users in the current period. i p(Q) is the integer value of the popularity weight of the basic information of the i-th user in the current period. i ) for Q i The probability of occurrence Let the standard deviation of the popularity weights of n users in the current period be denoted as .
[0140] The distribution determination unit is used to determine the distribution of the rounded values of the popularity weights of the basic information of all users as follows: if Skew(Q) is less than 0, the distribution is left-skewed; if Skew(Q) is equal to 0, the distribution is normal; if Skew(Q) is greater than 0, the distribution is right-skewed.
[0141] The heat value determination unit is used to determine the heat value of basic information in the current period based on the distribution.
[0142] The prediction unit 54 is specifically used to predict the popularity value of basic information in the next period based on the popularity value of basic information in the current period and the popularity value of basic information in the previous period.
[0143] The popularity value setting unit is used to set three popularity values R1, R2 and R3. R1 is the highest popularity value and R3 is the lowest popularity value. The popularity value is positively correlated with the number of times it is accessed.
[0144] The popularity value determination unit is specifically used to determine the popularity value if the distribution of the rounded values of the popularity weights of all users' basic information is left-skewed, and then, according to the first mapping relationship... Determine the popularity value of basic information in the current period;
[0145] If the distribution of the integer values of the popularity weights of all users' basic information is right-skewed, then according to the second mapping relationship... Determine the popularity value of basic information in the current period;
[0146] If the distribution of the integer values of the popularity weights of all users' basic information follows a normal distribution, then according to the third mapping relationship... Determine the popularity value of basic information in the current period;
[0147] Among them, Q i Let be the integer value of the popularity weight of the i-th user's basic information in the current period, M be the value that appears most frequently among the popularity weights of the n users, Qmax be the largest value among the popularity weights of the n users, Qmin be the smallest value among the popularity weights of the n users, and σ be the standard deviation of the popularity weights of the n users in the current period.
[0148] Prediction unit 54, specifically based on the prediction relation r i(t+1) =R i(t) +b i(t) Predict the popularity of basic information in the next period;
[0149] Where, r i(t+1) R is the predicted popularity value of the basic information of the i-th user in period t+1. i(t) Let b be the actual popularity value of the basic information of the i-th user in period t. i(t) Let b be the predicted change value of the basic information of the i-th user in period t. i(t) =β[α(R i(t) -R i(t-1) )+(1-α)b i(t-1) ]+(1-β)b i(t-1) b i(t-1) Let α be the predicted change value of the basic information of the i-th user in period t-1, where α is the discount coefficient (0 < α < 1) and β is the learning efficiency (0 < β < 1).
[0150] The predicted heat value determination unit is used to determine the predicted heat value of basic information in the next period as R1, R2 or R3;
[0151] The storage unit 55 is specifically used to save the basic information of the predicted heat value R1 for the next period to memory, the basic information of the predicted heat value R2 for the next period to solid-state drive, and the basic information of the predicted heat value R3 for the next period to mechanical hard drive.
[0152] Figure 6 A schematic diagram of a data storage device provided by the present invention, the data storage device comprising:
[0153] Memory 61 is used to store computer programs;
[0154] Processor 62 is used to implement the above-described method for data storage when executing a computer program.
[0155] The description of the data storage device provided in this application is given in the above embodiments and will not be repeated here.
[0156] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0157] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0158] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for data storage, characterized in that, include: Determine the customer's basic information and the number of times the basic information is accessed within the current period. The basic information includes one or more combinations of customer number, customer type, customer name, ID number, ID type, and customer status. The business data of the customer and the number of times the business data is accessed in the current period are determined. The business data includes one or more combinations of customer capital information, family member information, financial information, risk rating information, credit information, collateral information, loan information, contract information and permission information. The number of times the business data is accessed is related to the number of times the basic information is accessed. The popularity weight of the basic information in the current period is determined based on the basic information, the number of times the basic information is accessed in the current period, the business data, and the number of times the business data is accessed in the current period. The popularity weight is positively correlated with the number of times the basic information is accessed. Based on the popularity weight of the basic information in the current period and the popularity weight of the basic information in the previous period, the popularity weight of the basic information in the next period is predicted. The basic information is saved to the storage location corresponding to the heat weight of the next period according to the heat weight of the next period; After determining the popularity weight of the basic information in the current period based on the basic information, the number of times the basic information is accessed in the current period, the business data, and the number of times the business data is accessed in the current period, the method further includes: Based on the relationship between heat weight and distribution skewness value Determine the skewness value of the heat weight distribution The skewness value of the heat weight distribution represents the distribution of the rounded values of the heat weights of the basic information of all users; in, Let n be the mathematical expectation of the popularity weights of the n users in the current period. Let be the integer value of the popularity weight of the basic information of the i-th user in the current period. for The probability of occurrence The standard deviation of the popularity weights of the n users in the current period. ; like If the value is less than 0, then the distribution of the integer values of the popularity weights of all users' basic information is left-skewed. If the value equals 0, then the distribution of the integer values of the popularity weights of all users' basic information follows a normal distribution. If the value is greater than 0, then the distribution of the rounded values of the popularity weights of all users' basic information is determined to be a right-skewed distribution. Determine the popularity value of the basic information in the current period based on the distribution; Based on the popularity weight of the basic information in the current period and the popularity weight of the basic information in the previous period, the popularity weight of the basic information in the next period is predicted, including: Based on the popularity value of the basic information in the current period and the popularity value of the basic information in the previous period, the popularity value of the basic information in the next period is predicted; Determining the popularity value of the basic information in the current period based on the distribution includes: Three popularity values, R1, R2, and R3, are set, with R1 having the highest popularity value and R3 having the lowest. The popularity value is positively correlated with the number of visits. If the distribution of the integer values of the popularity weights of all users' basic information is left-skewed, then according to the first mapping relationship... Determine the popularity value of the aforementioned basic information in the current period; If the distribution of the integer values of the popularity weights of all users' basic information is right-skewed, then according to the second mapping relationship... Determine the popularity value of the aforementioned basic information in the current period; If the distribution of the integer values of the popularity weights of all users' basic information follows a normal distribution, then according to the third mapping relationship... Determine the popularity value of the aforementioned basic information in the current period; in, Let M be the integer value of the popularity weight of the i-th user's basic information in the current period, and M be the value that appears most frequently among the popularity weights of the n users. The largest value among the popularity weights of the n users. The minimum value among the popularity weights of the n users. The standard deviation of the popularity weights of the n users in the current period.
2. The data storage method as described in claim 1, characterized in that, Determine the customer's basic information and the number of times that basic information has been accessed during the current period, including: Determine the basic information set consisting of the basic information of n customers. , , This contains the basic information of the i-th user; The number of times the basic information was accessed within the current period was determined as follows: , for The number of times it has been accessed within the current period; Determining the customer's business data and the number of times the business data was accessed within the current period includes: Determine the set of m types of business data for each customer. , , It is a set of the j-th type of business data for n customers; Determine the number of times the business data is accessed within the current period. , This represents the total number of times the j-th type of business data for the i-th user is accessed.
3. The data storage method as described in claim 2, characterized in that, The popularity weight of the basic information in the current period is determined based on the basic information, the number of times the basic information is accessed in the current period, the business data, and the number of times the business data is accessed in the current period, including: Based on the basic information, the number of times the basic information was accessed in the current period, the business data, the number of times the business data was accessed in the current period, and the relationship between the popularity weight formula. Determine the popularity weight of basic information in the current period; in, Let be the integer value of the popularity weight of the basic information of the i-th user in the current period. Main influencing factor , to For related influencing factors, , to The sum of is 1.
4. The data storage method as described in claim 3, characterized in that, After determining the popularity weight of the basic information in the current period based on the basic information, the number of times the basic information is accessed in the current period, the business data, and the number of times the business data is accessed in the current period, the method further includes: The data statistics are performed using the heat weight as the horizontal axis and the probability density as the vertical axis. The probability density is the ratio of the number of basic information corresponding to each heat weight to the total number of basic information. If the statistical results of the data are left-skewed, then it is determined that there are more hot data than cold data in the basic information set. If the statistical results of the data are right-skewed, then it is determined that there is more cold data than hot data in the basic information set.
5. The data storage method as described in claim 1, characterized in that, Based on the popularity value of the basic information in the current period and the popularity value of the basic information in the previous period, predict the popularity value of the basic information in the next period, including: According to the predictive relation Predict the popularity value of the aforementioned basic information in the next period; in, Let the basic information of the i-th user be the predicted popularity value in period t+1. Let be the actual popularity value of the basic information of the i-th user in period t. Let be the predicted change value of the basic information of the i-th user in period t. , Let be the predicted change value of the basic information of the i-th user in period t-1. This is the discount factor. , For learning efficiency, .
6. The data storage method as described in claim 5, characterized in that, The basic information is saved to the storage location corresponding to the heat weight of the next period according to the heat weight of the next period, including: Determine the predicted popularity value of the basic information in the next period as R1, R2, or R3; The basic information of the predicted popularity value R1 for the next period is saved to memory, the basic information of the predicted popularity value R2 for the next period is saved to solid-state drive, and the basic information of the predicted popularity value R3 for the next period is saved to hard disk drive.
7. A data storage system, characterized in that, include: The basic information access count determination unit is used to determine the customer's basic information and the number of times the basic information is accessed in the current period. The basic information includes one or more combinations of customer number, customer type, customer name, ID number, ID type, and customer status. The business data access frequency determination unit is used to determine the customer's business data and the number of times the business data is accessed in the current period. The business data includes one or more combinations of customer capital information, family member information, financial information, risk rating information, credit information, collateral information, loan information, contract information, and permission information. The number of times the business data is accessed is related to the number of times the basic information is accessed. A popularity weight determination unit is used to determine the popularity weight of the basic information in the current period based on the basic information, the number of times the basic information is accessed in the current period, the business data, and the number of times the business data is accessed in the current period. The popularity weight is positively correlated with the number of times the basic information is accessed. The prediction unit is used to predict the heat weight of the basic information in the next period based on the heat weight of the basic information in the current period and the heat weight of the basic information in the previous period. A storage unit is used to save the basic information to the storage location corresponding to the heat weight of the next period according to the heat weight of the next period. The heat weight distribution skewness value determination unit is used to determine the heat weight and distribution skewness value according to the relationship formula. Determine the skewness value of the heat weight distribution The skewness value of the heat weight distribution represents the distribution of the rounded values of the heat weights of the basic information of all users; in, Let n be the mathematical expectation of the popularity weights of the n users in the current period. Let be the integer value of the popularity weight of the basic information of the i-th user in the current period. for The probability of occurrence The standard deviation of the popularity weights of the n users in the current period. ; Distribution determination unit, used for if If the value is less than 0, then the distribution of the integer values of the popularity weights of all users' basic information is left-skewed. If the value equals 0, then the distribution of the integer values of the popularity weights of all users' basic information follows a normal distribution. If the value is greater than 0, then the distribution of the rounded values of the popularity weights of all users' basic information is determined to be a right-skewed distribution. A heat value determination unit is used to determine the heat value of the basic information in the current period based on the distribution. The prediction unit is specifically used to predict the popularity value of the basic information in the next period based on the popularity value of the basic information in the current period and the popularity value of the basic information in the previous period. The popularity value setting unit is used to set three popularity values R1, R2 and R3. R1 is the highest popularity value and R3 is the lowest popularity value. The popularity value is positively correlated with the number of times it is accessed. The heat value determination unit is specifically used to determine the heat value if the distribution of the rounded values of the heat weights of all users' basic information is left-skewed, and then, according to the first mapping relationship... Determine the popularity value of the aforementioned basic information in the current period; If the distribution of the integer values of the popularity weights of all users' basic information is right-skewed, then according to the second mapping relationship... Determine the popularity value of the aforementioned basic information in the current period; If the distribution of the integer values of the popularity weights of all users' basic information follows a normal distribution, then according to the third mapping relationship... Determine the popularity value of the aforementioned basic information in the current period; in, Let M be the integer value of the popularity weight of the i-th user's basic information in the current period, and M be the value that appears most frequently among the popularity weights of the n users. The largest value among the popularity weights of the n users. The minimum value among the popularity weights of the n users. The standard deviation of the popularity weights of the n users in the current period.
8. A data storage device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the data storage method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Data storage method and device and electronic equipment
CN109992210A
Data table popularity distinguishing method and device and related equipment
CN115203195A