Data cold and hot state prediction method and device, electronic equipment and storage medium
By acquiring the access frequency sequence and business attributes of the data, and using a Hidden Markov Model to train a target prediction model, the problem of accuracy in predicting the hot and cold status of data was solved, and more accurate prediction of the hot and cold status was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies have low accuracy in predicting the hot and cold states of data, and a single hot and cold prediction model cannot fit the patterns of all the data to be predicted.
By obtaining the target access frequency sequence of the data to be predicted, its business attributes are determined. A hidden Markov model is trained using sample data with the same business attributes to establish a target prediction model and perform hot and cold prediction on the access frequency sequence.
It improves the accuracy of predicting the hot and cold status of data, ensures the precision of the prediction results for the hot and cold status of the data to be predicted for each business attribute, and combines the Hidden Markov Model to mine the changing patterns of the hot and cold status of the data.
Smart Images

Figure CN116522158B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer application technology, and in particular to a method, apparatus, electronic device, and storage medium for predicting the hot and cold status of data. Background Technology
[0002] In typical data access processes, some data is frequently accessed within a certain period, while other data is rarely accessed. Frequently accessed data is called "hot data," and rarely accessed data is called "cold data." To improve data query speed and reduce storage costs, hot data is usually stored in storage media with faster query speeds, while cold data is stored in storage media with larger capacity and / or lower cost. However, due to continuous changes in business scenarios or data validity periods, the hot / cold status of data can also change. Therefore, it is necessary to predict the hot / cold status of data to store it accordingly.
[0003] In existing technologies, the method for predicting the hot or cold status of data is based on the short-term access frequency of the data and uses a deep learning model to predict the hot or cold status of the data. However, the prediction results are often incorrect, so the accuracy of the predicted hot or cold status of the data is low. Summary of the Invention
[0004] This invention provides a method, apparatus, electronic device, and storage medium for predicting the hot and cold state of data, in order to solve the technical problem of low accuracy in predicting the hot and cold state of data.
[0005] According to one aspect of the present invention, a method for predicting the hot or cold state of data is provided, wherein the method includes:
[0006] Obtain the data to be predicted and determine the target access frequency sequence corresponding to the data to be predicted, wherein the target access frequency sequence is the time series corresponding to the access volume of the data to be predicted in at least one time period to be predicted.
[0007] Determine the business attributes of the data to be predicted, and determine the target prediction model corresponding to the data to be predicted based on the business attributes;
[0008] The target prediction model is used to perform hot and cold prediction on the input target access frequency sequence to obtain the hot and cold prediction results corresponding to the data to be predicted.
[0009] The target prediction model is a hot / cold prediction model trained on a hidden Markov model using sample data with the same business attributes as the data to be predicted.
[0010] According to another aspect of the present invention, a data hot / cold state prediction apparatus is provided, wherein the apparatus comprises:
[0011] The data processing module is used to acquire the data to be predicted and determine the target access frequency sequence corresponding to the data to be predicted, wherein the target access frequency sequence is a time series corresponding to the access volume of the data to be predicted in at least one time period to be predicted.
[0012] The model determination module is used to determine the business attributes of the data to be predicted, and to determine the target prediction model corresponding to the data to be predicted based on the business attributes.
[0013] The hot / cold prediction module is used to perform hot / cold prediction on the input target access frequency sequence through the target prediction model, and obtain the hot / cold prediction result corresponding to the data to be predicted.
[0014] The target prediction model is a hot / cold prediction model trained on a hidden Markov model using sample data with the same business attributes as the data to be predicted.
[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the data hot / cold state prediction method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data hot / cold state prediction method according to any embodiment of the present invention.
[0020] The technical solution of this invention involves acquiring data to be predicted, determining the target access frequency sequence corresponding to the data to be predicted, determining the business attributes of the data to be predicted, and determining a target prediction model corresponding to the data to be predicted based on the business attributes. The target prediction model is then used to perform hot / cold prediction on the input target access frequency sequence to obtain the hot / cold prediction result corresponding to the data to be predicted. The target prediction model is a hot / cold prediction model trained on a Hidden Markov Model using sample data with the same business attributes as the data to be predicted. This solves the problem that a single hot / cold prediction model cannot fit the patterns of all data to be predicted. By using a hot / cold prediction model corresponding to the business attributes, the accuracy of the hot / cold prediction results for each business attribute is ensured. Furthermore, the invention incorporates a Hidden Markov Model to mine the hot / cold state change patterns of the data to be predicted, further improving the accuracy of the hot / cold prediction results based on the prediction of the hot / cold state of the data to be predicted according to the business attributes.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a method for predicting the hot and cold state of data according to Embodiment 1 of the present invention;
[0024] Figure 2 This is a flowchart of a method for predicting the hot and cold state of data according to Embodiment 2 of the present invention;
[0025] Figure 3 This is an overall flowchart of a data hot / cold state prediction method provided by an embodiment of the present invention;
[0026] Figure 4 This is a schematic diagram of the structure of a data hot / cold state prediction device according to Embodiment 3 of the present invention;
[0027] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the data hot / cold state prediction method of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Example 1
[0031] Figure 1 This is a flowchart illustrating a method for predicting the hot and cold states of data according to Embodiment 1 of the present invention. This embodiment is applicable to data state prediction. The method can be executed by a data hot and cold state prediction device, which can be implemented in hardware and / or software and can be configured within computer software. Figure 1 As shown, the method includes:
[0032] S110. Obtain the data to be predicted and determine the target access frequency sequence corresponding to the data to be predicted.
[0033] The data to be predicted can be understood as data indicating a hot or cold state to be predicted. Optionally, the data to be predicted can be network access data. In this embodiment of the invention, the data to be predicted can be set according to scenario requirements, and is not specifically limited here. For example, the data to be predicted can be financial data, medical data, or educational data, etc.
[0034] The target access frequency sequence is a time series corresponding to the access volume of the data to be predicted within at least one prediction time period. The prediction time period can be understood as the time period during which the data to be predicted is to be predicted. In this embodiment of the invention, the prediction time period can be preset according to scenario requirements and is not specifically limited here. Optionally, the prediction time period can be April, May, and June, etc. Specifically, if it is necessary to predict the hot / cold prediction results of the data to be predicted in April, May, and June, the target access frequency sequence can be a time series generated based on the access volume of the data to be predicted in April, May, and June.
[0035] S120. Determine the business attributes of the data to be predicted, and determine the target prediction model corresponding to the data to be predicted based on the business attributes.
[0036] The business attribute can be understood as an attribute characterizing the business features of the data to be predicted. In this embodiment of the invention, the business attribute can be preset according to scenario requirements, and is not specifically limited here. Optionally, in the scenario where the data to be predicted is financial data, the business attribute can be deposit, loan, foreign exchange and / or savings, etc.
[0037] The target prediction model can be understood as a model determined based on the business attributes of the data to be predicted, used to predict the hot / cold status of the data to be predicted. Specifically, the target prediction model is a hot / cold prediction model trained on a Hidden Markov Model using sample data with the same business attributes as the data to be predicted.
[0038] It's important to understand that Hidden Markov Models (HMMs) are a type of Markov chain, typically used to uncover patterns of change in things related to historical states. An HMM mainly consists of the following components: an observation sequence, a hidden state sequence, a state transition probability matrix, and an emission probability function. In an HMM, each time step corresponds to a hidden state and an observation, each hidden state has a corresponding emission probability function, and each observation is generated based on the emission probability function corresponding to the current hidden state. The hidden state at a later time step depends on the current hidden state and the state transition probability matrix. This model is typically used to solve three types of problems: evaluation problems, i.e., given the model parameters (including the state transition probability matrix, the initial state probability matrix, and the emission probability function), calculating the probability of a given observation sequence; decoding problems, i.e., given the model parameters and the observation sequence, calculating the most likely hidden state sequence; and learning problems, i.e., given the observation sequence, finding the model parameters that maximize the probability of that observation sequence occurring.
[0039] Therefore, optionally, the Hidden Markov Model includes an observation sequence, a hidden state sequence, a state transition probability matrix, and an emission probability function. In this embodiment of the invention, the observation sequence is a target access frequency sequence, and the hidden state sequence is a hot / cold state sequence. Further, the hidden states are set as the hot / cold states of the data to be predicted, and the access volume of the data to be predicted is set as the observed data.
[0040] Optionally, the step of performing cold / hot prediction on the input data using the target prediction model to obtain the cold / hot prediction result corresponding to the data to be predicted includes:
[0041] The target prediction model is used to perform hot / cold prediction on the input target access frequency sequence to obtain the hot / cold state sequence corresponding to the data to be predicted, and the hot / cold prediction result corresponding to the data to be predicted is obtained based on the hot / cold state sequence.
[0042] The hot and cold state sequence can be understood as the time series corresponding to the hot and cold state of the data to be predicted in each time period to be predicted.
[0043] S130. The target access frequency order of the input target is predicted by the target prediction model to obtain the cold and hot prediction result corresponding to the data to be predicted.
[0044] The cold / hot prediction result can be understood as the prediction result of the data to be predicted. In this embodiment of the invention, the cold / hot prediction result can be preset according to scenario requirements, and is not specifically limited here. Optionally, the cold / hot prediction result may include a cold state and a hot state.
[0045] In this embodiment of the invention, predicting the hot or cold state of the data to be predicted allows for appropriate storage of the data. For example, data predicted as hot is stored in a storage medium with fast query speed, such as memory or a cache; data predicted as hot is stored in a storage medium with large capacity and / or low cost, such as a large-capacity hard disk.
[0046] The technical solution of this invention involves acquiring data to be predicted, determining the target access frequency sequence corresponding to the data to be predicted, determining the business attributes of the data to be predicted, and determining a target prediction model corresponding to the data to be predicted based on the business attributes. The target prediction model is then used to perform hot / cold prediction on the input target access frequency sequence to obtain the hot / cold prediction result corresponding to the data to be predicted. The target prediction model is a hot / cold prediction model trained on a Hidden Markov Model using sample data with the same business attributes as the data to be predicted. This solves the problem that a single hot / cold prediction model cannot fit the patterns of all data to be predicted. By using a hot / cold prediction model corresponding to the business attributes, the accuracy of the hot / cold prediction results for each business attribute is ensured. Furthermore, the invention incorporates a Hidden Markov Model to mine the hot / cold state change patterns of the data to be predicted, further improving the accuracy of the hot / cold prediction results based on the prediction of the hot / cold state of the data to be predicted according to the business attributes.
[0047] Example 2
[0048] Figure 2 This is a flowchart of a data hot / cold state prediction method provided in Embodiment 2 of the present invention. This embodiment adds to the above embodiment by determining the business attributes of the data to be predicted and determining the target prediction model corresponding to the data to be predicted based on the business attributes. Figure 2 As shown, the method includes:
[0049] S210. Obtain the data to be predicted and determine the target access frequency sequence corresponding to the data to be predicted.
[0050] S220. Obtain multiple sample data, determine at least one sample data group corresponding to the multiple sample data, and determine the sample sequence group corresponding to each sample data group.
[0051] The sample data can be understood as the data used to train the Hidden Markov Model. It is understood that the sample data can be of the same type as the data to be predicted. In this embodiment of the invention, the sample data and the data to be predicted can be the same or different. Optionally, the sample data can be network access data. In this embodiment of the invention, the sample data can be set according to scenario requirements, and is not specifically limited here. For example, the sample data can be financial data, medical data, or educational data, etc.
[0052] The sample data group can be understood as a data group obtained by dividing the sample data. The sample sequence group can be understood as the sequence group corresponding to the sample data group.
[0053] Optionally, determining at least one sample data group corresponding to the plurality of sample data, and determining a sample sequence group corresponding to each sample data group, includes:
[0054] Obtain transaction prediction data, determine the similarity of business attributes between each sample data based on the transaction prediction data, and determine at least one sample data group based on the similarity, wherein each sample data group includes at least one sample data.
[0055] Obtain transaction log data, and for each sample data group, determine the sample access frequency sequence of each sample data in the sample data group based on the transaction log data.
[0056] The sample access frequency sequence corresponding to each sample data in each sample data group is taken as the sample sequence group corresponding to the sample data group.
[0057] The transaction prediction data can be understood as data characterizing the business attributes of each sample data. Optionally, in the scenario where the data to be predicted is financial data, the transaction prediction data can be the transaction data corresponding to the sample data. In this embodiment of the invention, the business attributes corresponding to the sample data can be determined based on the transaction data of the sample data.
[0058] The similarity can be understood as the degree of similarity of business attributes between the various sample data.
[0059] The transaction log data can be understood as data representing the access frequency of each sample data within the sample data group.
[0060] The sample access frequency sequence is a time series corresponding to the access volume of the sample data within at least one historical time period. The historical time period can be understood as the historical time period of the sample data. In this embodiment of the invention, the historical time period can be preset according to scenario requirements and is not specifically limited here. Optionally, the historical time period can be January, February, and March, etc. Specifically, if the Hidden Markov Model needs to be trained based on the access volume of the sample data in January, February, and March, then the sample access frequency sequence can be a time series generated based on the access volume of the sample data in January, February, and March.
[0061] Optionally, determining the similarity of business attributes among the various sample data based on the transaction prediction data includes:
[0062] The overlap between the various sample data is determined based on the transaction prediction data, and the similarity of business attributes between the various sample data is determined based on the overlap.
[0063] The overlap can be understood as the degree of overlap between the various sample data.
[0064] It should be understood that a single sample data point often participates in multiple transactions, making it difficult to quantify and categorize its business attributes. Therefore, in this embodiment of the invention, the similarity of business attributes among the sample data points is determined based on the overlap of the transactions in which the sample data points participate. Specifically, exemplarily, the formula for calculating the overlap between the first sample and the second sample can be:
[0065]
[0066] Where u represents the first sample, v represents the second sample, set(u) represents the set of transactions in which the first sample participates, set(v) represents the set of transactions in which the second sample participates, and S uv S represents the degree of overlap in transactions between the first and second samples in the quantification. vu V represents the degree of overlap in the transactions participated in by the second sample and the first sample in the quantification, and V represents the sample dataset.
[0067] Understandably, the higher the overlap between two sample data, the more similar their business attributes.
[0068] Optionally, determining at least one sample data group based on the similarity includes:
[0069] The undirected graph corresponding to the sample data is determined based on the similarity.
[0070] Based on the undirected graph, the sample data are grouped using a community detection algorithm to obtain at least one sample data group.
[0071] The undirected graph can be understood as an undirected weighted network constructed for each sample data group, with each sample data as a node and the similarity of business attributes between the sample data as the weight of the edge.
[0072] The community detection algorithm can be understood as an algorithm that groups the sample data into groups. It's important to understand that community detection algorithms are commonly used for clustering analysis of nodes in large networks. Its input is an undirected weighted network consisting of nodes and edge weights, where the edge weights represent the strength of the relationship between two nodes, i.e., the similarity of their business attributes. Specifically, the community detection algorithm can group the sample data in the following ways:
[0073]
[0074] Where Q represents the optimization objective, W represents the sum of all edge weights in the undirected graph, u represents the first node, v represents the second node, and W... uv K represents the edge weight between the first node and the second node. u K represents the sum of the edge weights connected to the first node. v δ(c) represents the sum of the edge weights connected to the second node. u ,c v The parameter indicates whether the first node and the second node belong to the same community.
[0075] Specifically, if the first node and the second node belong to the same community, then δ(c u ,c v ) = 1, otherwise δ(c) u ,c v = 0. The community detection algorithm continuously adjusts the node clustering results to maximize the optimization objective and outputs the node clustering results, i.e., a set of at least one sample data group, for example, C = {C1, C2, ..., C}. m Let} represent m sample data groups. Nodes with closer relationships are more likely to be grouped into the same community; that is, sample data with higher similarity are more likely to be grouped into the same sample data group. Since sample data in the same sample data group usually participate in the same transactions, their hot / cold status changes typically follow the same pattern.
[0076] Optionally, determining the sample access frequency sequence of each sample data based on the transaction log data includes:
[0077] Based on the transaction log data, obtain the access volume of each of the sample data;
[0078] The access volume of each sample data is sorted by time to obtain the sample access frequency sequence corresponding to each sample data.
[0079] The access volume can be understood as the number of times the sample data is accessed. Optionally, the access volume can be the access volume over a preset time period, which can be preset according to scenario requirements and is not specifically limited here. For example, the preset time period can be a month or a day, etc. Specifically, the access volume can be the number of times the sample data is accessed within a month. Further, for each sample data, the access volume is sorted by time to obtain a sample access frequency sequence corresponding to each sample data. It can be understood that the sample access frequency sequence can characterize the access frequency of the sample data in different time periods.
[0080] S230. For each of the sample data groups, a hidden Markov model is trained based on the sample sequence group to obtain a hot / cold prediction model corresponding to the sample data group for each business attribute.
[0081] Specifically, for each set of sample data, a Hidden Markov Model is trained based on the sample sequence set. The observed value sequence is set as the target access frequency sequence, and the hidden state sequence is set as the hot / cold state sequence. There are two hidden states, corresponding to the hot / cold state of the sample data respectively. In summary, the hot / cold state of the sample data is related to both the historical hot / cold state and the current access frequency, which is consistent with the actual situation.
[0082] It is important to understand that, regarding the aforementioned Hidden Markov Model:
[0083] Hidden state: using Q i ∈{C,H} represents the hot / cold status of the sample data in the i-th month. Where, Q i Let C represent the cold or hot state of the sample data at time i, where C represents the cold state and H represents the hot state.
[0084] The state transition matrix P is defined as follows:
[0085]
[0086] P[m,n]=P(Q) i =n|Q i-i =m), m, n∈{C, H}, β 0,1 ,β 1,0 ∈[0,1]
[0087] Where P represents the state transition matrix, P[m,n] represents the probability of transitioning from state m at time i-1 to state n at time i, and β 0,1 β represents the probability of transitioning from a cold state to a hot state. 1,0 Q represents the probability of transitioning from a hot state to a cold state. i Q represents the hot / cold state of the sample data at time i. i-1 This indicates the hot or cold state of the sample data at time i-1.
[0088] It is important to understand that, due to the existence of the state transition matrix P, the hot and cold states Q of the sample data at time i are... i The hot and cold state Q at time i-1 i-1 This mechanism ensures that changes in the hot / cold status of sample data follow the inherent patterns of data state changes, rather than solely depending on the frequency of current access.
[0089] Emission probability function: Changes in the hot / cold state of sample data lead to different access frequencies. Frequent access indicates that data may be in a hot state, and vice versa. In a Hidden Markov Model, the emission probability function determines the frequency of access in the hidden state Q. i The corresponding access frequency R is generated below. i The probability of R, thus making R i Affected by Q i The impact of this. In this embodiment of the invention, a normal function is chosen as the emission probability function, defined as:
[0090]
[0091] Among them, R i μ represents the frequency of access to the sample data at time i. c σ represents the first parameter of the emission probability function in the cold state. 2 c μ represents the square of the second parameter of the emission probability function in the cold state. H σ represents the first parameter of the emission probability function under thermal conditions. 2 H Q represents the square of the second parameter of the emission probability function under thermal conditions. i Let C represent the cold or hot state of the sample data at time i, where C represents the cold state and H represents the hot state.
[0092] Furthermore, specifically, based on the sample access frequency sequence group, the Baum-Welch algorithm is used to train the model parameters λ for each sample data group. k λ, k = 1, 2, 3, ..., m. λ contains the following parts:
[0093] λ={β 0,1 ,β 1,0 μ C μ H , σ C , σ H},
[0094] Where λ represents the model parameters, β 0,1 β represents the probability of transitioning from a cold state to a hot state. 1,0 μ represents the probability of transitioning from a hot state to a cold state. c σ represents the first parameter of the emission probability function in the cold state. c μ represents the second parameter of the emission probability function in the cold state. H σ represents the first parameter of the emission probability function under thermal conditions. H The second parameter represents the emission probability function under thermal conditions.
[0095] It is understandable that the model parameters can be the same or different for different sets of sample data.
[0096] What needs to be understood is that, specifically, β 0,1 ,β 1,0 β represents the probability of state transition in a system where the access frequency changes steadily. 0,1 ,β 1,0 Smaller values mean that the hot / cold state doesn't change frequently. μ and σ represent the emission probability function parameters in different hidden states. For data from frequently used systems, μ is typically larger, indicating higher access frequency, while for less frequently used systems, μ is typically smaller. Since the hidden Markov model for each sample data set is trained using completely different sets of sample frequency sequences, λ... k Only representing sample data group C k The model parameters are independent of other sample data groups.
[0097] Furthermore, specifically, based on the trained Hidden Markov Model (HMM) corresponding to each group of data to be predicted, i.e., the hot / cold prediction model, and the target access frequency sequence, the Viterbi algorithm is used to generate the hidden state sequence Q = {Q} for each group of data to be predicted. i |i=1,2,3,...,T}, where Q represents the generated hidden state sequence, Q i This represents the hot / cold state of the data to be predicted at different time periods. It can be understood that the hot / cold state of the last time period can be used as the current hot / cold state of the data to be predicted, i.e., the hot / cold prediction result corresponding to the data to be predicted. In the embodiment of this invention, the solution of the hidden Markov model uses the Baum-Welch algorithm and the Viterbi algorithm, where the application of the maximum likelihood estimation idea can make the determined hot / cold prediction result more accurate.
[0098] S240. Determine the business attributes of the data to be predicted, and determine the target prediction model corresponding to the data to be predicted based on the business attributes.
[0099] S250. The target access frequency sequence is used to perform hot and cold prediction on the input target access frequency sequence through the target prediction model to obtain the hot and cold prediction result corresponding to the data to be predicted.
[0100] The technical solution of this invention involves acquiring multiple sample data sets, determining at least one sample data group corresponding to each sample data set, and determining a sample sequence group corresponding to each sample data group. For each sample data group, a Hidden Markov Model (HMM) is trained based on the sample sequence group to obtain a hot / cold prediction model corresponding to each business attribute of the sample data group. By training the HMM based on the business attributes of the sample data group, the problem that a single hot / cold prediction model cannot fit the patterns of all data to be predicted is solved, and the accuracy of the HMM is improved when performing hot / cold prediction on data to be predicted for different business attributes.
[0101] Optional, Figure 3 This is an overall flowchart of a data hot / cold state prediction method provided by an embodiment of the present invention, as shown below. Figure 3 As shown, the overall process of the data's hot / cold state prediction method can be as follows:
[0102] 1. Quantify the similarity of business attributes among sample data. Determine the similarity of business attributes among various sample data based on transaction prediction data.
[0103] 2. Divide the sample data into groups. Based on the similarity of business attributes, the sample data is grouped using a community detection algorithm, generating multiple sample data groups with different business attributes.
[0104] 3. Generate sample access frequency sequences separately. Generate the sample access frequency sequence for each sample in the sample data group based on the transaction log data.
[0105] 4. Hidden Markov Models (HMMs) are used to uncover patterns in the hot and cold states of the data. An HMM is built for each data set based on the frequency of access to the samples to uncover patterns in the changes in the hot and cold states of the data.
[0106] 5. Predict the hot / cold state of the data to be predicted. Based on the changing patterns of the hot / cold state of the data, predict the hot / cold state of the data to be predicted in different time periods. The hot / cold state of the last time period can be regarded as the current hot / cold state of the data to be predicted.
[0107] The technical solution of this invention combines a community detection algorithm with a Hidden Markov Model (HMM). First, based on the similarity of business attributes, the sample data is divided into multiple sample data groups using the community detection algorithm. Then, an HMM is trained for each sample data group. This solves the problem that a single model cannot fit all patterns, combining the advantages of the high data partitioning degree of the community detection algorithm and the high data fitting degree of the HMM, thereby improving the accuracy of predicting hot and cold data for the data to be predicted.
[0108] Example 3
[0109] Figure 4 This is a schematic diagram of a data hot / cold state prediction device provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes: a data processing module 310, a model determination module 320, and a hot / cold prediction module 330.
[0110] The data processing module 310 is used to acquire the data to be predicted and determine the target access frequency sequence corresponding to the data to be predicted, wherein the target access frequency sequence is a time series corresponding to the access volume of the data to be predicted within at least one time period to be predicted; the model determination module 320 is used to determine the business attributes of the data to be predicted and determine the target prediction model corresponding to the data to be predicted based on the business attributes; the hot / cold prediction module 330 is used to perform hot / cold prediction on the input target access frequency sequence through the target prediction model to obtain the hot / cold prediction result corresponding to the data to be predicted; wherein the target prediction model is a hot / cold prediction model trained on a hidden Markov model using sample data with the same business attributes as the data to be predicted.
[0111] The technical solution of this invention involves acquiring data to be predicted, determining the target access frequency sequence corresponding to the data to be predicted, determining the business attributes of the data to be predicted, and determining a target prediction model corresponding to the data to be predicted based on the business attributes. The target prediction model is then used to perform hot / cold prediction on the input target access frequency sequence to obtain the hot / cold prediction result corresponding to the data to be predicted. The target prediction model is a hot / cold prediction model trained on a Hidden Markov Model using sample data with the same business attributes as the data to be predicted. This solves the problem that a single hot / cold prediction model cannot fit the patterns of all data to be predicted. By using a hot / cold prediction model corresponding to the business attributes, the accuracy of the hot / cold prediction results for each business attribute is ensured. Furthermore, the invention incorporates a Hidden Markov Model to mine the hot / cold state change patterns of the data to be predicted, further improving the accuracy of the hot / cold prediction results based on the prediction of the hot / cold state of the data to be predicted according to the business attributes.
[0112] Optionally, the data hot / cold state prediction device further includes a sample determination module and a model training module.
[0113] The sample determination module is used to acquire multiple sample data, determine at least one sample data group corresponding to the multiple sample data, and determine a sample sequence group corresponding to each sample data group before determining the business attributes of the data to be predicted and determining the target prediction model corresponding to the data to be predicted based on the business attributes.
[0114] The model training module is used to train a hidden Markov model based on the sample sequence group for each sample data group, so as to obtain a hot / cold prediction model corresponding to the sample data group for each business attribute.
[0115] Optionally, the sample determination module includes: a sample data group determination submodule, a transaction log processing submodule, and a sample sequence group determination submodule.
[0116] The sample data group determination submodule is used to acquire transaction prediction data, determine the similarity of business attributes between each sample data based on the transaction prediction data, and determine at least one sample data group according to the similarity, wherein each sample data group includes at least one sample data.
[0117] The transaction log processing submodule is used to acquire transaction log data and, for each sample data group, determine the sample access frequency sequence of each sample data in the sample data group based on the transaction log data, wherein the sample access frequency sequence is the time series corresponding to the access volume of the sample data in at least one historical time period.
[0118] The sample sequence group determination submodule is used to take the sample access frequency sequence corresponding to each sample data in each sample data group as the sample sequence group corresponding to the sample data group.
[0119] Optionally, the sample data group determination submodule includes: a similarity determination unit.
[0120] The similarity determination unit is used to determine the overlap between each of the sample data based on the transaction prediction data, and to determine the similarity of business attributes between each of the sample data based on the overlap.
[0121] Optionally, the sample data group determination submodule includes: an undirected graph determination unit and a sample data group determination unit.
[0122] The undirected graph determination unit is used to determine the undirected graph corresponding to the sample data based on the similarity.
[0123] The sample data group determination unit is used to group each of the sample data based on the undirected graph using a community detection algorithm to obtain at least one sample data group.
[0124] Optionally, the transaction log processing submodule includes: an access volume determination unit and a sample access frequency sequence determination unit.
[0125] The access volume determination unit is used to obtain the access volume of each sample data based on the transaction log data.
[0126] The sample access frequency sequence determination unit is used to sort the access volume of each sample data according to time to obtain the sample access frequency sequence corresponding to each sample data.
[0127] Optionally, the hidden Markov model includes an observation sequence, a hidden state sequence, a state transition probability matrix, and an emission probability function, wherein the observation sequence is a target access frequency sequence, and the hidden state sequence is a hot / cold state sequence.
[0128] Optional, the hot / cold prediction module 330 is used for:
[0129] The target prediction model is used to perform hot / cold prediction on the input target access frequency sequence to obtain the hot / cold state sequence corresponding to the data to be predicted, and the hot / cold prediction result corresponding to the data to be predicted is obtained based on the hot / cold state sequence.
[0130] The data hot / cold state prediction device provided in the embodiments of the present invention can execute the data hot / cold state prediction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0131] Example 4
[0132] Figure 5 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0133] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0134] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0135] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods for predicting the hot or cold state of data.
[0136] In some embodiments, the data hot / cold state prediction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data hot / cold state prediction method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data hot / cold state prediction method by any other suitable means (e.g., by means of firmware).
[0137] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0138] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0139] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0141] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0142] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0143] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0144] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method of predicting a cold and hot state of data, characterized by, The method comprises: obtaining to-be-predicted data, determining a target access frequency sequence corresponding to the to-be-predicted data, wherein the target access frequency sequence is a time sequence corresponding to the access amount of the to-be-predicted data in at least one to-be-predicted time period; determining a service attribute of the to-be-predicted data, and determining a target prediction model corresponding to the to-be-predicted data based on the service attribute, wherein the service attribute is an attribute representing a service feature of the to-be-predicted data; performing cold and hot prediction on the input target access frequency sequence by using the target prediction model to obtain a cold and hot prediction result corresponding to the to-be-predicted data; wherein the target prediction model is a cold and hot prediction model obtained by training a hidden Markov model based on sample data having the same service attribute as the to-be-predicted data; before determining the service attribute of the to-be-predicted data and determining the target prediction model corresponding to the to-be-predicted data based on the service attribute, the method comprises: obtaining a plurality of sample data, determining at least one sample data group corresponding to the plurality of sample data, and determining a sample sequence group corresponding to each sample data group; for each sample data group, training a hidden Markov model based on the sample sequence group to obtain a cold and hot prediction model corresponding to each sample data group for each service attribute; the method of determining at least one sample data group corresponding to the plurality of sample data and determining a sample sequence group corresponding to each sample data group comprises: obtaining transaction prediction data, determining the similarity of the service attributes between each sample data based on the transaction prediction data, and determining at least one sample data group according to the similarity, wherein each sample data group comprises at least one sample data; obtaining transaction log data, and for each sample data group, determining a sample access frequency sequence of each sample data in the sample data group according to the transaction log data, wherein the sample access frequency sequence is a time sequence corresponding to the access amount of the sample data in at least one historical time period; taking the sample access frequency sequence corresponding to each sample data in each sample data group as the sample sequence group corresponding to the sample data group; the method of determining the similarity of the service attributes between each sample data based on the transaction prediction data comprises: determining the coincidence degree of participating in transactions between each sample data based on the transaction prediction data, and determining the similarity of the service attributes between each sample data based on the coincidence degree.
2. The method of claim 1, wherein, the method of determining at least one sample data group according to the similarity comprises: determining an undirected graph corresponding to the sample data according to the similarity; grouping each sample data by using a community discovery algorithm based on the undirected graph to obtain at least one sample data group.
3. The method of claim 1, wherein, the method of determining a sample access frequency sequence of each sample data based on the transaction log data comprises: obtaining the access amount of each sample data based on the transaction log data; sorting the access amount of each sample data according to time to obtain a sample access frequency sequence corresponding to each sample data.
4. The method of claim 1, wherein, The hidden Markov model comprises an observation value sequence, a hidden state sequence, a state transition probability matrix, and an emission probability function, wherein the observation value sequence is a target access frequency sequence, and the hidden state sequence is a hot and cold state sequence.
5. The method of claim 4, wherein, The cold and hot prediction of the target prediction model on the input target access frequency sequence is performed to obtain a cold and hot prediction result corresponding to the to-be-predicted data. The cold and hot prediction of the target prediction model on the input target access frequency sequence is performed to obtain a cold and hot prediction result corresponding to the to-be-predicted data.
6. A cold and hot state prediction apparatus of data, characterized by, Comprise: The data processing module is configured to obtain to-be-predicted data and determine a target access frequency sequence corresponding to the to-be-predicted data, wherein the target access frequency sequence is a time sequence corresponding to the access amount of the to-be-predicted data in at least one to-be-predicted time period; The model determination module is configured to determine a business attribute of the to-be-predicted data, and determine a target prediction model corresponding to the to-be-predicted data based on the business attribute, wherein the business attribute is an attribute representing a business feature of the to-be-predicted data; The cold and hot prediction module is configured to perform cold and hot prediction on the input target access frequency sequence through the target prediction model to obtain a cold and hot prediction result corresponding to the to-be-predicted data; The target prediction model is a cold and hot prediction model obtained by training a hidden Markov model based on sample data having the same business attribute as the to-be-predicted data; The cold and hot state prediction device for the data further comprises a sample determination module and a model training module; The sample determination module is configured to obtain a plurality of sample data, determine a plurality of sample data groups corresponding to the plurality of sample data, and determine a sample sequence group corresponding to each sample data group before determining the business attribute of the to-be-predicted data and determining the target prediction model corresponding to the to-be-predicted data based on the business attribute; The model training module is configured to train a hidden Markov model based on the sample sequence group for each sample data group to obtain a cold and hot prediction model corresponding to each sample data group of each business attribute; The sample determination module comprises a sample data group determination sub-module, a transaction log processing sub-module, and a sample sequence group determination sub-module; The sample data group determination sub-module is configured to obtain transaction prediction data, determine the similarity of business attributes between each sample data based on the transaction prediction data, and determine at least one sample data group based on the similarity, wherein each sample data group comprises at least one sample data; The transaction log processing sub-module is configured to obtain transaction log data, and determine a sample access frequency sequence of each sample data in the sample data group based on the transaction log data for each sample data group, wherein the sample access frequency sequence is a time sequence corresponding to the access amount of the sample data in at least one historical time period; The transaction log processing sub-module is configured to obtain transaction log data, and determine a sample access frequency sequence of each sample data in the sample data group based on the transaction log data for each sample data group, wherein the sample access frequency sequence is a time sequence corresponding to the access amount of the sample data in at least one historical time period; The sample sequence group determination sub-module is configured to take the sample access frequency sequence corresponding to each sample data in each sample data group as a sample sequence group corresponding to the sample data group. The sample data group determination sub-module comprises a similarity determination unit. The similarity determination unit is configured to determine the coincidence degree of transaction participation between each sample data based on the transaction prediction data, and determine the similarity of business attributes between each sample data based on the coincidence degree.
7. A computer readable storage medium characterized in that, The computer readable storage medium stores computer instructions, and the computer instructions are used to make the processor execute the data cold and hot state prediction method in any one of claims 1-5.
Citation Information
Patent Citations
Data processing method and training method of data state prediction model
CN110968564A
Railway data state evaluation method and system
CN111079827A