Data storage method and system based on big data analysis and storage medium
By classifying real-time data in big data storage twice, combining data information and high-frequency data, the problem of potential hot data in the prior art may be ignored, and a more accurate and reliable data storage is achieved.
Patent Information
- Application Number
- CN202510169615.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art only makes a single judgment on the target data in big data storage, which may ignore potential hot data, resulting in an error in the final storage location.
By collecting data information of multiple target data, dividing it into hot data or cold data, creating a data classification model and training the model, combining the classification results of real-time data with high-frequency data for secondary classification, and finally determining the storage location of real-time data.
Improve the accuracy and reliability of real-time data classification, ensuring that hot and cold data are correctly stored in the appropriate storage medium.
Smart Images

Figure CN120123337A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and particularly relates to a data storage method, system and storage medium based on big data analysis. Background Art
[0002] Big data storage is to persistently store data sets in a computer. Big data usually refers to those data sets that are huge in quantity, difficult to collect, process, and analyze, and also refers to the data that has been stored in traditional infrastructure for a long time.
[0003] Chinese Patent with the publication number CN116150184A discloses a method, device, equipment and medium for separating hot and cold data based on big data. The cluster hot and cold data sets are divided based on the heat index to obtain a cluster hot data set and a cluster warm data set; the cluster hot data set and the cluster warm data set are respectively sharded to obtain a hot data shard set and a warm data shard set; the data in the second database cluster is separately stored based on the hot data shard set and the warm data shard set. However, in the prior art, only a single judgment is made on the target data, which may cause potential hot data to be ignored, resulting in incorrect final storage locations. Summary of the Invention
[0004] The object of the present invention is to address the problems in the background art and propose a data storage method, system and storage medium based on big data analysis.
[0005] The technical solution of the present invention:
[0006] On the one hand, the present application provides a data storage method based on big data analysis, including:
[0007] Collect data information of multiple target data, and divide the multiple target data into hot data or cold data according to the data information;
[0008] Create a data classification model, input the target data into the data classification model to train the data classification model, and obtain a trained data classification model;
[0009] Collect real-time data, and input the real-time data into the trained data classification model to obtain a classification result of the real-time data. Combine the classification result of the real-time data with the high-frequency data of the real-time data to perform secondary classification on the real-time data;
[0010] Divide the storage location of the real-time data in combination with the secondary classification result of the real-time data.
[0011] Preferably, collecting data information of multiple target data and dividing the multiple target data into hot data or cold data includes:
[0012] Create a data information table;
[0013] Collect data information of multiple target data; the data information includes data, data access frequency, and data access mode; the data access frequency refers to the ratio of the number of times the data is accessed to the time, and the data access mode refers to the way and rule in which the data is accessed in the storage system, which reflects the usage and demand characteristics of the data. The data access mode includes online access, DAO mode, and offline mode;
[0014] Sort all the target data based on the data access frequency;
[0015] Record the first N target data as hot data;
[0016] Record the remaining target data as cold data.
[0017] Preferably, create a data classification model, input the target data into the data classification model to train the data classification model, and obtain the trained data classification model, including:
[0018] Create a data classification model;
[0019] Divide all the target data into a training set and a test set according to a random ratio;
[0020] Input the training set into the data classification model to train the data classification model, and obtain the trained data classification model;
[0021] Input the test set into the trained data classification model to verify whether the trained data classification model is trained successfully.
[0022] Preferably, input the training set into the data classification model to train the data classification model, and obtain the trained data classification model, including:
[0023] Randomly select a target data from the training set, and obtain the data access frequency and classification result of the target data;
[0024] Establish a coupling relationship of target data - data access frequency - classification result, and use the target data, data access frequency, classification result, and the coupling relationship of target data - data access frequency - classification result as a training sample to input into the data classification model to train the data classification model;
[0025] Return to randomly select a target data from the training set until all the target data in the training set are selected, and obtain the trained data classification model; the trained data classification model has the ability to automatically classify the target data according to the data information of the input target data.
[0026] Preferably, before collecting real-time data, inputting the real-time data into the trained data classification model to obtain the classification result of the real-time data, and performing secondary classification on the real-time data by combining the classification result of the real-time data with the high-frequency data of the real-time data, it includes:
[0027] Traverse each sub-data in the target data, perform a hash operation on each data, and put the data with the result of i into file i;
[0028] Set a high-frequency threshold;
[0029] Successively judge for each sub-data whether the corresponding i of the sub-data is greater than or equal to the high-frequency threshold;
[0030] If the corresponding i of the sub-data is greater than or equal to the high-frequency threshold, mark the sub-data as standard high-frequency data.
[0031] Preferably, collecting real-time data, inputting the real-time data into the trained data classification model to obtain the classification result of the real-time data, and performing secondary classification on the real-time data by combining the classification result of the real-time data with the high-frequency data of the real-time data, it includes:
[0032] Collect real-time data and input the real-time data into the trained data classification model;
[0033] Obtain the predicted classification output by the trained data classification model;
[0034] Judge whether the predicted classification of the real-time data belongs to cold data;
[0035] If the predicted classification of the real-time data belongs to cold data, perform a hash operation on the real-time data and extract the high-frequency data of the real-time data;
[0036] Judge whether the high-frequency data of the real-time data contains standard high-frequency data;
[0037] If the high-frequency data of the real-time data contains standard high-frequency data, mark the real-time data as hot data.
[0038] Preferably, dividing the storage location of the real-time data by combining the secondary classification result of the real-time data includes:
[0039] Obtain the classification of the real-time data;
[0040] If the real-time data is hot data, store the real-time data in a high-performance storage medium;
[0041] If the real-time data is cold data, store the real-time data in a low-performance storage medium.
[0042] On the other hand, the present application also provides a data storage system based on big data analysis, including a data collection module, a data storage module, and a data processing module;
[0043] Collect data information through the data collection module;
[0044] Store data information through the data storage module;
[0045] Process data information through the data processing model.
[0046] Preferably, the data storage module includes a cold data module and a hot data module. The hot data module is used to store hot data, and the cold data module is used to store cold data.
[0047] On the other hand, the present application also provides a data storage medium based on big data analysis. The data storage medium based on big data analysis stores computer-readable instructions, and the computer-readable instructions can execute the data storage method based on big data analysis as described in any one of the foregoing.
[0048] Compared with the prior art, the above technical solution of the present invention has the following beneficial technical effects:
[0049] By collecting the data information of multiple target data, dividing the multiple target data into hot data or cold data according to the data information, then creating a data classification model, inputting the target data into the data classification model to train the data classification model, obtaining the trained data classification model, then collecting real-time data, and inputting the real-time data into the trained data classification model to obtain the classification result of the real-time data, combining the classification result of the real-time data with the high-frequency data of the real-time data to perform secondary classification on the real-time data, and combining the secondary classification result of the real-time data to divide the storage location of the real-time data. The present application performs two data classifications on the real-time data, the first time is divided according to the data information of the real-time data, and the second time is divided by the high-frequency data of the real-time data, thereby improving the accuracy and reliability of the classification of the real-time data. Description of the Drawings
[0050] Figure 1 It is a schematic flowchart of a data storage method based on big data analysis proposed by the present invention;
[0051] Figure 2 It is a schematic structural diagram of a data storage system based on big data analysis proposed by the present invention;
[0052] Reference numerals: 100, data collection module; 200, data storage module; 201, cold data module;
[0053] 202, hot data module; 300, data processing module. Detailed Embodiments
[0054] Example 1, asFigure 1 As shown in the figure, a data storage method based on big data analysis proposed by the present invention includes:
[0055] S100, collecting data information of multiple target data, and dividing the multiple target data into hot data or cold data according to the data information;
[0056] S200, creating a data classification model, inputting the target data into the data classification model to train the data classification model, and obtaining the trained data classification model;
[0057] S300, collecting real-time data, inputting the real-time data into the trained data classification model, obtaining the classification result of the real-time data, and performing secondary classification on the real-time data by combining the classification result of the real-time data with the high-frequency data of the real-time data;
[0058] S400, dividing the storage location of the real-time data by combining the secondary classification result of the real-time data.
[0059] In the present invention, by collecting data information of multiple target data, dividing the multiple target data into hot data or cold data according to the data information, then creating a data classification model, inputting the target data into the data classification model to train the data classification model, obtaining the trained data classification model, then collecting real-time data, inputting the real-time data into the trained data classification model, obtaining the classification result of the real-time data, performing secondary classification on the real-time data by combining the classification result of the real-time data with the high-frequency data of the real-time data, and dividing the storage location of the real-time data by combining the secondary classification result of the real-time data, the present application performs two data classifications on the real-time data, the first time is divided according to the data information of the real-time data, and the second time is divided by the high-frequency data of the real-time data, so as to improve the accuracy and reliability of the classification of the real-time data.
[0060] In an optional embodiment, the S100 includes:
[0061] S110, creating a data information table;
[0062] S120, collecting data information of multiple target data; the data information includes data, data access frequency, and data access mode;
[0063] Specifically, the data access frequency refers to the ratio of the number of times the data is accessed to the time, and the data access mode refers to the way and rule in which the data is accessed in the storage system, which reflects the usage situation and demand characteristics of the data. The data access mode includes online access, DAO mode, and offline mode;
[0064] S130, sorting all the target data based on the data access frequency;
[0065] S140, Denote the first N target data as hot data;
[0066] S150, Denote the remaining target data as cold data.
[0067] It should be noted that this application collects the data information of multiple target data for the training of the data classification model. Since hot data is data with a higher access frequency and cold data is data with a lower access frequency, when dividing, as long as the target data with a high access frequency is set as hot data and the target data with a low access frequency is set as cold data according to the data access frequency.
[0068] In an optional embodiment, the S200 includes:
[0069] S210, Create a data classification model;
[0070] S220, Divide all the target data into a training set and a test set according to a random ratio;
[0071] Specifically, the division ratio of the training set is greater than that of the test set;
[0072] S230, Input the training set into the data classification model to train the data classification model, and obtain the trained data classification model;
[0073] S240, Input the test set into the trained data classification model to verify whether the trained data classification model is trained successfully.
[0074] It should be noted that by creating and training the data classification model, the data classification model can automatically classify the target data according to the data information of the input target data after training, thereby improving the efficiency of data classification and facilitating the subsequent storage of data into the corresponding storage unit.
[0075] In an optional embodiment, the S230 includes:
[0076] S231, Randomly select a target data from the training set, and obtain the data access frequency and classification result of this target data;
[0077] S232, Establish the coupling relationship of target data - data access frequency - classification result, and input the target data, data access frequency, classification result, and the coupling relationship of target data - data access frequency - classification result as a training sample into the data classification model to train the data classification model;
[0078] S233, returning to randomly select a target data from the training set until all the target data in the training set are selected, and obtaining a trained data classification model; the trained data classification model has the ability to automatically classify the target data according to the data information of the input target data;
[0079] Optionally, since the judgment of cold data and hot data is not only based on the data access frequency, when training the data classification model, hot data and cold data can also be distinguished according to the time dimension. For example, target data with an update time of less than 30 minutes can be classified as hot data, and target data with an update time of more than 30 minutes can be classified as cold data.
[0080] It should be noted that the present application uses the target data and its data information as training samples, and inputs the classification results of the target information into the data classification model, so that the data classification model continuously learns the coupling relationship between data access frequency and classification results.
[0081] In an optional embodiment, before S300, the following steps are included:
[0082] K100, traverses each sub-data in the target data, performs hash operation on each data, and puts the sub-data with result i into file i;
[0083] Specifically, the hash operations include create, insert, find, update and delete operations; in a hash table, data is mapped to a specific location in the table through a hash function, which allows fast data access and operation
[0084] K110, set high frequency threshold;
[0085] K120, determining for each sub-data in turn whether i corresponding to the sub-data is greater than or equal to the high frequency threshold;
[0086] K130, if i corresponding to the sub-data is greater than or equal to the high-frequency threshold, the sub-data is recorded as standard high-frequency data.
[0087] It should be noted that in this embodiment, high-frequency sub-data is extracted from the target data. Since the high-frequency sub-data can reflect the importance of a piece of data, the high-frequency sub-data is also used as a judgment criterion when subsequently dividing the data types.
[0088] In an optional embodiment, the 300 includes:
[0089] S310, collecting real-time data and inputting the real-time data into the trained data classification model;
[0090] S320, obtaining the predicted classification output by the trained data classification model;
[0091] S330, determine whether the predicted classification of the real-time data belongs to cold data;
[0092] S340, if the predicted classification of the real-time data belongs to cold data, perform a hashing operation on the real-time data to extract the high-frequency data of the real-time data;
[0093] S350, determine whether the high-frequency data of the real-time data contains standard high-frequency data;
[0094] S360, if the high-frequency data of the real-time data contains standard high-frequency data, mark the real-time data as hot data;
[0095] Furthermore, the data type of the real-time data can be further determined by performing the following steps:
[0096] K140, perform a hashing operation on the real-time data to extract the high-frequency data of the real-time data;
[0097] K150, determine whether the high-frequency data of the real-time data contains standard high-frequency data;
[0098] K160, if the high-frequency data of the real-time data contains standard high-frequency data, count the number of high-frequency data in the real-time data;
[0099] K170, set a quantity threshold, and determine whether the number of high-frequency data in the real-time data is greater than or equal to the quantity threshold;
[0100] K180, if the number of high-frequency data in the real-time data is greater than or equal to the quantity threshold, mark the real-time data as hot data;
[0101] In this embodiment, in addition to determining whether the high-frequency data of the real-time data contains standard high-frequency data, it is also necessary to further determine the quantity of the included standard high-frequency data, so as to improve the accuracy of the determination of the real-time data type.
[0102] It should be noted that after the real-time data is input into the trained data classification model, the data classification of the real-time data can be obtained. If the real-time data is classified as cold data, a second check can be performed on the real-time data. By determining whether the real-time data contains standard high-frequency data, it can be determined whether the real-time data is cold data.
[0103] Since the real-time data is checked twice in this application, the first time is to use the data information of the real-time data as the judgment criterion, such as data access frequency or data update time. After the real-time data is classified as cold data for the first time, a second judgment is made, that is, to determine whether the real-time data contains standard high-frequency data. Through the two checks, misjudgment of the real-time data can be avoided.
[0104] In an optional embodiment, the S400 includes:
[0105] S410, obtaining the classification of real-time data;
[0106] S420, if the real-time data is hot data, storing the real-time data in a high-performance storage medium;
[0107] S430, if the real-time data is cold data, storing the real-time data in a low-performance storage medium.
[0108] It should be noted that after obtaining the classification result of the real-time data, the real-time data is stored in the corresponding storage location according to the classification result of the real-time data.
[0109] As Figure 2 shown, the present application further provides a data storage system based on big data analysis, including a data acquisition module 100, a data storage module 200, and a data processing module 300. The data information is acquired through the data acquisition module 100, the data information is stored through the data storage module 200, and the data information is processed through the data processing module 300.
[0110] It should be noted that since hot data is frequently accessed, hot data needs to be stored in a storage medium that can perform fast read and write operations. While the access frequency of cold data is relatively low, so cold data can be stored in a medium with a slower read and write speed but lower cost, thereby meeting the storage efficiency requirements for different types of target data on the premise of reducing costs.
[0111] In an optional embodiment, the data storage module 200 includes a cold data module 201 and a hot data module 202. The hot data module 202 is used to store hot data, and the cold data module 201 is used to store cold data.
[0112] It should be noted that the hot data module 202 can be an SSD, and the cold data module 201 can be an HDD.
[0113] The present application further provides a data storage medium based on big data analysis. The data storage medium based on big data analysis stores computer-readable instructions, and the computer-readable instructions can execute the data storage method based on big data analysis as described in any one of the first embodiments.
[0114] It should be noted that the computer-readable instructions can execute the data storage method based on big data analysis as described in any one of the first embodiments.
[0115] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited thereto, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those skilled in the relevant technical field.
Claims
1. A data storage method based on big data analysis, characterized in that: include: Collect data information of multiple target data, and classify the multiple target data into hot data or cold data according to the data information; Creating a data classification model, inputting target data into the data classification model to train the data classification model, and obtaining a trained data classification model; Collect real-time data and input the real-time data into the trained data classification model to obtain the classification results of the real-time data. Combine the classification results of the real-time data with the high-frequency data of the real-time data to perform secondary classification on the real-time data. The storage location of the real-time data is divided according to the secondary classification results of the real-time data.
2. A data storage method based on big data analysis according to claim 1, characterized in that: Collect data information of multiple target data, and classify the multiple target data into hot data or cold data according to the data information, including: Create a data information table; Collect data information of multiple target data; the data information includes data, data access frequency and data access mode; Sort all target data based on data access frequency; The first N target data are recorded as hot data; The remaining target data is recorded as cold data.
3. The data storage method based on big data analysis according to claim 2 is characterized in that: Create a data classification model, input the target data into the data classification model to train the data classification model, and obtain the trained data classification model, including: Create data classification models; Divide all target data into training sets and test sets according to random proportions; Inputting the training set into the data classification model to train the data classification model, thereby obtaining a trained data classification model; Input the test set into the trained data classification model to verify whether the trained data classification model is trained.
4. The data storage method based on big data analysis according to claim 3 is characterized in that: The training set is input into the data classification model to train the data classification model, and the trained data classification model is obtained, including: Randomly select a target data from the training set and obtain the data access frequency and classification result of the target data; Establishing a coupling relationship of target data-data access frequency-classification result, and inputting the target data, data access frequency, classification result, and the coupling relationship of target data-data access frequency-classification result into a data classification model as a training sample to train the data classification model; Return to randomly select a target data from the training set until all the target data in the training set are selected, and obtain the trained data classification model; the trained data classification model has the ability to automatically classify the target data according to the data information of the input target data.
5. The data storage method based on big data analysis according to claim 4 is characterized in that: Before collecting real-time data and inputting the real-time data into the trained data classification model to obtain the classification result of the real-time data, and combining the classification result of the real-time data with the high-frequency data of the real-time data to perform secondary classification on the real-time data, the following steps are included: Traverse each sub-data in the target data, perform hash operation on each data, and put the data with result i into file i; Set high frequency threshold; For each sub-data, determine in turn whether the i corresponding to the sub-data is greater than or equal to the high-frequency threshold; If the i corresponding to the sub-data is greater than or equal to the high-frequency threshold, the sub-data is recorded as standard high-frequency data.
6. A data storage method based on big data analysis according to claim 5, characterized in that: Collect real-time data and input the real-time data into the trained data classification model to obtain the classification results of the real-time data. Combine the classification results of the real-time data with the high-frequency data of the real-time data to perform secondary classification on the real-time data, including: Collect real-time data and input the real-time data into the trained data classification model; Get the predicted classification output by the trained data classification model; Determine whether the predicted classification of real-time data belongs to cold data; If the predicted classification of real-time data belongs to cold data, a hash operation is performed on the real-time data to extract the high-frequency data of the real-time data; Determine whether the high-frequency data of the real-time data contains standard high-frequency data; If the high-frequency data of the real-time data includes the standard high-frequency data, the real-time data is recorded as hot data.
7. The data storage method based on big data analysis according to claim 6 is characterized in that: Combine the secondary classification results of real-time data to divide the storage location of real-time data, including: Get real-time data classification; If the real-time data is hot data, the real-time data is stored in a high-performance storage medium; If the real-time data is cold data, the real-time data is stored in a low-performance storage medium.
8. A data storage system based on big data analysis, applied to the data storage method based on big data analysis according to claim 1 to claim 7, characterized in that: The data storage system based on big data analysis includes: A data acquisition module, through which data information is collected; A data storage module, through which data information is stored; A data processing module is used to process data information.
9. The data storage system based on big data analysis according to claim 8, characterized in that: The data storage module includes a cold data module and a hot data module. The hot data module is used to store hot data, and the cold data module is used to store cold data.
10. A data storage medium based on big data analysis, characterized in that: The data storage medium based on big data analysis stores computer-readable instructions, and the computer-readable instructions can execute the data storage method based on big data analysis as described in any one of claims 1 to claim 7.
Citation Information
Patent Citations
Cold and hot data separation method and device based on big data, equipment and medium
CN116150184A