Data processing method, device, electronic device and computer readable storage medium
By obtaining the data generation path and accessed records in the database, classifying the data and compressing and storing or deleting it according to its category, the problems of high data repetition rate and waste of storage space in hot and cold data processing in the prior art are solved, and efficient storage management and data quality improvement are achieved.
Patent Information
- Application Number
- CN202011529450.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-12-22
AI Technical Summary
In the prior art, hot and cold data processing has problems such as high data repetition rate and wasting storage space, and it has failed to effectively consider the attributes and correlations of data.
By obtaining the data generation path and accessed records in the database, the data is classified and its categories are determined. When the data is cold data, it is compressed and stored or deleted based on the generation path.
It realizes saving storage space, ensuring timely response to applications, improving data quality, and reducing cache pollution.
Smart Images

Figure CN114610760B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular, to a data processing method, device, electronic device, and computer-readable storage medium. Background Art
[0002] In recent years, big data has emerged and has become a hot topic that almost everyone knows. In new fields such as machine learning, deep learning, and big data, data is an important part of solving problems and even plays a certain guiding role. The value of data lies in the degree to which it is used, that is, the frequency of being queried or updated, which gives rise to the concepts of cold data and hot data.
[0003] In the processing of hot and cold data by existing technologies, generally, the access frequency of data is recorded, and some weight mechanisms are used to effectively divide hot / cold data, and finally the hot data is retained. However, these methods do not take into account the ownership and relevance of data, and the data duplication rate is high, resulting in space waste. Alternatively, the hot and cold data are distinguished by data creation time or data access heat. However, the creation time method ignores the value of data, that is, the frequency of being queried or updated. Although the access heat method takes into account the value of data, it does not pay attention to the quality of data. In this way, even if the hot data is screened out, the data duplication rate is too high.
[0004] It can be seen that the processing of hot and cold data in the existing technology has technical problems such as high data duplication rate and waste of storage space, which needs to be improved. Summary of the invention
[0005] The purpose of the present disclosure is to solve at least one of the above-mentioned technical defects, especially the technical problems of high data duplication and waste of storage space in the processing of hot and cold data in the prior art.
[0006] In a first aspect, a data processing method is provided, the method comprising:
[0007] Obtaining the generation path of data in the database and the access records of the data;
[0008] Classifying the data based on access records of the data to determine the category of the data;
[0009] When the data is cold data, the data is compressed, stored or deleted based on a generation path of the data.
[0010] In a second aspect, a data processing device is provided, the device comprising:
[0011] A data acquisition module, used to acquire the generation path of data in the database and the access records of the data;
[0012] A data classification module, used to classify the data based on the access records of the data and determine the category of the data;
[0013] The data processing module is used to compress, store or delete the data based on the generation path of the data when the data is cold data.
[0014] In a third aspect, an electronic device is provided, the electronic device comprising:
[0015] processor, memory, and bus;
[0016] The bus is used to connect the processor and the memory;
[0017] The memory is used to store operation instructions;
[0018] The processor is used to execute the above-mentioned data processing method by calling the operation instruction.
[0019] In a fourth aspect, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned data processing method.
[0020] The disclosed embodiment classifies data based on accessed records. When data is determined to be cold data, the data is compressed, stored or deleted based on the generation path of the cold data. At the same time, hot data can be updated and saved in real time, which not only saves storage space but also ensures timely response to applications. While stripping hot data and cold data, only hot data in the link is retained in consideration of the repeatability of the data, thereby improving data quality, and in particular, reducing cache pollution to a greater extent in occasional and periodic batch operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for describing the embodiments of the present disclosure are briefly introduced below.
[0022] Figure 1 A flowchart of a data processing method provided by an embodiment of the present disclosure;
[0023] Figure 2 A schematic diagram of a data generation path provided by an embodiment of the present disclosure;
[0024] Figure 3 A flowchart of a data classification method provided by an embodiment of the present disclosure;
[0025] Figure 4A flowchart of a data access method provided by an embodiment of the present disclosure;
[0026] Figure 5 A flowchart of a data compression storage or deletion method provided by an embodiment of the present disclosure;
[0027] Figure 6 A flowchart of a method for decompressing or reconstructing access data provided by an embodiment of the present disclosure;
[0028] Figure 7 A schematic diagram of the structure of a data processing device provided in an embodiment of the present disclosure;
[0029] Figure 8 A structural diagram of a data classification module provided in an embodiment of the present disclosure;
[0030] Fig. 9 A structural diagram of a data access module provided in an embodiment of the present disclosure;
[0031] Fig.10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure.
[0032] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale. DETAILED DESCRIPTION
[0033] Embodiments of the present disclosure are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present disclosure, and cannot be interpreted as limiting the present disclosure.
[0034] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of the features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0035] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.
[0036] The data processing method, device, electronic device and computer-readable storage medium provided by the present disclosure are intended to solve the above technical problems in the prior art.
[0037] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present invention will be described below in conjunction with the accompanying drawings.
[0038] The present disclosure provides a data processing method, such as Figure 1 As shown, the method includes:
[0039] Step S101, obtaining the generation path of data in the database and the access records of the data;
[0040] In the disclosed embodiment, the data in the database may be any data to be stored. For a piece of data, the data generation path refers to the generation path of the data, which is used to represent the blood relationship of the path. When necessary, the data can be reconstructed according to the path.
[0041] Step S102, classifying the data based on the access records of the data to determine the category of the data;
[0042] The access record of data includes the number of times the data is accessed, the time of access, and the type of access, etc., wherein the type of access includes adding, retrieving, updating, and deleting data, etc. In the embodiment of the present disclosure, data can be classified based on the access record of data, and the categories that can be distinguished include hot data and cold data, wherein hot data refers to data that is accessed more times within a certain period of time, and cold data refers to data that is accessed less times within a certain period of time.
[0043] Step S103: When the data is cold data, compress and store or delete the data based on the generation path of the data.
[0044] For a piece of data, when it is determined that the data is cold data, the data is compressed and stored or deleted based on the generation path of the data to retrieve the storage memory of the data. Among them, whether to compress and store the cold data or delete it can be selected according to a preset rule, wherein the preset rule can be whether the time for reconstructing the cold data is greater than a preset value. If it is greater than the preset value, the cold data is compressed and stored, and if it is less than the preset value, the cold data is deleted. Of course, the above are all optional implementations in the embodiments of the present disclosure, and other implementations based on the central idea of the above embodiments of the present disclosure are also within the protection scope of the present disclosure.
[0045] For the convenience of description, a specific embodiment is taken as an example. Figure 2 As shown, the database includes data A, B, and C, wherein the relationship between data A, B, and C is that data A can generate data B, and data B can generate data C, then the generated data of data A is A, the generated path of data B is AB, and the generated path of data C is ABC; wherein, the access record of data A is queried 64 times in the past 12 hours, the access record of data B is queried 85 times in the past 12 hours, and the access record of data C is queried 1 time in the past 12 hours. Of course, the access record can be not only queried, but also created, updated, and deleted, which is for the convenience of explanation here. Based on the access records of the above A, B, and C data, based on the preset rules, such as the number of accesses in the past 12 hours exceeding 20 times is hot data, and the number of accesses less than 20 times is cold data, then it can be determined that data A and data B are hot data, and data C is cold data. When processing data C, it can be determined based on the time when data A or data B generates data C to compress, store, or delete data C. Optionally, when the time for generating data C according to data B exceeds 1 second, data C is compressed and stored, and when the time for generating data C according to data B does not exceed 1 second, data C is deleted. In the embodiment of the present disclosure, each data can have multiple data generated, and one data can also generate multiple data, and the data generation path is also longer, but the data processing method is the same as the above embodiment, and all belong to the protection scope of the present disclosure.
[0046] The disclosed embodiment classifies data based on accessed records. When data is determined to be cold data, the data is compressed, stored or deleted based on the generation path of the cold data. At the same time, hot data can be updated and saved in real time, which not only saves storage space but also ensures timely response to applications. While stripping hot data and cold data, only hot data in the link is retained in consideration of the repeatability of the data, thereby improving data quality, and in particular, reducing cache pollution to a greater extent in occasional and periodic batch operations.
[0047] The present disclosure provides a possible implementation method, such as Figure 3 As shown, in this implementation, classifying the data based on the accessed records of the data to determine the category of the data includes:
[0048] Step S301, when the number of accesses to the data is less than a preset threshold, determining that the data is cold data;
[0049] Step S302, when the number of accesses to the data is greater than the preset threshold, the data is added to a queue to be classified;
[0050] Step S303, determining the time point when all the data in the queue to be classified were last accessed, taking a preset number of data with the earliest last accessed time point in the queue to be classified as cold data, and determining the data in the queue to be classified except the cold data as hot data.
[0051] In the embodiment of the present disclosure, when determining the category of data, it is necessary to determine it based on the access records of the data and the preset determination rules, wherein the preset determination rules can be set by the user, such as when the number of accesses within a specified time reaches a threshold, it is identified as hot data. The access records of data refer to the access records that continue until the data type needs to be confirmed.
[0052] For the embodiments of the present disclosure, for the convenience of explanation, taking a specific embodiment as an example, there are data a, b, c, d, and e in the database. When it is necessary to confirm the type of data a, b, c, d, and e, the access records of each data are obtained, wherein the access records of each data are as follows: data a is queried 65 times and updated 32 times, data b is queried 32 times and updated 16 times, data c is queried 36 times and updated 8 times, data d is queried 48 times and updated 85 times, and data e is queried 6 times. Based on preset rules, when the number of accesses to data is less than the preset number, it is determined that the data is cold data. If the threshold is 10 times, the corresponding data e is cold data, and data a, b, c, and d are all hot data; if the threshold is 100 times, the corresponding data a, b, c, and e are all cold data, and data d is hot data. After the cold data is determined, the hot data is added to the queue to be classified, wherein all the data in the queue to be classified are data whose access times are greater than a preset threshold. For example, the queue to be classified originally includes data f, g, and h. After data d is added, the data are arranged in order based on the time when each data was last accessed, and the access time point of each data is recorded in the next preset time period. Each time the data is accessed, the data is reordered according to the access time point. After the preset time is reached, the data of the first preset number of data whose last access time point in the queue to be classified is closest to the preset time point is regarded as hot data, and the other data is still regarded as cold data. As a specific embodiment of the present disclosure, if the preset time point is 12:00 p.m., the time points at which the data in the queue to be classified were last accessed are as follows: data d was last accessed at 11:58 p.m., data f was last accessed at 6:00 p.m., data g was last accessed at 8:37 a.m., and data h was last accessed at 10:16 p.m. If the data of the first two numbers whose last accessed time points in the queue to be classified are closest to the preset time point are taken as hot data, then data d and data h are hot data, and data f and data g are cold data.
[0053] The disclosed embodiment classifies data according to the data access records. When the total number of accesses to the data is lower than a threshold, the data is treated as cold data. When the total number of accesses to the data exceeds the threshold, the number of accesses to the data within a preset time period is used to confirm whether the data is hot data again, thereby ensuring the accuracy of data confirmation and preventing randomness of the data.
[0054] The present disclosure provides a possible implementation method, such as Figure 4 As shown, in this implementation, compressing, storing or deleting the data based on the generation path of the data includes:
[0055] Step S401, determining upper layer data of the data based on the generation path of the data;
[0056] Step S402, determining target data in upper layer data of the data, the target data being data closest to the data on a generation path;
[0057] Step S403, determining the time required for generating the data from the target data, and compressing and storing or deleting the data according to the relationship between the time and a preset threshold.
[0058] In the embodiments of the present disclosure, when it is necessary to determine the processing of cold data, it is necessary to determine it based on the upper-layer data of the cold data. The upper-layer data refers to the data existing in the path of generating the cold data. For the convenience of explanation, taking the aforementioned embodiment as an example, the generation path of data C is ABC, then the upper-layer data of data C is data A and data B.
[0059] For the embodiments of the present disclosure, for the convenience of explanation, taking a specific embodiment as an example, the generation path of data D is ABCD, wherein data C and data D are cold data, and data A and data B are hot data. When determining the processing method of data D, the upper-layer data A, B, and C of data D are obtained, and the target data in data A, B, and C are determined, wherein the target data is the data closest to the data D on the generation path, which is data C in the embodiment of the present disclosure. The time required to generate data D based on data C is calculated. When the time exceeds the preset time, the data D is compressed and stored. When the time does not exceed the preset time, the data D is deleted. As another embodiment of the present disclosure, data A and data B generate data C, and data C and data D generate data E, wherein E is cold data. When E needs to be processed, it is necessary to calculate the time it takes for data C and data D to be spliced together to form data E, and compare the time with a preset threshold. When the time it takes for data C and data D to be spliced together to form data E exceeds the preset threshold, data E is compressed and stored in its own node. When the time it takes for data C and data D to be spliced together to form data E does not exceed the preset threshold, data E is directly deleted.
[0060] The disclosed embodiment determines the upper layer data of cold data and the target data in the upper layer data of the cold data. When the time for generating the cold data based on the target data exceeds a preset threshold, the cold data is compressed and stored. When the time for generating the cold data based on the target data exceeds the preset threshold, the data is deleted. This can effectively reduce the memory occupied by the data and ensure the timely response of the data to the application.
[0061] The present disclosure provides a possible implementation method, in which: Figure 5 As shown, the data is compressed, stored or deleted according to the size relationship between the duration and the preset threshold, including:
[0062] Step S501, when the duration is greater than the preset threshold, compressing the data into the target data for storage;
[0063] Step S502: when the duration is not greater than the preset threshold, delete the data.
[0064] Based on the aforementioned embodiment, if the time required to generate the cold data based on the target data does not exceed the preset threshold, deleting the cold data can effectively reduce the memory occupied by the data. When the data needs to be accessed, the data can be regenerated, and the time required to generate the data is short and the response speed is also very fast.
[0065] For the embodiments of the present disclosure, for the convenience of explanation, taking a specific embodiment as an example, the generation path of data D is ABCD, wherein data C and data D are cold data, and data A and data B are hot data. When determining the processing method of data D, the upper-layer data A, B, and C of data D are obtained, and the target data in data A, B, and C are determined, wherein the target data is the data closest to data D on the generation path, which is data C in the embodiment of the present disclosure. The time required to generate data D based on data C is calculated. When the time exceeds the preset time, data D is compressed and stored. When data C needs to be processed, similarly, the upper-layer data of data C is data A and data B. If the time to generate data C based on data B exceeds the preset threshold, data C is directly compressed and stored in its own node.
[0066] For the embodiments of the present disclosure, for the convenience of explanation, taking a specific embodiment as an example, the generation path of data D is ABCD, wherein data C and data D are cold data, and data A and data B are hot data. When determining the processing method of data D, the upper-layer data A, B, and C of data D are obtained, and the target data in data A, B, and C are determined, wherein the target data is the data closest to the data D on the generation path, which is data C in the embodiment of the present disclosure. The time required to generate data D based on data C is calculated. When the time exceeds the preset time, the data D is compressed and stored.
[0067] In the disclosed embodiment, when the time duration for generating cold data based on target data exceeds a preset threshold, the cold data is directly deleted, which can effectively reduce the memory occupied by the data. When the data needs to be accessed, the data can be regenerated, and the time required to generate the data is short and the response speed is also very fast.
[0068] The present disclosure provides a possible implementation method, such as Figure 5 As shown, in this implementation, after compressing and storing or deleting the data based on the generation path of the data, it also includes:
[0069] Step S601, receiving a data access request and determining the category of the accessed data;
[0070] Step S602, when the category of the accessed data is cold data, determining a generation path of the accessed data, and determining the target data in upper-layer data of the accessed data based on the generation path of the accessed data;
[0071] Step S603: determine the time required for the target data to generate the data, and obtain the accessed data according to the relationship between the time and the preset threshold.
[0072] In an embodiment of the present disclosure, when a data access request is received, the corresponding response is different based on the type of data accessed by the access request, wherein the access request may be sent by a user through a terminal or by a server, and the access request carries at least the data requested to be accessed, and may also carry the category of the data requested to be accessed, and based on the category, the target data in the upper-layer data of the accessed data is determined, and the target data is decompressed to obtain the accessed data.
[0073] For the embodiments of the present disclosure, for the convenience of explanation, taking a specific embodiment as an example, when a data access request is received, the category of the data requested to be accessed in the data access request is determined. For example, the data requested to be accessed is G, where data G is cold data, then it is necessary to determine the upper-layer data of data G, for example, the upper-layer data of data G is data E and data F, and determine that data E and data F are target data. Corresponding to the aforementioned embodiment, if the time length for generating data G based on data E and F exceeds a preset threshold, data G is stored by compression. When data G needs to be accessed, data G can be first decompressed to obtain data G, and data G can be accessed. If the time length for generating data G based on data E and F does not exceed the preset threshold, data G has been deleted, and data G can be regenerated based on data E and F to access data G.
[0074] The disclosed embodiment obtains the category of the data requested to be accessed in the data access request, and when the data is cold data, determines the target data in the upper-layer data of the cold data, and decompresses or reconstructs the target data according to the relationship between the time length of generating the accessed data based on the target data and the threshold value to obtain the accessed data, with a fast response speed.
[0075] The embodiment of the present disclosure provides a possible implementation method, in which the acquiring of the accessed data according to the magnitude relationship between the duration and the preset threshold value includes:
[0076] According to the time length being greater than a preset threshold, decompressing the target data to obtain the accessed data;
[0077] According to the time length being no greater than a preset threshold, based on the target data and a generation path of the accessed data, the accessed data is restored and called.
[0078] In an embodiment of the present disclosure, when a data access request is received, the corresponding response is different based on the type of data accessed by the access request, wherein the access request may be sent by a user through a terminal or by a server, and the access request carries at least the data requested to be accessed, and may also carry the category of the data requested to be accessed, and based on the category, the target data in the upper-layer data of the accessed data is determined, and the accessed data is regenerated for the target data, and the accessed data is called.
[0079] For the embodiments of the present disclosure, for the convenience of explanation, a specific embodiment is taken as an example. When a data access request is received, the category of the data requested to be accessed in the data access request is determined. For example, the data requested to be accessed is X, where data X is cold data, then the upper layer data of data X needs to be determined. For example, the upper layer data of data X is data Y and data Z, and data Y and data Z are determined to be target data. Corresponding to the aforementioned embodiment, if data X has been deleted, data X is regenerated based on data Y and data Z, and data X is called.
[0080] The disclosed embodiment obtains the category of the data requested to be accessed in the data access request, determines the target data in the upper layer data of the cold data if the data is cold data, and regenerates the accessed data based on the target data, thereby achieving a fast response speed.
[0081] The disclosed embodiment provides a possible implementation method, in which, when the category of the accessed data is hot data, the accessed data is called. When a data access request is received, the corresponding response is different based on the type of data accessed by the access request, wherein the access request may be sent by a user through a terminal or by a server, and the access request may contain at least the data requested to be accessed and may also contain the category of the data requested to be accessed. When the category of the accessed data is hot data, the accessed data may be directly called, and the response speed is fast.
[0082] The disclosed embodiment classifies data based on accessed records. When data is determined to be cold data, the data is compressed, stored or deleted based on the generation path of the cold data. At the same time, hot data can be updated and saved in real time, which not only saves storage space but also ensures timely response to applications. While stripping hot data and cold data, only hot data in the link is retained in consideration of the repeatability of the data, thereby improving data quality, and in particular, reducing cache pollution to a greater extent in occasional and periodic batch operations.
[0083] The present disclosure provides a data processing device, such as Figure 7 As shown, the data processing device 70 may include: a data acquisition module 710, a data classification module 720 and a data processing module 730, wherein:
[0084] The data acquisition module 710 is used to acquire the generation path of data in the database and the access records of the data;
[0085] A data classification module 720, configured to classify the data based on access records of the data to determine a category of the data;
[0086] The data processing module 730 is used to compress, store or delete the data based on the generation path of the data when the data is cold data.
[0087] Further, such as Figure 8 As shown, the data classification module also includes: a threshold comparison unit 810, a queue adding unit 820 and a classification unit 830, wherein:
[0088] A threshold comparison unit 810, configured to determine that the data is cold data when the number of accesses to the data is less than a preset threshold;
[0089] A queue adding unit 820, configured to add the data to a queue to be classified when the number of accesses to the data is greater than the preset threshold;
[0090] The classification unit 830 is used to determine the time point when all the data in the queue to be classified was last accessed, and to determine a preset number of data with the earliest time point of the last access in the queue to be classified as cold data, and to determine the data in the queue to be classified except the cold data as hot data.
[0091] In the disclosed embodiment, when determining the category of data, the data classification module needs to determine it according to the access records of the data and the preset determination rules, wherein the preset determination rules can be set by the user, such as when the number of accesses within a specified time reaches a threshold, it is identified as hot data. The access records of data refer to the access records that continue until the data type needs to be confirmed.
[0092] For the embodiments of the present disclosure, for the convenience of explanation, taking a specific embodiment as an example, there are data a, b, c, d, and e in the database. When it is necessary to confirm the type of data a, b, c, d, and e, the access records of each data are obtained, wherein the access records of each data are as follows: data a is queried 65 times and updated 32 times, data b is queried 32 times and updated 16 times, data c is queried 36 times and updated 8 times, data d is queried 48 times and updated 85 times, and data e is queried 6 times. Based on the preset rules, when the number of accesses to data is less than the preset number, it is determined that the data is cold data. If the threshold is 10 times, the corresponding data e is cold data, and data abcd are all hot data; if the threshold is 100 times, the corresponding data abce are all cold data, and data d is hot data. After the cold data is determined, the hot data is added to the queue to be classified, wherein all the data in the queue to be classified are data whose access times are greater than a preset threshold. For example, the queue to be classified originally includes data fgh. After data d is added, the data are arranged in order based on the time when each data was last accessed, and the access time point of each data is recorded in the next preset time period. Each time the data is accessed, the data is reordered according to the access time point. After the preset time is reached, the data of the first preset number of data whose last access time point in the queue to be classified is closest to the preset time point is regarded as hot data, and the other data is still regarded as cold data. As a specific embodiment of the present disclosure, if the preset time point is 12:00 p.m., the time points at which the data in the queue to be classified were last accessed are as follows: data d was last accessed at 11:58 p.m., data f was last accessed at 6:00 p.m., data g was last accessed at 8:37 a.m., and data h was last accessed at 10:16 p.m. If the data of the first two numbers whose last accessed time points in the queue to be classified are closest to the preset time point are taken as hot data, then data d and data h are hot data, and data f and data g are cold data.
[0093] The disclosed embodiment classifies data according to the data access records. When the total number of accesses to the data is lower than a threshold, the data is treated as cold data. When the total number of accesses to the data exceeds the threshold, the number of accesses to the data within a preset time period is used to confirm whether the data is hot data again, thereby ensuring the accuracy of data confirmation and preventing randomness of the data.
[0094] Optionally, when compressing, storing or deleting the data based on the generation path of the data, the data processing module 730 may be used to:
[0095] Determine upper layer data of the data based on a generation path of the data;
[0096] Determine target data in upper layer data of the data, the target data being data closest to the data on a generation path;
[0097] When the time required for generating the data from the target data exceeds a preset threshold, the data is compressed into the target data for storage.
[0098] Optionally, when compressing, storing or deleting the data based on the generation path of the data, the data processing module 730 may be used to:
[0099] When the time length required for generating the data from the target data does not exceed a preset threshold, the data is deleted.
[0100] Optional, such as Fig. 9 As shown, the data processing device further includes a data access module 740, which is used to:
[0101] Receive data access requests and determine the category of data to be accessed;
[0102] When the category of the accessed data is cold data, determining the target data in upper layer data of the accessed data based on the generation path of the accessed data;
[0103] The target data is decompressed to obtain the accessed data.
[0104] Optionally, the data access module 740 may also be used to:
[0105] Receive data access requests and determine the category of data to be accessed;
[0106] When the category of the accessed data is cold data, determining the target data in upper layer data of the accessed data based on the generation path of the accessed data;
[0107] Based on the target data and the generation path of the accessed data, the accessed data is restored and the accessed data is called.
[0108] Optionally, the data access module 740 may also be used to:
[0109] When the category of the accessed data is hot data, the accessed data is called.
[0110] The disclosed embodiment classifies data based on accessed records. When data is determined to be cold data, the data is compressed, stored or deleted based on the generation path of the cold data. At the same time, hot data can be updated and saved in real time, which not only saves storage space but also ensures timely response to applications. While stripping hot data and cold data, only hot data in the link is retained in consideration of the repeatability of the data, thereby improving data quality, and in particular, reducing cache pollution to a greater extent in occasional and periodic batch operations.
[0111] The data processing device of the embodiment of the present disclosure can execute the data processing method shown in the above embodiment of the present disclosure. The implementation principle is similar and will not be repeated here.
[0112] Reference below Fig.10 , which shows a schematic diagram of the structure of an electronic device 1000 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig.10 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0113] The electronic device includes: a memory and a processor, wherein the processor here may be referred to as the processing device 1001 described below, and the memory may include at least one of the read-only memory (ROM) 1002, the random access memory (RAM) 1003, and the storage device 1008 described below, as shown below:
[0114] like Fig.10 As shown, the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are also stored. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0115] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange data. Although Fig.10 The electronic device 1000 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0116] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 1009, or installed from a storage device 1008, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0117] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0118] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0119] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0120] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains at least two Internet Protocol addresses; sends a node evaluation request including the at least two Internet Protocol addresses to a node evaluation device, wherein the node evaluation device selects an Internet Protocol address from the at least two Internet Protocol addresses and returns it; receives the Internet Protocol address returned by the node evaluation device; wherein the obtained Internet Protocol address indicates an edge node in a content distribution network.
[0121] Alternatively, the computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device: obtains the data generation path and the access records of the data in the database; classifies the data based on the access records of the data to determine the category of the data; when the data is cold data, compresses, stores or deletes the data based on the data generation path.
[0122] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0123] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0124] The modules or units involved in the embodiments described in the present disclosure may be implemented in software or hardware.
[0125] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0126] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0127] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0128] The above description is only a partial implementation mode of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present disclosure. These improvements and modifications should also be regarded as the protection scope of the present disclosure.
Claims
1. A data processing method, characterized in that: include: Obtaining the generation path of data in the database and the access records of the data; Classifying the data based on access records of the data to determine the category of the data; When the data is cold data, compressing and storing or deleting the data based on the generation path of the data; The compressing, storing or deleting the data based on the generation path of the data includes: Determine upper layer data of the data based on a generation path of the data; Determine target data in upper layer data of the data, the target data being data closest to the data on a generation path; Determine the time required for the target data to generate the data, and compress and store or delete the data according to the relationship between the time and a preset threshold; The compressing, storing or deleting the data according to the size relationship between the duration and a preset threshold value includes: When the duration is greater than the preset threshold, compressing and storing the data; When the duration is not greater than the preset threshold, the data is deleted.
2. The data processing method according to claim 1, characterized in that: The classifying the data based on the accessed records of the data to determine the category of the data includes: When the number of accesses to the data is less than a preset threshold, determining that the data is cold data; When the number of accesses to the data is greater than the preset threshold, the data is added to a queue to be classified; Determine the time point when all the data in the queue to be classified were last accessed, take a preset number of data with the earliest last accessed time point in the queue to be classified as cold data, and determine the data in the queue to be classified except the cold data as hot data.
3. The data processing method according to claim 1, characterized in that: After compressing and storing or deleting the data based on the generation path of the data, the method further includes: Receive data access requests and determine the category of data to be accessed; When the category of the accessed data is cold data, determining a generation path of the accessed data, and determining the target data in upper-layer data of the accessed data based on the generation path of the accessed data; The time required for the target data to generate the data is determined, and the accessed data is acquired according to a relationship between the time and the preset threshold.
4. The data processing method according to claim 3, characterized in that: The acquiring the accessed data according to the magnitude relationship between the duration and the preset threshold comprises: According to the time length being greater than a preset threshold, the accessed data is decompressed to obtain the accessed data.
5. The data processing method according to claim 3, characterized in that: The acquiring the accessed data according to the magnitude relationship between the duration and the preset threshold comprises: According to the time length being no greater than a preset threshold, based on the target data and a generation path of the accessed data, the accessed data is restored and called.
6. The data processing method according to claim 3, characterized in that: The determining of the category of the accessed data further comprises: When the category of the accessed data is hot data, the accessed data is called.
7. A data processing device, characterized in that: include: A data acquisition module, used to acquire the generation path of data in the database and the access records of the data; A data classification module, used to classify the data based on the access records of the data and determine the category of the data; A data processing module, configured to compress, store or delete the data based on a generation path of the data when the data is cold data; The compressing, storing or deleting the data based on the generation path of the data includes: Determine upper layer data of the data based on a generation path of the data; Determine target data in upper layer data of the data, the target data being data closest to the data on a generation path; Determine the time required for the target data to generate the data, and compress and store or delete the data according to the relationship between the time and a preset threshold; The compressing, storing or deleting the data according to the size relationship between the duration and a preset threshold value includes: When the duration is greater than the preset threshold, compressing and storing the data; When the duration is not greater than the preset threshold, the data is deleted.
8. An electronic device, characterized in that: It includes: processor, memory, and bus; The bus is used to connect the processor and the memory; The memory is used to store operation instructions; The processor is used to execute the data processing method described in any one of claims 1 to 6 by calling the operation instruction.
9. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the data processing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Processing method and system for cold data in hdfs
CN107861999A
distributed cold and hot data separation method based on Redis and HBase
CN109871367A