Classification Method, Device, Computer Equipment and Storage Medium of Data Types
By carefully analyzing and reordering the access data set, a more detailed access frequency queue is generated, which solves the problem of high server pressure when the number of IOs in the prior art, and realizes more efficient data classification and storage resource utilization.
Patent Information
- Application Number
- CN202310901502.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-07-21
AI Technical Summary
The current technical solution for classification of hot and cold data is high when the number of IOs is high, resulting in a sharp increase in server processing pressure and network paralysis.
By obtaining the data set to be accessed, the number of accesses for each data to be accessed is determined, and the subset of data to be accessed is reordered according to the data with the same number of accesses and the number of targets, a more detailed access frequency queue is generated, and adaptive data classification is performed.
It realizes more detailed data classification, makes full use of storage resources, improves system efficiency, reduces server processing pressure, and reduces the probability of network crashes.
Smart Images

Figure CN116910327B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to a method, apparatus, computer device, and storage medium for classifying data types. Background Art
[0002] In a storage system, cold data generally refers to data that has not been accessed for a long time, and hot data refers to data that is frequently accessed. In a distributed storage system, storage is often divided into a cache layer composed of high-speed and expensive SSD (Solid State Disk) and RAM (Random Access Memory), and a permanent storage layer composed of low-speed and inexpensive HDD.
[0003] Usually, only data with a high access frequency is stored in the cache layer to obtain higher access efficiency, and data with a low access frequency is stored in the permanent storage layer to obtain better economic benefits. Therefore, in order to improve the utilization rate of different types of data, it is usually necessary to classify hot and cold data, and then based on the classified hot data or cold data, use them in applicable scenarios respectively to achieve the maximum utilization of data.
[0004] In the current technical solutions for classifying hot and cold data, data is often determined to be cold data or hot data through thresholds or simple LRU (Least Recently Used) methods. This requires frequent processing of accessed data, and when the number of IOs (Input / Output) is high, it causes a sharp increase in the processing pressure of the server and leads to problems such as network paralysis. Summary of the Invention
[0005] In view of this, the present disclosure provides a method, apparatus, computer device, and storage medium for classifying data types to solve the problem that when the number of IOs is high in the current technical solutions for classifying hot and cold data, it causes a sharp increase in the processing pressure of the server and leads to problems such as network paralysis.
[0006] In a first aspect, the present disclosure provides a method for classifying data types, the method including:
[0007] Obtain a set of data to be accessed, and determine the access times of each piece of data to be accessed in the set of data to be accessed;
[0008] Obtain first data to be accessed with the same access times and count the target number of the first data to be accessed;
[0009] According to the target number, determine a subset of data to be accessed to be reordered, and add the reordered subset of data to be accessed to the set of data to be accessed to obtain an access history queue;
[0010] Determine a first-frequency access queue and a second-frequency access queue in the access history queue according to the reordering times;
[0011] Reorder all the to-be-accessed data in the access history queue according to the first-frequency access queue and the second-frequency access queue, and regard the to-be-accessed data located at the first preset position in the queue as first-type data, and regard the other to-be-accessed data outside the first preset position as second-type data.
[0012] In the embodiments of the present disclosure, by obtaining a set of to-be-accessed data, determine the access times of each to-be-accessed data in the set of to-be-accessed data; obtain first to-be-accessed data with the same access times and count the target number of the first to-be-accessed data; determine a subset of to-be-accessed data to be reordered according to the target number, and add the reordered subset of to-be-accessed data to the set of to-be-accessed data to obtain an access history queue; determine a first-frequency access queue and a second-frequency access queue in the access history queue according to the reordering times; reorder all the to-be-accessed data in the access history queue according to the first-frequency access queue and the second-frequency access queue, and regard the to-be-accessed data located at the first preset position in the queue as first-type data, and regard the other to-be-accessed data outside the first preset position as second-type data. In this way, the embodiments of the present disclosure can generate a first-frequency access queue and a second-frequency access queue with more detailed division based on access frequencies according to the access times of each access data and the subset of to-be-accessed data to be reordered, and then complete adaptive data classification for each to-be-accessed data, so as to make full use of storage resources, improve system efficiency, reduce the processing pressure on the server, and reduce the probability of network crashes.
[0013] In an alternative embodiment, adding the reordered subset of to-be-accessed data to the set of to-be-accessed data to obtain an access history queue includes:
[0014] Sort the to-be-accessed data according to the access times of each to-be-accessed data to obtain an initial access history queue;
[0015] Sort the subset of to-be-accessed data and place the sorted subset of to-be-accessed data at the second preset position in the initial access history queue to obtain an access history queue.
[0016] In the embodiments of the present disclosure, by sorting the subset of to-be-accessed data that needs to be reordered and adding it to the second preset position in the initial access history queue, the data with the most frequent access frequency is replaced in real time to the second preset position to achieve adaptive data classification.
[0017] In an alternative embodiment, determining a subset of to-be-accessed data to be reordered according to the target number includes:
[0018] Determine the corresponding standard deviation according to the number of targets;
[0019] Obtain the frequency density of accessing the data to be accessed according to the standard deviation and the first preset threshold;
[0020] Determine the subset of the data to be accessed to be re - sorted according to the frequency density.
[0021] In the embodiments of the present disclosure, the frequency density of the data to be accessed is determined according to the standard deviation corresponding to the number of targets and the set first preset threshold, which facilitates determining which data to be accessed in the data set to be accessed needs to be re - sorted.
[0022] In an alternative embodiment, determining the subset of the data to be accessed to be re - sorted according to the frequency density includes:
[0023] In the case where the frequency density is determined to be the first frequency density type, maintain the current sorting of the data to be accessed in the initial access history queue;
[0024] In the case where the frequency density is determined to be the second frequency density type, determine the corresponding mean value according to the number of targets;
[0025] Obtain the first subset of the data to be accessed to be re - sorted according to the number of targets, the mean value and the standard deviation, and set the data type of the first subset of the data to be accessed to the preset type;
[0026] In the case where the frequency density is determined to be the third frequency density type, obtain the second subset of the data to be accessed to be re - sorted according to the number of targets, the mean value and the standard deviation, and set the data type of the second subset of the data to be accessed to the preset type.
[0027] In the embodiments of the present disclosure, by determining the corresponding subset of the data to be accessed according to the type of the frequency density, and then classifying the data according to the subset of the data to be accessed under each type of the frequency density, the storage resources are fully utilized.
[0028] In an alternative embodiment, determining the first frequency access queue and the second frequency access queue in the access history queue according to the number of re - sorts includes:
[0029] When the number of re - sorts is greater than the second preset threshold, fill the third data to be accessed whose number of times classified into the preset type is the first number into the first frequency access queue to generate the first frequency access queue, and fill the third data to be accessed whose number of times classified into the preset type is the second number into the second frequency access queue to generate the second frequency access queue.
[0030] In the embodiments of the present disclosure, the second access data with the data type set to the preset type is determined according to the target number, and then the corresponding data to be accessed is filled into the first frequency access queue and the second frequency access queue respectively by combining the number of reordering times with the number of times classified into the preset type. In this way, the embodiments of the present disclosure increase the activity type of the data to be accessed with the first frequency access queue and the second frequency access queue, improving the subsequent classification accuracy.
[0031] In an alternative embodiment, all the data to be accessed in the access history queue is reordered according to the first frequency access queue and the second frequency access queue. The data to be accessed at the first preset position in the queue is used as the first type of data, and the other data to be accessed outside the first preset position is used as the second type of data, including:
[0032] Obtain the first weights assigned to each piece of data to be accessed in the first frequency access queue;
[0033] Obtain the second weights assigned to each piece of data to be accessed in the second frequency access queue;
[0034] According to the first weights and the access times of each piece of data to be accessed in the first frequency access queue, obtain the access times of each piece of data to be accessed in the first frequency access queue. According to the second weights and the access times of each piece of data to be accessed in the second frequency access queue, obtain the access times of each piece of data to be accessed in the second frequency access queue;
[0035] Reorder all the data to be accessed in the access history queue according to the access times of each piece of data to be accessed in the first frequency access queue and the access times of each piece of data to be accessed in the second frequency access queue;
[0036] Discharge the data to be accessed at the first preset position in the queue, and use the discharged data to be accessed as the first type of data, and use the data to be accessed outside the first preset position as the second type of data.
[0037] In the embodiments of the present disclosure, based on the weights set for different frequency access queues and the access times of each piece of data to be accessed in different frequency access queues, the access times of each piece of data to be accessed in different frequency access queues are obtained, and then reordering is performed based on the access times of each piece of data to be accessed in different frequency access queues to obtain the final data classification result, saving time and improving the classification accuracy.
[0038] In an alternative embodiment, the access times of each data to be accessed in the queue are accessed according to the first weight and the first frequency, and the access times of each data to be accessed in the first-frequency access queue are obtained. The access times of each data to be accessed in the second-frequency access queue are accessed according to the second weight and the second frequency, and the access times of each data to be accessed in the second-frequency access queue are obtained, including:
[0039] Obtain the access times of each data to be accessed in the first-frequency access queue, sort each data to be accessed in sequence according to the access time, and obtain the access times of the most recent access to each data to be accessed relative to the current time and the access times at the current time;
[0040] According to the first weight, the access times of the most recent access, and the access times at the current time, obtain the access times of each data to be accessed in the first-frequency access queue;
[0041] Obtain the access times of each data to be accessed in the second-frequency access queue, sort each data to be accessed in sequence according to the access time, and obtain the access times of the most recent access to each data to be accessed relative to the current time and the access times at the current time;
[0042] According to the second weight, the access times of the most recent access, and the access times at the current time, obtain the access times of each data to be accessed in the second-frequency access queue.
[0043] In the embodiments of the present disclosure, according to the weights of each frequency access queue, the access times of the most recent access, and the access times at the current time, the access times of each data to be accessed in each frequency access queue are obtained, and then these access times are used for the subsequent sorting of the data to be accessed, improving the overall processing efficiency of the system.
[0044] In a second aspect, the present disclosure provides a classification device for data types, the device including:
[0045] A first acquisition module, configured to acquire a set of data to be accessed and determine the access times of each data to be accessed in the set of data to be accessed;
[0046] A second acquisition module, configured to acquire first data to be accessed with the same access times and count the target number of the first data to be accessed;
[0047] A first determination module, configured to determine a subset of data to be accessed to be re-sorted according to the target number, and add the re-sorted subset of data to be accessed to the set of data to be accessed to obtain an access history queue;
[0048] A second determination module, configured to determine a first-frequency access queue and a second-frequency access queue in the access history queue according to the number of re-sorting times;
[0049] A classification module is configured to re - sort all the to - be - accessed data in the access history queue according to the first frequency access queue and the second frequency access queue, and regard the to - be - accessed data located at the first preset position in the queue as the first type of data, and regard the other to - be - accessed data outside the first preset position as the second type of data.
[0050] In a third aspect, the present disclosure provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the classification method of the data type in the first aspect or any corresponding embodiment thereof.
[0051] In a fourth aspect, the present disclosure provides a computer - readable storage medium, on which computer instructions are stored. The computer instructions are used to cause a computer to execute the classification method of the data type in the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following - described drawings are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0053] Figure 1 is a flowchart of the classification method of the data type according to the embodiment of the present disclosure;
[0054] Figure 2 is a schematic diagram of queue division according to the embodiment of the present disclosure;
[0055] Figure 3 is a schematic diagram of the overall process of the classification of the data type according to the embodiment of the present disclosure;
[0056] Figure 4 is a structural block diagram of the classification device of the data type according to the embodiment of the present disclosure;
[0057] Figure 5 is a schematic diagram of the hardware structure of the computer device according to the embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0059] In the current technical solutions for classifying hot and cold data, data is often determined to be cold or hot data through thresholds or simple LRU methods. This requires frequent processing of accessed data. When the number of I / Os is high, it causes a sharp increase in the processing pressure of the server, resulting in problems such as network paralysis. To solve the above problems, according to the embodiments of the present disclosure, a method embodiment for classifying data types is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0060] In this embodiment, a method for classifying data types is provided. Figure 1 It is a flowchart of the method for classifying data types according to the embodiments of the present disclosure, as Figure 1 shown. This method can be applied to the server side. The method process includes the following steps:
[0061] Step S101, obtain a set of data to be accessed, and determine the number of accesses of each piece of data to be accessed in the set of data to be accessed.
[0062] Optionally, in the embodiments of the present disclosure, as Figure 2 shown, all newly accessed data is inserted into the head of the queue, thereby obtaining an initial access history queue with a length of m. The upper layer issues an I / O request, and fills the initial access history queue according to the request until there are n pieces of data in the queue. At this time, the initial access history queue filled with request information is used as the set of data to be accessed.
[0063] Then, from the time when the initial access history queue is full to a preset time, such as within T time, count the number of accesses of each piece of data to be accessed that receives an I / O request during this period.
[0064] Step S102, obtain the first piece of data to be accessed with the same number of accesses and count the target number of the first piece of data to be accessed.
[0065] Optionally, obtain the first data to be accessed with the same access count from the access counts of each data to be accessed. For example, the access count of data A to be accessed is 3 times, the access count of data B to be accessed is 5 times, and the access count of data C to be accessed is 3 times. At this time, the access counts of data A and data C to be accessed are the same. Both data A and data C to be accessed are called the first data to be accessed at this time. At the same time, count the target number of the first data to be accessed. The current target number is 2 (i.e., data A and data C to be accessed).
[0066] Step S103: Determine the subset of data to be accessed to be re-sorted according to the target number, and add the re-sorted subset of data to be accessed to the set of data to be accessed to obtain an access history queue.
[0067] Optionally, since there are many access counts, such as 1 time, 2 times, 3 times,..., i times, etc., at this time, the access count can be represented by x i The total access count is represented by N, and the target number of all the first data to be accessed with the same access count is represented by n i At this time, count the target number n i of the data with the same access count x i The proportion of the total accessed data N is obtained. After obtaining the normal distribution function It is possible to fit the normal distribution function according to the obtained data to determine the subset of data to be accessed to be re-sorted in the set of data to be accessed.
[0068] Then add the re-sorted subset of data to be accessed to the set of data to be accessed to obtain an access history queue. Here, the access history queue is formed by adding the re-sorted subset of data to be accessed to the set of data to be accessed, that is, after adding it to the initial access history queue.
[0069] It should be noted that the re-sorted subset of data to be accessed can be placed in the second preset position (such as the head of the queue) of the initial access history queue, and then the access history queue is obtained.
[0070] Step S104: Determine the first frequency access queue and the second frequency access queue in the access history queue according to the number of re-sorts.
[0071] Optionally, after obtaining the subset of data to be accessed that needs to be re-sorted, the number of re-sorts can be obtained. According to the number of re-sorts, the access history queue can be divided into a first frequency access queue and a second frequency access queue. Among them, the first frequency access queue can be Figure 2 the medium frequency access queue in Figure 2 and the second frequency access queue can be the high frequency access queue in
[0072] Step S105, re - sort all the to - be - accessed data in the access history queue according to the first - frequency access queue and the second - frequency access queue. Take the to - be - accessed data at the first preset position in the queue as the first - type data, and take the other to - be - accessed data outside the first preset position as the second - type data.
[0073] Optionally, since the first - frequency access queue and the second - frequency access queue are set in the access history queue, at this time, all the to - be - accessed data in the first - frequency access queue and the second - frequency access queue can be re - sorted, so as to obtain the access history queue after re - sorting all the to - be - accessed data in the first - frequency access queue and the second - frequency access queue.
[0074] After that, take the to - be - accessed data that exceeds the queue at the first preset position (such as the end of the queue) in the queue as the first - type data, for example, as cold data; then take the other to - be - accessed data located inside the queue outside the first preset position as the second - type data, for example, as hot data.
[0075] In the embodiment of the present disclosure, by obtaining the set of to - be - accessed data, determining the access times of each to - be - accessed data in the set of to - be - accessed data; obtaining the first to - be - accessed data with the same access times and counting the target number of the first to - be - accessed data; determining the subset of to - be - accessed data to be re - sorted according to the target number, and adding the re - sorted subset of to - be - accessed data to the set of to - be - accessed data to obtain the access history queue; determining the first - frequency access queue and the second - frequency access queue in the access history queue according to the number of re - sorting times; re - sorting all the to - be - accessed data in the access history queue according to the first - frequency access queue and the second - frequency access queue, taking the to - be - accessed data at the first preset position in the queue as the first - type data, and taking the other to - be - accessed data outside the first preset position as the second - type data. In this way, the embodiment of the present disclosure can generate a more detailed first - frequency access queue and second - frequency access queue based on the access frequency according to the access times of each access data and the subset of to - be - accessed data to be re - sorted, and then complete the adaptive data classification of each to - be - accessed data, so as to make full use of the storage resources, improve the system efficiency, reduce the processing pressure of the server, and reduce the probability of network crashes.
[0076] In some alternative embodiments, determining the subset of to - be - accessed data to be re - sorted according to the target number includes:
[0077] Determine the corresponding standard deviation according to the target number;
[0078] Obtain the frequency density of accessing the to - be - accessed data according to the standard deviation and the first preset threshold;
[0079] Determine the subset of to - be - accessed data to be re - sorted according to the frequency density.
[0080] Optionally, fit a normal distribution function according to the obtained data, and determine the corresponding standard deviation σ according to the target number of the first data to be accessed with the same number of all accesses. Then compare the standard deviation σ with a preset first preset threshold to obtain the frequency density of the data to be accessed.
[0081] It should be noted that the first preset threshold here can be a single value or multiple values; the frequency density includes concentrated, relatively dispersed, dispersed, etc.
[0082] After that, determine the subset of the data to be accessed to be reordered according to the frequency density.
[0083] In the embodiments of the present disclosure, determining the frequency density of the data to be accessed according to the standard deviation corresponding to the target number and the set first preset threshold facilitates determining which data to be accessed in the data set to be accessed need to be reordered.
[0084] In some alternative embodiments, determining the subset of the data to be accessed to be reordered according to the frequency density includes:
[0085] When determining that the frequency density is the first frequency density type, keep the current sorting of the data to be accessed in the initial access history queue;
[0086] When determining that the frequency density is the second frequency density type, determine the corresponding mean according to the target number;
[0087] Obtain the first subset of the data to be accessed to be reordered according to the target number, the mean, and the standard deviation, and set the data type of the first subset of the data to be accessed to the preset type;
[0088] When determining that the frequency density is the third frequency density type, obtain the second subset of the data to be accessed to be reordered according to the target number, the mean, and the standard deviation, and set the data type of the second subset of the data to be accessed to the preset type.
[0089] Optionally, determine the corresponding mean μ according to the target number of the first data to be accessed with the same number of all accesses, and directly obtain the normal distribution probability density function according to the mean μ and the standard deviation σ:
[0090]
[0091] Analyze the normal distribution probability density function, and obtain the frequency density according to the mean, the standard deviation, and multiple first preset thresholds.
[0092] If σ is less than the first preset threshold k, where k can be 1, then the frequency density is determined to be the first frequency density type (i.e., the access frequencies are concentrated), and the current sorting of the data to be accessed in the initial access history queue is maintained without further update.
[0093] If σ is greater than a first preset threshold k1 and less than another first preset threshold k2, where k1 can be 1 and k2 can be 10, then the frequency density is determined to be the second frequency density type (i.e., the data volumes with different access frequencies are relatively dispersed). According to the target number, mean, and standard deviation, the set of data where x > (μ + σ) is used as the first subset of data to be accessed for re - sorting, and then it is re - sorted by access frequency and placed at the head of the initial history queue. Also, the data type of the data to be accessed in the first subset of data to be accessed is set to the preset type, i.e., marked as active data, and the remaining data remains unchanged.
[0094] If σ is greater than the first preset threshold k2, where k2 can be 10, then the frequency density is determined to be the third frequency density type (i.e., the data volumes with different access frequencies are distributed dispersedly). According to the target number, mean, and standard deviation, the set of data where x > (μ + 2σ) is used as the second subset of data to be accessed for re - sorting, and then it is re - sorted by access frequency and placed at the head of the initial history queue. Also, the data type of the data to be accessed in the second subset of data to be accessed is set to the preset type, i.e., marked as active data, and the remaining data remains unchanged.
[0095] In the embodiments of the present disclosure, by determining the corresponding subsets of data to be accessed according to the type of frequency density, and then implementing data classification based on the subsets of data to be accessed under each type of frequency density, the storage resources are fully utilized.
[0096] In some alternative embodiments, determining the first frequency access queue and the second frequency access queue in the access history queue according to the number of re - sortings includes:
[0097] When the number of re - sortings is greater than the second preset threshold, the third subset of data to be accessed whose number of times classified into the preset type is the first number is filled into the first frequency access queue to generate the first frequency access queue, and the third subset of data to be accessed whose number of times classified into the preset type is the second number is filled into the second frequency access queue to generate the second frequency access queue.
[0098] Optionally, in the above embodiments, after obtaining the first subset of data to be accessed and the second subset of data to be accessed that need to be re-sorted, the current re-sorting times can be determined. When the re-sorting times are greater than the second preset threshold (such as 5), at this time, the data types are classified into preset types, that is, the data that is an active data for the first number of times (such as any number of times from 3 to 5) is filled into the first frequency access queue to generate the first frequency access queue, and the data that is an active data for the second number of times (such as 5 times or more than 5 times) is filled into the second frequency access queue to generate the second frequency access queue.
[0099] In the embodiments of the present disclosure, the second accessed data with the data type set to the preset type is determined according to the target number, and then, in combination with the number of times classified into the preset type by the re-sorting times, the corresponding data to be accessed is filled into the first frequency access queue and the second frequency access queue respectively. In this way, the embodiments of the present disclosure increase the activity type of the data to be accessed with the first frequency access queue and the second frequency access queue, improving the subsequent classification accuracy.
[0100] In some alternative embodiments, all the data to be accessed in the access history queue are re-sorted according to the first frequency access queue and the second frequency access queue. The data to be accessed at the first preset position in the queue is used as the first type of data, and the other data to be accessed outside the first preset position is used as the second type of data, including:
[0101] Obtain the first weights assigned to each data to be accessed in the first frequency access queue;
[0102] Obtain the second weights assigned to each data to be accessed in the second frequency access queue;
[0103] According to the first weights and the access times of each data to be accessed in the first frequency access queue, obtain the access times of each data to be accessed in the first frequency access queue. According to the second weights and the access times of each data to be accessed in the second frequency access queue, obtain the access times of each data to be accessed in the second frequency access queue;
[0104] Re-sort all the data to be accessed in the second access history queue according to the access times of each data to be accessed in the first frequency access queue and the access times of each data to be accessed in the second frequency access queue;
[0105] Discharge the data to be accessed at the first preset position in the queue, and use the discharged data to be accessed as the first type of data, and use the data to be accessed outside the first preset position as the second type of data.
[0106] Optionally, weights are assigned to the data in the first frequency access queue and the second frequency access queue respectively. For example, a first weight (such as the first weight is 1.5) is assigned to each data to be accessed in the first frequency access queue, and a second weight (such as the second weight is 2) is assigned to each data to be accessed in the second frequency access queue. Then, according to the first weight and the access time of each data to be accessed in the first frequency access queue, the access times of each data to be accessed in the first frequency access queue are calculated; according to the second weight and the access time of each data to be accessed in the second frequency access queue, the access times of each data to be accessed in the second frequency access queue are calculated.
[0107] Then, the access times of each data to be accessed in the first frequency access queue are sorted from largest to smallest, and the access times of each data to be accessed in the second frequency access queue are sorted from largest to smallest, so as to re - sort all the data to be accessed in the access history queue.
[0108] After that, the data to be accessed at the first preset position (i.e., the end of the queue) in the access history queue is discharged and eliminated as cold data, and the data to be accessed in the queue is used as hot data to continue to perform data classification - related operations.
[0109] In the embodiments of the present disclosure, based on the weights set for different frequency access queues and the access times of each data to be accessed in different frequency access queues, the access times of each data to be accessed in different frequency access queues are obtained. Furthermore, based on the access times of each data to be accessed in different frequency access queues, re - sorting is performed to obtain the final data classification result, saving time and improving classification accuracy.
[0110] In some alternative embodiments, obtaining the access times of each data to be accessed in the first frequency access queue according to the first weight and the access time of each data to be accessed in the first frequency access queue, and obtaining the access times of each data to be accessed in the second frequency access queue according to the second weight and the access time of each data to be accessed in the second frequency access queue includes:
[0111] Obtain the access times of each data to be accessed in the first frequency access queue, sort each data to be accessed in sequence according to the access time, and obtain the access times of the nearest previous access to the current time and the access times at the current time for each data to be accessed;
[0112] According to the first weight, the access times of the nearest previous access and the access times at the current time, obtain the access times of each data to be accessed in the first frequency access queue;
[0113] Obtain the access time of each data to be accessed in the second-frequency access queue, sort each data to be accessed in sequence according to the access time, and obtain the access times of the last access closest to the current time and the access time at the current time corresponding to each data to be accessed;
[0114] According to the second weight, the number of times of the last access, and the number of times of access at the current time, obtain the access times of each data to be accessed in the second-frequency access queue.
[0115] Optionally, arranging the access times of each data to be accessed in the first-frequency access queue in the order of access occurrence can obtain the access times corresponding to the last access time and the access time corresponding to the current access time of each data to be accessed. Then multiply the first weight by the access times corresponding to the last access time, and add it to the access times at the current time to obtain the access times of each data to be accessed in the first-frequency access queue;
[0116] Similarly, arranging the access times of each data to be accessed in the second-frequency access queue in the order of access occurrence can obtain the access times corresponding to the last access time and the access time corresponding to the current access time of each data to be accessed. Then multiply the second weight by the access times corresponding to the last access time, and add it to the access times at the current time to obtain the access times of each data to be accessed in the second-frequency access queue.
[0117] In the embodiments of the present disclosure, according to the weight of each frequency access queue, the number of times of the last access, and the number of times of access at the current time, obtain the access times of each data to be accessed in each frequency access queue, and then use these access times for the subsequent sorting of the data to be accessed, improving the overall processing efficiency of the system.
[0118] In some alternative embodiments, as Figure 3 shown, Figure 3 is a schematic overall flow diagram of the classification of data types according to the embodiments of the present disclosure. The specific process is as follows:
[0119] Set up a data queue;
[0120] Issue IO to fill the access history queue;
[0121] Count the data access frequency;
[0122] Fit the normal distribution function;
[0123] Perform data classification and rearrangement;
[0124] Fill the medium-frequency and high-frequency access queues;
[0125] Determine whether the number of rearrangement times is greater than 5 times;
[0126] If it is greater than, rearrange the data according to the weight; obtain the classification result of hot and cold data; if it is not greater than, loop to execute the statistics of data access frequency and subsequent steps.
[0127] In this embodiment, a classification device for data types is further provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0128] This embodiment provides a classification device for data types, as Figure 4 shown, including:
[0129] The first acquisition module 401 is used to acquire the set of data to be accessed and determine the access times of each piece of data to be accessed in the set of data to be accessed;
[0130] The second acquisition module 402 is used to acquire the first data to be accessed with the same access times and count the target number of the first data to be accessed;
[0131] The first determination module 403 is used to determine the subset of data to be accessed to be reordered according to the target number, and add the reordered subset of data to be accessed to the set of data to be accessed to obtain the access history queue;
[0132] The second determination module 404 is used to determine the first frequency access queue and the second frequency access queue in the access history queue according to the reordering times;
[0133] The classification module 405 is used to reorder all the data to be accessed in the access history queue according to the first frequency access queue and the second frequency access queue, and use the data to be accessed at the first preset position in the queue as the first type of data, and use the other data to be accessed outside the first preset position as the second type of data.
[0134] In some alternative implementation manners, the first determination module 403 includes:
[0135] The first sorting unit is used to sort the data to be accessed according to the access times of each piece of data to be accessed to obtain the initial access history queue;
[0136] The second sorting unit is used to sort the subset of data to be accessed and place the sorted subset of data to be accessed at the second preset position in the initial access history queue to obtain the access history queue.
[0137] In some alternative embodiments, the first determination module 403 includes:
[0138] A first determination unit, configured to determine a corresponding standard deviation according to the target number;
[0139] A first obtaining unit, configured to obtain the frequency density of accessing the data to be accessed according to the standard deviation and a first preset threshold;
[0140] A second determination unit, configured to determine a subset of the data to be accessed to be re - sorted according to the frequency density.
[0141] In some alternative embodiments, the second determination unit includes:
[0142] A retention sub - module, configured to retain the current sorting of the data to be accessed in the initial access history queue when it is determined that the frequency density is of a first frequency density type;
[0143] A first determination sub - module, configured to determine a corresponding mean according to the target number when it is determined that the frequency density is of a second frequency density type;
[0144] A first obtaining sub - module, configured to obtain a first subset of the data to be accessed to be re - sorted according to the target number, the mean, and the standard deviation, and set the data type of the first subset of the data to be accessed to a preset type;
[0145] A second obtaining sub - module, configured to obtain a second subset of the data to be accessed to be re - sorted according to the target number, the mean, and the standard deviation, and set the data type of the second subset of the data to be accessed to a preset type when it is determined that the frequency density is of a third frequency density type.
[0146] In some alternative embodiments, the second determination module 404 includes:
[0147] A generating unit, configured to, when the number of re - sortings is greater than a second preset threshold, fill the second data to be accessed with the number of times classified into the preset type being the first number into the first - frequency access queue to generate a first - frequency access queue, and fill the third data to be accessed with the number of times classified into the preset type being the second number into the second - frequency access queue to generate a second - frequency access queue.
[0148] In some alternative embodiments, the classification module 405 includes:
[0149] A first obtaining unit, configured to obtain a first weight assigned to each data to be accessed in the first - frequency access queue;
[0150] A second obtaining unit, configured to obtain a second weight assigned to each data to be accessed in the second - frequency access queue;
[0151] A second obtaining unit, configured to access the access times of each data to be accessed in the queue according to the first weight and the first frequency, obtain the access times of each data to be accessed in the first-frequency access queue, access the access times of each data to be accessed in the queue according to the second weight and the second frequency, and obtain the access times of each data to be accessed in the second-frequency access queue;
[0152] A third sorting unit, configured to re-sort all the data to be accessed in the access history queue according to the access times of each data to be accessed in the first-frequency access queue and the access times of each data to be accessed in the second-frequency access queue;
[0153] A setting unit, configured to discharge the data to be accessed located at the first preset position in the queue, and use the discharged data to be accessed as the first type of data, and use the data to be accessed outside the first preset position as the second type of data.
[0154] In some alternative embodiments, the second obtaining unit includes:
[0155] A first obtaining sub-module, configured to obtain the access times of each data to be accessed in the first-frequency access queue, sort each data to be accessed in sequence according to the access time, and obtain the access times of the most recent access to each data to be accessed relative to the current time and the access times at the current time;
[0156] A third obtaining sub-module, configured to obtain the access times of each data to be accessed in the first-frequency access queue according to the first weight, the most recent access times, and the access times at the current time;
[0157] A second obtaining sub-module, configured to obtain the access times of each data to be accessed in the second-frequency access queue, sort each data to be accessed in sequence according to the access time, and obtain the access times of the most recent access to each data to be accessed relative to the current time and the access times at the current time;
[0158] A fourth obtaining sub-module, configured to obtain the access times of each data to be accessed in the second-frequency access queue according to the second weight, the most recent access times, and the access times at the current time.
[0159] The classification device for data types in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0160] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above embodiments, and will not be elaborated here.
[0161] This embodiment of the present disclosure further provides a computer device having the aboveFigure 4 Classification device for the data type shown
[0162] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present disclosure. As Figure 5 shown, the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 5 Taking one processor 10 as an example in
[0163] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.
[0164] Among them, the memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.
[0165] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device presented by a small program landing page, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and a combination thereof.
[0166] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid state drive; the memory 20 may further include a combination of the above types of memory.
[0167] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.
[0168] Embodiments of the present disclosure also provide a computer-readable storage medium. The methods according to the embodiments of the present disclosure can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading via a network the original computer code stored in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk or a solid state drive, etc.; further, the storage medium can also include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor or the hardware, the methods shown in the above embodiments are implemented.
[0169] Although the embodiments of the present disclosure are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A classification method for a data type, characterized in that, The method includes: Obtaining a set of data to be accessed, and determining the access times of each piece of data to be accessed in the set of data to be accessed; Obtaining first pieces of data to be accessed with the same access times and counting the target number of the first pieces of data to be accessed; Determining a subset of data to be accessed to be re-sorted according to the target number, and adding the re-sorted subset of data to be accessed to the set of data to be accessed to obtain an access history queue; Determining a first-frequency access queue and a second-frequency access queue in the access history queue according to the number of re-sorting times, wherein the first-frequency access queue is obtained by filling second pieces of data to be accessed whose number of times classified into a preset type is the first number into the first-frequency access queue, and the second-frequency access queue is obtained by filling third pieces of data to be accessed whose number of times classified into the preset type is the second number into the second-frequency access queue, and the preset type is a data type whose number of re-sorting times is greater than a second preset threshold; Re-sorting all the data to be accessed in the access history queue according to the first-frequency access queue and the second-frequency access queue, taking the data to be accessed at the first preset position in the queue as first-type data, and taking the other data to be accessed outside the first preset position as second-type data, wherein the re-sorting all the data to be accessed in the access history queue according to the first-frequency access queue and the second-frequency access queue, taking the data to be accessed at the first preset position in the queue as first-type data, and taking the other data to be accessed outside the first preset position as second-type data includes: obtaining first weights assigned to each piece of data to be accessed in the first-frequency access queue; obtaining second weights assigned to each piece of data to be accessed in the second-frequency access queue; obtaining the access times of each piece of data to be accessed in the first-frequency access queue according to the first weights and the access times of each piece of data to be accessed in the first-frequency access queue, and obtaining the access times of each piece of data to be accessed in the second-frequency access queue according to the second weights and the access times of each piece of data to be accessed in the second-frequency access queue; re-sorting all the data to be accessed in the access history queue according to the access times of each piece of data to be accessed in the first-frequency access queue and the access times of each piece of data to be accessed in the second-frequency access queue; discharging the data to be accessed at the first preset position in the queue and taking the discharged data to be accessed as first-type data, and taking the data to be accessed outside the first preset position as second-type data, wherein the first preset position is the end position of the queue.
2. The method according to claim 1, wherein The adding the re-sorted subset of data to be accessed to the set of data to be accessed to obtain an access history queue includes: Sorting the data to be accessed according to the access times of each piece of data to be accessed to obtain an initial access history queue; Sort the subset of data to be accessed, and place the sorted subset of data to be accessed at the second preset position in the initial access history queue to obtain the access history queue.
3. The method according to claim 2, wherein The determining the subset of data to be accessed that needs to be re-sorted according to the target number includes: Determine the corresponding standard deviation according to the target number; Obtain the frequency density of accessing the data to be accessed according to the standard deviation and the first preset threshold; Determine the subset of data to be accessed that needs to be re-sorted according to the frequency density.
4. The method according to claim 3, wherein The determining the subset of data to be accessed that needs to be re-sorted according to the frequency density includes: In the case where the frequency density is determined to be the first frequency density type, keep the current sorting of the data to be accessed in the initial access history queue; In the case where the frequency density is determined to be the second frequency density type, determine the corresponding mean according to the target number; Obtain the first subset of data to be accessed that needs to be re-sorted according to the target number, the mean, and the standard deviation, and set the data type of the first subset of data to be accessed to the preset type; In the case where the frequency density is determined to be the third frequency density type, obtain the second subset of data to be accessed that needs to be re-sorted according to the target number, the mean, and the standard deviation, and set the data type of the second subset of data to be accessed to the preset type.
5. The method according to claim 1, wherein The determining the first frequency access queue and the second frequency access queue in the access history queue according to the number of re-sorting times includes: When the number of re-sorting times is greater than the second preset threshold, fill the second subset of data to be accessed with the number of times classified into the preset type being the first number into the first frequency access queue to generate the first frequency access queue, and fill the third subset of data to be accessed with the number of times classified into the preset type being the second number into the second frequency access queue to generate the second frequency access queue.
6. The method according to claim 1, characterized in that, The obtaining the access times of each piece of data to be accessed in the first frequency access queue according to the first weight and the access time of each piece of data to be accessed in the first frequency access queue, and obtaining the access times of each piece of data to be accessed in the second frequency access queue according to the second weight and the access time of each piece of data to be accessed in the second frequency access queue includes: Obtain the access time of each piece of data to be accessed in the first frequency access queue, sort each piece of data to be accessed in sequence according to the access time, and obtain the access times of the nearest previous access to the current time and the access time at the current time corresponding to each piece of data to be accessed; Obtain the access times of each piece of data to be accessed in the first frequency access queue according to the first weight, the nearest previous access times, and the access time at the current time; Obtain the access time of each piece of data to be accessed in the second frequency access queue, sort each piece of data to be accessed in sequence according to the access time, and obtain the access times of the nearest previous access to the current time and the access time at the current time corresponding to each piece of data to be accessed; Based on the second weight, the number of visits in the most recent access, and the number of visits at the current time, obtain the number of visits for each data to be accessed in the second-frequency access queue.
7. A classification device for a data type, characterized in that, The device includes: A first acquisition module, configured to acquire a set of data to be accessed, and determine the number of visits for each data to be accessed in the set of data to be accessed; A second acquisition module, configured to acquire first data to be accessed with the same number of visits and count the target number of the first data to be accessed; A first determination module, configured to determine a subset of data to be accessed to be re-sorted according to the target number, and add the re-sorted subset of data to be accessed to the set of data to be accessed to obtain an access history queue; A second determination module, configured to determine a first-frequency access queue and a second-frequency access queue in the access history queue according to the number of re-sorts, where the first-frequency access queue is obtained by filling second data to be accessed that are classified into a preset type and have a first number of times into the first-frequency access queue, and the second-frequency access queue is obtained by filling third data to be accessed that are classified into the preset type and have a second number of times into the second-frequency access queue, and the preset type is a data type with the number of re-sorts greater than a second preset threshold; A classification module, configured to re-sort all data to be accessed in the access history queue according to the first-frequency access queue and the second-frequency access queue, use the data to be accessed at the first preset position in the queue as first-type data, and use the other data to be accessed outside the first preset position as second-type data. The classification module includes: a first acquisition unit, configured to acquire a first weight assigned to each data to be accessed in the first-frequency access queue; a second acquisition unit, configured to acquire a second weight assigned to each data to be accessed in the second-frequency access queue; a second obtaining unit, configured to obtain the number of visits for each data to be accessed in the first-frequency access queue according to the first weight and the access time of each data to be accessed in the first-frequency access queue, and obtain the number of visits for each data to be accessed in the second-frequency access queue according to the second weight and the access time of each data to be accessed in the second-frequency access queue; a third sorting unit, configured to re-sort all data to be accessed in the access history queue according to the number of visits for each data to be accessed in the first-frequency access queue and the number of visits for each data to be accessed in the second-frequency access queue; a setting unit, configured to discharge the data to be accessed at the first preset position in the queue, and use the discharged data to be accessed as first-type data, and use the data to be accessed outside the first preset position as second-type data, where the first preset position is the end position of the queue.
8. A computer device, characterized in that, It includes: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the data type classification method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the classification method of the data type described in any one of claims 1 to 6.
Citation Information
Patent Citations
Data classified storage method, device and system
CN112506433A
Method and device for processing access request
CN113783924A