Data clustering processing method, flash memory device and computer readable storage medium
By using the clustering method based on the current clustering center and the preset clustering algorithm in the solid-state drive shunt/grouping scheme, the classification error problem under the influence of historical data is solved and the accuracy of data clustering is improved.
Patent Information
- Application Number
- CN202411981657.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-31
AI Technical Summary
During the clustering or classification process of existing solid-state drive diversion/grouping schemes, it is greatly affected by historical data information, and it is difficult to accurately reflect recent input/output characteristics, resulting in classification errors.
By clustering the data points to be clustered based on the current clustering center of multiple clustering groups, the clustering group corresponding to the data points to be clustered is determined, and the current clustering center point of the clustering group is updated through the preset clustering algorithm to obtain the updated clustering center point.
Eliminate the adverse impact of historical data on the clustering center, update the initial clustering center point, so that it more accurately reflects the characteristics of the logical unit, and improves the accuracy of data clustering.
Smart Images

Figure CN119989021A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of data processing technology, and in particular to a data clustering processing method, a flash memory device, and a computer-readable storage medium. Background Art
[0002] In flash memory devices, since the storage medium uses flash memory technology, the way of writing data is different from that of traditional hard disks. When writing to a flash memory unit, the original data must be erased first, and the erase operation is performed in units of larger units (i.e., storage blocks) rather than single data pages. This writing mechanism results in the amount of data actually written to the storage medium often being greater than the original requested write amount, a phenomenon known as write amplification. In order to alleviate the problem of write amplification within the SSD, technologies such as Open Channel (OC), Multi Stream (MS), Zoned Namespace (ZNS), and Flexible Data Placement (FDP) have emerged. These technologies need to be compatible with the upper software stack including applications, file systems, and device drivers, and are difficult to promote. In order to address this issue and achieve application-unaware input / output data placement optimization, many intelligent SSD diversion / grouping solutions have been proposed.
[0003] At present, the existing SSD diversion / grouping scheme averages all historical data points in the cluster through an averaging algorithm. The dynamic incremental averaging takes into account both the incoming data points and the outgoing data points. In the clustering or classification process of the existing scheme, each group is greatly affected by the historical data information, and it is difficult to accurately reflect the recent input / output characteristics, resulting in classification errors. Taking the common K-means clustering as an example, the cluster center is calculated by the mean of all historical data points in the cluster. If the access pattern of a logical block changes significantly, it may move from one cluster to another. If the block continues to contribute to the calculation of the center point of the original cluster, the cluster center of the original cluster may be offset, resulting in inaccurate clustering results. Therefore, the existing technical solutions have the defect of low clustering accuracy. Summary of the invention
[0004] The embodiments of the present application provide a data clustering processing method, a flash memory device and a computer-readable storage medium. The data points to be clustered are clustered based on the current cluster centers of multiple cluster groups to determine the cluster groups corresponding to the data points to be clustered; the data points to be clustered are added to the cluster groups, and the current cluster center points of the cluster groups are updated through a preset clustering algorithm to obtain updated cluster center points. The present application can eliminate the adverse effects of historical data on the cluster centers, update the initial cluster center points, and obtain updated cluster center points, so that the updated cluster centers can more accurately reflect the characteristics of the logical units and improve the accuracy of data clustering.
[0005] The embodiments of the present application provide the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides a data clustering processing method, which is applied to a flash memory device, the flash memory device includes a plurality of logic units, and the method includes:
[0007] Acquire historical feature data corresponding to the logical unit, and determine an initial cluster center point corresponding to the historical feature data, wherein the historical feature data includes a plurality of historical data points, the historical data points are feature vectors, and the feature vectors are used to characterize the hotness or coldness of the historical feature data points;
[0008] According to the initial cluster center, multiple cluster groups are determined, wherein one initial cluster center corresponds to one cluster group;
[0009] According to the initial cluster center point, the historical feature data is clustered, the initial cluster center point corresponding to each historical data point is determined, and the historical feature data with the same initial cluster center are divided into the same cluster group, wherein the hotness and coldness of the historical feature data points in each cluster group are the same or similar;
[0010] Acquire sample data to be clustered, wherein the sample data to be clustered includes a plurality of data points to be clustered;
[0011] Based on the current cluster centers of the multiple cluster groups, cluster the data points to be clustered, and determine the cluster groups corresponding to the data points to be clustered;
[0012] The data points to be clustered are added to the cluster grouping, and the current cluster center point of the cluster grouping is updated through a preset clustering algorithm to obtain an updated cluster center point.
[0013] In some embodiments, clustering the data points to be clustered based on the current cluster centers of the plurality of cluster groups, and determining the cluster groups corresponding to the data points to be clustered, includes:
[0014] Traverse the data points to be clustered in the sample data to obtain the current data points to be clustered;
[0015] Calculate the Euclidean distance between the current data point to be clustered and multiple current cluster center points to obtain multiple distance results;
[0016] According to the multiple distance results, determine the minimum value among the distance results;
[0017] According to the minimum value, determine the current cluster center point corresponding to the minimum value;
[0018] According to the current cluster center point corresponding to the minimum value, the cluster group corresponding to the current data point to be clustered is determined, and the group size of the cluster group is obtained, wherein the group size is the number of logical addresses of the logical units in the cluster group.
[0019] In some embodiments, the sample data includes a timestamp of each sample point, and traversing the data points to be clustered in the sample data to obtain the current data points to be clustered includes:
[0020] Get the timestamp of the data points to be clustered;
[0021] According to the timestamps of the data points to be clustered, determining the time series corresponding to the sample data to be clustered, wherein the data points to be clustered in the time series are sorted from small to large according to the timestamps;
[0022] Traverse the time series in ascending order of timestamps and obtain the current data points to be clustered from the time series.
[0023] In some embodiments, the preset clustering algorithm includes a first clustering algorithm, a second clustering algorithm, or a third clustering algorithm. The current cluster center point of the cluster grouping is updated by the preset clustering algorithm to obtain the updated cluster center point, including:
[0024] By using the first clustering algorithm, the current cluster center point of the cluster grouping is updated to obtain an updated cluster center point;
[0025] or,
[0026] By using the second clustering algorithm, the current cluster center point of the cluster grouping is updated to obtain an updated cluster center point;
[0027] or,
[0028] The current cluster center point of the cluster grouping is updated through the third clustering algorithm to obtain an updated cluster center point.
[0029] In some embodiments, the current cluster center point of the cluster grouping is updated by the first clustering algorithm to obtain the updated cluster center point, including:
[0030] According to the data points to be clustered, the current cluster center point and the group size of the cluster group, the updated cluster center point is calculated, which specifically includes:
[0031]
[0032] Among them, μ new is the updated cluster center point, N is the group size of the cluster group, x in is the data point to be clustered, and μ1 is the current cluster center.
[0033] In some embodiments, the method further comprises:
[0034] Get the current cluster center point of the cluster grouping;
[0035] When there are removed data points in the cluster grouping, the current cluster center point is updated according to the removed data points, the current cluster center point and the grouping size, and the updated cluster center point is calculated, which specifically includes:
[0036]
[0037] Among them, μ new is the updated cluster center point, N is the group size of the cluster group, x out is the data point to be removed, and μ2 is the current cluster center.
[0038] In some embodiments, the current cluster center point of the cluster grouping is updated by the second clustering algorithm to obtain the updated cluster center point, including:
[0039] Setting a moving window and a window size of the moving window, wherein the window size is less than or equal to the group size of the cluster grouping;
[0040] Determine the historical data point to be removed according to the timestamp of the historical data point, wherein the historical data point to be removed is the historical data point with the smallest timestamp in the moving window;
[0041] According to the removed historical data points, the current cluster center point and the current data points to be clustered, the current cluster center point is updated to obtain the updated cluster center point, specifically including:
[0042] mean new =mean+(x k -x k-w ) / w,
[0043] Among them, mean new is the updated cluster center, mean is the current cluster center, x k is the data point to be clustered, x k-wis the historical data point to be removed, w is the window size, where the current cluster center point = the sum of all historical data points in the moving window / window size.
[0044] In some embodiments, the current cluster center point of the cluster grouping is updated by a third clustering algorithm to obtain an updated cluster center point, including:
[0045] Set the smoothing factor;
[0046] According to the smoothing coefficient, the current cluster center point and the current sample point, the current cluster center point is updated to obtain the updated cluster center point, which specifically includes:
[0047] EMA t =αx t +(1-α)EMA t-1 ,
[0048] Among them, EMA t is the updated cluster center point, α is the smoothing coefficient, x t is the data point to be clustered, EMA t-1 is the current cluster center.
[0049] In some embodiments, setting a smoothing coefficient includes:
[0050] According to the group size, the smoothing coefficient is determined, including:
[0051] α=2 / (γ*N+1),
[0052] Among them, α is the smoothing coefficient, γ is the grouping adjustment coefficient, and N is the grouping size.
[0053] In a second aspect, an embodiment of the present application provides a flash memory device, including:
[0054] A processor and a memory, the processor is used to execute the executable program code in the memory, and when the executable program code is executed, the processor executes the instructions of the data clustering processing method of the first aspect.
[0055] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, it implements the data clustering processing method as described in any one of the first aspects.
[0056] The beneficial effects of the embodiments of the present application are as follows: Different from the prior art, the embodiments of the present application provide a data clustering processing method, which is applied to a flash memory device, and the flash memory device includes a plurality of logical units. The method includes: obtaining historical feature data corresponding to the logical unit, and determining an initial cluster center point corresponding to the historical feature data, wherein the historical feature data includes a plurality of historical data points, and the historical data points are feature vectors, and the feature vectors are used to characterize the hotness and coldness of the historical feature data points; determining a plurality of cluster groups according to the initial cluster center, wherein one initial cluster center corresponds to one cluster group; clustering the historical feature data according to the initial cluster center point; Classify, determine the initial cluster center point corresponding to each historical data point, and divide the historical feature data with the same initial cluster center into the same cluster group, wherein the hotness or coldness of the historical feature data points in each cluster group is the same or similar; obtain sample data to be clustered, wherein the sample data to be clustered includes multiple data points to be clustered; cluster the data points to be clustered based on the current cluster centers of the multiple cluster groups, and determine the cluster groups corresponding to the data points to be clustered; add the data points to be clustered to the cluster groups, and update the current cluster center points of the cluster groups through a preset clustering algorithm to obtain updated cluster center points.
[0057] By clustering the data points to be clustered based on the current cluster centers of multiple cluster groups, the cluster groups corresponding to the data points to be clustered are determined; the data points to be clustered are added to the cluster groups, and the current cluster center points of the cluster groups are updated through a preset clustering algorithm to obtain updated cluster center points. The present application can eliminate the adverse effects of historical data on the cluster centers, update the initial cluster center points, and obtain updated cluster center points, so that the updated cluster centers can more accurately reflect the characteristics of the logical units and improve the accuracy of data clustering. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] One or more embodiments are exemplarily described by corresponding drawings, which do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and the figures in the drawings do not constitute proportional limitations unless otherwise stated.
[0059] Figure 1 It is a schematic diagram of a process of intelligent traffic distribution of a solid state disk based on a preset clustering algorithm provided in an embodiment of the present application;
[0060] Figure 2 It is a flow chart of a data clustering processing method provided in an embodiment of the present application;
[0061] Figure 3 yes Figure 2 A detailed flow chart of step S205 in FIG.
[0062] Figure 4 yes Figure 3 A detailed flow chart of step S251 in FIG.
[0063] Figure 5 yes Figure 2 A detailed flow chart of step S206 in FIG.
[0064] Figure 6 yes Figure 5 A detailed flow chart of step S261 in FIG.
[0065] Figure 7 It is a schematic diagram of a flow chart of calculating updated cluster center points when there are removed data points in cluster grouping provided by an embodiment of the present application;
[0066] Figure 8 It is a schematic diagram of an overall process of updating cluster centers based on a dynamic incremental averaging method provided in an embodiment of the present application;
[0067] Fig. 9 yes Figure 5 A detailed flow chart of step S262 in FIG.
[0068] Fig.10 It is a schematic diagram of an overall process of updating cluster centers based on a moving sliding window average method provided in an embodiment of the present application;
[0069] Fig.11 yes Figure 5 A detailed flow chart of step S263 in FIG.
[0070] Fig.12 yes Fig.11 A detailed flow chart of step S2631 in FIG.
[0071] Fig.13 It is a schematic diagram of an overall process of updating cluster centers based on an exponential moving average method provided in an embodiment of the present application;
[0072] Fig.14 It is a structural schematic diagram of a flash memory device provided in an embodiment of the present application.
[0073] Description of Figure Numbers:
[0074] Label name Label name 140 Flash memory devices 141 processor 142 Memory DETAILED DESCRIPTION
[0075] In order to facilitate the understanding of the present application, the present application is described in more detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that when an element is described as "fixed to" another element, it can be directly on the other element, or there can be one or more centered elements therebetween. When an element is described as "connected to" another element, it can be directly connected to the other element, or there can be one or more centered elements therebetween. The terms "vertical", "horizontal", "left", "right" and similar expressions used in this specification are for illustrative purposes only.
[0076] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification and in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.
[0077] The technical solution of the present application is described in detail below in conjunction with the accompanying drawings:
[0078] See also Figure 1 , Figure 1 It is a schematic diagram of a process of intelligent traffic distribution of a solid state disk based on a preset clustering algorithm provided in an embodiment of the present application;
[0079] like Figure 1 As shown, the process of intelligent traffic distribution of solid state drives based on a preset clustering algorithm includes:
[0080] Step (1): Extract the feature vector of the logical block. It is necessary to extract the feature vector that can represent the characteristics of each logical block of the SSD. These feature vectors can contain a variety of information such as the size of the logical block, the read and write frequency, and the access mode, depending on the needs and goals of the diversion strategy.
[0081] Step (2): Obtain clustering results by using a preset clustering algorithm, and use the preset clustering algorithm to cluster the extracted feature vectors. The preset clustering algorithms include a first clustering algorithm, a second clustering algorithm, and a third clustering algorithm, wherein the first clustering algorithm is an algorithm for updating cluster centers based on a dynamic incremental average method, the second clustering algorithm is an algorithm for updating cluster centers based on a moving sliding window average method, and the third clustering algorithm is an algorithm for updating cluster centers based on an exponential moving average method. In the clustering process, the current cluster center point is updated by accepting new data points. The algorithm will iteratively redistribute data points to the nearest cluster center and update the position of the cluster center until the convergence condition is reached (such as the cluster center no longer changes or the preset number of iterations is reached), and obtain the clustering results, wherein the clustering results include multiple cluster groups and the updated cluster center point corresponding to each cluster group.
[0082] Step (3): According to the clustering results, the data is grouped / divided and placed. According to the clustering results obtained by the preset clustering algorithm, the logic blocks are grouped / divided and placed. Specifically, the logic blocks belonging to the same cluster can be placed in the same area of the SSD or on the same type of storage medium to optimize the read and write performance, extend the life of the SSD, or achieve other goals.
[0083] Step (4): After the diversion placement is completed, it is necessary to continuously monitor the write operation of the SSD to determine whether there is new data written. If there is new data written, it is necessary to repeat the above process (starting from extracting the feature vector of the logical block) to dynamically adjust the diversion strategy according to the latest data distribution and access pattern; if there is new data written, the data clustering process ends.
[0084] Please refer to Figure 2 , Figure 2 It is a flow chart of a data clustering processing method provided in an embodiment of the present application;
[0085] like Figure 2 As shown, the process of the data clustering processing method includes:
[0086] Step S201: Acquire historical feature data corresponding to the logic unit, and determine the initial cluster center point corresponding to the historical feature data;
[0087] Specifically, the flash memory device includes multiple logical units, and the types of logical units include but are not limited to logical blocks or logical pages. Historical feature data corresponding to the logical units are obtained from the flash memory device, wherein the historical feature data includes multiple historical data points, and the historical data points are feature vectors. The feature vectors are used to characterize the hotness or coldness of the historical feature data points. It can be understood that when the logical unit is a logical block, one historical data point corresponds to the historical feature data of a logical block, and when the logical unit is a logical page, one historical data point corresponds to the historical feature data of a logical page.
[0088] In some embodiments, by calculating the covariance matrix of historical data points, the covariance matrix is eigendecomposed to extract features of the historical data points, and eigenvalues and corresponding eigenvectors are obtained, wherein the eigenvalue represents the variance of the data projected onto the corresponding eigenvector, and the eigenvector defines the direction of the new coordinate axis, thereby obtaining the eigenvector corresponding to each historical data point, wherein the eigenvector includes but is not limited to the size of the logical unit, the read and write frequency, the access mode, and other information that can reflect the hot and cold characteristics of the logical unit.
[0089] Step S202: determining multiple cluster groups according to the initial cluster center point;
[0090] Specifically, k historical data points are randomly selected from the historical feature data as initial cluster centers, or areas with higher density of data points are determined through density estimation algorithms (such as kernel density estimation, DBSCAN), and the center points of these areas are selected as initial cluster centers to determine the initial cluster centers, and then multiple cluster groups are determined based on the initial cluster center points, where one initial cluster center point corresponds to one cluster group, that is, the number of initial cluster center points is equal to the number of cluster groups.
[0091] Step S203: clustering the historical feature data according to the initial cluster center point, determining the initial cluster center point corresponding to each historical data point, and dividing the historical feature data with the same initial cluster center into the same cluster group;
[0092] Specifically, the clustering algorithm includes a K-Means algorithm. Based on the K-Means algorithm, k historical data points are randomly selected from the historical data as initial clustering centers. For each historical data point, the Euclidean distance between the historical data point and the k initial clustering centers is calculated to obtain k Euclidean distances corresponding to each historical data point. The minimum value of the Euclidean distance is determined from the k Euclidean distances, and the initial clustering center corresponding to the minimum value of the Euclidean distance is determined as the initial clustering center point corresponding to the historical data point.
[0093] For example, assuming that there are initial clustering centers A and B, as well as historical data points C and D, for historical data point C, first calculate the Euclidean distance between the feature vector corresponding to historical data point C and the initial clustering center A, recorded as the first Euclidean distance, then calculate the Euclidean distance between the feature vector corresponding to historical data point C and the initial clustering center B, recorded as the second Euclidean distance, compare the first Euclidean distance and the second Euclidean distance to determine the minimum value of the first Euclidean distance and the second Euclidean distance, and determine the initial clustering center corresponding to the minimum value as the clustering result of the historical data point C; for historical data point D, first calculate the Euclidean distance between the feature vector corresponding to historical data point D and the initial clustering center A, recorded as the third Euclidean distance, then calculate the Euclidean distance between the feature vector corresponding to historical data point C and the initial clustering center B, recorded as the fourth Euclidean distance, compare the third Euclidean distance and the fourth Euclidean distance to determine the minimum value of the third Euclidean distance and the fourth Euclidean distance, and determine the initial clustering center corresponding to the minimum value as the clustering result of the historical data point D.
[0094] In the embodiment of the present application, the number of k can be determined by the "elbow rule", which refers to plotting the relationship between the total sum of squared errors (SSE) and k, and selecting the k value corresponding to the inflection point (elbow). In a multi-stream scenario, k is generally set to the maximum number of streams / groups supported by the flash device.
[0095] Step S204: obtaining sample data to be clustered;
[0096] Specifically, sample data to be clustered is obtained, wherein the sample data to be clustered includes a plurality of data points to be clustered. It is understandable that the data points to be clustered have not yet been divided into any cluster group.
[0097] Step S205: clustering the data points to be clustered based on the current cluster centers of the multiple cluster groups, and determining the cluster groups corresponding to the data points to be clustered;
[0098] For details, please refer to Figure 3 , Figure 3 yes Figure 2 A detailed flow chart of step S205 in FIG.
[0099] like Figure 3 As shown, step S205: clustering the data points to be clustered based on the current cluster centers of the multiple cluster groups, and determining the cluster groups corresponding to the data points to be clustered, including:
[0100] Step S251: traverse the data points to be clustered in the sample data to obtain the current data points to be clustered;
[0101] For details, please refer to Figure 4, Figure 4 yes Figure 3 A detailed flow chart of step S251 in FIG.
[0102] like Figure 4 As shown, step S251: traverse the data points to be clustered in the sample data to obtain the current data points to be clustered, including:
[0103] Step S2511: Obtain the timestamp of the data point to be clustered;
[0104] Specifically, the sample data to be clustered includes the timestamp of each data point to be clustered. Based on the sample data, the timestamp of each data point to be clustered is obtained. For example, assuming that the data point to be clustered is the feature vector corresponding to the write request of the solid-state drive, then the timestamp corresponding to the sample data is the time when the solid-state drive receives the write request.
[0105] Step S2512: determining the time series corresponding to the sample data to be clustered according to the timestamps of the data points to be clustered;
[0106] Specifically, according to the timestamp of each data point to be clustered, the data points to be clustered are sorted in ascending order of timestamp to determine the time series corresponding to the sample data to be clustered, wherein the data points to be clustered in the time series are sorted in ascending order of timestamp. For example, assuming that the sample data to be clustered includes data point A to be clustered, data point B to be clustered and data point C to be clustered, the timestamp corresponding to data point A to be clustered>the timestamp corresponding to data point B to be clustered>the timestamp corresponding to data point C to be clustered, then the time series corresponding to the sample data can be expressed as {C, B, A}.
[0107] Step S2513: traverse the time series in ascending order of timestamps, and obtain the current data point to be clustered from the time series;
[0108] Specifically, the time series corresponding to the sample data to be clustered is traversed in order of timestamps from small to large, and the current sample point to be clustered is obtained from the time series. For example, assuming that the time series corresponding to the sample data to be clustered is {C, B, A}, then during the first traversal, the sample point C to be clustered is taken out from the time series, and the current sample point is determined as the sample point C to be clustered. During the second traversal, the sample point B to be clustered is taken out from the time series, and the current sample point is determined as the sample point B to be clustered. During the third traversal, the sample point A to be clustered is taken out from the time series, and the current sample point is determined as the sample point A to be clustered.
[0109] Step S252: Calculate the Euclidean distance between the current data point to be clustered and multiple current cluster center points to obtain multiple distance results;
[0110] Specifically, the data points to be clustered are obtained. It can be understood that the data points to be clustered have not been divided into any cluster group. By calculating the Euclidean distance between the data points to be clustered and multiple current cluster center points, multiple distance results are obtained.
[0111] In some embodiments, the Euclidean distance between each current cluster center point and the data point to be clustered can be calculated according to the coordinates of the current cluster center point and the coordinates of the data point to be clustered, and multiple distance results can be obtained, wherein the distance results include the Euclidean distances between multiple current cluster center points and the data point to be clustered, and the calculation formula of the Euclidean distance is as follows:
[0112] d = sqrt((x2-x1) 2 +(y2-y1) 2 ),
[0113] Among them, d is the Euclidean distance between the current cluster center and the data point to be clustered, x1 is the horizontal coordinate of the current cluster center, y2 is the vertical coordinate of the current cluster center, x2 is the horizontal coordinate of the data point to be clustered, and y2 is the vertical coordinate of the data point to be clustered.
[0114] Step S253: determining a minimum value among the distance results according to the multiple distance results;
[0115] Specifically, the Euclidean distance between the data point to be clustered and multiple current cluster center points is calculated, and after multiple distance results are obtained, the minimum value among the distance results is determined by comparing the multiple cluster results.
[0116] For example, assuming that there are current cluster centers A and B, and data points C and D to be clustered, for the data point C to be clustered, first calculate the Euclidean distance between the feature vector corresponding to the data point C to be clustered and the current cluster center A, recorded as the first Euclidean distance, then calculate the Euclidean distance between the feature vector corresponding to the data point C to be clustered and the current cluster center B, recorded as the second Euclidean distance, compare the first Euclidean distance and the second Euclidean distance to determine the minimum value of the first Euclidean distance and the second Euclidean distance; for the data point D to be clustered, first calculate the Euclidean distance between the feature vector corresponding to the data point D to be clustered and the current cluster center A, recorded as the third Euclidean distance, then calculate the Euclidean distance between the feature vector corresponding to the data point D to be clustered and the current cluster center B, recorded as the fourth Euclidean distance, compare the third Euclidean distance and the fourth Euclidean distance to determine the minimum value of the third Euclidean distance and the fourth Euclidean distance.
[0117] Step S254: determining the current cluster center point corresponding to the minimum value according to the minimum value;
[0118] Specifically, each current cluster center point corresponds to a cluster group, and according to the minimum value, the current cluster center point corresponding to the minimum value is determined, and then the cluster group corresponding to the current cluster center point is determined through the minimum value.
[0119] Step S255: determining the cluster group corresponding to the current data point to be clustered according to the current cluster center point corresponding to the minimum value, and obtaining the group size of the cluster group;
[0120] Specifically, each data point to be clustered corresponds to a logical block and a logical block address, the data point to be clustered is added to the cluster group corresponding to the second cluster center point, and the group size of the cluster group is obtained, wherein the group size of the cluster group is the number of logical block addresses in the cluster group.
[0121] For example, assuming that if there is a data point C to be clustered and its current cluster center point is A, then the data point C to be clustered is divided into the cluster group corresponding to the current cluster center point A. If there is a data point D to be clustered and its current cluster center point is B, then the data point D to be clustered is divided into the cluster group corresponding to the current cluster center point B, and two cluster groups are obtained. The first cluster group contains the current cluster center point A and the data point C to be clustered, and the second cluster group contains the current cluster center point B and the data point D to be clustered.
[0122] In some embodiments, when the logical unit is a logical page, the grouping size of the cluster grouping can also be determined as the number of logical pages contained in the cluster grouping. It can be understood that the grouping size of each cluster grouping can be adjusted according to actual needs and system configuration.
[0123] Generally speaking, cluster groups with a high access frequency, that is, the first group of hot data, may be relatively small in size because they contain relatively less data and need to be accessed frequently; while cluster groups with a low access frequency, that is, the size of cold data cluster groups, may be relatively large because cold data usually does not need to be accessed frequently and occupies more storage space. It can be understood that hot data refers to data with a high access frequency and has been frequently accessed recently; while cold data refers to data with a low access frequency and has not been accessed for a long time. Once the data is identified as hot data or cold data, the system will assign them to different cluster groups, which are usually called hot data groups, warm data groups, and cold data groups.
[0124] Step S206: adding the data points to be clustered to the cluster grouping, and updating the current cluster center point of the cluster grouping by using a preset clustering algorithm to obtain an updated cluster center point;
[0125] For details, please refer to Figure 5 , Figure 5 yes Figure 2 A detailed flow chart of step S206 in FIG.
[0126] like Figure 5 As shown, step S206: adding the data points to be clustered to the cluster group, and updating the current cluster center point of the cluster group through a preset clustering algorithm to obtain an updated cluster center point, including:
[0127] Step S261: updating the current cluster center point of the cluster grouping by using the first clustering algorithm to obtain an updated cluster center point;
[0128] For details, please refer to Figure 6 , Figure 6 yes Figure 5 A detailed flow chart of step S261 in FIG.
[0129] like Figure 6 As shown, step S261: using the first clustering algorithm, the current cluster center point of the cluster grouping is updated to obtain the updated cluster center point, including:
[0130] Step S2611: Calculate the updated cluster center point according to the data points to be clustered, the current cluster center point and the group size of the cluster group;
[0131] Specifically, the first clustering algorithm includes an algorithm for updating cluster centers based on a dynamic incremental average method, and the calculation formula of the algorithm is as follows:
[0132]
[0133] Among them, μ new is the updated cluster center point, N is the group size of the cluster group, x in is the data point to be clustered, and μ1 is the current cluster center.
[0134] Please refer to Figure 7 , Figure 7 It is a schematic diagram of a flow chart of calculating updated cluster center points when there are removed data points in cluster grouping provided by an embodiment of the present application;
[0135] like Figure 7 As shown in FIG. 1 , when there are removed data points in the cluster grouping, the process of calculating the updated cluster center point includes:
[0136] Step S701: obtaining the current cluster center point of the cluster grouping;
[0137] Specifically, when there are removed data points in the cluster grouping, the current cluster center point of the cluster grouping is obtained.
[0138] Step S702: When there are removed data points in the cluster grouping, the current cluster center point is updated according to the removed data points, the current cluster center point and the grouping size, and the updated cluster center point is calculated;
[0139] Specifically, the removed data points are obtained, and the current cluster center points are updated according to the removed data points, the current cluster center points and the group size, and the updated cluster center points are calculated. The calculation formula is as follows:
[0140]
[0141] Among them, μ new is the updated cluster center point, N is the group size of the cluster group, x out is the data point to be removed, and μ2 is the current cluster center.
[0142] Please refer to Figure 8 , Figure 8 It is a schematic diagram of an overall process of updating cluster centers based on a dynamic incremental averaging method provided in an embodiment of the present application;
[0143] like Figure 8 As shown, step (1): initialize the cluster center and the counter, wherein the counter is used to record the number of received data points, for example, the initial values of the cluster center and the counter are set to zero.
[0144] Step (2): Every time a new data point is received, the cluster center is updated based on the dynamic incremental averaging method to obtain the updated cluster center. Specifically, the cluster center update formula is as follows: mean_new = mean + (x_k-mean) / (n+1), where mean_new represents the updated cluster center, mean is the cluster center before updating, x_k is the new data point, and n represents the number of received data points.
[0145] Step (3): Increase the counter once and use the updated cluster center as the current cluster center.
[0146] Step (4): Determine whether all data points have been accepted. If not, return to step (2) and accept the next data point to update the cluster center. If not, end the data clustering process.
[0147] In the embodiment of the present application, the dynamic incremental averaging method considers both the incoming data points and the outgoing data points. Assume that the mean is μ, the counter is N, and the initial value is 0. When new data x enters the group, N=N+1. When old data is removed from the group, N=N-1.
[0148] Step S262: updating the current cluster center point of the cluster grouping by a second clustering algorithm to obtain an updated cluster center point;
[0149] For details, please refer to Fig. 9 , Fig. 9 yes Figure 5 A detailed flow chart of step S262 in FIG.
[0150] like Fig. 9 As shown, step S262: using a second clustering algorithm, the current cluster center point of the cluster grouping is updated to obtain an updated cluster center point, including:
[0151] Step S2621: Setting the moving window and the window size of the moving window;
[0152] Specifically, a moving window and a window size corresponding to the moving window are set, wherein the window size (W) determines the number of data points considered in the clustering process. It should be noted that the window size of the moving window can be set according to actual needs. This application does not limit the window size. For example, if the data points are densely distributed, a smaller W can be selected to capture data changes more carefully. If the data points are sparsely distributed, a larger W can be selected to find clustering structures in a larger range.
[0153] Step S2622: determining the historical data point to be removed according to the timestamp of the historical data point;
[0154] Specifically, according to the window size of the moving window, determine the first number of historical data points, where the first number is equal to the window size. It can be understood that the first number of historical data points can be randomly selected historical data points, or can be selected according to the size of the timestamp. This application does not limit this. For example, if the window size is set to 100, the first number is 100, and 100 historical data points need to be obtained. Obtain the timestamps of all historical data points in the moving window, and obtain the historical data point corresponding to the minimum timestamp based on the timestamps of all historical data points in the moving window. It can be understood that among all historical data points in the current moving window, this historical data point is the earliest historical data point to enter the moving window. When a new data point enters the moving window, the earliest historical data point to enter the moving window needs to be removed.
[0155] Step S2623: updating the current cluster center point according to the removed historical data points, the current cluster center point and the current data points to be clustered, to obtain an updated cluster center point;
[0156] Specifically, the current cluster center point corresponding to the historical data is determined according to the historical data points and the window size, where the current cluster center point = the sum of the historical data points in the moving window / window size. For example, assuming the window size is w, the number of historical data points in the moving window is w, where the historical data points are x_1, x_2, ..., x_w, respectively. The first cluster center is the mean of all historical data points in the moving window, and its calculation formula is mean = (x_1+x_2+...+x_w) / w, where mean represents the first cluster center, w represents the window size, and x_1, x_2, ..., x_w represent the w historical data points in the moving window.
[0157] According to the removed historical data points, the current cluster center point and the current data points to be clustered, the current cluster center point is updated to obtain the updated cluster center point, specifically including:
[0158] mean new =mean+(x k -x k-w ) / w,
[0159] Among them, mean new is the updated cluster center, mean is the current cluster center, x k is the data point to be clustered, x k-w is the historical data point to be removed, w is the window size, where the current cluster center point = the sum of all historical data points in the moving window / window size.
[0160] In some embodiments, in addition to calculating the mean of historical data points in the moving window, a density estimation algorithm (such as kernel density estimation, DBSCAN) can be used to determine areas with higher density of data points, and the center points of these areas can be selected as the first cluster center to determine the first cluster center.
[0161] Fig.10 It is a schematic diagram of an overall process of updating cluster centers based on a moving sliding window average method provided in an embodiment of the present application;
[0162] like Fig.10 As shown in FIG. 1 , the process of updating the cluster center based on the moving sliding window average method includes:
[0163] Step (1): Set a moving window, initialize the window size of the moving window to w, obtain the first w historical data points, and calculate the average value of the first w historical data points: mean = (x_1+x_2+…+x_w) / w, where mean represents the current initial cluster center, that is, the first cluster center point.
[0164] Step (2): Every time a new data point x_k is accepted, the moving window slides to the right to allow the new data point to enter the moving window, and the earliest accepted data point in the moving window is removed; the cluster center is updated according to the following calculation formula: mean_new = mean + (x_k-x_{kw}) / w, where mean_new represents the updated cluster center, mean represents the current cluster center, x_k represents the new data point, x_{kw} represents the removed data point, and w represents the size of the moving window.
[0165] Step (3): Use the updated cluster center as the current cluster center.
[0166] Step (4): Determine whether all data points have been accepted. If not, return to step (2) and accept the next data point to update the cluster center. If not, end the data clustering process.
[0167] In the embodiment of the present application, in order to reduce memory overhead, the present application further proposes cluster center update based on moving sliding window average. This method maintains a fixed-size window for each cluster. When new data x comes, it is added to the window. If the number of data points in the window exceeds W, the earliest data point in the window is removed to ensure that the window always contains only W recent data points. Then, the calculation formula of the moving sliding window average can be derived:
[0168] Step S263: updating the current cluster center point of the cluster grouping by using the third clustering algorithm to obtain an updated cluster center point;
[0169] For details, please refer to Fig.11 , Fig.11 yes Figure 5 A detailed flow chart of step S263 in FIG.
[0170] like Fig.11 As shown, step S263: using the third clustering algorithm, the current cluster center point of the cluster grouping is updated to obtain the updated cluster center point, including:
[0171] Step S2631: setting a smoothing coefficient;
[0172] For details, please refer to Fig.12 , Fig.12 yes Fig.11 A detailed flow chart of step S2631 in FIG.
[0173] like Fig.12 As shown, step S2631: setting the smoothing coefficient includes:
[0174] Step S6311: Determine a smoothing coefficient according to the group size;
[0175] Specifically, the smoothing coefficient is determined according to the group size of the cluster grouping, wherein the calculation formula of the smoothing coefficient is as follows:
[0176] α=2 / (γ*W+1),
[0177] Among them, α is the smoothing coefficient, γ is the grouping adjustment coefficient, and W is the grouping size.
[0178] Step S2632: updating the current cluster center point according to the smoothing coefficient, the current cluster center point and the current sample point to obtain an updated cluster center point;
[0179] Specifically, according to the smoothing coefficient, the second cluster center point and the current sample point, the calculation formula for determining the updated cluster center point is as follows:
[0180] EMA t =αx t +(1-α)EMA t-1 ,
[0181] Among them, EMA t is the updated cluster center point, x t is the current sample point, EMA t-1 is the second cluster center, α is the smoothing coefficient, and the smoothing coefficient ranges from (0,1). When α is large, the weight decays faster, so the EMA's "memory" is shorter and the influence of historical data decreases rapidly. When α is small, the EMA's memory is longer and the influence of historical data lasts longer.
[0182] For example, based on the calculation formula for determining the updated cluster center point, the specific calculation process is as follows: Assume that the time series is [10, 20, 30, 40, 50], the smoothing coefficient is 0.5, initialize the second cluster center point EMA1 = 10, obtain the current sample point from the time series in the order of timestamp from small to large, and then use the calculation formula EMA of the updated cluster center point t =αx t +(1-α)EMA t-1 , calculate the updated cluster center points:
[0183] When t=2, the current sample point obtained from the time series is 20, then EMA2=0.5*20+(1-0.5)*10=10+5=15;
[0184] When t=3, the current sample point obtained from the time series is 30, then EMA3=0.5*30+(1-0.5)*15=15+7.5=22.5;
[0185] When t=4, the current sample point obtained from the time series is 40, then EMA4=0.5*40+(1-0.5)*22.5=20+11.25=31.25;
[0186] t=n, and so on, continuously obtain the next sample point in the time series, and then continuously update the second cluster center point until all sample points in the sample data are traversed to obtain the updated cluster center point.
[0187] In some embodiments, before acquiring sample data, the grouping adjustment coefficient can be set to an initial value. Preferably, the initial value of the grouping adjustment coefficient can be set to 1. It can be understood that when the grouping adjustment coefficient is 1, the smoothing coefficient calculation formula α=2 / (γ*W+1) in γ*W is equal to the group size of the first group. At this time, the group size is in an unadjusted state. In the subsequent clustering process, the grouping adjustment coefficient can be adjusted to make the algorithm clustering effect better, so that the final clustering result can better reflect the changing trend of the new data.
[0188] In an embodiment of the present application, a differentiated forgetting mechanism can be implemented by adjusting the size of the smoothing coefficient through the grouping adjustment coefficient. The differentiated forgetting mechanism means that the forgetting rate of historical data should be treated differently. When the value of the smoothing coefficient increases, the weight of the current sample point increases, which means that the influence of the new sample data on the updated cluster center point increases, that is, a faster forgetting rate is set for the historical data. Conversely, when the value of the smoothing coefficient decreases, the weight of the current sample point decreases, which means that the influence of the new sample data on the updated cluster center point decreases, that is, a slower forgetting rate is set for the historical data.
[0189] Please refer to Fig.13 , Fig.13 It is a schematic diagram of an overall process of updating cluster centers based on an exponential moving average method provided in an embodiment of the present application;
[0190] like Fig.13 As shown in the figure, the overall process of updating the cluster center based on the exponential moving average method includes:
[0191] Step (1): Set the smoothing coefficient and the group size of the cluster grouping, initialize the smoothing coefficient to initialize the smoothing coefficient to α=2 / (N+1), where α is the smoothing coefficient and N is the group size of the cluster grouping, obtain the historical data point with the smallest timestamp among the historical data points, and determine the historical data point as the first historical data point, and initialize the second cluster center point of the current cluster grouping as the first historical data point.
[0192] Step (2): Every time a new sample point x_k is accepted, the second cluster center is updated according to the following calculation formula: mean_new = α*x_k+(1-α)*mean, where mean_new represents the updated cluster center, mean represents the second cluster center, x_k represents the new sample point, and α represents the smoothing coefficient.
[0193] Step (3): Use the updated cluster center as the current second cluster center.
[0194] Step (4): Determine whether all sample points have been accepted. If not, return to step (2) and accept the next sample point, and then update the second cluster center point. If not, obtain the updated cluster center point and end the data clustering process.
[0195] In the embodiment of the present application, in order to minimize memory overhead, the present application updates the second cluster center by means of exponential moving average (EMA), which gives higher weight to the latest data and gradually reduces the weight to the historical data, thereby more sensitively reflecting the recent changes in the data. i and the second cluster center EMA t-1 , determine the updated cluster center EMA t , where the updated cluster center point = smoothing coefficient * current sample point + (1-smoothing coefficient) * second cluster center point.
[0196] In an embodiment of the present application, the present application provides a data clustering processing method, which is applied to a flash memory device, and the flash memory device includes a plurality of logical units. The method includes: obtaining historical feature data corresponding to the logical unit, and determining an initial cluster center point corresponding to the historical feature data, wherein the historical feature data includes a plurality of historical data points, and the historical data points are feature vectors, and the feature vectors are used to characterize the hotness and coldness of the historical feature data points; according to the initial cluster center, determining a plurality of cluster groups, wherein one initial cluster center corresponds to one cluster group; according to the initial cluster center point, clustering the historical feature data to determine each historical feature data point; The initial cluster center point corresponding to the data point, and the historical feature data with the same initial cluster center are divided into the same cluster group, wherein the hotness or coldness of the historical feature data points in each cluster group is the same or similar; the sample data to be clustered is obtained, wherein the sample data to be clustered includes multiple data points to be clustered; based on the current cluster centers of the multiple cluster groups, the data points to be clustered are clustered to determine the cluster groups corresponding to the data points to be clustered; the data points to be clustered are added to the cluster groups, and the current cluster center points of the cluster groups are updated through a preset clustering algorithm to obtain updated cluster center points.
[0197] By clustering the data points to be clustered based on the current cluster centers of multiple cluster groups, the cluster groups corresponding to the data points to be clustered are determined; the data points to be clustered are added to the cluster groups, and the current cluster center points of the cluster groups are updated through a preset clustering algorithm to obtain updated cluster center points. The present application can eliminate the adverse effects of historical data on the cluster centers, update the initial cluster center points, and obtain updated cluster center points, so that the updated cluster centers can more accurately reflect the characteristics of the logical units and improve the accuracy of data clustering.
[0198] Please refer to Fig.14 , Fig.14 A schematic diagram of the structure of a flash memory device provided in an embodiment of the present application;
[0199] like Fig.14 As shown, the flash memory device 140 includes one or more processors 141 and a memory 142. Fig.14 A processor 141 is taken as an example.
[0200] The processor 141 and the memory 142 may be connected via a bus or other means. Fig.14 The example of connecting through bus is taken in the following.
[0201] Processor 141 is used to provide computing and control capabilities to control the sending end to perform corresponding tasks, for example, to control the sending end to perform the above Figure 2The data clustering processing method in the invention is applied to a flash memory device, the flash memory device includes a plurality of logic units, and the method includes: obtaining historical feature data corresponding to the logic unit, and determining an initial cluster center point corresponding to the historical feature data, wherein the historical feature data includes a plurality of historical data points, the historical data points are feature vectors, and the feature vectors are used to characterize the degree of hotness and coldness of the historical feature data points; determining a plurality of cluster groups according to the initial cluster center, wherein one initial cluster center corresponds to one cluster group; clustering the historical feature data according to the initial cluster center point, and determining an initial cluster center corresponding to each historical data point; The method comprises the following steps: obtaining the sample data to be clustered, wherein the sample data to be clustered includes a plurality of data points to be clustered; clustering the data points to be clustered based on the current cluster centers of the plurality of cluster groups, and determining the cluster groups corresponding to the data points to be clustered; adding the data points to be clustered to the cluster groups, and updating the current cluster center points of the cluster groups through a preset clustering algorithm, and obtaining the updated cluster center points.
[0202] By clustering the data points to be clustered based on the current cluster centers of multiple cluster groups, the cluster groups corresponding to the data points to be clustered are determined; the data points to be clustered are added to the cluster groups, and the current cluster center points of the cluster groups are updated through a preset clustering algorithm to obtain updated cluster center points. The present application can eliminate the adverse effects of historical data on the cluster centers, update the initial cluster center points, and obtain updated cluster center points, so that the updated cluster centers can more accurately reflect the characteristics of the logical units and improve the accuracy of data clustering.
[0203] The processor 141 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a hardware chip or any combination thereof; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The above-mentioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0204] The memory 142, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as program instructions / modules corresponding to the data clustering processing method in the embodiment of the present application. The processor 141 can implement the data clustering processing method in any of the above method embodiments by running the non-transitory software programs, instructions and modules stored in the memory 142. Specifically, the memory 142 may include a volatile memory (VM), such as a random access memory (RAM); the memory 142 may also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD) or other non-transitory solid-state storage device; the memory 142 may also include a combination of the above types of memories.
[0205] The memory 142 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 142 may optionally include a memory remotely arranged relative to the processor 141, and these remote memories may be connected to the processor 141 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0206] One or more modules are stored in the memory 142, and when executed by one or more processors 141, the data clustering processing method in any of the above method embodiments is executed, for example, the data clustering processing method described above is executed. Figure 2 The steps shown.
[0207] In the embodiment of the present application, the flash memory device 140 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The flash memory device 140 may also include other components for realizing device functions, which will not be described in detail here.
[0208] The embodiment of the present application also provides a non-volatile computer-readable storage medium, such as a memory including a program code, and the program code can be executed by a processor to complete the data clustering processing method in the above embodiment. For example, the non-volatile computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), a magnetic tape, a floppy disk, and an optical data storage device.
[0209] The embodiment of the present application also provides a computer program product, which includes one or more program codes, and the program codes are stored in a non-volatile computer-readable storage medium. The processor of the electronic device reads the program code from the non-volatile computer-readable storage medium, and the processor executes the program code to complete the method steps of the data clustering processing method provided in the above embodiment.
[0210] A person skilled in the art will appreciate that all or part of the steps for implementing the above embodiments may be accomplished by hardware or by hardware associated with a program code, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0211] Through the description of the above implementation methods, a person of ordinary skill in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, can also be implemented by hardware. A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the non-volatile computer-readable storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Under the concept of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes in different aspects of the present application as mentioned above, which are not provided in detail for the sake of simplicity. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features can be replaced by equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A data clustering processing method, characterized in that: Applied to a flash memory device, the flash memory device includes a plurality of logic units, and the method includes: Acquire historical feature data corresponding to the logic unit, and determine an initial cluster center point corresponding to the historical feature data, wherein the historical feature data includes a plurality of historical data points, the historical data points are feature vectors, and the feature vectors are used to characterize the hotness or coldness of the historical feature data points; Determine a plurality of cluster groups according to the initial cluster center, wherein one initial cluster center corresponds to one cluster group; According to the initial cluster center point, the historical feature data is clustered to determine the initial cluster center point corresponding to each historical data point, and the historical feature data with the same initial cluster center point are divided into the same cluster group, wherein the hotness or coldness of the historical feature data points in each cluster group is the same or similar; Acquire sample data to be clustered, wherein the sample data to be clustered includes a plurality of data points to be clustered; Based on the current cluster centers of the plurality of cluster groups, clustering the data points to be clustered, and determining the cluster groups corresponding to the data points to be clustered; The data points to be clustered are added to the cluster grouping, and the current cluster center point of the cluster grouping is updated through a preset clustering algorithm to obtain an updated cluster center point.
2. The method according to claim 1, characterized in that The clustering of the data points to be clustered based on the current cluster centers of the plurality of cluster groups to determine the cluster groups corresponding to the data points to be clustered includes: Traversing the data points to be clustered in the sample data to obtain the current data points to be clustered; Calculating the Euclidean distance between the current data point to be clustered and the multiple current cluster center points to obtain multiple distance results; Determine a minimum value among the distance results according to the multiple distance results; According to the minimum value, determining a current cluster center point corresponding to the minimum value; According to the current cluster center point corresponding to the minimum value, the cluster group corresponding to the current data point to be clustered is determined, and the group size of the cluster group is obtained, wherein the group size is the number of logical addresses of the logical units in the cluster group.
3. The method according to claim 2, characterized in that The sample data includes a timestamp of each sample point, and traversing the data points to be clustered in the sample data to obtain the current data points to be clustered includes: Get the timestamp of the data points to be clustered; Determine the time series corresponding to the sample data to be clustered according to the timestamps of the data points to be clustered, wherein the data points to be clustered in the time series are sorted from small to large according to the timestamps; The time series is traversed in the order of timestamps from small to large, and the data points to be clustered are obtained from the time series.
4. The method according to claim 3, characterized in that The preset clustering algorithm includes a first clustering algorithm, a second clustering algorithm, or a third clustering algorithm. The method of updating the current cluster center point of the cluster grouping by the preset clustering algorithm to obtain the updated cluster center point includes: By using the first clustering algorithm, the current cluster center point of the cluster grouping is updated to obtain an updated cluster center point; or, By using the second clustering algorithm, the current cluster center point of the cluster grouping is updated to obtain an updated cluster center point; or, The current cluster center point of the cluster grouping is updated by using the third clustering algorithm to obtain an updated cluster center point.
5. The method according to claim 4, characterized in that By using the first clustering algorithm, the current cluster center point of the cluster grouping is updated to obtain an updated cluster center point, including: Calculating an updated cluster center point according to the data points to be clustered, the current cluster center point and the group size of the cluster grouping includes: Among them, μ new is the updated cluster center point, N is the group size of the cluster group, x in is the data point to be clustered, and μ1 is the current cluster center point.
6. The method according to claim 5, characterized in that The method further comprises: Obtaining the current cluster center point of the cluster grouping; When there are removed data points in the cluster grouping, the current cluster center point is updated according to the removed data points, the current cluster center point and the grouping size, and the updated cluster center point is calculated, which specifically includes: Among them, μ new is the updated cluster center point, N is the group size of the cluster group, x out is the data point to be removed, and μ2 is the current cluster center.
7. The method according to claim 4, characterized in that The updating of the current cluster center point of the cluster grouping by the second clustering algorithm to obtain an updated cluster center point includes: Setting a moving window and a window size of the moving window, wherein the window size is less than or equal to a group size of the cluster grouping; Determine the historical data point to be removed according to the timestamp of the historical data point, wherein the historical data point to be removed is the historical data point with the smallest timestamp in the moving window; The method of updating the current cluster center point according to the removed historical data point, the current cluster center point and the current data point to be clustered to obtain an updated cluster center point specifically includes: mean new =mean+(x k -x k-w ) / w, Among them, mean new is the updated cluster center, mean is the current cluster center, x k is the data point to be clustered, x k-w is the historical data point to be removed, w is the window size, wherein the current cluster center point = the sum of all historical data points in the moving window / the window size.
8. The method according to claim 4, characterized in that The updating of the current cluster center point of the cluster grouping by the third clustering algorithm to obtain the updated cluster center point includes: Set the smoothing factor; The current cluster center point is updated according to the smoothing coefficient, the current cluster center point and the current sample point to obtain an updated cluster center point, specifically including: EMA t =αx t +(1-α)EMA t-1 , Among them, EMA t is the updated cluster center point, α is the smoothing coefficient, x t is the data point to be clustered, EMA t-1 is the current cluster center.
9. The method according to claim 8, characterized in that The setting of the smoothing coefficient includes: According to the group size, a smoothing coefficient is determined, specifically including: α=2 / (γ*N+1), Among them, α is the smoothing coefficient, γ is the grouping adjustment coefficient, and N is the grouping size.
10. A flash memory device, characterized in that: include: A processor and a memory, wherein the processor is used to execute an executable program code in the memory, and when the executable program code is executed, the processor executes instructions of the data clustering processing method as described in any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the data clustering processing method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Hot data recognizing method of solid state disk by fusing various machine learning algorithms
CN106874213A
Clustering method and device, computer equipment and storage medium
CN115344692A
Energy-optimizing placement of resources in data centers
US11875191B1
Systems and methods for visualizing a pattern in a dataset
US20180225416A1
Clustering method and device, and storage medium
WO2018196673A1