Data clustering method and device

By obtaining multiple feature values ​​of data in the data clustering method and subdividing it, the problem of inability to effectively distinguish differences in similar data sets in the prior art is solved, and more efficient data clustering and compression effects are achieved, reducing storage costs.

CN120234633APending Publication Date: 2025-07-01HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410361769.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-03-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing data clustering methods cannot effectively distinguish the differences between different data blocks in similar data sets, resulting in poor clustering effects and unable to improve data compression efficiency.

Method used

By obtaining multiple eigenvalues ​​of the data and dividing data with the same eigenvalue into the same similar data set, further subdividing it according to the total number of eigenvalues, distinguishing data with different redundancy levels, and improving clustering accuracy and compression rate.

Benefits of technology

It improves the accuracy and efficiency of data clustering, enhances the effect of data compression, and reduces storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234633A_ABST
    Figure CN120234633A_ABST
Patent Text Reader

Abstract

The invention discloses a data clustering method and device, and belongs to the technical field of storage. The method comprises the following steps: acquiring multiple pieces of data to be clustered; obtaining a plurality of first feature values of target data, wherein the target data is any one of the plurality of data; and dividing at least two pieces of data with a plurality of correspondingly same first feature values into the same similar data set. According to the method and the device, the data with different similarities can be divided into different similar data sets, so that the accuracy of clustering the similar data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202311850193.5 and the invention title "A Method, Device and Other Equipment for Data Processing" filed on December 28, 2023, the entire content of which is incorporated herein by reference. Technical Field

[0002] This application relates to the field of storage technology, and particularly to a data clustering method and device. Background Art

[0003] Data deduplication and compression technology is the most effective and direct method to reduce storage costs. The deduplication and compression technology based on similar data clustering can cluster similar data distributed in time and space in the system together to obtain a set of similar data, and then compress the data in the set of similar data.

[0004] Currently, similar data is usually clustered based on the feature values of data blocks. Before clustering the data blocks, it is necessary to obtain the feature values of each data block. When multiple data blocks have the same feature value, these multiple data blocks are aggregated together to form a set of similar data.

[0005] However, this way of clustering data is too rough and cannot distinguish the differences between different data blocks in the set of similar data, resulting in poor clustering effect. Summary of the Invention

[0006] This application provides a data clustering method and device. This application can divide data with different similarities into different sets of similar data, improving the accuracy of clustering similar data. The technical solutions provided by this application are as follows:

[0007] In a first aspect, this application provides a data clustering method. The data clustering method includes: obtaining a plurality of data to be clustered; obtaining a plurality of first feature values of a target data, where the target data is any one of the plurality of data; and dividing at least two data with corresponding identical plurality of first feature values into the same set of similar data.

[0008] In this way, at least two data in the same similar data set not only have the same eigenvalue, but also the total number of the same first eigenvalues of the at least two data is equal. Therefore, based on the fact that the data have the same eigenvalue, this solution can further screen according to whether different data have corresponding same eigenvalues, can further subdivide the data according to the similarity degree between the data, treat the data with different redundancies differently, so that there is a large redundancy between the data in the same similar data set, and there is a large difference in the redundancy of the data in different similar data sets, realizing the distinction of the data with different redundancies, being able to divide the data with different similarities into different similar data sets, improving the accuracy of clustering similar data, and improving the clustering effect of clustering the data. Moreover, this solution also helps to improve the efficiency of performing operations on the data based on the clustering results of the data. For example, when compressing the data based on the clustering results, since the higher the similar redundancy of the similar data set, the better the compression effect on the similar data in the similar data set, therefore, this application can utilize the greater redundancy between the data in the same similar data set to improve the compression ratio of compressing the data.

[0009] In a possible implementation manner, before dividing at least two data with corresponding same multiple first eigenvalues into the same similar data set, it includes: counting the total number of corresponding same first eigenvalues of each two data among the multiple data to obtain multiple total values. Correspondingly, dividing at least two data with corresponding same multiple first eigenvalues into the same similar data set includes: performing a clustering process on the multiple data in the order from large to small of the multiple total values, where performing the clustering process according to the i-th total value among the multiple total values includes: among the multiple data to be clustered in the multiple data, dividing the multiple data to be clustered with corresponding same i-th total value of first eigenvalues into a similar data set. In this way, the data can be divided into the same similar data set with other data that has the most corresponding same first eigenvalues with it, which helps to maximize the accuracy of clustering similar data.

[0010] In a possible implementation manner, obtaining multiple first eigenvalues of the target data includes: obtaining multiple data shards from the target data; obtaining the second eigenvalue of the target data shard, where the target data shard is any one of the multiple data shards; processing the second eigenvalues of the multiple data shards to obtain multiple first eigenvalues of the target data, where any one first eigenvalue of the target data is obtained based on the second eigenvalues of some of the multiple data shards, and any two first eigenvalues of the target data are obtained based on the second eigenvalues of different data shards.

[0011] In a possible implementation, obtaining a second eigenvalue of a target data shard includes: obtaining a plurality of data segments from the target data shard; obtaining a modulus value of the target data shard modulo each of the plurality of data segments; and obtaining a second eigenvalue of the target data shard based on the modulus values corresponding to the plurality of data segments.

[0012] In a possible implementation, in response to the target data shard, the data segments, and the modulus value being represented as strings, obtaining a second eigenvalue of the target data shard based on the modulus values corresponding to the plurality of data segments includes: concatenating the strings representing the modulus values corresponding to the plurality of data segments in the order in which the strings representing the plurality of data segments appear in the string representing the target data shard to obtain the second eigenvalue of the target data shard.

[0013] In the implementation of obtaining the second eigenvalue of a data shard in this application, since the hash value of a data shard is obtained based on the modulus value of the data shard modulo a data segment, its computational complexity is low and the amount of computation is small, reducing the computational overhead of obtaining the eigenvalue and ensuring the speed of obtaining the eigenvalue. For example, using the implementation of calculating the second eigenvalue of a data shard in this application, compared with the implementation of randomly selecting a hash value using the Robin hash algorithm, the computational overhead can be reduced from a computational complexity of O(N*N) to O(N), reducing the computational overhead.

[0014] In a possible implementation, before obtaining a second eigenvalue of the target data shard based on the modulus values corresponding to the plurality of data segments, the data clustering method further includes: obtaining an optimization coefficient of a target data segment, where the target data segment is any one of the plurality of data segments; and eliminating a hash conflict of the modulus value corresponding to the target data segment based on the optimization coefficient of the target data segment to obtain an optimized modulus value of the target data segment. By using the optimization coefficient of the data segment to eliminate the hash conflict of the modulus value corresponding to the data segment, the correlation between the plurality of data segments in the data shard can be reduced or even eliminated, thereby reducing or even avoiding the local interference caused by the correlation between the data segments to the clustering result, and further improving the accuracy of the clustering result.

[0015] In a possible implementation, in response to the modulus value and the optimization coefficient being represented as strings, the optimized modulus value is obtained by concatenating the string representing the optimization coefficient and the string representing the modulus value.

[0016] In a possible implementation, in response to the target data shard being represented as a string, multiple data segments are obtained from the target data shard, including: respectively intercepting multiple first character segments from the string used to represent the target data shard to obtain multiple data segments, where each first character segment represents a data segment, the first character segment includes one character or multiple consecutively arranged characters, and the positions of the characters in any two first character segments among the multiple first character segments in the string used to represent the target data shard are different from each other.

[0017] In a possible implementation, in response to the second eigenvalue and the first eigenvalue being represented as strings, the first eigenvalue is obtained by character concatenation based on the string used to represent the multiple second eigenvalues.

[0018] In a possible implementation, before processing the second eigenvalues of multiple data shards to obtain multiple first eigenvalues of the target data, the data clustering method further includes: performing a de-correlation process on the second eigenvalues of the multiple data shards to obtain the second eigenvalues of the multiple data shards after the de-correlation process, and the de-correlation process includes: sorting according to the numerical magnitude. By performing the de-correlation process on the second eigenvalues of the multiple data shards, the correlation relationship between the eigenvalues of the multiple data shards can be reduced or even eliminated, thereby reducing or even avoiding the local interference on the clustering result caused by the correlation relationship between the data shards, and further improving the accuracy of the clustering result. When the second eigenvalues of the multiple data shards are sorted according to the numerical magnitude, the correlation relationship carried by the second eigenvalues of the multiple data shards due to the positions of the multiple data shards in the target data can be broken, so that the local interference on the clustering result caused by the correlation relationship due to the positions of the multiple data shards can be reduced or even avoided.

[0019] In a possible implementation, in response to the target data being represented as a string, multiple data shards are obtained from the target data, including: respectively intercepting multiple second character segments from the string used to represent the target data to obtain multiple data shards, where each second character segment represents a data shard, the second character segment includes multiple consecutively arranged characters, and the positions of the characters in any two second character segments among the multiple second character segments in the string used to represent the target data are different from each other.

[0020] In a second aspect, the present application provides a data clustering device. The data clustering device includes: an acquisition module, configured to acquire multiple data to be clustered; the acquisition module is further configured to acquire multiple first eigenvalues of the target data, where the target data is any one of the multiple data; a clustering module, configured to divide at least two data with corresponding identical multiple first eigenvalues into the same similar data set.

[0021] In a possible implementation manner, the clustering module is specifically configured to: count the total number of pairs of data among multiple data that have the same corresponding first eigenvalue, obtaining multiple total values; perform a clustering process on the multiple data in descending order of the multiple total values, where performing the clustering process according to the i-th total value among the multiple total values includes: among the multiple data to be clustered in the multiple data, dividing the multiple data to be clustered that have the same corresponding i-th total value of first eigenvalues into a similar data set.

[0022] In a possible implementation manner, the obtaining module is specifically configured to: obtain multiple data shards from the target data; obtain the second eigenvalue of the target data shard, where the target data shard is any one of the multiple data shards; process the second eigenvalues of the multiple data shards to obtain multiple first eigenvalues of the target data, where any one first eigenvalue of the target data is obtained based on the second eigenvalues of some of the multiple data shards, and any two first eigenvalues of the target data are obtained based on the second eigenvalues of different data shards.

[0023] In a possible implementation manner, the obtaining module is specifically configured to: obtain multiple data segments from the target data shard; obtain the modulus value of the target data shard modulo each of the multiple data segments; obtain the second eigenvalue of the target data shard based on the modulus values corresponding to the multiple data segments.

[0024] In a possible implementation manner, in response to the target data shard, data segments, and modulus values being represented by strings, the obtaining module is specifically configured to: splice the strings representing the modulus values corresponding to the multiple data segments character by character in the order of the strings representing the multiple data segments in the string representing the target data shard to obtain the second eigenvalue of the target data shard.

[0025] In a possible implementation manner, the obtaining module is further configured to: obtain the optimization coefficient of the target data segment, where the target data segment is any one of the multiple data segments; eliminate the hash conflict of the modulus value corresponding to the target data segment based on the optimization coefficient of the target data segment to obtain the optimized modulus value of the target data segment.

[0026] In a possible implementation manner, in response to the modulus value and the optimization coefficient being represented by strings, the optimized modulus value is obtained by splicing the string representing the optimization coefficient and the string representing the modulus value.

[0027] In a possible implementation, in response to the target data shard being represented as a string, the acquisition module is specifically configured to: respectively intercept a plurality of first character segments from the string representing the target data shard to obtain a plurality of data segments, where each first character segment represents a data segment, the first character segment includes one character or a plurality of continuously arranged characters, and the positions of the characters in any two first character segments among the plurality of first character segments in the string representing the target data shard are different from each other.

[0028] In a possible implementation, in response to the second eigenvalue and the first eigenvalue being represented as strings, the first eigenvalue is obtained by character splicing based on the string representing the plurality of second eigenvalues.

[0029] In a possible implementation, the acquisition module is further configured to: perform a de-correlation process on the second eigenvalues of the plurality of data shards to obtain the second eigenvalues of the plurality of data shards after the de-correlation process, and the de-correlation process includes: sorting according to the numerical size.

[0030] In a possible implementation, in response to the target data being represented as a string, the acquisition module is specifically configured to: respectively intercept a plurality of second character segments from the string representing the target data to obtain a plurality of data shards, where each second character segment represents a data shard, the second character segment includes a plurality of continuously arranged characters, and the positions of the characters in any two second character segments among the plurality of second character segments in the string representing the target data are different from each other.

[0031] In a third aspect, the present application provides a computing device, including a memory and a processor, the memory stores program instructions, and the processor runs the program instructions to execute the method provided in the first aspect of the present application and any of its possible implementations.

[0032] In a fourth aspect, the present application provides a computing device cluster, including a plurality of computing devices, the plurality of computing devices include a plurality of processors and a plurality of memories, the plurality of memories store program instructions, and the plurality of processors run the program instructions, so that the computing device cluster executes the method provided in the first aspect of the present application and any of its possible implementations.

[0033] In a fifth aspect, the present application provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium, and the computer-readable storage medium includes program instructions. When the program instructions run on a computing device, the computing device is caused to execute the method provided in the first aspect of the present application and any of its possible implementations.

[0034] Sixth aspect, the present application provides a computer program product containing instructions. When the computer program product runs on a computer, it causes the computer to execute the method provided in the first aspect of the present application and any possible implementation manner thereof. Description of the Drawings

[0035] Figure 1 is a schematic diagram of an implementation environment related to a data clustering method provided in an embodiment of the present application;

[0036] Figure 2 is a schematic diagram of an implementation environment related to another data clustering method provided in an embodiment of the present application;

[0037] Figure 3 is a schematic diagram of the architecture of a storage system provided in an embodiment of the present application;

[0038] Figure 4 is a flowchart of a data clustering method provided in an embodiment of the present application;

[0039] Figure 5 is a flowchart of obtaining multiple first eigenvalue of target data provided in an embodiment of the present application;

[0040] Figure 6 is a schematic diagram of obtaining multiple data shards provided in an embodiment of the present application;

[0041] Figure 7 is a flowchart of obtaining a second eigenvalue of a target data shard provided in an embodiment of the present application;

[0042] Figure 8 is another schematic diagram of obtaining a second eigenvalue of a target data shard provided in an embodiment of the present application;

[0043] Figure 9 is another flowchart of obtaining multiple first eigenvalue of target data provided in an embodiment of the present application;

[0044] Figure 10 is a flowchart of another data clustering method provided in an embodiment of the present application;

[0045] Figure 11 is a schematic diagram of a clustering result of the present application provided in an embodiment of the present application;

[0046] Figure 12 is a schematic diagram of a clustering result of related art provided in an embodiment of the present application;

[0047] Figure 13 is a kind of based on provided in an embodiment of the present application Figure 11 and Figure 12Schematic diagram for comparing the compression effects of the differential compression and deep compression schemes implemented on the shown clustering results;

[0048] Figure 14 This is a kind provided by an embodiment of the present application Figure 11 and Figure 12 Schematic diagram for comparing the compression effects of the merge compression scheme implemented on the shown clustering results;

[0049] Figure 15 Schematic diagram of the structure of a storage system provided by an embodiment of the present application;

[0050] Figure 16 Schematic diagram of the process of applying the data clustering method of the present application during data writing provided by an embodiment of the present application;

[0051] Figure 17 This is a kind provided by an embodiment of the present application Figure 16 Schematic diagram of the data reading process corresponding to the data writing process;

[0052] Figure 18 Schematic diagram of the structure of a data clustering device provided by an embodiment of the present application;

[0053] Figure 19 Schematic diagram of the structure of a computing device provided by an embodiment of the present application;

[0054] Figure 20 Schematic diagram of the structure of a computing device cluster provided by an embodiment of the present application;

[0055] Figure 21 Schematic diagram of the structure of a computing device cluster provided by an embodiment of the present application. Detailed implementation manners

[0056] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0057] For ease of understanding, the technologies and backgrounds involved in the embodiments of the present application will be explained first below.

[0058] With the continuous expansion of the scale of storage (such as cloud storage) systems, and the higher performance requirements put forward by some new applications and services, this has led to two significant changes in storage systems. One is the rapid expansion of storage capacity, and it is necessary to reduce storage costs to enhance the core competitiveness of storage systems. The other is that applications and services with higher performance requirements need to use storage devices with better performance to meet the performance requirements of such applications and services, which has led to a multiple increase in storage costs. These two aspects of changes both require finding effective methods to reduce the storage costs of storage systems, thereby enhancing the core competitiveness of storage systems.

[0059] Data deduplication and compression technology is an effective and direct method to reduce storage costs. There are two key technical aspects: data deduplication and data compression. By data deduplication and data compression, redundant data in the storage system can be reduced, achieving the effect of data reduction and significantly reducing the storage cost of the storage system. In some scenarios, the data that the storage system needs to process includes a lot of similar data, and the distribution of these similar data in time and space is very discrete and uneven. Considering this characteristic, data deduplication and compression technology can obtain the feature values of the data, and cluster the similar data that is very discrete and unevenly distributed in time and space according to the feature values to obtain a set of similar data, and then use hierarchical data reduction technology to compress the data in the set of similar data. The hierarchical data reduction technology, on the one hand, finds exactly the same data in the set of similar data for deduplication, and on the other hand, compresses the data in the set of similar data through a differential compression scheme to obtain differential blocks, and compresses the differential blocks through a deep compression algorithm to achieve the effect of further improving the reduction rate.

[0060] In this way, an important technical point of data deduplication and compression technology lies in how to cluster similar data. The implementation method of clustering similar data will affect the reduction effect and compression overhead of the deduplication and compression scheme. Currently, when clustering similar data based on the feature values of data blocks, as long as multiple data blocks have the same feature value, these multiple data blocks will be gathered together to form a set of similar data. However, this way of clustering data is too rough to distinguish the differences between different data blocks in the set of similar data, resulting in a poor clustering effect.

[0061] Based on this, an embodiment of the present application provides a data clustering method. This method can obtain a plurality of data to be clustered and obtain a plurality of first feature values of a target data, and then divide at least two data with corresponding identical plurality of first feature values into the same similar data set. The target data is any one of the plurality of data to be clustered. In this way, at least two data in the same similar data set not only have the same feature values, but also the total number of the same first feature values possessed by the at least two data is equal. Therefore, this solution can further screen according to whether different data have corresponding identical feature values on the basis that the data have the same feature values, can further subdivide the data according to the similarity degree between the data, treat data with different redundancies differently, so that the data in the same similar data set have a large redundancy, and the redundancies of the data in different similar data sets have a large difference, realizing the distinction of data with different redundancies, can divide data with different similarities into different similar data sets, improving the accuracy of clustering similar data and the clustering effect of clustering data. Moreover, this solution also helps to improve the efficiency of operating on data based on the clustering result of the data. For example, when compressing data based on the clustering result, since the higher the similarity redundancy of the similar data set, the better the compression effect on the similar data in the similar data set, the present application can utilize the greater redundancy between the data in the same similar data set to improve the compression rate of data compression.

[0062] First, the implementation environment related to a data clustering method provided by an embodiment of the present application will be described below.

[0063] Figure 1 It is a schematic diagram of the implementation environment related to a data clustering method provided by an embodiment of the present application. As Figure 1 shown, the implementation environment includes: a computing device 10. The computing device 10 can execute the data clustering method provided by an embodiment of the present application.

[0064] In one implementation, the data clustering method provided by the embodiments of the present application can be implemented by a computing device 10 running an executable program. For example, the executable program of the data clustering method can be optionally presented in the form of an application installation package. After the application installation package is installed in the computing device 10, the data clustering method can be implemented by running the executable program. At this time, the computing device 10 can be a terminal. The terminal can be a computer, a personal computer, a portable mobile terminal, a multimedia player, an e-book reader, or a wearable device, etc. For example, after the application installation package of the executable program of the data clustering method is installed in the computing device 10, when the computing device 10 needs to obtain the similarity of multiple images, the computing device 10 can run the executable program to implement the data clustering method, and cluster the data representing the multiple images based on the data clustering method, and then obtain the similarity of the multiple images according to the clustering result. For example, multiple images divided into the same set have a large similarity, and multiple images divided into different sets have a small similarity.

[0065] Figure 2 is a schematic diagram of an implementation environment related to another data clustering method provided by the embodiments of the present application. As Figure 2 shown, the implementation environment may further include: a client 20. The client 20 can establish a communication connection with the computing device 10. For example, a communication connection can be established between the client 20 and the computing device 10 through a network. Optionally, the network can be a local area network, the Internet, or other networks, which are not limited in the embodiments of the present application.

[0066] In one possible implementation, the client 20 can be a desktop computer, a laptop computer, a mobile phone, a smart phone, a tablet computer, a multimedia player, a smart home appliance, an artificial intelligence device, a smart wearable device, an e-reader, a smart vehicle-mounted device, or an Internet of Things device, etc.

[0067] The client 20 is used for a user to interact with the computing device 10. In one implementation, the client 20 is used to send an instruction to the computing device 10 according to the user's instruction, and the operation indicated by the instruction depends on the execution of the data clustering method provided in the embodiments of the present application. The computing device 10 is used to execute the operation indicated by the instruction sent by the client 20, and the data clustering method provided in the embodiments of the present application needs to be executed during the execution of this operation. For example, when the computing device 10 is used as a storage node, the computing device 10 is used to store data. When the user needs to store data in the computing device 10, a data write instruction is sent to the computing device 10 through the client 20 to indicate writing data to the computing device 10. After receiving the write instruction, the computing device 10 first obtains the data to be written from the write instruction, then divides the data to be written into multiple data blocks, clusters the multiple data blocks by using the data clustering method provided in the embodiments of the present application to obtain multiple similar data sets, then compresses the data in the similar data sets, and stores the compressed data. At this time, the data clustering method provided in the embodiments of the present application is used in the data compression scenario. Another example is that the client 20 is used to send a data acquisition instruction to the computing device 10 according to the user's instruction to indicate acquiring data from the computing device 10. After receiving the data acquisition instruction, the computing device 10 acquires the data that the client 20 needs to acquire based on the data acquisition instruction, then divides the data into multiple data blocks, clusters the multiple data blocks by using the data clustering method provided in the embodiments of the present application to obtain multiple similar data sets, then compresses the data in the similar data sets, and sends the compressed data to the client 20. At this time, the data clustering method provided in the embodiments of the present application is used in the data transmission scenario.

[0068] In a possible implementation, Figure 2 The shown implementation environment includes a computing device cluster, and the computing device cluster includes multiple computing devices 10. At this time, the computing device 10 can be a server (such as a cloud server), the computing device cluster is a server cluster composed of several servers, or a cloud computing service center is implemented. Among them, a large number of basic resources owned by a cloud service provider are deployed in the cloud computing service center. For example, computing resources, storage resources, network resources, etc. are deployed in the cloud computing service center. The cloud computing service center can use the large number of basic resources to implement the data clustering method provided in the embodiments of the present application.

[0069] When a computing device cluster is implemented through a cloud computing service center, the functions that the computing device cluster can provide by executing the data clustering method provided by the embodiments of the present application can be abstracted by a cloud service provider into a clustering cloud service on a cloud platform. A user can access the cloud platform through a client 20, purchase the clustering cloud service on the cloud platform, and use the clustering cloud service provided by the computing device cluster through the cloud platform. Optionally, the cloud platform can be a cloud platform of a central cloud, a cloud platform of an edge cloud, or a cloud platform including a central cloud and an edge cloud, and the embodiments of the present application do not make specific limitations on it. Moreover, the clustering cloud service can be provided by the cloud platform as a separate cloud service, or the clustering cloud service can be provided as an additional cloud service of other cloud services. For example, the computing device cluster can be optionally a storage system, and the storage system is used to provide a storage cloud service, and the clustering cloud service can be optionally provided as an additional cloud service of the storage cloud service. At this time, if a customer purchases the clustering cloud service as an additional cloud service of the storage cloud service when purchasing the storage cloud service, after the client 20 instructs to store data in the computing device cluster, before storing the data in the computing device cluster, the computing device cluster first clusters the data by using the clustering cloud service, and then stores the clustered data by using the storage cloud service.

[0070] Figure 3 It is a schematic diagram of the architecture of a storage system provided by the embodiments of the present application. Optionally, the storage system can be a distributed storage system. As Figure 3 shown, the storage system includes: a service layer, an index layer, and a persistence layer.

[0071] The service layer is used to provide a unified interface protocol service to users. The services provided by the service layer can include: elastic volume service (EVS, also known as cloud disk), object storage service (OBS), scalable file service (SFS), data lake insight (DLI) service, data warehouse service (DWS). Moreover, in order to ensure service performance, a cache layer is also configured in the service layer.

[0072] The index layer is used to provide metadata management services for the distributed storage system. The index layer can run a database (DB) and perform deduplication and compression processing, and interact with the service layer through object guidance and file guidance. Among them, the database can be a key-value database (KVDB).

[0073] The persistence layer is used to provide persistent storage services for a distributed system. The persistence layer can achieve write-optimized and read-optimized through the ishard mode and the PLOG mode. The ishard mode and the PLOG mode can share a storage pool. The storage pool can be a storage pool for data function virtualisation (DFV). And the storage medium of the storage pool can be non-volatile memory (NVM), solid state drive (SSD), hard disk drive (HDD), optical storage medium, etc.

[0074] It should be understood that the above content is an exemplary description of the application scenario of the data clustering method provided by the embodiments of the present application, and does not constitute a limitation on the application scenario of the data clustering method. Those of ordinary skill in the art know that with the change of business requirements, its application scenario can be adjusted according to application requirements. For example, the data clustering method provided by the embodiments of the present application can also be applied to the field of network transmission. By clustering similar data in the data to be transmitted and then compressing the clustered similar data, the amount of data transmitted over the network can be reduced and the network transmission rate can be improved. Another example is that the data clustering method provided by the embodiments of the present application can also be applied to various compression fields, such as the field of image compression, the field of video compression, and the field of dedicated database compression, etc., so as to first cluster the data to be compressed through the clustering function provided by the embodiments of the present application, thereby reducing the time delay of data compression and decompression, improving the efficiency of data compression and decompression, and reducing the resource consumption of compression and decompression. In addition, when the method is applied to other scenarios, it is not limited to using Figure 1 or Figure 2 the implementation environment shown. For example, when the method is applied to a data transmission scenario, the implementation environment may optionally include multiple computing devices that can transmit data to each other. The embodiments of the present application do not make specific limitations on this.

[0075] Next, taking the data clustering method provided by the embodiments of the present application applied to Figure 2 the application scenario shown as an example, the implementation process of the data clustering method provided by the embodiments of the present application will be described. Figure 4 is a flowchart of a data clustering method provided by the embodiments of the present application. As Figure 4 shown, the data clustering method includes the following steps:

[0076] Step 401, obtain multiple data to be clustered.

[0077] In this application, the data to be clustered can be various data with clustering requirements. Clustering multiple data refers to the process of dividing multiple data into different similar data sets according to the similarity between the multiple data. Multiple data divided into the same similar data set have greater similarity. Multiple data divided into different similar data sets have less similarity.

[0078] The multiple data to be clustered can be optionally obtained by dividing the data recorded in a data record. At this time, the multiple data to be clustered belong to different parts of the data recorded in the one data record. In a possible implementation manner, the multiple data to be clustered can be optionally obtained by dividing the data to be compressed. The data to be compressed can be optionally data that needs to perform operations such as storage or transmission. The data that needs to be operated during a compression or transmission operation performed by a computing device can be divided into multiple data, and the multiple data are the multiple data to be clustered. For example, before a storage system stores a certain data to be written, deduplication and compression processing can be performed on the data to be written. During the deduplication and compression processing of the data to be written, the storage system can divide the data to be written into multiple data blocks, cluster the multiple data blocks, and then perform deduplication and compression on the data in units of the similar data sets obtained by clustering. Before performing these operations on the data, clustering the data can optimize the execution process of these operations according to the clustering results. For example, by clustering the data to be stored and then compressing the clustered data, the amount of data to be stored can be reduced, thereby reducing the storage resources occupied by the data to be stored and improving the utilization rate of the storage resources. By clustering the data to be transmitted and then compressing the clustered data, the amount of data to be transmitted can be reduced, thereby reducing the amount of data to be transmitted and improving the transmission rate. It should be noted that the computing device can optionally divide the data to be compressed in a fixed-length or non-fixed-length manner to obtain the multiple data to be clustered. Or, the computing device can also use other methods to divide the data recorded in a data record to obtain the multiple data to be clustered, and the embodiments of the present application do not make specific limitations on this.

[0079] Or, the multiple data to be clustered can also be data that does not belong to the data recorded in a data record. For example, the multiple data can be optionally independent data, and there is no relationship between the multiple data. By way of example, the multiple data can represent the data of multiple images for which the similarity between the images needs to be obtained. For example, when obtaining the similarity between multiple images based on the clustering results, if the clustering results indicate that multiple images divided into the same set have greater similarity and multiple images divided into different sets have less similarity, then the multiple data can be optionally the data used to represent the multiple images, and the data used to represent any one image is one data to be clustered.

[0080] Step 402: Obtain multiple first eigenvalues of the target data, where the target data is any one of multiple data to be clustered.

[0081] In a possible implementation, the computing device may optionally obtain multiple first eigenvalues of the data by sampling. For example, the computing device first obtains multiple data shards of the target data, then respectively obtains second eigenvalues of the multiple data shards, and obtains multiple first eigenvalues of the target data based on the second eigenvalues of the multiple data shards. As Figure 5 shown, the implementation process of this step 402 includes:

[0082] Step 4021: Obtain multiple data shards from the target data.

[0083] The computing device may optionally obtain partial data of the target data from the target data in multiple times, and use the partial data obtained each time as a data shard of the target data. In a possible implementation, in response to the target data being represented as a string, obtaining multiple data shards from the target data includes: respectively intercepting multiple second character segments from the string representing the target data to obtain multiple data shards. Wherein, each second character segment represents a data shard. The second character segment includes multiple continuously arranged characters, and the positions of the characters in any two second character segments in the string representing the target data are different. The positions of the characters in any two second character segments in the string representing the target data are different, including: any two second character segments do not include characters at the same position. Optionally, the computing device may optionally intercept multiple character segments of the same length in the string representing the target data at a fixed step length, and each character segment represents a data shard. As Figure 6 shown, the multiple data to be clustered include data block V1 and data block V2. When the computing device obtains multiple data shards from data block V1, it intercepts multiple character segments C1, C2, C3, and C4 of the same length in data block V1 at a fixed step length, and the multiple character segments of the same length are the multiple data shards obtained from data block V1. When the computing device obtains multiple data shards from data block V2, it intercepts multiple character segments C21, C22, C23, and C24 of the same length in data block V2 at a fixed step length, and the multiple character segments of the same length are the multiple data shards obtained from data block V2.

[0084] Step 4022: Obtain the second eigenvalue of the target data shard, where the target data shard is any one of the multiple data shards of the target data.

[0085] In a possible implementation, the computing device may optionally obtain the second eigenvalue of the data shard by sampling. For example, the computing device first obtains multiple data segments of the target data shard, and then obtains the second eigenvalue of the target data shard based on the multiple data segments. As Figure 7 shown, the implementation process of step 4022 includes:

[0086] Step 4022a: Obtain multiple data segments from the target data shard.

[0087] The computing device may optionally obtain a part of the target data shard in multiple times and use the part of the data as a data segment of the target data shard. In a possible implementation, in response to the target data shard being represented by a string, obtaining multiple data segments from the target data shard includes: respectively intercepting multiple first character segments from the string representing the target data shard to obtain multiple data segments. Wherein, each first character segment represents a data segment, the first character segment includes one character or multiple continuously arranged characters, and the positions of the characters in any two first character segments in the string representing the target data shard are different. The positions of the characters in any two first character segments in the string representing the target data shard are different, including: any two first character segments do not include characters at the same position. Optionally, the computing device may optionally intercept multiple character segments of the same length in the string representing the target data shard at a fixed step size, and each character segment represents a data segment. As Figure 6 shown, in data block V1 and data block V2. When the computing device obtains multiple data segments from data shard C3, multiple character segments of the same length are intercepted in data shard C3 at a fixed step size (such as Figure 6 the boxes filled with slashes in), and the multiple character segments of the same length are the multiple data segments obtained from data shard C3.

[0088] Step 4022b: Obtain the modulus value of the target data shard modulo each data segment in the multiple data segments.

[0089] In the field of computers, taking the modulus of a by b is used to find the remainder of the division of a by b. In this application, taking the modulus of a data shard with respect to a data segment means taking the modulus using the asic code value of the data. That is, taking the modulus of a data shard with respect to a data segment is to take the modulus of the asic code value of the data shard with the asic code value of the data segment. For example, assuming the data shard is represented as abcdefg and the data segment is abc, the modulus value obtained by taking the modulus of the data shard abcdefg with respect to the data segment abc is defg. After the target data shard takes the modulus with respect to each data segment, a modulus value can be obtained, and this modulus value is the modulus value corresponding to that data segment. Then, after the target data shard takes the modulus with respect to each of the multiple data segments, multiple modulus values corresponding one-to-one to the multiple data segments can be obtained.

[0090] Step 4022c: Based on the modulus values corresponding to the multiple data segments, obtain the second feature value of the target data shard.

[0091] After obtaining the multiple modulus values of the target data shard taking the modulus with respect to the multiple data segments, the computing device can obtain the second feature value of the target data shard based on these multiple modulus values. For example, the computing device performs a combination process on the multiple modulus values to obtain the second feature value of the second data shard. In a possible implementation manner, in response to the target data shard, data segment, and modulus value being represented as strings, based on the modulus values corresponding to the multiple data segments, obtaining the second feature value of the target data shard includes: arranging the strings representing the modulus values corresponding to the multiple data segments in the order of the strings representing the multiple data segments in the string representing the target data shard for character concatenation to obtain the second feature value of the target data shard. For example, assuming the target data shard is represented as abcdefg, according to the left-to-right arrangement order of the characters in the string representing the target data shard, multiple data segments are obtained in sequence, and the modulus values corresponding to the multiple data segments are defg, efg, and aeg respectively. Then, arranging the strings representing the modulus values corresponding to the multiple data segments in the order of the strings representing the multiple data segments in the string representing the target data shard for character concatenation, the second feature value of the target data shard obtained is defgefgaeg.

[0092] Optionally, before obtaining the second feature value of the target data shard based on the modulus values corresponding to the multiple data segments, it is optional to first optimize the modulus values corresponding to the multiple data segments. At this time, the modulus value used to obtain the second feature value of the target data shard is the optimized modulus value. That is Figure 8 As shown, the implementation process of this step 4022c includes: Step 4022c1: Based on the optimized modulus values corresponding to the multiple data segments, obtain the second feature value of the target data shard.

[0093] Exemplarily, such as Figure 8As shown, before obtaining the second eigenvalue of the target data shard based on the modulus values corresponding to multiple data segments, the method further includes step 4022d and step 4022e.

[0094] Step 4022d: Obtain the optimization coefficient of the target data segment, where the target data segment is any one of the multiple data segments of the target data shard.

[0095] Optimize the modulus value corresponding to the target data segment based on the optimization coefficient of the target data segment, aiming to eliminate the hash conflict of the modulus value corresponding to the target data segment. That is, the optimization coefficient of the target data segment is the coefficient used to eliminate the hash conflict of the modulus value corresponding to the target data segment. There are various ways to obtain this optimization coefficient. In a possible implementation, when using an existing hash algorithm (such as the Robin hash algorithm), the optimization coefficient is the value used to eliminate the hash conflict of the hash value of the data in this hash algorithm. Then, the optimization coefficient can be obtained by executing this existing hash algorithm on the target data segment and determining the value used to eliminate the hash conflict of the hash value of the target data segment in this existing hash algorithm as the optimization coefficient of the target data segment. Optionally, before executing step 4022d, the computing device can calculate the optimization coefficient of each data in this way for a large number of preset data in advance and establish the corresponding relationship between the data and its optimization coefficient. When executing this step 4022d, optionally query this corresponding relationship according to the target data segment to obtain the optimization coefficient corresponding to the target data segment. And for a certain data, there may be multiple values of the optimization coefficient calculated for this data during multiple calculations. Then, the computing device can establish the corresponding relationship between the optimization coefficient and the data based on the optimization coefficient used most frequently during these multiple calculations. In this way, by pre-establishing the corresponding relationship between the optimization coefficient and the data, it is not necessary to calculate the optimization coefficient corresponding to the target data segment online when executing step 4022d, which can reduce the computational complexity and computational load of the data clustering process and help improve the efficiency of clustering the data.

[0096] Step 4022e: Based on the optimization coefficient of the target data segment, eliminate the hash conflict of the modulus value corresponding to the target data segment to obtain the optimized modulus value corresponding to the target data segment.

[0097] After obtaining the optimization coefficient of the target data segment, the modulus value of the target data segment can be processed using this optimization coefficient to eliminate the hash conflict of the modulus value corresponding to the target data segment, and obtain the optimized modulus value corresponding to the target data segment. In a possible implementation manner, in response to the modulus value and the optimization coefficient being represented by strings, the optimized modulus value is obtained by character splicing based on the string representing the optimization coefficient and the string representing the modulus value. Exemplarily, the optimized modulus value is the string obtained by splicing the string representing the optimization coefficient after the string representing the modulus value. For example, assuming that the modulus value corresponding to the target data segment is defg and the optimization coefficient of the target data segment is ace, the optimized modulus value corresponding to the target data segment is defgace.

[0098] When multiple data segments of a data shard are obtained by intercepting data in the data shard, due to the correlation between characters at different positions in the data shard, multiple data segments of the data shard may also have a correlation. Therefore, by using the optimization coefficient of the data segment to eliminate the hash conflict of the modulus value corresponding to the data segment, the correlation between multiple data segments in the data shard can be reduced or even eliminated, thereby reducing or even avoiding the local interference caused by the correlation between data segments to the clustering result, and further improving the accuracy of the clustering result.

[0099] The implementation process of step 4022 can be regarded as a process of obtaining the hash value of the data shard, and the implementation manner of this step 4022 can be regarded as an implementation manner of a hash algorithm provided by an embodiment of the present application. In this implementation manner, since the hash value of the data shard is obtained according to the modulus value obtained by taking the modulus of the data shard with respect to the data segment, its calculation complexity is low and the amount of calculation is small, reducing the calculation overhead of obtaining the eigenvalue and ensuring the speed of obtaining the eigenvalue. For example, using the implementation manner of calculating the second eigenvalue of the data shard in the present application, compared with the implementation manner of calculating the hash value using the Robin hash algorithm, the calculation overhead can be reduced from the calculation complexity of 0(N*N) to 0(N), reducing the calculation overhead.

[0100] Step 4023: Process the second eigenvalues of multiple data shards to obtain multiple first eigenvalues of the target data. Any one of the first eigenvalues of the target data is obtained based on the second eigenvalues of some of the multiple data shards, and any two of the first eigenvalues of the target data are obtained based on the second eigenvalues of different data shards.

[0101] During the execution of step 4023, the second eigenvalues of multiple data shards can be grouped in advance to obtain multiple eigenvalue groups, and each eigenvalue group includes multiple second eigenvalues. The multiple eigenvalue groups correspond one-to-one with the multiple first eigenvalues, that is, each first eigenvalue among the multiple first eigenvalues is obtained based on the multiple second eigenvalues in the eigenvalue group corresponding to this first eigenvalue. In the embodiments of the present application, there are various optional principles for grouping the second eigenvalues of multiple data shards, which are determined according to application requirements. By way of example, when the positions of multiple data shards in the target data are different, after obtaining the second eigenvalues of multiple data shards, the second eigenvalues of multiple data shards can be sorted according to this position, and then starting from the second eigenvalue of the data shard ranked first, the second eigenvalues of every specified number of data shards in the sorting result are divided into one group, so as to obtain multiple eigenvalue groups. For example, assuming that one eigenvalue of each of the 12 data shards is obtained through the foregoing steps, according to the positions of these 12 data shards in the target data, a sorting of 12 eigenvalues can be obtained, and then starting from the second eigenvalue ranked first, the second eigenvalues of every 4 second eigenvalues in the sorting result are divided into one group, so as to obtain 3 eigenvalue groups. Among them, the first eigenvalue group includes the second eigenvalues ranked 1st to 4th, the second eigenvalue group includes the second eigenvalues ranked 5th to 8th, and the third eigenvalue group includes the second eigenvalues ranked 9th to 12th.

[0102] In a possible implementation manner, in response to the second eigenvalue and the first eigenvalue being represented by strings, the first eigenvalue can be optionally obtained by character splicing based on the string representing the multiple second eigenvalues in the eigenvalue group corresponding to this first eigenvalue. Continuing with the above example, the first eigenvalue group includes the second eigenvalues ac, ab, bc, and cd, then the first eigenvalue corresponding to this first eigenvalue group is acabccd. The second eigenvalue group includes the second eigenvalues ac, abc, bcd, and cd, then the first eigenvalue corresponding to this first eigenvalue group is acabcbcdcd. The third eigenvalue group includes the second eigenvalues abc, ab, abc, and cd, then the first eigenvalue corresponding to this first eigenvalue group is abcababccd.

[0103] Optionally, the second eigenvalue used when obtaining the multiple first eigenvalues of the target data can also be the second eigenvalue after de-correlation processing. Then as Figure 9 shown, this step 4023 includes: step 4023a, processing the second eigenvalues after de-correlation processing of multiple data shards to obtain multiple first eigenvalues of the target data. For the implementation process, please refer to the relevant description in step 4023 accordingly, and details are not described here again.

[0104] By way of example, as Figure 9As shown, before processing the second eigenvalues of multiple data shards to obtain multiple first eigenvalues of the target data, the method further includes: Step 4024, performing a de-correlation process on the second eigenvalues of the multiple data shards to obtain the de-correlated second eigenvalues of the multiple data shards.

[0105] There are multiple implementation manners for performing the de-correlation process on the second eigenvalues of the multiple data shards. This application illustrates it by taking one implementation manner as an example. For example, the de-correlation process includes: sorting according to the numerical size.

[0106] When multiple data segments of a data shard are obtained by intercepting data in the data shard, due to the correlation relationship between characters at different positions in the data shard, the multiple second eigenvalues of the multiple data shards may also have a correlation relationship. Therefore, by performing a de-correlation process on the second eigenvalues of the multiple data shards, the correlation relationship between the eigenvalues of the multiple data shards can be reduced or even eliminated, thereby reducing or even avoiding the local interference caused by the correlation relationship between the data shards to the clustering result, and further improving the accuracy of the clustering result. When the second eigenvalues of the multiple data shards are sorted according to the numerical size, the correlation relationship carried by the positions of the multiple data shards in the target data for the second eigenvalues of the multiple data shards can be broken, so that the local interference caused by the correlation relationship due to the positions of the multiple data shards to the clustering result can be reduced or even avoided.

[0107] Step 403, partitioning at least two data with corresponding identical multiple first eigenvalues into the same similar data set.

[0108] That at least two data have corresponding identical multiple first eigenvalues means that for any two data among the at least two data, the two data are data a and data b, and the multiple first eigenvalues of data a and the multiple first eigenvalues of data b are in one-to-one correspondence and equal. For example, data a and data b have 3 corresponding identical eigenvalues. Suppose the 3 eigenvalues of data a are a1, a2, and a3 respectively, and the 3 eigenvalues of data b are b1, b2, and b3 respectively. Then that data a and data b have 3 corresponding identical eigenvalues means that eigenvalue a1 is equal to eigenvalue b1, eigenvalue a2 is equal to eigenvalue b2, and eigenvalue a3 is equal to eigenvalue b3.

[0109] By dividing at least two data with corresponding identical multiple first eigenvalues into the same similar data set, at least two data in the same similar data set not only have the same eigenvalues, but also the total number of the same first eigenvalues of the at least two data is equal. Therefore, based on the fact that the data have the same eigenvalues, this solution can further screen according to whether different data have corresponding identical eigenvalues, further subdivide the data according to the similarity degree between the data, differentially treat data with different redundancies, so that there is a large redundancy between the data in the same similar data set, and there are large differences in the redundancies of the data in different similar data sets, realizing the distinction of data with different redundancies, being able to divide data with different similarities into different similar data sets, and improving the accuracy of clustering similar data.

[0110] Optionally, this step 403 may be executed sequentially according to the total number of corresponding identical first eigenvalues of different data. For example, as Figure 10 shown, before this step 403, the method further includes: step 404, counting the total number of corresponding identical first eigenvalues of every two data among multiple data to obtain multiple total values.

[0111] The implementation process of this step 404 includes: comparing the multiple first eigenvalues of every two data among multiple data. For any two data, when it is determined each time that the two data have a pair of identical eigenvalues, adding one to the total number of corresponding identical first eigenvalues of the two data until the comparison of all the first eigenvalues of the two data is completed. By analogy, the computing device can obtain the total number of corresponding identical first eigenvalues of every two data among multiple data according to this implementation logic.

[0112] When the data clustering method of this application further includes step 404, as Figure 10 shown, the implementation manner of step 403 includes: step 4031, performing a clustering process on multiple data in the order from largest to smallest of the multiple total values. Among them, performing a clustering process according to the i-th total value among the multiple total values includes: among the multiple data to be clustered among multiple data, dividing the multiple data to be clustered with corresponding identical i-th total value of first eigenvalues into one similar data set. In this way, the data can be divided into the same similar data set as the other data with which it has the most corresponding identical first eigenvalues, which helps to maximize the accuracy of clustering similar data.

[0113] Exemplarily, assume that multiple data are respectively data block 1, data block 2, data block 3, and data block 4. Data block 1 has 5 first eigenvalue, which are SFP1, SFP2, SFP3, SFP4, and SFP5 respectively. Data block 2 has 5 first eigenvalue, which are SFP1, SFP2, SFP3, SFP6, and SFP7 respectively. Data block 3 has 2 first eigenvalue, which are SFP1 and SFP8 respectively. Data block 4 has 2 first eigenvalue, which are SFP9 and SFP1 respectively. Then it can be obtained that the total number of corresponding identical first eigenvalue between data block 1 and data block 2 is 3, the total number of identical first eigenvalue between data block 1 and data block 3 is 0, the total number of identical first eigenvalue between data block 1 and data block 4 is 1, the total number of identical first eigenvalue between data block 2 and data block 3 is 1, and the total number of identical first eigenvalue between data block 2 and data block 4 is 1. Thus, when multiple total values are 3 and 1 respectively, in the process of step 4031, it is necessary to first perform the clustering process according to the total value of 3, and then perform the clustering process according to the total value of 1. When performing the clustering process according to the total value of 3, the computing device searches for multiple data blocks with corresponding identical 3 first eigenvalue among multiple data blocks, and obtains that data block 1 and data block 2 have corresponding identical 3 first eigenvalue, and the 3 first eigenvalue are SFP1, SFP2, and SFP3 respectively. Then data block 1 and data block 2 are divided into the same similar data set. When performing the clustering process according to the total value of 1, since data block 1 and data block 2 have completed clustering, the computing device searches for multiple data blocks with identical 1 first eigenvalue among the remaining data blocks 3 and 4 to be clustered, and obtains that data block 3 and data block 4 have corresponding identical 1 first eigenvalue, and the 1 first eigenvalue are SFP1 respectively. Then data block 3 and data block 4 are divided into the same similar data set. Figure 11 is the clustering result of this clustering process. As Figure 11 shown, similar data set 1 includes data block 1 and data block 2, and data block 1 and data block 2 have corresponding identical three first eigenvalue. Similar data set 2 includes data block 3 and data block 4, and data block 3 and data block 4 have identical one first eigenvalue.

[0114] In the related art, when clustering data block pairs, as long as multiple data blocks have the same eigenvalue, these multiple data blocks will be gathered together to form a similar data set. Then, after clustering the above data block 1, data block 2, data block 3, data block 4, and data block 5 using this related art, since these 5 data blocks have the first eigenvalue SFP1, data block 1, data block 2, data block 3, data block 4, and data block 5 are divided into the same similar data set.

[0115] The Figure 11 and Figure 12Compared with the clustering result of Figure 11 the clustering result of Figure 12 subdivides one similar data set in Figure 12 into two similar data sets. The number of data blocks in these two similar data sets with the same corresponding first eigenvalue is different. It is equivalent to dividing Figure 12 a similar data set into two similar data sets with different similar redundancies. At this time, when compressing data blocks based on the clustering result, compared with compressing based on Figure 11 the similar data set in

[0116] Figure 13 shows Figure 11 and Figure 12 the comparison of the compression effects of differential compression and deep compression implemented on the shown clustering results. Since Figure 11 the clustering result of Figure 12 can effectively divide the similar data set into multiple similar data sets with different data redundancies according to the data redundancy, it can significantly improve the compression effect of the compression scheme. When performing differential compression based on Figure 11 the clustering result of Figure 14 shown, it is very likely that only the four data blocks represented by SFP1 can be aggregated and compressed together, resulting in a very limited compression effect. In contrast, when performing differential compression based on

[0117] Further, before performing a clustering process on multiple data in descending order of multiple total values, it is optional to first perform a rough screening on the multiple data to obtain multiple data sets, and then perform a clustering process on the multiple data in each data set in descending order of multiple total values. By way of example, it is optional to first divide multiple data with the same eigenvalue into the same data set, and then perform a clustering process on the multiple data in each data set in descending order of multiple total values. Continuing with the above example, since data blocks 1 to 5 all have the eigenvalue SFP1, data blocks 1 to 5 can be first divided into a data set, and then for data blocks 1 to 5 included in this data set, a clustering process is performed on the multiple data in descending order of multiple total values.

[0118] In summary, in the data clustering method of the present application, the method can obtain multiple data to be clustered, and obtain multiple first eigenvalues of the target data, and then divide at least two data with corresponding identical multiple first eigenvalues into the same similar data set. The target data is any one of the multiple data to be clustered. In this way, at least two data in the same similar data set not only have the same eigenvalue, but also the total number of the same first eigenvalues of the at least two data is equal. Therefore, this solution can further screen according to whether different data have corresponding identical eigenvalues on the basis that the data have the same eigenvalue, can further subdivide the data according to the similarity degree between the data, differentially treat data with different redundancies, so that the data in the same similar data set have a large redundancy, and the redundancies of the data in different similar data sets have a large difference, realizing the distinction of data with different redundancies, and can divide data with different similarities into different similar data sets, improving the accuracy of clustering similar data and the clustering effect of clustering data. Moreover, this solution also helps to improve the efficiency of performing operations on data based on the clustering result of the data. For example, when compressing data based on the clustering result, since the higher the similarity redundancy of the similar data set, the better the compression effect on the similar data in the similar data set, the present application can utilize the greater redundancy between the data in the same similar data set to improve the compression ratio of data compression.

[0119] In addition, in some special scenarios (such as memory pools, vector databases, key-value databases (KVDB), data warehouses, and network transmissions, etc.), there are also problems of huge amounts of redundant data. The embodiments of the present application can also be used in these special scenarios. On the one hand, it can improve the compression effect based on the clustering effect of the present application, and on the other hand, it can reduce the overhead of compression and decompression.

[0120] Exemplarily, when applying the data clustering method of the present application to the scenario of writing data to a storage system, the storage system can use the data clustering method of the present application to cluster the data to be written, and then perform data deduplication and compression based on the clustering result. For example, Figure 15 is a schematic structural diagram of a storage system provided by an embodiment of the present application. As Figure 15 shown, the storage system is a distributed storage system. The storage system may include a metadata management module, a compression module, and a persistence module, and the metadata management module, the compression module, and the persistence module are implemented by different computing devices. The metadata management module is used to receive a user's input / output (I / O) request and perform a management process on metadata based on the I / O request. The compression module is used to perform merge compression and / or de-merge compression on data. The persistence module is used to store the received data to achieve the persistence of the data. The metadata management module includes a metadata management unit (such as LunMap) and a metadata clustering unit (such as DusClient). The metadata management unit is used to record the metadata of all data stored on the storage medium of the storage system, and the metadata includes index information of the storage medium. The metadata clustering unit is used to interface with the metadata management unit, the compression module, and the persistence module, calculate the fingerprint of each data block, and calculate the eigenvalue of each data block by using the data clustering method provided by the present application, then store the obtained fingerprint and eigenvalue in the persistence module, and is responsible for combining the obtained fingerprint and eigenvalue with other metadata information of the data block and sending it to the compression module. The compression module includes an aggregation unit (such as an opportunistic table (OpTable)), an analysis unit (such as a postdata analysis (PDA) unit), a task allocation unit (such as task mag), a compression unit (such as a postdata reduce (PDR) unit), and an index information storage unit (such as FPTable).

[0121] As Figure 15 and Figure 16As shown, during the process of writing data to the storage system, the aggregation unit is used to receive fingerprints, feature values, etc. sent by the metadata processing unit in the metadata management module, and cluster the data according to the feature values using the data clustering method provided in this application to obtain aggregated similar data, provide the aggregated similar data to the compression unit, and provide the physical address and logical index number of the similar data to the compression unit. The analysis unit is used to regularly analyze the aggregated similar data provided by the aggregation unit, and assemble the similar data scattered throughout the storage system into a similar data chain. The task allocation unit is used to receive the similar data chain sent by the analysis unit, form different compression tasks according to the characteristics of the dispersion of the feature values in the storage system, and allocate compression tasks to the compression unit. The compression unit is used to perform data deduplication compression according to the allocated compression tasks, persist the compressed data set to the persistence module, generate index information for each data block in the compressed data set, and provide the index information to the index information storage unit. Optionally, before performing compression processing on the data, the compression unit can first find the identical data on the similar data chain through fingerprint comparison, keep only one copy of multiple identical data, and then perform compression processing on the remaining data. The index information storage unit is used to receive the index information for each data block in the data set sent by the compression unit, persist the index information, and update the index information recorded in the metadata management unit based on it.

[0122] When it is necessary to read a target data block from Figure 15 the storage system as shown, the storage system can first obtain the index information of the target data block according to the data read request, then obtain the target data block from the target data set according to the index information, and then feedback the target data block to the data read request. For example, as Figure 15 and Figure 17 shown, during the process of reading data from the storage system, after receiving the data read request, the metadata management module can obtain the index information of the target data block indicated by the data read request and the target data set where it is located from the metadata management unit, and provide the index information to the compression module. The compression module can obtain the compressed target data set where the target data block is located from the storage medium according to the index information, perform decompression processing on the target data set, then obtain the target data block from the decompressed target data set, and feedback the target data block to the metadata management module, thereby completing the process of reading the target data block. However, when the target data set cannot be obtained based on the index information obtained from the metadata management unit, as Figure 17 shown, the metadata management module can obtain the new index information of the target data block and the target data set where it is located from the index information storage unit, and then obtain the target data block according to the new index information.

[0123] It should be noted that the order of the steps of the data clustering method provided in the embodiments of the present application can be appropriately adjusted, and the steps can also be increased or decreased accordingly according to the situation. Any person skilled in the art in the technical field disclosed in the present application can easily think of a changed method, which should be covered by the protection scope of the present application, so it will not be elaborated here.

[0124] The data clustering method of the embodiments of the present application is introduced above. Corresponding to the above method, the embodiments of the present application also provide a data clustering device. Figure 18 It is a schematic structural diagram of a data clustering device provided by an embodiment of the present application. Based on Figure 18 the following multiple modules shown, the Figure 18 data clustering device shown can execute all or part of the operations shown above Figure 4 、 Figure 5 and Figure 10 . It should be understood that the data clustering device may include more additional modules than the modules shown or omit some of the modules shown, and the embodiments of the present application do not limit this. Optionally, the data clustering device can be configured on a cloud platform. As Figure 18 shown, the data clustering device 180 includes:

[0125] An acquisition module 1801, configured to acquire a plurality of data to be clustered.

[0126] The acquisition module 1801 is further configured to acquire a plurality of first feature values of target data, where the target data is any one of the plurality of data.

[0127] A clustering module 1802, configured to divide at least two data with corresponding identical plurality of first feature values into the same similar data set.

[0128] In a possible implementation manner, the clustering module 1802 is specifically configured to: count the total number of corresponding identical first feature values for each two of the plurality of data to obtain a plurality of total values; perform a clustering process on the plurality of data in descending order of the plurality of total values, where performing the clustering process according to the i-th total value among the plurality of total values includes: among the plurality of data to be clustered in the plurality of data, dividing the plurality of data to be clustered with the corresponding identical i-th total value of first feature values into a similar data set.

[0129] In a possible implementation, the obtaining module 1801 is specifically configured to: obtain multiple data shards from target data; obtain a second eigenvalue of a target data shard, where the target data shard is any one of the multiple data shards; process the second eigenvalues of the multiple data shards to obtain multiple first eigenvalues of the target data, where any one of the first eigenvalues of the target data is obtained based on the second eigenvalues of some of the multiple data shards, and any two of the first eigenvalues of the target data are obtained based on the second eigenvalues of different data shards.

[0130] In a possible implementation, the obtaining module 1801 is specifically configured to: obtain multiple data segments from a target data shard; obtain a modulus value of the target data shard modulo each of the multiple data segments; and obtain a second eigenvalue of the target data shard based on the modulus values corresponding to the multiple data segments.

[0131] In a possible implementation, in response to the target data shard, the data segments, and the modulus value being represented by strings, the obtaining module 1801 is specifically configured to: splice the strings representing the modulus values corresponding to the multiple data segments character by character in the order of the strings representing the multiple data segments in the string representing the target data shard to obtain the second eigenvalue of the target data shard.

[0132] In a possible implementation, the obtaining module 1801 is further configured to: obtain an optimization coefficient of a target data segment, where the target data segment is any one of the multiple data segments; and eliminate a hash conflict of the modulus value corresponding to the target data segment based on the optimization coefficient of the target data segment to obtain an optimized modulus value of the target data segment.

[0133] In a possible implementation, in response to the modulus value and the optimization coefficient being represented by strings, the optimized modulus value is obtained by splicing the string representing the optimization coefficient and the string representing the modulus value.

[0134] In a possible implementation, in response to the target data shard being represented by a string, the obtaining module 1801 is specifically configured to: respectively intercept multiple first character segments from the string representing the target data shard to obtain multiple data segments, where each first character segment represents a data segment, the first character segment includes one character or multiple continuously arranged characters, and the positions of the characters in any two of the multiple first character segments in the string representing the target data shard are different.

[0135] In a possible implementation, in response to the second eigenvalue and the first eigenvalue being represented by strings, the first eigenvalue is obtained by splicing the strings representing the multiple second eigenvalues.

[0136] In a possible implementation, the obtaining module 1801 is further configured to: perform a de-correlation process on the second eigenvalue of multiple data shards to obtain the second eigenvalue of the multiple data shards after the de-correlation process, where the de-correlation process includes: sorting according to the numerical value size.

[0137] In a possible implementation, in response to the target data being represented by a string, the obtaining module 1801 is specifically configured to: respectively intercept multiple second character segments from the string used to represent the target data to obtain multiple data shards, where each second character segment represents a data shard, the second character segment includes multiple continuously arranged characters, and the positions of the characters in any two second character segments in the string used to represent the target data are different.

[0138] Here, for the detailed working processes of the obtaining module 1801 and the clustering module 1802, please refer to the descriptions in the foregoing method embodiments. For example, the obtaining module 1801 obtains multiple data to be clustered by using the foregoing step 401, and obtains multiple first eigenvalues of the target data by using the foregoing step 402. The clustering module 1802 divides at least two data with corresponding identical multiple first eigenvalues into the same similar data set by using the foregoing step 403. The embodiments of the present application will not be described repeatedly herein.

[0139] Among them, both the obtaining module 1801 and the clustering module 1802 can be implemented by software or can be implemented by hardware. Exemplarily, next, taking the obtaining module 1801 as an example, the implementation manner of the obtaining module 1801 will be introduced. Similarly, the implementation manner of the clustering module 1802 can refer to the implementation manner of the obtaining module 1801.

[0140] As an example of a software functional unit, the obtaining module 1801 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the foregoing computing instance may be one or more. For example, the obtaining module 1801 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running this code may be distributed in the same region, or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone (AZ), or may be distributed in different AZs, and each AZ includes one cloud data center or multiple geographically close cloud data centers. Usually, one region may include multiple AZs.

[0141] Similarly, multiple hosts / virtual machines / containers for running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Usually, one VPC is set up within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set up within each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0142] As an example of a hardware functional unit, the acquisition module 1801 may include at least one computing device, such as a server. Alternatively, the acquisition module 1801 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0143] The multiple computing devices included in the acquisition module 1801 can be distributed in the same region or in different regions. The multiple computing devices included in the acquisition module 1801 can be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the acquisition module 1801 can be distributed within the same VPC or across multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0144] It should be noted that in other embodiments, either the acquisition module 1801 or the clustering module 1802 can be used to execute any step in the data clustering method. The steps to be implemented by the acquisition module 1801 and the clustering module 1802 can be specified as needed. The full function of the data clustering device is achieved by implementing different steps in the data clustering method through the acquisition module 1801 and the clustering module 1802 respectively.

[0145] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described respective components can refer to the corresponding content in the foregoing method embodiments and will not be elaborated herein.

[0146] The basic hardware structure involved in the embodiments of the present application will be exemplified below.

[0147] The present application also provides a computing device 1900. As Figure 19 shown, the computing device 1900 includes: a bus 1902, a processor 1904, a memory 1906, and a communication interface 1908. The processor 1904, the memory 1906, and the communication interface 1908 communicate with each other through the bus 1902. The computing device 1900 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1900.

[0148] The bus 1902 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 19 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 1902 may include a path for transmitting information between various components of the computing device 1900 (for example, the memory 1906, the processor 1904, and the communication interface 1908).

[0149] The processor 1904 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0150] The memory 1906 may include a volatile memory, such as a random access memory (RAM). The processor 1904 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0151] The executable program code is stored in the memory 1906, and the processor 1904 executes the executable program code to implement the functions of the aforementioned acquisition module 1801 and clustering module 1802 respectively, so as to implement the data clustering method of the present application. That is, the memory 1906 stores instructions for executing the data clustering method of the present application.

[0152] The communication interface 1908 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 1900 and other devices or communication networks.

[0153] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0154] As Figure 20 shown, the computing device cluster includes at least one computing device 1900. The memory 1906 in one or more computing devices 1900 in the computing device cluster can store the same instructions for executing the data clustering method of the present application.

[0155] In some possible implementation manners, the memory 1906 of one or more computing devices 1900 in the computing device cluster can also store partial instructions for executing the data clustering method of the present application respectively. In other words, the combination of one or more computing devices 1900 can jointly execute the instructions for executing the data clustering method of the present application.

[0156] It should be noted that the memories 1906 in different computing devices 1900 in the computing device cluster can store different instructions, which are respectively used to execute partial functions of the data clustering device of the present application. That is, the instructions stored in the memories 1906 of different computing devices 1900 can implement the functions of one or more modules among the acquisition module 1801 and the clustering module 1802.

[0157] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. Among them, the network can be a wide area network or a local area network, etc. Figure 21 Shows a possible implementation manner. As Figure 21As shown, two computing devices 1900A and 1900B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the instructions for implementing the functions of the acquisition module 1801 are stored in the memory 1906 of the computing device 1900A. At the same time, the instructions for implementing the functions of the clustering module 1802 are stored in the memory 1906 of the computing device 1900B.

[0158] Figure 21 The connection method between the computing device clusters shown can be considered. Since the data clustering method provided in this application requires a large amount of data storage, it is considered to hand over the functions implemented by the clustering module 1802 to the computing device 1900B for execution.

[0159] It should be understood that Figure 21 the functions of the computing device 1900A shown in can also be completed by multiple computing devices 1900. Similarly, the functions of the computing device 1900B can also be completed by multiple computing devices 1900.

[0160] The embodiments of this application also provide another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to Figure 20 and Figure 21 the connection method of the computing device cluster. The difference is that the memory 1906 in one or more computing devices 1900 in this computing device cluster may store the same instructions for executing the data clustering method.

[0161] In some possible implementation manners, the memory 1906 of one or more computing devices 1900 in this computing device cluster may also separately store partial instructions for executing the data clustering method. In other words, a combination of one or more computing devices 1900 can jointly execute the instructions for executing the data clustering method.

[0162] The embodiments of this application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the data clustering method.

[0163] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the data clustering method or direct the computing device to execute the data clustering method.

[0164] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.

[0165] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the present application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions. For example, the original data and executable code involved in the present application are obtained under full authorization.

[0166] In the embodiments of the present application, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. The term "at least one" means one or more, and the term "a plurality" means two or more, unless otherwise clearly defined.

[0167] The term "and / or" in the present application is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data clustering method, characterized in that: The method comprises: Obtain multiple data to be clustered; Acquire multiple first feature values ​​of target data, where the target data is any one of the multiple data; At least two data having corresponding identical first feature values ​​are divided into the same similar data set.

2. The method according to claim 1, characterized in that Before dividing at least two data having the same corresponding multiple first feature values ​​into the same similar data set, the method includes: Counting the total number of first characteristic values ​​corresponding to each two data in the plurality of data to obtain a plurality of total number values; The step of dividing at least two data having the same corresponding multiple first feature values ​​into the same similar data set includes: A clustering process is performed on the multiple data in a descending order of the multiple total values, wherein the clustering process is performed according to the i-th total value among the multiple total values, including: among the multiple data to be clustered in the multiple data, the multiple data to be clustered having corresponding to the same i-th total value of the first eigenvalue are divided into a similar data set.

3. The method according to claim 1 or 2, characterized in that The step of obtaining a plurality of first characteristic values ​​of the target data includes: Acquire multiple data slices from the target data; Obtaining a second characteristic value of a target data slice, where the target data slice is any one of the multiple data slices; The second characteristic values ​​of the multiple data slices are processed to obtain multiple first characteristic values ​​of the target data, wherein any first characteristic value of the target data is obtained based on the second characteristic values ​​of some data slices among the multiple data slices, and any two first characteristic values ​​of the target data are obtained based on the second characteristic values ​​of different data slices.

4. The method according to claim 3, characterized in that The obtaining of the second characteristic value of the target data slice includes: Acquire multiple data segments from the target data slice; Obtaining a modulus value of the target data slice modulo each data segment of the multiple data segments; Based on the module values ​​corresponding to the multiple data segments, a second characteristic value of the target data slice is obtained.

5. The method according to claim 4, characterized in that In response to the target data slice, the data segment, and the modulus value being represented by a character string, obtaining a second characteristic value of the target data slice based on the modulus values ​​corresponding to the multiple data segments includes: According to the order of the character strings representing the multiple data segments in the character string representing the target data slice, the character strings representing the module values ​​corresponding to the multiple data segments are concatenated to obtain the second characteristic value of the target data slice.

6. The method according to claim 4 or 5, characterized in that Before acquiring the second characteristic value of the target data slice based on the modulus values ​​corresponding to the multiple data segments, the method further includes: Obtaining an optimization coefficient of a target data segment, where the target data segment is any one of the multiple data segments; Based on the optimization coefficient of the target data segment, the hash conflict of the module value corresponding to the target data segment is eliminated to obtain the optimized module value of the target data segment.

7. The method according to claim 6, characterized in that In response to the modulus value and the optimization coefficient being represented by a character string, the optimized modulus value is obtained by concatenating characters of the character string used to represent the optimization coefficient and the character string used to represent the modulus value.

8. The method according to any one of claims 4 to 6, characterized in that: In response to the target data slice being represented by a character string, obtaining a plurality of data segments from the target data slice includes: A plurality of first character segments are respectively extracted from a character string used to represent the target data slice to obtain the plurality of data segments, wherein each first character segment represents a data segment, the first character segment includes a character or a plurality of characters arranged continuously, and the positions of the characters in any two first character segments among the plurality of first character segments in the character string used to represent the target data slice are different.

9. The method according to any one of claims 3 to 8, characterized in that: In response to the second feature value and the first feature value being represented by a character string, the first feature value is obtained by concatenating characters of the character string used to represent a plurality of the second feature values.

10. The method according to any one of claims 3 to 9, characterized in that: Before processing the second characteristic values ​​of the plurality of data slices to obtain the plurality of first characteristic values ​​of the target data, the method further includes: A decorrelation process is performed on the second eigenvalues ​​of the multiple data slices to obtain the decorrelation-processed second eigenvalues ​​of the multiple data slices, wherein the decorrelation process includes: sorting according to numerical values.

11. The method according to any one of claims 3 to 10, characterized in that: In response to the target data being represented by a character string, obtaining a plurality of data slices from the target data includes: A plurality of second character segments are respectively extracted from the character string used to represent the target data to obtain the plurality of data slices, wherein each second character segment represents a data slice, the second character segment includes a plurality of characters arranged continuously, and the positions of the characters in any two of the plurality of second character segments in the character string used to represent the target data are different.

12. A data clustering device, characterized in that: The device comprises: An acquisition module, used for acquiring multiple data to be clustered; The acquisition module is further used to acquire multiple first characteristic values ​​of target data, where the target data is any one of the multiple data; The clustering module is used to classify at least two data having corresponding identical first characteristic values ​​into the same similar data set.

13. The device according to claim 12, characterized in that The clustering module is specifically used for: Counting the total number of first characteristic values ​​corresponding to each two data in the plurality of data to obtain a plurality of total number values; A clustering process is performed on the multiple data in a descending order of the multiple total values, wherein the clustering process is performed according to the i-th total value among the multiple total values, including: among the multiple data to be clustered in the multiple data, the multiple data to be clustered having corresponding to the same i-th total value of the first eigenvalue are divided into a similar data set.

14. The device according to claim 12 or 13, characterized in that The acquisition module is specifically used for: Acquire multiple data slices from the target data; Obtaining a second characteristic value of a target data slice, where the target data slice is any one of the multiple data slices; The second characteristic values ​​of the multiple data slices are processed to obtain multiple first characteristic values ​​of the target data, wherein any first characteristic value of the target data is obtained based on the second characteristic values ​​of some data slices among the multiple data slices, and any two first characteristic values ​​of the target data are obtained based on the second characteristic values ​​of different data slices.

15. The device according to claim 14, characterized in that The acquisition module is specifically used for: Acquire multiple data segments from the target data slice; Obtaining a modulus value of the target data slice modulo each data segment of the multiple data segments; Based on the module values ​​corresponding to the multiple data segments, a second characteristic value of the target data slice is obtained.

16. The device according to claim 15, characterized in that In response to the target data slice, the data segment and the modulus value being represented by a character string, the acquisition module is specifically configured to: According to the order of the character strings representing the multiple data segments in the character string representing the target data slice, the character strings representing the module values ​​corresponding to the multiple data segments are concatenated to obtain the second characteristic value of the target data slice.

17. The device according to claim 15 or 16, characterized in that Get module, also used for: Obtaining an optimization coefficient of a target data segment, where the target data segment is any one of the multiple data segments; Based on the optimization coefficient of the target data segment, the hash conflict of the module value corresponding to the target data segment is eliminated to obtain the optimized module value of the target data segment.

18. The device according to claim 17, characterized in that In response to the modulus value and the optimization coefficient being represented by a character string, the optimized modulus value is obtained by concatenating characters of the character string used to represent the optimization coefficient and the character string used to represent the modulus value.

19. The device according to any one of claims 15 to 17, characterized in that: In response to the target data slice being represented by a character string, the acquisition module is specifically configured to: A plurality of first character segments are respectively extracted from a character string used to represent the target data slice to obtain the plurality of data segments, wherein each first character segment represents a data segment, the first character segment includes a character or a plurality of characters arranged continuously, and the positions of the characters in any two first character segments among the plurality of first character segments in the character string used to represent the target data slice are different.

20. The device according to any one of claims 14 to 19, characterized in that In response to the second feature value and the first feature value being represented by a character string, the first feature value is obtained by concatenating characters of the character string used to represent a plurality of the second feature values.

21. The device according to any one of claims 14 to 20, characterized in that The acquisition module is further used for: A decorrelation process is performed on the second eigenvalues ​​of the multiple data slices to obtain the decorrelation-processed second eigenvalues ​​of the multiple data slices, wherein the decorrelation process includes: sorting according to numerical values.

22. The device according to any one of claims 14 to 21, characterized in that In response to the target data being represented by a character string, the acquisition module is specifically configured to: A plurality of second character segments are respectively extracted from the character string used to represent the target data to obtain the plurality of data slices, wherein each second character segment represents a data slice, the second character segment includes a plurality of characters arranged continuously, and the positions of the characters in any two of the plurality of second character segments in the character string used to represent the target data are different.

23. A computing device cluster, characterized in that: The method comprises a plurality of computing devices, wherein the plurality of computing devices comprises a plurality of processors and a plurality of memories, wherein program instructions are stored in the plurality of memories, and the plurality of processors execute the program instructions, so that the computing device cluster executes any one of the methods described in claims 1 to 11.

24. A computer-readable storage medium, characterized in that: The method comprises program instructions, and when the program instructions are executed on a computing device, the computing device is caused to execute the method according to any one of claims 1 to 11.

25. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 11.