Data clustering method and device
By obtaining multiple feature values of data in the data clustering method and subdividing it, the problem of inability to effectively distinguish differences in similar data sets in the prior art is solved, the clustering accuracy and compression rate are improved, and the storage cost is reduced.
Patent Information
- Application Number
- PCT/CN2024/118054
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2024-09-10
- Publication Date
- 2025-07-03
AI Technical Summary
Existing data clustering methods cannot effectively distinguish the differences between different data blocks in similar data sets, resulting in poor clustering effects and affecting data compression and storage efficiency.
By obtaining multiple eigenvalues of the data and dividing data with the same eigenvalue into the same similar data set, further subdividing is done according to the number and similarity of the eigenvalues, distinguishing data with different redundancy degrees, and improving clustering accuracy and compression rate.
It improves the accuracy and efficiency of data clustering, enhances data compression effect, and reduces storage costs.
Smart Images

Figure CN2024118054_03072025_PF_FP_ABST
Abstract
Description
Data clustering method and device
[0001] This application claims priority to Chinese patent application No. 202311850193.5 filed on December 28, 2023, with invention name “A method, device and other equipment for data processing”, and priority to Chinese patent application No. 202410361769.X filed on March 27, 2024, with invention name “Data clustering method and device”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of storage technology, and in particular to a data clustering method and device. Background Art
[0003] Data deduplication and compression technology is the most effective and direct way to reduce storage costs. Deduplication and compression technology, based on clustering similar data, can cluster similar data distributed across time and space in the system to form similar data sets, which are then compressed.
[0004] Currently, similar data is usually clustered based on the feature values of data blocks. Before clustering data blocks, the feature values of each data block need to be obtained. When multiple data blocks have the same feature values, the multiple data blocks are clustered together to form a similar data set.
[0005] However, this method of clustering data is too rough and cannot distinguish the differences between different data blocks in similar data sets, resulting in poor clustering effect.
[0006] Summary of the Invention
[0007] This application provides a data clustering method and device. This application can divide data with different similarities into different similar data sets, thereby improving the accuracy of clustering similar data. The technical solutions provided by this application are as follows:
[0008] In a first aspect, the present application provides a data clustering method. The data clustering method includes: obtaining multiple data to be clustered; obtaining multiple first eigenvalues of target data, where the target data is any one of the multiple data; and grouping at least two data having the same corresponding multiple first eigenvalues into the same similar data set.
[0009] In this way, at least two data in the same similar data set not only have the same eigenvalue, but also the total number of the same first eigenvalues of the at least two data is equal. Therefore, the scheme can, on the basis of the data having the same eigenvalue, further screen according to whether different data have the corresponding same eigenvalue, and can further subdivide the data according to the degree of similarity between the data, and treat data with different redundancies differently, so that the data in the same similar data set have a larger redundancy, and the redundancy of the data in different similar data sets has a large difference, thereby achieving the distinction between data with different redundancies, and being able to divide data with different similarities into different similar data sets, thereby improving the accuracy of clustering similar data and improving the clustering effect of clustering data. In addition, the scheme also helps to improve the efficiency of performing operations on data based on the clustering results of the data. For example, when compressing data based on the clustering results, since the higher the similarity redundancy of the similar data set, the better the compression effect of the similar data in the similar data set, the present application can utilize the greater redundancy between data in the same similar data set to improve the compression rate of the data.
[0010] In one possible implementation, before dividing at least two data with corresponding identical multiple first eigenvalues into the same similar data set, the method includes: counting the total number of each two data in the multiple data with corresponding identical first eigenvalues to obtain multiple total values. Accordingly, dividing at least two data with corresponding identical multiple first eigenvalues into the same similar data set includes: performing a clustering process on the multiple data in descending order of the multiple total values, wherein the clustering process is performed according to the i-th total value in the multiple total values, including: among the multiple data to be clustered in the multiple data, the multiple data to be clustered with the corresponding i-th total value of the first eigenvalue are divided into a similar data set. In this way, the data and other data with which it has the most corresponding identical first eigenvalues can be divided into the same similar data set, which helps to maximize the accuracy of clustering similar data.
[0011] In one possible implementation, obtaining multiple first characteristic values of target data includes: obtaining multiple data slices from the target data; obtaining a second characteristic value of the target data slice, where the target data slice is any one of the multiple data slices; processing the second characteristic values of the multiple data slices to obtain multiple first characteristic values of the target data, wherein any first characteristic value of the target data is obtained based on the second characteristic values of some data slices in the multiple data slices, and any two first characteristic values of the target data are obtained based on the second characteristic values of different data slices.
[0012] In one possible implementation, obtaining the second characteristic value of the target data shard includes: obtaining multiple data segments from the target data shard; obtaining a modulus value of the target data shard modulo each of the multiple data segments; and obtaining the second characteristic value of the target data shard based on the modulus values corresponding to the multiple data segments.
[0013] In one possible implementation, in response to the target data slice, data segment and modulus value being represented by a string, a second characteristic value of the target data slice is obtained based on the modulus values corresponding to multiple data segments, including: according to the order of the strings used to represent the multiple data segments in the string used to represent the target data slice, the strings used to represent the modulus values corresponding to the multiple data segments are concatenated to obtain the second characteristic value of the target data slice.
[0014] In the implementation method of obtaining the second eigenvalue of the data shard in the present application, since the hash value of the data shard is obtained based on the modulus value of the data shard modulo the data segment, its computational complexity is low and the amount of computation is small, which reduces the computational overhead of obtaining the eigenvalue and can ensure the speed of obtaining the eigenvalue. For example, the implementation method of calculating the second eigenvalue of the data shard using the present application can reduce the computational overhead from the computational complexity of 0(N*N) to 0(N) compared to the implementation method of using the Robin hash algorithm to select the hash value, thereby reducing the computational overhead.
[0015] In one possible implementation, before obtaining the second eigenvalue of a target data segment based on the modulus values corresponding to the multiple data segments, the data clustering method further includes: obtaining an optimization coefficient for the target data segment, where the target data segment is any one of the multiple data segments; and eliminating hash conflicts of the modulus values corresponding to the target data segment based on the optimization coefficient of the target data segment, thereby obtaining an optimized modulus value for the target data segment. By using the optimization coefficient of the data segment to eliminate hash conflicts of the modulus values corresponding to the data segment, the correlation between the multiple data segments in the data segment can be reduced or even eliminated, thereby reducing or even avoiding local interference on the clustering results caused by the correlation between the data segments, thereby improving the accuracy of the clustering results.
[0016] In a possible implementation, in response to the modulus value and the optimization coefficient being represented by a character string, the optimized modulus value is obtained by concatenating the character string used to represent the optimization coefficient and the character string used to represent the modulus value.
[0017] In one possible implementation, in response to the target data slice being represented by a character string, multiple data segments are obtained from the target data slice, including: respectively intercepting multiple first character segments from the character string used to represent the target data slice to obtain multiple data segments, wherein each first character segment represents a data segment, the first character segment includes one character or multiple characters arranged continuously, and the positions of the characters in any two of the multiple first character segments in the character string used to represent the target data slice are different.
[0018] In a possible implementation, in response to the second eigenvalue and the first eigenvalue being represented by character strings, the first eigenvalue is obtained by concatenating characters in character strings used to represent a plurality of second eigenvalues.
[0019] In one possible implementation, before processing the second eigenvalues of multiple data slices to obtain multiple first eigenvalues of target data, the data clustering method further includes: performing decorrelation processing on the second eigenvalues of multiple data slices to obtain the second eigenvalues of multiple data slices after decorrelation processing, and the decorrelation processing includes: sorting according to the size of the numerical value. By performing decorrelation processing on the second eigenvalues of multiple data slices, the correlation between the eigenvalues of multiple data slices can be reduced or even eliminated, thereby reducing or even avoiding the local interference of the clustering results caused by the correlation between the data slices, thereby improving the accuracy of the clustering results. After sorting the second eigenvalues of multiple data slices according to the size of the numerical value, the correlation carried by the second eigenvalues of multiple data slices due to the positions of the multiple data slices in the target data can be broken, thereby reducing or even avoiding the local interference of the clustering results caused by the correlation between the positions of the multiple data slices due to the positions of the multiple data slices.
[0020] In one possible implementation, in response to the target data being represented by a string, multiple data slices are obtained from the target data, including: respectively intercepting multiple second character segments from the string used to represent the target data to obtain multiple data slices, wherein each second character segment represents a data slice, the second character segment includes multiple characters arranged continuously, and the positions of the characters in any two of the multiple second character segments in the string used to represent the target data are different.
[0021] In a second aspect, the present application provides a data clustering device. The data clustering device includes: an acquisition module for acquiring multiple data to be clustered; the acquisition module is further configured to acquire multiple first eigenvalues of target data, where the target data is any one of the multiple data; and a clustering module for grouping at least two data having corresponding identical first eigenvalues into the same similar data set.
[0022] In one possible implementation, the clustering module is specifically used to: count the total number of data with the same first eigenvalue corresponding to each two data in multiple data to obtain multiple total values; perform a clustering process on the multiple data in descending order of the multiple total values, wherein the clustering process is performed according to the i-th total value in the multiple total values, including: among the multiple data to be clustered in the multiple data, the multiple data to be clustered with the same i-th total value first eigenvalue are divided into a similar data set.
[0023] In one possible implementation, the acquisition module is specifically used to: obtain multiple data slices from the target data; obtain the second characteristic value of the target data slice, where the target data slice is any one of the multiple data slices; process the second characteristic values of the multiple data slices to obtain multiple first characteristic values of the target data, where any first characteristic value of the target data is obtained based on the second characteristic values of some data slices in the multiple data slices, and any two first characteristic values of the target data are obtained based on the second characteristic values of different data slices.
[0024] In one possible implementation, the acquisition module is specifically used to: obtain multiple data segments from the target data slice; obtain the modulus value of the target data slice modulo each data segment in the multiple data segments; and obtain the second eigenvalue of the target data slice based on the modulus values corresponding to the multiple data segments.
[0025] In one possible implementation, in response to the target data slice, data segment and modulus value being represented by a string, the acquisition module is specifically used to: according to the order of the strings used to represent the multiple data segments in the string used to represent the target data slice, perform character concatenation on the strings used to represent the modulus values corresponding to the multiple data segments to obtain the second characteristic value of the target data slice.
[0026] In one possible implementation, the acquisition module is also used to: obtain the optimization coefficient of the target data segment, where the target data segment is any one of the multiple data segments; based on the optimization coefficient of the target data segment, eliminate the hash conflict of the module value corresponding to the target data segment to obtain the optimized module value of the target data segment.
[0027] In a possible implementation, in response to the modulus value and the optimization coefficient being represented by a character string, the optimized modulus value is obtained by concatenating the character string used to represent the optimization coefficient and the character string used to represent the modulus value.
[0028] In one possible implementation, in response to the target data slice being represented by a string, the acquisition module is specifically used to: respectively intercept multiple first character segments from the string used to represent the target data slice to obtain multiple data segments, wherein each first character segment represents a data segment, the first character segment includes one character or multiple characters arranged continuously, and the positions of the characters in any two first character segments among the multiple first character segments in the string used to represent the target data slice are different.
[0029] In a possible implementation, in response to the second eigenvalue and the first eigenvalue being represented by character strings, the first eigenvalue is obtained by concatenating characters in character strings used to represent a plurality of second eigenvalues.
[0030] In a possible implementation, the acquisition module is further used to: perform decorrelation processing on the second eigenvalues of the multiple data slices to obtain the decorrelation-processed second eigenvalues of the multiple data slices, and the decorrelation processing includes: sorting according to numerical value size.
[0031] In one possible implementation, in response to the target data being represented by a string, the acquisition module is specifically used to: respectively intercept multiple second character segments from the string used to represent the target data to obtain multiple data fragments, wherein each second character segment represents a data fragment, the second character segment includes multiple characters arranged continuously, and the positions of the characters in any two second character segments in the multiple second character segments in the string used to represent the target data are different.
[0032] In a third aspect, the present application provides a computing device comprising a memory and a processor, wherein the memory stores program instructions, and the processor runs the program instructions to execute the method provided in the first aspect of the present application and any possible implementation thereof.
[0033] In a fourth aspect, the present application provides a computing device cluster, comprising multiple computing devices, wherein the multiple computing devices include multiple processors and multiple memories, wherein program instructions are stored in the multiple memories, and the multiple processors execute the program instructions, so that the computing device cluster executes the method provided in the first aspect of the present application and any possible implementation thereof.
[0034] In a fifth aspect, the present application provides a computer-readable storage medium, which is a non-volatile computer-readable storage medium. The computer-readable storage medium includes program instructions. When the program instructions are executed on a computing device, the computing device executes the method provided in the first aspect of the present application and any possible implementation thereof.
[0035] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when run on a computer, enables the computer to execute the method provided in the first aspect of the present application and any possible implementation thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] FIG1 is a schematic diagram of an implementation environment involved in a data clustering method provided in an embodiment of the present application;
[0037] FIG2 is a schematic diagram of an implementation environment involved in another data clustering method provided in an embodiment of the present application;
[0038] FIG3 is a schematic diagram of the architecture of a storage system provided in an embodiment of the present application;
[0039] FIG4 is a flow chart of a data clustering method provided in an embodiment of the present application;
[0040] FIG5 is a flow chart of obtaining multiple first characteristic values of target data provided by an embodiment of the present application;
[0041] FIG6 is a schematic diagram of obtaining multiple data slices provided in an embodiment of the present application;
[0042] FIG7 is a flow chart of obtaining a second characteristic value of a target data slice according to an embodiment of the present application;
[0043] FIG8 is a schematic diagram of another embodiment of the present application for obtaining a second characteristic value of a target data slice;
[0044] FIG9 is a flowchart of another embodiment of the present application for obtaining multiple first characteristic values of target data;
[0045] FIG10 is a flow chart of another data clustering method provided in an embodiment of the present application;
[0046] FIG11 is a schematic diagram of a clustering result of the present application provided in an embodiment of the present application;
[0047] FIG12 is a schematic diagram of a clustering result of a related technology provided in an embodiment of the present application;
[0048] FIG13 is a schematic diagram showing a comparison of compression effects of differential compression and deep compression schemes based on the clustering results shown in FIG11 and FIG12 , provided by an embodiment of the present application;
[0049] FIG14 is a schematic diagram showing a comparison of compression effects of a combined compression scheme implemented on the clustering results shown in FIG11 and FIG12 according to an embodiment of the present application;
[0050] FIG15 is a schematic structural diagram of a storage system provided in an embodiment of the present application;
[0051] FIG16 is a schematic diagram of a process of applying a data clustering method of the present application to data writing provided by an embodiment of the present application;
[0052] FIG17 is a schematic diagram of a data reading process corresponding to the data writing process of FIG16 provided in an embodiment of the present application;
[0053] FIG18 is a schematic structural diagram of a data clustering device provided in an embodiment of the present application;
[0054] FIG19 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0055] FIG20 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0056] FIG21 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0057] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0058] To facilitate understanding, the technology and background involved in the embodiments of this application are explained below.
[0059] The continuous expansion of storage systems (such as cloud storage) and the increasing performance demands of some new applications and services have led to two significant changes in storage systems. First, the rapid expansion of storage capacity has necessitated the reduction of storage costs to enhance the core competitiveness of storage systems. Second, applications and services with higher performance requirements require the use of higher-performing storage devices to meet these demands, leading to an exponential increase in storage costs. Both of these changes require the development of effective methods to reduce storage costs and thereby enhance the core competitiveness of storage systems.
[0060] Data deduplication and compression technology is an effective and direct method for reducing storage costs. It involves two key technical aspects: data deduplication and data compression. Deduplication and compression can reduce redundant data in storage systems, resulting in a data reduction effect and significantly lowering storage costs. In some scenarios, the data a storage system needs to process includes a large amount of similar data, and this data is distributed very discretely and unevenly in time and space. Taking this into account, data deduplication and compression technology can extract data eigenvalues and cluster similar data that is distributed very discretely and unevenly in time and space based on these eigenvalues to form similar data sets. This data is then compressed using a layered data reduction technique. Layered data reduction technology not only identifies identical data within similar data sets for deduplication but also compresses data within similar data sets using a differential compression scheme to generate differential blocks. These differential blocks are then compressed using a deep compression algorithm to further improve the data reduction rate.
[0061] Therefore, a key technical aspect of data deduplication and compression technology lies in how to cluster similar data. The implementation method for clustering similar data affects the deduplication and compression scheme's reduction effectiveness and compression overhead. Currently, when clustering similar data based on data block feature values, if multiple data blocks share the same feature value, they are grouped together to form a similar data set. However, this data clustering method is too coarse and cannot distinguish between different data blocks within a similar data set, resulting in poor clustering results.
[0062] Based on this, an embodiment of the present application provides a data clustering method. This method can obtain multiple data to be clustered, obtain multiple first eigenvalues of target data, and then divide at least two data with corresponding identical multiple first eigenvalues into the same similar data set. The target data is any one of the multiple data to be clustered. In this way, at least two data in the same similar data set not only have the same eigenvalues, but also have the same total number of identical first eigenvalues. Therefore, this solution can further filter different data based on whether they have corresponding identical eigenvalues, based on the data having the same eigenvalues. It can further subdivide the data based on the degree of similarity between the data, and treat data with different redundancies differently, so that the data in the same similar data set have greater redundancy, and the data in different similar data sets have greater redundancy, thereby achieving the distinction between data with different redundancies, and being able to divide data with different similarities into different similar data sets, thereby improving the accuracy of clustering similar data and improving the clustering effect of clustering data. In addition, this solution also helps to improve the efficiency of performing operations on data based on the clustering results of the data. For example, when compressing data based on clustering results, since the higher the similarity redundancy of similar data sets, the better the compression effect of similar data in similar data sets, the present application can utilize the greater redundancy between data in the same similar data set to improve the compression rate of data.
[0063] The following first describes the implementation environment involved in a data clustering method provided in an embodiment of the present application.
[0064] Figure 1 is a schematic diagram of an implementation environment involved in a data clustering method provided in an embodiment of the present application. As shown in Figure 1, the implementation environment includes: a computing device 10. The computing device 10 is capable of executing the data clustering method provided in an embodiment of the present application.
[0065] In one implementation, the data clustering method provided in the embodiment of the present application can be implemented by running an executable program on the computing device 10. For example, the executable program of the data clustering method can be optionally presented in the form of an application installation package. After the application installation package is installed in the computing device 10, the data clustering method can be implemented by running the executable program. In this case, the computing device 10 can be a terminal. The terminal can be a computer, a personal computer, a portable mobile terminal, a multimedia player, an e-book reader, or a wearable device. For example, after the application installation package of the executable program of the data clustering method is installed in the computing device 10, when the computing device 10 needs to obtain the similarity of multiple images, the computing device 10 can run the executable program to implement the data clustering method, and cluster the data used to represent the multiple images based on the data clustering method, and then obtain the similarity of the multiple images based on the clustering result, such as multiple images divided into the same set have a greater similarity, and multiple images divided into different sets have a smaller similarity.
[0066] FIG2 is a schematic diagram of an implementation environment involved in another data clustering method provided in an embodiment of the present application. As shown in FIG2 , the implementation environment may further include a client 20. The client 20 is capable of establishing a communication connection with the computing device 10. For example, the communication connection between the client 20 and the computing device 10 may be established via a network. Optionally, the network may be a local area network, the Internet, or other network, and the present embodiment does not limit this.
[0067] In one implementation, the client 20 may be a desktop computer, a laptop computer, a mobile phone, a smart phone, a tablet computer, a multimedia player, a smart home appliance, an artificial intelligence device, a smart wearable device, an e-reader, a smart vehicle-mounted device, or an Internet of Things device.
[0068] The client 20 is used for the user to interact with the computing device 10. In one implementation, the client 20 is used to send instructions to the computing device 10 according to the user's instructions, and the operation indicated by the instruction depends on the execution of the data clustering method provided by the embodiment of the present application. The computing device 10 is used to execute the operation indicated by the instruction sent by the client 20, and the data clustering method provided by the embodiment of the present application needs to be executed in the process of executing the operation. For example, when the computing device 10 is used as a storage node, the computing device 10 is used to store data. When the user needs to store data in the computing device 10, the user sends a data write instruction to the computing device 10 through the client 20 to instruct the computing device 10 to write data. After the computing device 10 receives the write instruction, it first obtains the data to be written from the write instruction, and then divides the data to be written into multiple data blocks, clusters the multiple data blocks using the data clustering method provided by the embodiment of the present application, obtains multiple similar data sets, and then compresses the data in the similar data sets, and stores the compressed data. At this time, the data clustering method provided by the embodiment of the present application is used in a data compression scenario. For another example, the client 20 is configured to send a data acquisition instruction to the computing device 10 according to the user's instructions, instructing the computing device 10 to acquire data from the computing device 10. After receiving the data acquisition instruction, the computing device 10 acquires the data that the client 20 needs to acquire based on the data acquisition instruction, then divides the data into multiple data blocks, clusters the multiple data blocks using the data clustering method provided in the embodiment of the present application, obtains multiple similar data sets, and then compresses the data in the similar data sets, and sends the compressed data to the client 20. In this case, the data clustering method provided in the embodiment of the present application is used in a data transmission scenario.
[0069] In one possible implementation, the implementation environment shown in FIG2 includes a computing device cluster, which includes multiple computing devices 10. In this case, the computing device 10 can be a server (such as a cloud server), and the computing device cluster can be a server cluster composed of several servers, or a cloud computing service center. Among them, a large number of basic resources owned by the cloud service provider are deployed in the cloud computing service center. For example, computing resources, storage resources, and network resources are deployed in the cloud computing service center. The cloud computing service center can use this large number of basic resources to implement the data clustering method provided in the embodiment of the present application.
[0070] When the computing device cluster is implemented through a cloud computing service center, the computing device cluster performs the functions that can be provided by the data clustering method provided in the embodiment of the present application, and can be abstracted by the cloud service provider into a clustering cloud service on the cloud platform. The user can access the cloud platform through the client 20, purchase the clustering cloud service on the cloud platform, and use the clustering cloud service provided by the computing device cluster through the cloud platform. Optionally, the cloud platform can be a cloud platform of a central cloud, a cloud platform of an edge cloud, or a cloud platform including a central cloud and an edge cloud, and the embodiment of the present application does not make specific limitations on it. In addition, the clustering cloud service can be optionally provided by the cloud platform as a separate cloud service, or the clustering cloud service can be provided as an additional cloud service to other cloud services. For example, the computing device cluster can be optionally a storage system, which is used to provide a storage cloud service, and the clustering cloud service can be optionally provided as an additional cloud service to the storage cloud service. At this time, if the customer purchases the clustering cloud service as an additional cloud service when purchasing the storage cloud service, after the client 20 instructs to store data in the computing device cluster, the computing device cluster first uses the clustering cloud service to cluster the data before storing the data, and then uses the storage cloud service to store the clustered data.
[0071] FIG3 is a schematic diagram of the architecture of a storage system provided in an embodiment of the present application. Optionally, the storage system may be a distributed storage system. As shown in FIG3 , the storage system includes a service layer, an index layer, and a persistence layer.
[0072] The service layer provides users with unified interface protocol services. Services provided by the service layer include: Elastic Volume Service (EVS, also known as cloud hard disk), Object Storage Service (OBS), Scalable File Service (SFS), Data Lake Insight (DLI) service, and Data Warehouse Service (DWS). To ensure service performance, the service layer also includes a cache layer.
[0073] The boot layer provides metadata management services for the distributed storage system. It can run a database (DB) and perform deduplication and compression, interacting with the service layer through object and file bootstrapping. The database can be a key-value database (KVDB).
[0074] The persistence layer provides persistent storage services for distributed systems. The persistence layer can achieve write-optimization and read-optimization through ishard mode and PLOG mode. Both ishard and PLOG modes can share a storage pool. This storage pool can be a data function virtualization (DFV) storage pool. The storage media in this storage pool can include non-volatile memory (NVM), solid-state drives (SSDs), hard disk drives (HDDs), and optical storage media.
[0075] It should be understood that the above content is an exemplary description of the application scenario of the data clustering method provided in the embodiment of the present application, and does not constitute a limitation on the application scenario of the data clustering method. It is known to those skilled in the art that as business needs change, its application scenario can be adjusted according to application needs. For example, the data clustering method provided in the embodiment of the present application can also be applied to the field of network transmission. By clustering similar data in the data to be transmitted and then compressing the clustered similar data, the amount of data transmitted over the network can be reduced and the network transmission rate can be improved. For another example, the data clustering method provided in the embodiment of the present application can also be applied to a variety of compression fields, such as the image compression field, the video compression field, and the dedicated database compression field, so that the compressed data can be clustered by the clustering function provided in the embodiment of the present application, thereby reducing the delay of data compression and decompression, improving the efficiency of data compression and decompression, and reducing the resource consumption of compression and decompression. In addition, when the method is applied to other scenarios, it is not limited to the implementation environment shown in Figure 1 or Figure 2. For example, when the method is applied to a data transmission scenario, its implementation environment can optionally include multiple computing devices, and data can be transmitted between the multiple computing devices. The embodiment of the present application does not specifically limit it.
[0076] The following describes the implementation process of the data clustering method provided in the embodiment of the present application by applying the data clustering method provided in the embodiment of the present application to the application scenario shown in Figure 2. Figure 4 is a flow chart of a data clustering method provided in the embodiment of the present application. As shown in Figure 4, the data clustering method includes the following steps:
[0077] Step 401: Acquire multiple data to be clustered.
[0078] In this application, the data to be clustered can be any data that has clustering requirements. Clustering multiple data refers to the process of dividing multiple data into different similar data sets based on the similarities between the multiple data. Multiple data divided into the same similar data set have greater similarity. Multiple data divided into different similar data sets have less similarity.
[0079] The multiple data to be clustered may be obtained by dividing the data recorded in a data record. In this case, the multiple data to be clustered belong to different parts of the data recorded in the data record. In a possible implementation, the multiple data to be clustered may be obtained by dividing the data to be compressed. The data to be compressed may be data that needs to be stored or transmitted. The data that needs to be operated during a compression or transmission operation performed by a computing device may be divided into multiple data, and the multiple data are the multiple data to be clustered. For example, before a storage system stores a certain data to be written, it may perform deduplication and compression on the data to be written. In the process of performing deduplication and compression on the data to be written, the storage system may divide the data to be written into multiple data blocks, cluster the multiple data blocks, and then dedupe and compress the data using the similar data sets obtained by clustering as units. Before performing these operations on the data, clustering the data can optimize the execution process of these operations based on the clustering results. For example, by clustering the data to be stored and then compressing the clustered data, the amount of data to be stored can be reduced, thereby reducing the storage resources occupied by the data to be stored and improving the utilization rate of storage resources. By clustering the data to be transmitted and then compressing the clustered data, the amount of data to be transmitted can be reduced, thereby reducing the amount of data to be transmitted and improving the transmission rate. It should be noted that the computing device can optionally divide the compressed data into fixed-length or non-fixed-length segments to obtain multiple data to be clustered. Alternatively, the computing device can also use other methods to divide the data recorded in a data record to obtain multiple data to be clustered, which is not specifically limited in the embodiments of the present application.
[0080] Alternatively, the multiple data to be clustered may also be data that does not belong to a data record. For example, the multiple data may be selected as independent data, and there is no relationship between the multiple data. For example, the multiple data may represent data of multiple images for which similarity between the images needs to be obtained. For example, when obtaining the similarity between multiple images based on the clustering result, the clustering result indicates that multiple images divided into the same set have a greater similarity, and multiple images divided into different sets have a smaller similarity, then the multiple data may be selected as the data used to represent the multiple images, and the data used to represent any one image is one data to be clustered.
[0081] Step 402: Acquire multiple first eigenvalues of target data, where the target data is any one of the multiple data to be clustered.
[0082] In one possible implementation, the computing device may optionally obtain multiple first eigenvalues of the data using a sampling method. For example, the computing device first obtains multiple data slices of the target data, then obtains the second eigenvalues of each of the multiple data slices, and obtains multiple first eigenvalues of the target data based on the second eigenvalues of the multiple data slices. As shown in FIG5 , the implementation process of step 402 includes:
[0083] Step 4021: Obtain multiple data shards from the target data.
[0084] The computing device may optionally obtain partial data of the target data from the target data in multiple times, and use the partial data obtained each time as a data slice of the target data. In one possible implementation, in response to the target data being represented by a character string, multiple data slices are obtained from the target data, including: intercepting multiple second character segments from the character string used to represent the target data to obtain multiple data slices. Each second character segment represents a data slice. The second character segment includes multiple characters arranged continuously, and the positions of the characters in any two of the multiple second character segments in the character string used to represent the target data are different. The positions of the characters in any two second character segments in the character string used to represent the target data are different, including: any two second character segments do not include characters located at the same position. Optionally, the computing device may optionally intercept multiple character segments of the same length in the character string used to represent the target data according to a fixed step size, and each character segment represents a data slice. As shown in Figure 6, the multiple data to be clustered include data block V1 and data block V2. When the computing device obtains multiple data fragments from data block V1, it extracts multiple character segments C1, C2, C3, and C4 of the same length from data block V1 according to a fixed step size. These multiple character segments of the same length are the multiple data fragments obtained from data block V1. When the computing device obtains multiple data fragments from data block V2, it extracts multiple character segments C21, C22, C23, and C24 of the same length from data block V2 according to a fixed step size. These multiple character segments of the same length are the multiple data fragments obtained from data block V2.
[0085] Step 4022: Obtain a second characteristic value of a target data slice, where the target data slice is any one of multiple data slices of the target data.
[0086] In one possible implementation, the computing device may optionally obtain the second characteristic value of the data slice using a sampling method. For example, the computing device first obtains multiple data segments of the target data slice, and then obtains the second characteristic value of the target data slice based on the multiple data segments. As shown in FIG7 , the implementation process of step 4022 includes:
[0087] Step 4022a: Obtain multiple data segments from the target data shard.
[0088] The computing device may optionally obtain partial data of the target data slice from the target data slice multiple times and use the partial data as a data segment of the target data slice. In one possible implementation, in response to the target data slice being represented by a string, obtaining multiple data segments from the target data slice includes: intercepting multiple first character segments from the string used to represent the target data slice to obtain multiple data segments. Each first character segment represents a data segment, the first character segment includes one character or multiple characters arranged continuously, and the positions of the characters in any two of the multiple first character segments in the string used to represent the target data slice are different. The positions of the characters in any two first character segments in the string used to represent the target data slice are different, including: any two first character segments do not include characters located at the same position. Optionally, the computing device may optionally intercept multiple character segments of the same length from the string used to represent the target data slice according to a fixed step size, each character segment representing a data segment. As shown in Figure 6, in data block V1 and data block V2. When the computing device obtains multiple data segments from the data slice C3, it intercepts multiple character segments of the same length in the data slice C3 according to a fixed step size (such as the boxes filled with slashes in Figure 6). The multiple character segments of the same length are the multiple data segments obtained from the data slice C3.
[0089] Step 4022b: Obtain the modulus value of each data segment in the multiple data segments of the target data fragment.
[0090] In the computer field, a modulo b is used to find the remainder of the division of two a and b. In this application, a data slice modulo a data segment refers to using the ASIC code value of the data to take the modulo. That is, a data slice modulo a data segment is using the ASIC code value of the data slice to take the modulo of the ASIC code value of the data segment. For example, assuming that the data slice is represented as abcdefg and the data segment is abc, the modulus value of the data slice abcdefg modulo the data segment abc is defg. After the target data slice takes the modulo of each data segment, a modulus value corresponding to the data segment can be obtained. Then, after the target data slice takes the modulo of each data segment in the multiple data segments, multiple modulo values corresponding one to one to the multiple data segments can be obtained.
[0091] Step 4022c: Based on the module values corresponding to the multiple data segments, obtain the second characteristic value of the target data segment.
[0092] After obtaining multiple modulus values of the target data slice modulo multiple data segments, the computing device can obtain the second characteristic value of the target data slice based on the multiple modulus values. For example, the computing device combines and processes the multiple modulus values to obtain the second characteristic value of the second data slice. In one possible implementation, in response to the target data slice, the data segment and the modulus value being represented by a string, the second characteristic value of the target data slice is obtained based on the modulus values corresponding to the multiple data segments, including: according to the order of the strings used to represent the multiple data segments in the string used to represent the target data slice, the strings used to represent the modulus values corresponding to the multiple data segments are concatenated to obtain the second characteristic value of the target data slice. For example, assuming that the target data slice is represented as abcdefg, multiple data segments are obtained in sequence according to the order of the characters from left to right in the string used to represent the target data slice, and the modulus values corresponding to the multiple data segments are defg, efg and aeg respectively. Then, according to the order of the character strings used to represent the multiple data segments in the character string used to represent the target data fragment, after character concatenation of the character strings used to represent the modulus values corresponding to the multiple data segments, the second characteristic value of the target data fragment can be obtained as defgefgaeg.
[0093] Optionally, before obtaining the second eigenvalue of the target data slice based on the modulus values corresponding to the multiple data segments, the modulus values corresponding to the multiple data segments may be optimized. In this case, the modulus value used to obtain the second eigenvalue of the target data slice is the optimized modulus value. Specifically, as shown in FIG8 , the implementation process of step 4022c includes: step 4022c1, obtaining the second eigenvalue of the target data slice based on the optimized modulus values corresponding to the multiple data segments.
[0094] For example, as shown in FIG8 , before obtaining the second characteristic value of the target data segment based on the module values corresponding to the multiple data segments, the method further includes step 4022d and step 4022e.
[0095] Step 4022d: Obtain the optimization coefficient of the target data segment, where the target data segment is any one of the multiple data segments of the target data shard.
[0096] The modulus value corresponding to the target data segment is optimized based on the optimization coefficient of the target data segment, with the aim of eliminating hash conflicts of the modulus value corresponding to the target data segment. That is, the optimization coefficient of the target data segment is the coefficient used to eliminate hash conflicts of the modulus value corresponding to the target data segment. There are multiple ways to obtain this optimization coefficient. In one possible implementation, the optimization coefficient is the value used in the hash algorithm to eliminate hash conflicts of the hash value of the data when using an existing hash algorithm (such as the Robin hash algorithm). The optimization coefficient can be determined by executing the existing hash algorithm on the target data segment and using the value used in the existing hash algorithm to eliminate hash conflicts of the hash value of the target data segment as the optimization coefficient of the target data segment. Optionally, before executing step 4022d, the computing device can pre-calculate the optimization coefficient of each data in this manner for a large amount of preset data and establish a corresponding relationship between the data and its optimization coefficient. When executing step 4022d, the corresponding relationship can be optionally queried according to the target data segment to obtain the optimization coefficient corresponding to the target data segment. Furthermore, for a particular piece of data, the optimization coefficient calculated for that data may have multiple values during multiple calculations. The computing device can then establish a correspondence between the optimization coefficient and the data based on the optimization coefficient used most frequently during those multiple calculations. This pre-established correspondence between the optimization coefficient and the data eliminates the need to calculate the optimization coefficient corresponding to the target data segment online during step 4022d, reducing the computational complexity and load of the data clustering process and improving the efficiency of data clustering.
[0097] Step 4022e: Based on the optimization coefficient of the target data segment, eliminate the hash conflict of the modulus value corresponding to the target data segment to obtain the optimized modulus value corresponding to the target data segment.
[0098] After obtaining the optimization coefficient of the target data segment, the optimization coefficient can be used to process the modulus value of the target data segment to eliminate the hash conflict of the modulus value corresponding to the target data segment, and obtain the optimized modulus value corresponding to the target data segment. In one possible implementation, in response to the modulus value and the optimization coefficient being represented by a string, the optimized modulus value is obtained by concatenating the string used to represent the optimization coefficient and the string used to represent the modulus value. For example, the optimized modulus value is a string obtained by concatenating the string used to represent the modulus value with the string used to represent the optimization coefficient. For example, assuming that the modulus value corresponding to the target data segment is defg and the optimization coefficient of the target data segment is ace, the optimized modulus value corresponding to the target data segment is defgace.
[0099] When multiple data segments of a data shard are obtained by intercepting data from the data shard, the multiple data segments of the data shard may also be correlated due to the correlation between characters at different positions in the data shard. Therefore, by using the optimization coefficients of the data segments to eliminate hash conflicts of the module values corresponding to the data segments, the correlation between the multiple data segments in the data shard can be reduced or even eliminated, thereby reducing or even avoiding the local interference of the clustering results caused by the correlation between the data segments, thereby improving the accuracy of the clustering results.
[0100] The implementation process of step 4022 can be regarded as the process of obtaining the hash value of the data shard, and the implementation method of step 4022 can be regarded as an implementation method of a hash algorithm provided by an embodiment of the present application. In this implementation method, since the hash value of the data shard is obtained according to the modulus value of the data shard modulo the data segment, its computational complexity is low and the amount of computation is small, which reduces the computational overhead of obtaining the eigenvalue and can ensure the speed of obtaining the eigenvalue. For example, the implementation method of calculating the second eigenvalue of the data shard using the present application can reduce the computational overhead from the computational complexity 0(N*N) to 0(N) compared to the implementation method of calculating the hash value using the Robin hash algorithm, thereby reducing the computational overhead.
[0101] Step 4023: Process the second eigenvalues of multiple data slices to obtain multiple first eigenvalues of the target data. Any first eigenvalue of the target data is obtained based on the second eigenvalues of some data slices in the multiple data slices, and any two first eigenvalues of the target data are obtained based on the second eigenvalues of different data slices.
[0102] In the process of executing step 4023, the second eigenvalues of the plurality of data slices can be grouped in advance to obtain a plurality of eigenvalue groups, each eigenvalue group including a plurality of second eigenvalues. The plurality of eigenvalue groups correspond one-to-one to the plurality of first eigenvalues, that is, each first eigenvalue in the plurality of first eigenvalues is obtained based on the plurality of second eigenvalues in the eigenvalue group corresponding to the first eigenvalue. In an embodiment of the present application, there are multiple optional principles for grouping the second eigenvalues of the plurality of data slices, which are determined according to application requirements. For example, when the positions of the plurality of data slices in the target data are different, after obtaining the second eigenvalues of the plurality of data slices, the second eigenvalues of the plurality of data slices can be optionally sorted according to the position, and then starting from the second eigenvalue of the data slice ranked first, the second eigenvalues of each specified number of data slices in the sorting result are divided into a group, thereby obtaining a plurality of eigenvalue groups. For example, assuming that one eigenvalue is obtained for each of the 12 data slices through the aforementioned steps, then based on the positions of the 12 data slices in the target data, a ranking of the 12 eigenvalues can be obtained. Then, starting with the second eigenvalue ranked first, every four second eigenvalues in the ranking result are divided into a group, thereby obtaining three eigenvalue groups. Among them, the first eigenvalue group includes the second eigenvalues ranked from 1 to 4, the second eigenvalue group includes the second eigenvalues ranked from 5 to 8, and the third eigenvalue group includes the second eigenvalues ranked from 9 to 12.
[0103] In one possible implementation, in response to the second eigenvalue and the first eigenvalue being represented by a string, the first eigenvalue may be optionally obtained by character concatenation based on a string representing multiple second eigenvalues in the eigenvalue group corresponding to the first eigenvalue. Continuing with the above example, the first eigenvalue group includes the second eigenvalues ac, ab, bc, and cd, then the first eigenvalue corresponding to the first eigenvalue group is acabbccd. The second eigenvalue group includes the second eigenvalues ac, abc, bcd, and cd, then the first eigenvalue corresponding to the first eigenvalue group is acabcbcdcd. The third eigenvalue group includes the second eigenvalues abc, ab, abc, and cd, then the first eigenvalue corresponding to the first eigenvalue group is abcababccd.
[0104] Optionally, the second eigenvalue used when obtaining the multiple first eigenvalues of the target data may also be a second eigenvalue that has undergone decorrelation processing. As shown in FIG9 , step 4023 includes: step 4023a, processing the decorrelated second eigenvalues of the multiple data slices to obtain the multiple first eigenvalues of the target data. The implementation process is described in step 4023 and is not further elaborated here.
[0105] For example, as shown in FIG9 , before processing the second eigenvalues of multiple data slices to obtain multiple first eigenvalues of target data, the method further includes: step 4024, performing decorrelation processing on the second eigenvalues of multiple data slices to obtain the second eigenvalues of multiple data slices after decorrelation processing.
[0106] There are many ways to implement the de-correlation process on the second characteristic values of multiple data slices. This application uses one implementation as an example to illustrate it. For example, the de-correlation process includes: sorting by numerical value.
[0107] When multiple data segments of a data slice are obtained by intercepting data in the data slice, since the characters at different positions in the data slice have a correlation, the multiple second eigenvalues of the multiple data slices may also have a correlation. Therefore, by performing decorrelation processing on the second eigenvalues of multiple data slices, the correlation between the eigenvalues of multiple data slices can be reduced or even eliminated, thereby reducing or even avoiding the local interference of the clustering results caused by the correlation between the data slices, thereby improving the accuracy of the clustering results. After sorting the second eigenvalues of multiple data slices according to the size of the values, the correlation carried by the second eigenvalues of multiple data slices due to the positions of the multiple data slices in the target data can be broken, thereby reducing or even avoiding the local interference of the clustering results caused by the correlation between the positions of the multiple data slices.
[0108] Step 403: Divide at least two data with corresponding identical first eigenvalues into the same similar data set.
[0109] At least two data have the same multiple first eigenvalues, which means that for any two data among the at least two data, the arbitrary two data are data a and data b, the multiple first eigenvalues of data a and the multiple first eigenvalues of data b are equal in a one-to-one correspondence. For example, data a and data b have the same three eigenvalues. Assume that the three eigenvalues of data a are a1, a2, and a3, and the three eigenvalues of data b are b1, b2, and b3. Then, data a and data b have the same three eigenvalues, which means that eigenvalue a1 is equal to eigenvalue b1, eigenvalue a2 is equal to eigenvalue b2, and eigenvalue a3 is equal to eigenvalue b3.
[0110] By dividing at least two data with corresponding identical multiple first eigenvalues into the same similar data set, at least two data in the same similar data set not only have the same eigenvalues, but also the total number of identical first eigenvalues of the at least two data is equal. Therefore, the scheme can further filter different data based on whether they have corresponding identical eigenvalues on the basis of having identical eigenvalues, can further subdivide the data according to the degree of similarity between the data, and treat data with different redundancies differently, so that the data in the same similar data set have greater redundancy, and the redundancy of data in different similar data sets has greater differences, thereby realizing the distinction between data with different redundancies, and being able to divide data with different similarities into different similar data sets, thereby improving the accuracy of clustering similar data.
[0111] Optionally, step 403 may be performed sequentially according to the total number of different data having the same first eigenvalue. For example, as shown in FIG10 , before step 403 , the method further includes: step 404 , counting the total number of each two data having the same first eigenvalue among the multiple data, to obtain multiple total values.
[0112] The implementation process of step 404 includes comparing the multiple first feature values of each pair of data in the plurality of data. For any two data, each time it is determined that the two data have a pair of identical feature values, the total number of first feature values corresponding to the two data is increased by one until all first feature values of the two data are compared. Similarly, the computing device can obtain the total number of first feature values corresponding to each pair of data in the plurality of data according to this implementation logic.
[0113] When the data clustering method of the present application further includes step 404, as shown in FIG10 , the implementation of step 403 includes: step 4031, performing a clustering process on multiple data in descending order of multiple total values, wherein the clustering process is performed according to the i-th total value among the multiple total values, including: among the multiple data to be clustered in the multiple data, the multiple data to be clustered having the same i-th total value first eigenvalue are divided into a similar data set. In this way, the data and other data with which it has the most corresponding identical first eigenvalues can be divided into the same similar data set, which helps to maximize the accuracy of clustering similar data.
[0114] For example, assume that multiple data are data block 1, data block 2, data block 3, and data block 4. Data block 1 has five first characteristic values, namely SFP1, SFP2, SFP3, SFP4, and SFP5. Data block 2 has five first characteristic values, namely SFP1, SFP2, SFP3, SFP6, and SFP7. Data block 3 has two first characteristic values, namely SFP1 and SFP8. Data block 4 has two first characteristic values, namely SFP9 and SFP1. It can be seen that the total number of data blocks 1 and data block 2 that have the same first characteristic value is 3, the total number of data blocks 1 and data block 3 that have the same first characteristic value is 0, the total number of data blocks 1 and data block 4 that have the same first characteristic value is 1, the total number of data blocks 2 and data block 3 that have the same first characteristic value is 1, and the total number of data blocks 2 and data block 4 that have the same first characteristic value is 1. Thus, when the total values are 3 and 1, respectively, in step 4031, the clustering process is first performed according to the total value of 3, and then the clustering process is performed according to the total value of 1. When the clustering process is performed according to the total value of 3, the computing device searches for multiple data blocks with the same three first eigenvalues among the multiple data blocks. It is found that data block 1 and data block 2 have the same three first eigenvalues, which are SFP1, SFP2, and SFP3, respectively. Data blocks 1 and 2 are then classified into the same similar data set. When the clustering process is performed according to the total value of 1, since data blocks 1 and 2 have already been clustered, the computing device searches for multiple data blocks with the same first eigenvalue among the remaining data blocks 3 and 4 to be clustered. It is found that data blocks 3 and 4 have the same first eigenvalue, which is SFP1, respectively. Data blocks 3 and 4 are then classified into the same similar data set. Figure 11 shows the clustering results of this clustering process. As shown in Figure 11, similar data set 1 includes data block 1 and data block 2, which have the same three first eigenvalues. Similar data set 2 includes data block 3 and data block 4, which have the same one first eigenvalue.
[0115] In related technologies, when clustering data blocks, as long as multiple data blocks have the same eigenvalue, they are grouped together to form a similar data set. Using this related technology to cluster data blocks 1, 2, 3, 4, and 5, since these five data blocks have the first eigenvalue SFP1, they are classified into the same similar data set.
[0116] Comparing the clustering results of Figures 11 and 12, the clustering result of Figure 11 subdivides the one similar data set in Figure 12 into two similar data sets, and the data blocks in the two similar data sets have different numbers of corresponding identical first eigenvalues, which is equivalent to dividing the one similar data set in Figure 12 into two similar data sets with different similarity redundancies. At this time, when compressing the data blocks based on the clustering results, compared to compressing based on the similar data sets in Figure 12, due to the greater redundancy between the data in the similar data sets in Figure 11, the overall reduction rate of compressing the data including data block 1, data block 2, data block 3, data block 4 and data block 5 can be improved.
[0117] Figure 13 compares the compression effects of differential compression and deep compression based on the clustering results shown in Figures 11 and 12. Because the clustering results in Figure 11 effectively divide similar data sets into multiple similar data sets with different redundancies based on data redundancy, the compression effect of the compression scheme can be significantly improved. When performing differential compression based on the clustering results in Figure 12, it is very likely that only the four data blocks represented by SFP1 will be compressed together, resulting in very limited compression effect. In contrast, when performing differential compression based on the clustering results in Figure 11, the first eigenvalues SFP1, SPF2, and SFP3 can be used to aggregate data blocks 1 and 2 into different data sets based on similarity redundancy gradients. This can more effectively group similar data sets together, and when differential compression is performed, each differentially compressed data block achieves better compression effect due to the higher similarity between the data blocks. Similarly, as shown in Figure 14, when the two similar data clustering schemes are implemented in a merged compression scheme, the comparison of compression effects is also very obvious. The reason for this is that the compression effects of compression schemes such as merge compression and differential compression all depend on the data similarity of the entire similar data set. Therefore, the data clustering method provided by this application can help improve the compression rate of data.
[0118] Furthermore, before executing the clustering process on the multiple data in descending order of the multiple total values, the multiple data may be coarsely screened to obtain multiple data sets, and then the clustering process may be executed on the multiple data in descending order of the multiple total values in each data set. For example, the multiple data with the same characteristic value may be first divided into the same data set, and then the clustering process may be executed on the multiple data in descending order of the multiple total values in each data set. Continuing with the above example, since data blocks 1 to 5 all have the characteristic value SFP1, the data blocks 1 to 5 may be first divided into a data set, and then the clustering process may be executed on the multiple data in descending order of the multiple total values for the data blocks 1 to 5 included in the data set.
[0119] In summary, in the data clustering method of the present application, the method can obtain multiple data to be clustered, obtain multiple first eigenvalues of target data, and then divide at least two data with corresponding identical multiple first eigenvalues into the same similar data set. The target data is any one of the multiple data to be clustered. In this way, at least two data in the same similar data set not only have the same eigenvalues, but also have the same total number of identical first eigenvalues. Therefore, the scheme can further filter different data based on whether they have the same eigenvalues, further subdivide the data based on the degree of similarity between the data, and treat data with different redundancies differently, so that the data in the same similar data set have greater redundancy, and the data in different similar data sets have greater redundancy, thereby achieving the distinction between data with different redundancies, and being able to divide data with different similarities into different similar data sets, thereby improving the accuracy of clustering similar data and improving the clustering effect of clustering data. In addition, the scheme also helps to improve the efficiency of performing operations on data based on the clustering results of the data. For example, when compressing data based on clustering results, since the higher the similarity redundancy of similar data sets, the better the compression effect of similar data in similar data sets, the present application can utilize the greater redundancy between data in the same similar data set to improve the compression rate of data.
[0120] In addition, some special scenarios (such as memory pools, vector databases, key-value databases (KVDB), data warehouses and network transmission, etc.) also have the problem of huge amounts of redundant data. The embodiments of the present application can also be used in these special scenarios. On the one hand, it can improve the compression effect based on the clustering effect of the present application, and on the other hand, it can reduce the overhead of compression and decompression.
[0121] For example, when the data clustering method of the present application is applied to a scenario where data is written to a storage system, the storage system can use the data clustering method of the present application to cluster the data to be written, and then perform data deduplication and compression based on the clustering results. For example, Figure 15 is a structural schematic diagram of a storage system provided in an embodiment of the present application. As shown in Figure 15, the storage system is a distributed storage system. The storage system may include a metadata management module, a compression module, and a persistence module, and the metadata management module, the compression module, and the persistence module are implemented by different computing devices. The metadata management module is used to receive user input / output (I / O) requests and execute a metadata management process based on the I / O requests. The compression module is used to merge and compress data and / or decompress and merge data. The persistence module is used to store the received data and realize data persistence. The metadata management module includes a metadata management unit (such as LunMap) and a metadata clustering unit (such as DusClient). The metadata management unit is used to record the metadata of all data stored on the storage medium of the storage system, and the metadata includes index information of the storage medium. The metadata clustering unit is used to connect with the metadata management unit, the compression module, and the persistence module, calculate the fingerprint of each data block, and use the data clustering method provided in this application to calculate the characteristic value of each data block. The obtained fingerprint and characteristic value are then stored in the persistence module, and the unit is responsible for combining the obtained fingerprint and characteristic value with other metadata information of the data block and sending them to the compression module. The compression module includes an aggregation unit (such as an opportunistic table (OpTable)), an analysis unit (such as a post data analysis (PDA) unit), a task allocation unit (such as a task mag), a compression unit (such as a post data reduction (PDR) unit), and an index information storage unit (such as an FPTable).
[0122] As shown in Figures 15 and 16 , during the process of writing data to the storage system, the aggregation unit is configured to receive fingerprints and feature values sent by the metadata processing unit in the metadata management module, cluster the data based on the feature values using the data clustering method provided in this application, obtain aggregated data with similarity, provide this aggregated data with similarity to the compression unit, and provide the compression unit with the physical address and logical index number of the similar data. The analysis unit is configured to periodically analyze the aggregated data with similarity provided by the aggregation unit and assemble similar data scattered throughout the storage system into similarity data chains. The task allocation unit is configured to receive the similar data chains sent by the analysis unit, compose different compression tasks based on the distribution of the feature values in the storage system, and allocate the compression tasks to the compression unit. The compression unit is configured to perform data deduplication and compression based on the assigned compression tasks, persist the compressed data set to the persistence module, generate index information for each data block in the compressed data set, and provide this index information to the index information storage unit. Optionally, before performing the data compression process, the compression unit may first find identical data on similar data chains through fingerprint comparison, retain only one copy of multiple identical copies of data, and then perform the compression process on the remaining data. The index information storage unit is used to receive the index information of each data block in the data set sent by the compression unit, persist the index information, and update the index information recorded in the metadata management unit based on the index information.
[0123] When a target data block needs to be read from the storage system shown in FIG15 , the storage system can first obtain index information for the target data block based on a data read request, then retrieve the target data block from the target data set based on the index information, and then feedback the target data block to the data read request. For example, as shown in FIG15 and FIG17 , during the process of reading data from the storage system, after receiving a data read request, the metadata management module can obtain index information for the target data block and the target data set indicated by the data read request from the metadata management unit, and provide the index information to the compression module. The compression module can then obtain the compressed target data set containing the target data block from the storage medium based on the index information, decompress the target data set, and then retrieve the target data block from the decompressed target data set. The target data block can then be fed back to the metadata management module, thereby completing the target data block reading process. However, if the target data set cannot be retrieved based on the index information obtained from the metadata management unit, as shown in FIG17 , the metadata management module can obtain new index information for the target data block and the target data set it contains from the index information storage unit, and then retrieve the target data block based on the new index information.
[0124] It should be noted that the order of the steps of the data clustering method provided in the embodiments of the present application can be adjusted appropriately, and the steps can be increased or decreased accordingly. Any person skilled in the art who can easily conceive of a modified method within the technical scope disclosed in this application should be included in the scope of protection of this application, and therefore will not be described in detail.
[0125] The above introduces the data clustering method of the embodiment of the present application. Corresponding to the above method, the embodiment of the present application also provides a data clustering device. Figure 18 is a structural schematic diagram of a data clustering device provided by the embodiment of the present application. Based on the following multiple modules shown in Figure 18, the data clustering device shown in Figure 18 can perform all or part of the operations shown in Figures 4, 5 and 10 above. It should be understood that the data clustering device may include more additional modules than the modules shown or omit some of the modules shown therein, and the embodiment of the present application does not limit this. Optionally, the data clustering device can be configured on a cloud platform. As shown in Figure 18, the data clustering device 180 includes:
[0126] The acquisition module 1801 is used to acquire a plurality of data to be clustered.
[0127] The acquisition module 1801 is further configured to acquire multiple first feature values of target data, where the target data is any one of the multiple data.
[0128] The clustering module 1802 is configured to group at least two data items corresponding to the same plurality of first eigenvalues into the same similar data set.
[0129] In one possible implementation, the clustering module 1802 is specifically used to: count the total number of data with the same first eigenvalue corresponding to each two data in multiple data to obtain multiple total values; perform a clustering process on the multiple data in descending order of the multiple total values, wherein the clustering process is performed according to the i-th total value in the multiple total values, including: among the multiple data to be clustered in the multiple data, the multiple data to be clustered with the same i-th total value first eigenvalue are divided into a similar data set.
[0130] In one possible implementation, the acquisition module 1801 is specifically used to: obtain multiple data slices from the target data; obtain the second characteristic value of the target data slice, where the target data slice is any one of the multiple data slices; process the second characteristic values of the multiple data slices to obtain multiple first characteristic values of the target data, where any first characteristic value of the target data is obtained based on the second characteristic values of some data slices in the multiple data slices, and any two first characteristic values of the target data are obtained based on the second characteristic values of different data slices.
[0131] In one possible implementation, the acquisition module 1801 is specifically used to: obtain multiple data segments from the target data slice; obtain the modulus value of the target data slice modulo each data segment in the multiple data segments; and obtain the second eigenvalue of the target data slice based on the modulus values corresponding to the multiple data segments.
[0132] In one possible implementation, in response to the target data fragment, data segment and modulus value being represented by a string, the acquisition module 1801 is specifically used to: according to the order of the strings used to represent the multiple data segments in the string used to represent the target data fragment, perform character concatenation on the strings used to represent the modulus values corresponding to the multiple data segments to obtain the second characteristic value of the target data fragment.
[0133] In one possible implementation, the acquisition module 1801 is also used to: obtain the optimization coefficient of the target data segment, where the target data segment is any one of the multiple data segments; based on the optimization coefficient of the target data segment, eliminate the hash conflict of the module value corresponding to the target data segment to obtain the optimized module value of the target data segment.
[0134] In a possible implementation, in response to the modulus value and the optimization coefficient being represented by a character string, the optimized modulus value is obtained by concatenating the character string used to represent the optimization coefficient and the character string used to represent the modulus value.
[0135] In one possible implementation, in response to the target data slice being represented by a string, the acquisition module 1801 is specifically used to: respectively intercept multiple first character segments from the string used to represent the target data slice to obtain multiple data segments, wherein each first character segment represents a data segment, the first character segment includes one character or multiple characters arranged continuously, and the positions of the characters in any two first character segments in the multiple first character segments in the string used to represent the target data slice are different.
[0136] In a possible implementation, in response to the second eigenvalue and the first eigenvalue being represented by character strings, the first eigenvalue is obtained by concatenating characters in character strings used to represent a plurality of second eigenvalues.
[0137] In a possible implementation, the acquisition module 1801 is further configured to perform decorrelation processing on the second eigenvalues of the plurality of data slices to obtain decorrelation-processed second eigenvalues of the plurality of data slices, wherein the decorrelation processing includes sorting by numerical value.
[0138] In one possible implementation, in response to the target data being represented by a string, the acquisition module 1801 is specifically used to: respectively intercept multiple second character segments from the string used to represent the target data to obtain multiple data fragments, wherein each second character segment represents a data fragment, the second character segment includes multiple characters arranged continuously, and the positions of the characters in any two second character segments in the multiple second character segments in the string used to represent the target data are different.
[0139] The detailed working processes of acquisition module 1801 and clustering module 1802 are described in the previous method embodiment. For example, acquisition module 1801 uses step 401 to acquire multiple data to be clustered and uses step 402 to acquire multiple first eigenvalues of the target data. Clustering module 1802 uses step 403 to group at least two data sets that have the same multiple first eigenvalues into the same similar data set. This embodiment of the present application will not be repeated here.
[0140] Among them, both acquisition module 1801 and clustering module 1802 can be implemented by software or hardware. For example, the implementation of acquisition module 1801 is described below using acquisition module 1801 as an example. Similarly, the implementation of clustering module 1802 can refer to the implementation of acquisition module 1801.
[0141] As an example of a software functional unit, the acquisition module 1801 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the acquisition module 1801 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one cloud data center or multiple geographically close cloud data centers. Typically, a region may include multiple AZs.
[0142] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0143] As an example of a hardware functional unit, acquisition module 1801 may include at least one computing device, such as a server. Alternatively, acquisition module 1801 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0144] The multiple computing devices included in acquisition module 1801 can be distributed in the same region or in different regions. The multiple computing devices included in acquisition module 1801 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in acquisition module 1801 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0145] It should be noted that in other embodiments, either acquisition module 1801 or clustering module 1802 can be used to execute any step in the data clustering method. The steps that acquisition module 1801 and clustering module 1802 are responsible for implementing can be specified as needed. By having acquisition module 1801 and clustering module 1802 respectively implement different steps in the data clustering method, the full functionality of the data clustering device is realized.
[0146] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the various components described above can refer to the corresponding contents in the aforementioned method embodiments and will not be repeated here.
[0147] The following is an example of the basic hardware structure involved in the embodiments of the present application.
[0148] This application also provides a computing device 1900. As shown in Figure 19, computing device 1900 includes a bus 1902, a processor 1904, a memory 1906, and a communication interface 1908. Processor 1904, memory 1906, and communication interface 1908 communicate with each other via bus 1902. Computing device 1900 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1900.
[0149] Bus 1902 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG19 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1902 may include a path for transmitting information between various components of computing device 1900 (e.g., memory 1906, processor 1904, and communication interface 1908).
[0150] The processor 1904 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0151] The memory 1906 may include volatile memory, such as random access memory (RAM). The processor 1904 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0152] Memory 1906 stores executable program code, and processor 1904 executes the executable program code to implement the functions of the aforementioned acquisition module 1801 and clustering module 1802, thereby implementing the data clustering method of the present application. In other words, memory 1906 stores instructions for executing the data clustering method of the present application.
[0153] The communication interface 1908 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1900 and other devices or a communication network.
[0154] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0155] As shown in Figure 20, the computing device cluster includes at least one computing device 1900. The memory 1906 in one or more computing devices 1900 in the computing device cluster may store the same instructions for executing the data clustering method of the present application.
[0156] In some possible implementations, the memory 1906 of one or more computing devices 1900 in the computing device cluster may also store some instructions for executing the data clustering method of the present application. In other words, the combination of one or more computing devices 1900 can jointly execute the instructions for executing the data clustering method of the present application.
[0157] It should be noted that the memory 1906 in different computing devices 1900 in the computing device cluster can store different instructions, each for executing a portion of the functions of the data clustering apparatus of the present application. In other words, the instructions stored in the memory 1906 in different computing devices 1900 can implement the functions of one or more modules in the acquisition module 1801 and the clustering module 1802.
[0158] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG. 21 illustrates a possible implementation. As shown in FIG. 21 , two computing devices 1900A and 1900B are connected via a network. Specifically, the connection to the network is made via a communication interface in each computing device. In this type of possible implementation, the memory 1906 in the computing device 1900A stores instructions for executing the functions of the acquisition module 1801. Simultaneously, the memory 1906 in the computing device 1900B stores instructions for executing the functions of the clustering module 1802.
[0159] The connection method between the computing device clusters shown in Figure 21 can be considered to be that the data clustering method provided by this application requires a large amount of data storage, so it is considered to hand over the functions implemented by the clustering module 1802 to the computing device 1900B for execution.
[0160] It should be understood that the functionality of the computing device 1900A shown in FIG21 may also be implemented by multiple computing devices 1900. Similarly, the functionality of the computing device 1900B may also be implemented by multiple computing devices 1900.
[0161] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection relationship between computing device clusters in Figures 20 and 21. However, the memory 1906 in one or more computing devices 1900 in this computing device cluster can store the same instructions for executing the data clustering method.
[0162] In some possible implementations, the memory 1906 of one or more computing devices 1900 in the computing device cluster may also store some instructions for executing the data clustering method. In other words, the combination of one or more computing devices 1900 can jointly execute the instructions for executing the data clustering method.
[0163] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be executed on a computing device or stored in any available medium. When the computer program product is executed on at least one computing device, the at least one computing device executes a data clustering method.
[0164] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the data clustering method, or instruct the computing device to execute the data clustering method.
[0165] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0166] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, display, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the raw data and executable code involved in this application were obtained with full authorization.
[0167] In the embodiments of the present application, the terms "first," "second," and "third" are used for descriptive purposes only and should not be understood as indicating or implying relative importance. The term "at least one" refers to one or more, and the term "plurality" refers to two or more, unless otherwise expressly limited.
[0168] In this application, the term "and / or" simply describes an association between related objects, indicating that three possible relationships exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A data clustering method, characterized in that, The method includes: Obtaining a plurality of data to be clustered; Obtaining a plurality of first eigenvalues of target data, where the target data is any one of the plurality of data; Dividing at least two data with corresponding identical plurality of first eigenvalues into the same similar data set.
2. The method according to claim 1, wherein Before dividing at least two data with corresponding identical plurality of first eigenvalues into the same similar data set, it includes: Counting the total number of corresponding identical first eigenvalues of each two data among the plurality of data to obtain a plurality of total values; The dividing at least two data with corresponding identical plurality of first eigenvalues into the same similar data set includes: Performing a clustering process on the plurality of data in descending order of the plurality of total values. Wherein, performing the clustering process according to the i-th total value among the plurality of total values includes: among the plurality of data to be clustered in the plurality of data, dividing the plurality of data to be clustered with corresponding identical i-th total value of first eigenvalues into a similar data set.
3. The method according to claim 1 or 2, characterized in that The obtaining a plurality of first eigenvalues of target data includes: Obtaining a plurality of data shards from the target data; Obtaining a second eigenvalue of a target data shard, where the target data shard is any one of the plurality of data shards; Processing the second eigenvalues of the plurality of data shards to obtain a plurality of first eigenvalues of the target data. Any one first eigenvalue of the target data is obtained based on the second eigenvalues of some of the plurality of data shards, and any two first eigenvalues of the target data are obtained based on the second eigenvalues of different data shards.
4. The method according to claim 3, wherein The obtaining a second eigenvalue of a target data shard includes: Obtaining a plurality of data segments from the target data shard; Obtaining a modulus value of the target data shard modulo each data segment among the plurality of data segments; Based on the modulus values corresponding to the plurality of data segments, obtaining the second eigenvalue of the target data shard.
5. The method according to claim 4, characterized in that, In response to the target data shard, the data segment, and the modulus value being represented by strings, the obtaining the second eigenvalue of the target data shard based on the modulus values corresponding to the plurality of data segments includes: Concatenating the strings representing the modulus values corresponding to the plurality of data segments in the order of the strings representing the plurality of data segments in the string representing the target data shard to obtain the second eigenvalue of the target data shard.
6. The method according to claim 4 or 5, characterized in that Before obtaining the second eigenvalue of the target data shard based on the modulus values corresponding to the plurality of data segments, the method further includes: Obtaining an optimization coefficient of a target data segment, where the target data segment is any one of the plurality of data segments; Based on the optimization coefficient of the target data segment, eliminating the hash conflict of the modulus value corresponding to the target data segment to obtain the optimized modulus value of the target data segment.
7. The method according to claim 6, characterized in that, In response to the modulus value and the optimization coefficient being represented by strings, the optimized modulus value is obtained by concatenating the string representing the optimization coefficient and the string representing the modulus value.
8. The method according to any one of claims 4 to 6, characterized in that In response to the target data shard being represented by a string, obtaining a plurality of data segments from the target data shard includes: Respectively intercepting a plurality of first character segments from the string representing the target data shard to obtain the plurality of data segments, where each first character segment represents a data segment, the first character segment includes one character or a plurality of continuously arranged characters, and the positions of the characters in any two first character segments among the plurality of first character segments in the string representing the target data shard are different.
9. The method according to any one of claims 3 to 8, characterized in that In response to the second eigenvalue and the first eigenvalue being represented by a string, the first eigenvalue is obtained by character splicing based on the string representing the plurality of second eigenvalues.
10. The method according to any one of claims 3 to 9, characterized in that Before processing the second eigenvalues of the plurality of data shards to obtain the plurality of first eigenvalues of the target data, the method further includes: Performing a de-correlation process on the second eigenvalues of the plurality of data shards to obtain the second eigenvalues of the plurality of data shards after the de-correlation process, where the de-correlation process includes: sorting according to the numerical size.
11. The method according to any one of claims 3 to 10, characterized in that, In response to the target data being represented by a string, obtaining a plurality of data shards from the target data includes: Respectively intercepting a plurality of second character segments from the string representing the target data to obtain the plurality of data shards, where each second character segment represents a data shard, the second character segment includes a plurality of continuously arranged characters, and the positions of the characters in any two second character segments among the plurality of second character segments in the string representing the target data are different.
12. A data clustering device, characterized in that, The apparatus includes: An acquisition module, configured to acquire a plurality of data to be clustered; The acquisition module is further configured to acquire a plurality of first eigenvalues of target data, where the target data is any one of the plurality of data; A clustering module, configured to divide at least two data having corresponding identical plurality of first eigenvalues into the same similar data set.
13. The device according to claim 12, wherein Specifically, the clustering module is configured to: Count the total number of corresponding identical first eigenvalues for each pair of data among the plurality of data to obtain a plurality of total values; Perform a clustering process on the plurality of data in descending order of the plurality of total values, where performing the clustering process according to the i-th total value among the plurality of total values includes: among the plurality of data to be clustered in the plurality of data, dividing the plurality of data to be clustered having corresponding identical i-th total value of first eigenvalues into a similar data set.
14. The device according to claim 12 or 13, characterized in that, Specifically, the acquisition module is configured to: Obtain a plurality of data shards from the target data; Obtain a second eigenvalue of a target data shard, where the target data shard is any one of the plurality of data shards; Process the second eigenvalues of the plurality of data shards to obtain the plurality of first eigenvalues of the target data, where any one first eigenvalue of the target data is obtained based on the second eigenvalues of some of the plurality of data shards, and any two first eigenvalues of the target data are obtained based on the second eigenvalues of different data shards.
15. The device according to claim 14, wherein Specifically, the acquisition module is configured to: Obtain multiple data segments from the target data shard; Obtain the modulus value of the target data shard modulo each data segment among the multiple data segments; Based on the modulus values corresponding to the multiple data segments, obtain the second eigenvalue of the target data shard.
16. The device according to claim 15, characterized in that, In response to the target data shard, the data segment, and the modulus value being represented by strings, the obtaining module is specifically configured to: According to the order of the strings representing the multiple data segments in the string representing the target data shard, splice the strings representing the modulus values corresponding to the multiple data segments character by character to obtain the second eigenvalue of the target data shard.
17. The device according to claim 15 or 16, characterized in that, The obtaining module is further configured to: Obtain the optimization coefficient of the target data segment, where the target data segment is any one of the multiple data segments; Based on the optimization coefficient of the target data segment, eliminate the hash conflict of the modulus value corresponding to the target data segment to obtain the optimized modulus value of the target data segment.
18. The device according to claim 17, characterized in that, In response to the modulus value and the optimization coefficient being represented by strings, the optimized modulus value is obtained by splicing the string representing the optimization coefficient and the string representing the modulus value character by character.
19. The device according to any one of claims 15 to 17, characterized in that In response to the target data shard being represented by a string, the obtaining module is specifically configured to: Respectively intercept multiple first character segments from the string representing the target data shard to obtain the multiple data segments, where each first character segment represents a data segment, the first character segment includes one character or multiple continuously arranged characters, and the positions of the characters in any two first character segments among the multiple first character segments in the string representing the target data shard are different.
20. The device according to any one of claims 14 to 19, characterized in that In response to the second eigenvalue and the first eigenvalue being represented by strings, the first eigenvalue is obtained by splicing the strings representing multiple second eigenvalues character by character.
21. The device according to any one of claims 14 to 20, characterized in that, The obtaining module is further configured to: Perform a de-correlation process on the second eigenvalues of the multiple data shards to obtain the de-correlated second eigenvalues of the multiple data shards, where the de-correlation process includes: sorting according to the numerical size.
22. The device according to any one of claims 14 to 21, characterized in that, In response to the target data being represented by a string, the obtaining module is specifically configured to: Respectively intercept multiple second character segments from the string representing the target data to obtain the multiple data shards, where each second character segment represents a data shard, the second character segment includes multiple continuously arranged characters, and the positions of the characters in any two second character segments among the multiple second character segments in the string representing the target data are different.
23. A cluster of computing devices, characterized in that, Comprising a plurality of computing devices, the plurality of computing devices include a plurality of processors and a plurality of memories, program instructions are stored in the plurality of memories, and the plurality of processors run the program instructions to enable the computing device cluster to execute the method according to any one of claims 1 to 11.
24. A computer-readable storage medium, characterized in that, Comprising program instructions, when the program instructions run on a computing device, enabling the computing device to execute the method according to any one of claims 1 to 11.
25. A computer program product comprising instructions, characterized in that, When the instructions are run by the computing device cluster, the computing device cluster is caused to execute the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Data compression method and device
CN111061428A
Data block construction and comparison method and device, medium and equipment
CN112667144A
Data compression method, controller, device, medium and program product
CN115145467A
Data access method and device
CN116257180A
Using data similarity to select segments for garbage collection
CN116601596A