Method and device for determining backup node in global deduplication storage scenario

By merging backup tasks and selecting high-performance backup nodes in a global deduplication storage scenario using fingerprint feature vectors, the problem of low efficiency in backup node determination is solved, achieving efficient backup task management and storage resource optimization.

CN119356943BActive Publication Date: 2026-05-05ANHUI DINGJIA COMPUTER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI DINGJIA COMPUTER TECH CO LTD
Filing Date
2024-10-10
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In a global deduplication storage scenario, manually selecting backup nodes leads to low efficiency in determining backup nodes.

Method used

By determining the fingerprint feature vectors of multiple backup tasks, backup tasks with high similarity are merged, and target backup nodes are selected for data storage based on the performance information of backup nodes, while deleting duplicate data.

Benefits of technology

It improves the efficiency of backup node determination, simplifies the complexity of backup task scheduling, optimizes storage resource utilization, and reduces network bandwidth consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119356943B_ABST
    Figure CN119356943B_ABST
Patent Text Reader

Abstract

This application relates to a method and apparatus for determining backup nodes in a global deduplication storage scenario. The method includes: determining multiple backup tasks; each backup task includes data to be backed up; each backup task corresponds to a fingerprint feature vector determined based on the data to be backed up; merging the data to be backed up from backup tasks that meet merging conditions based on the feature distance between the fingerprint feature vectors corresponding to each backup task, generating new multiple backup tasks; determining the similarity between the new multiple backup tasks and multiple backup nodes based on the fingerprint feature vectors corresponding to the new multiple backup tasks and the fingerprint feature vectors of cached data in multiple backup nodes; and sequentially determining the target backup node for executing the backup tasks from the multiple backup nodes based on the similarity and the performance information of the backup nodes. This method can improve the efficiency of determining backup nodes in a global deduplication storage scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for determining backup nodes in a global deduplication storage scenario. Background Technology

[0002] With the rapid growth of data volume, enterprises may have dozens or even hundreds of backup nodes for the data to be backed up, which poses a significant challenge in managing these backup nodes. Data deduplication is a widely adopted technology, also known as deduplication. Deduplication refers to replacing redundant data in physical storage with logical references, achieving single-instance storage of data, significantly saving storage space and reducing user costs.

[0003] Distributed global deduplication is a technique that performs data deduplication across multiple storage locations (backup nodes). Traditional techniques typically require users to select a suitable backup node from multiple backup nodes for data backup. However, manually selecting a backup node leads to inefficient backup node determination in global deduplication storage scenarios. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for determining backup nodes in a global deduplication storage scenario, which can improve the efficiency of backup node determination in a global deduplication storage scenario, in order to address the above-mentioned technical problems.

[0005] Firstly, this application provides a method for determining backup nodes in a global deduplication storage scenario, including:

[0006] Multiple backup tasks are identified; each backup task includes data to be backed up; each backup task corresponds to a fingerprint feature vector determined based on the data to be backed up.

[0007] Based on the feature distance between the fingerprint feature vectors corresponding to each backup task, the backup data to be backed up in each backup task that meets the merging conditions is merged to generate multiple new backup tasks.

[0008] Based on the fingerprint feature vectors corresponding to the new multiple backup tasks and the fingerprint feature vectors of the cached data in the multiple backup nodes, the similarity between the new multiple backup tasks and the multiple backup nodes is determined;

[0009] Based on the similarity and performance information of the backup nodes, a target backup node for executing the backup task is sequentially determined from the plurality of backup nodes; the target backup node is used to store the backup data to be backed up for the backup task, and to delete duplicate data in the backup data and the cached data.

[0010] In one embodiment, the step of merging the backup data of backup tasks that meet the merging conditions based on the feature distance between the fingerprint feature vectors corresponding to each backup task to generate multiple new backup tasks includes:

[0011] Based on the feature distance between the fingerprint feature vectors corresponding to each of the backup tasks, determine the two backup tasks with the smallest feature distance;

[0012] If the total amount of data to be backed up in the two backup tasks is less than the preset total amount of data, then the two backup tasks are merged to obtain multiple new backup tasks, and the step of determining the two backup tasks with the smallest feature distance based on the feature distance between the fingerprint feature vectors corresponding to each backup task is returned, until the number of new multiple backup tasks is less than the preset number of categories.

[0013] In one embodiment, the preset total data volume is determined based on the total data volume of the plurality of backup tasks and the preset number of categories; the preset number of categories is determined based on the performance information and number of the plurality of backup nodes.

[0014] In one embodiment, determining the target backup node for performing the backup task sequentially from the plurality of backup nodes based on the similarity and performance information of the backup nodes includes:

[0015] For any of the new multiple backup tasks, a target backup node for executing the backup task is determined from the multiple backup nodes based on the similarity between the backup task and the multiple backup nodes and the performance information of the backup nodes.

[0016] If the target backup node completes any of the backup tasks, the other backup tasks besides the target backup node are taken as the new multiple backup tasks, and the other backup nodes besides the target backup node are taken as the new multiple backup nodes. Then, the step of determining the efficiency parameters of the multiple backup nodes based on the similarity between the target backup task and the multiple backup nodes and the performance information of the backup nodes is returned for any of the new multiple backup tasks.

[0017] In one embodiment, determining the target backup node for performing the backup task sequentially from the plurality of backup nodes based on the similarity and performance information of the backup nodes includes:

[0018] The storage cost and efficiency cost of the backup node are obtained; the storage cost is determined based on the deduplication savings of the backup node, which includes the amount of data saved by the backup node after deleting duplicate data in the data to be backed up and the cached data; the efficiency cost is determined based on the network performance and backup efficiency of the backup node.

[0019] Based on the similarity and the data movement cost of the backup node, the transmission cost of the backup node is determined;

[0020] Based on the storage cost, the efficiency cost, and the transmission cost, determine the efficiency parameters of the plurality of backup nodes;

[0021] Based on the efficiency parameters of the plurality of backup nodes, a target backup node for performing the backup task is sequentially determined from the plurality of backup nodes.

[0022] In one embodiment, before merging the backup data of backup tasks that meet the merging conditions based on the feature distance between the fingerprint feature vectors corresponding to each backup task, the method further includes:

[0023] Multiple data blocks to be backed up are determined from the data to be backed up in the backup task;

[0024] Determine multiple hash values ​​corresponding to any of the data blocks to be backed up;

[0025] The smallest hash value among the plurality of hash values ​​is used as the feature value of any of the data blocks to be backed up.

[0026] Based on the feature values ​​of each of the data blocks to be backed up, a fingerprint feature vector corresponding to the backup task is constructed.

[0027] In one embodiment, before determining the similarity between the new multiple backup tasks and the multiple backup nodes based on the fingerprint feature vectors corresponding to the new multiple backup tasks and the fingerprint feature vectors of cached data in the multiple backup nodes, the method further includes:

[0028] The cached data is determined based on the hot data in the backup node; the hot data includes data in the backup node whose access frequency is higher than a frequency threshold.

[0029] Multiple node data blocks are determined from the cached data;

[0030] The fingerprint feature vector of the cached data is determined based on the fingerprints of the multiple node data blocks.

[0031] Secondly, this application also provides a device for determining backup nodes in a global deduplication storage scenario, comprising:

[0032] A task module is used to determine multiple backup tasks; each backup task includes data to be backed up; each backup task corresponds to a fingerprint feature vector determined based on the data to be backed up.

[0033] The merging module is used to merge the backup data of backup tasks that meet the merging conditions in each backup task according to the feature distance between the fingerprint feature vectors corresponding to each backup task, and generate multiple new backup tasks.

[0034] The determination module is used to determine the similarity between the new multiple backup tasks and the multiple backup nodes based on the fingerprint feature vectors corresponding to the new multiple backup tasks and the fingerprint feature vectors of cached data in the multiple backup nodes; the backup nodes are used to store the data to be backed up and to delete duplicate data in the data to be backed up and the cached data;

[0035] The backup module is used to sequentially determine the target backup node for performing the backup task from the plurality of backup nodes based on the similarity and performance information of the backup nodes.

[0036] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0037] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0038] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0039] The method, apparatus, computer equipment, computer-readable storage medium, and computer program product for determining backup nodes in the aforementioned global deduplication storage scenario determine multiple backup tasks. Each backup task includes data to be backed up, and each backup task corresponds to a fingerprint feature vector determined based on the data to be backed up. Based on the feature distance between the fingerprint feature vectors corresponding to each backup task, the data to be backed up in the backup tasks that meet the merging conditions are merged to generate multiple new backup tasks. The similarity between the new backup tasks and the backup nodes is determined based on the fingerprint feature vectors corresponding to the new backup tasks and the fingerprint feature vectors of the cached data in the multiple backup nodes. Based on the similarity and the performance information of the backup nodes, backup tasks are sequentially selected from the multiple backup nodes to execute the backup tasks. The target backup node is used to store the backup data to be backed up for backup tasks and to delete duplicate data in the backup data and cached data. By analyzing the backup data in the backup tasks and determining its fingerprint feature vector, highly similar backup data is merged, thereby efficiently classifying backup tasks. By comparing the fingerprint feature vectors of newly generated backup tasks with the cached data in the backup nodes, the similarity analysis between each backup task and each backup node is realized. Based on the similarity and the performance information of the backup nodes, the target backup nodes for executing backup tasks are determined in sequence. This ensures that the backup data is assigned to backup nodes with high data similarity and excellent performance, simplifies the complexity of backup task scheduling, and improves the efficiency of backup node determination in global deduplication storage scenarios. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is an application environment diagram of a method for determining backup nodes in a global deduplication storage scenario, as shown in one embodiment.

[0042] Figure 2 This is a flowchart illustrating a method for determining backup nodes in a global deduplication storage scenario, as shown in one embodiment.

[0043] Figure 3 This is a schematic diagram of the structure of a backup node in one embodiment;

[0044] Figure 4 This is a flowchart illustrating a backup task merging process in one embodiment.

[0045] Figure 5 This is a schematic diagram illustrating the result of merging backup tasks in one embodiment;

[0046] Figure 6 This is a flowchart of a method for selecting a backup node in one embodiment;

[0047] Figure 7 This is a structural block diagram of a backup node determination device in a global deduplication storage scenario, as shown in one embodiment.

[0048] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0050] The method for determining backup nodes in a global deduplication storage scenario provided in this application embodiment can be applied to, for example... Figure 1 The application environment shown is a global deduplication storage scenario. This refers to a scenario where data deduplication technology is applied in a distributed system spanning multiple storage nodes, devices, or storage pools to reduce redundant data storage. This global deduplication storage scenario can include a backup system deployed through a distributed architecture that applies deduplication technology to deduplicate data globally (across multiple storage nodes) to improve storage efficiency and optimize resource utilization.

[0051] like Figure 1 As shown, the backup system may include a network device 100 and multiple servers. The network device 100 may connect to one or more servers, which may include computing servers, virtualization servers and storage servers, providing various cloud services and business applications. These servers, through connection with the network device 100, realize data management, data monitoring, data backup and data disaster recovery.

[0052] The backup system may include multiple servers, such as backup center server 102 and backup nodes 104. Multiple backup nodes 104 can form a deduplication storage cluster, which may include multiple collaborative storage nodes that manage and store data globally through deduplication technology.

[0053] Among them, network equipment 100 may include, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, portable wearable devices and other terminals, and may also include workstations within an enterprise.

[0054] Backup node 104 can be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.

[0055] Backup service center 102 can be used to create and assign backup tasks to backup node 104 based on the backup data uploaded by network device 100, and can monitor and manage the data backup process and status, thereby providing data recovery and business continuity support in the event of a disaster.

[0056] The backup center server 102 communicates with the network device 100 and the backup node 104 via the network. Specifically, the network device 100 can connect to the backup node 104 through a network connector and a caching service. The network device 100 can also connect to the backup center server 102 through a network connector.

[0057] Backup center server 102 determines multiple backup tasks; each backup task includes data to be backed up; each backup task corresponds to a fingerprint feature vector determined based on the data to be backed up; backup center server 102 merges the data to be backed up of backup tasks that meet the merging conditions based on the feature distance between the fingerprint feature vectors corresponding to each backup task, generating multiple new backup tasks; backup center server 102 determines the similarity between the multiple new backup tasks and multiple backup nodes 104 based on the fingerprint feature vectors corresponding to the multiple new backup tasks and the fingerprint feature vectors of the cached data in multiple backup nodes 104; backup center server 102 determines the target backup node for executing the backup tasks sequentially from the multiple backup nodes 104 based on the similarity and the performance information of the backup nodes 104; the target backup node is used to store the data to be backed up for the backup tasks and delete duplicate data in the data to be backed up and the cached data.

[0058] In one exemplary embodiment, such as Figure 2 As shown, a method for determining backup nodes in a global deduplication storage scenario is provided, and this method is applied to... Figure 1 Taking backup center server 102 as an example, the explanation includes:

[0059] Step S202: Identify multiple backup tasks.

[0060] The backup task includes the data to be backed up; each backup task corresponds to a fingerprint feature vector determined based on the data to be backed up.

[0061] In practice, the backup task includes a data set consisting of one or more files. Each backup task is responsible for backing up a specific data set. The backup center server can determine the fingerprint feature vector of the backup task based on the data to be backed up.

[0062] A fingerprint feature vector can be a unique identifier generated based on data content, used to describe the data characteristics in the backup task. A fingerprint feature vector can uniquely identify the data set of the backup task.

[0063] As an example, the backup center server can calculate a fingerprint set for the backup task before the backup task begins, and construct a fingerprint feature vector from the fingerprint set. The feature values ​​in the fingerprint feature vector can be fingerprints from the fingerprint set, thereby identifying the data to be backed up by the backup task through the fingerprint feature vector. Optionally, the fingerprint set can include multiple fingerprints. A fingerprint is a short, unique identifier or digest generated from data, used to characterize the features of that data.

[0064] As an example, the fingerprints in the fingerprint set can be hash fingerprints determined by a hash algorithm. A hash algorithm is an algorithm that converts an input of any length (such as text or data) into a fixed-length output through a specific algorithm. The output of the hash algorithm can be called a hash value, which can be used to uniquely identify the data content of the data to be backed up.

[0065] In one embodiment, the fingerprint set of the backup task can be a set of fingerprints corresponding to each data block in the data to be backed up. The size of the data block to be backed up can be determined according to the file structure and specifications of the data to be backed up. For example, if the data to be backed up includes large files such as virtual machines and databases, the data to be backed up can be divided into 64 kilobytes (KB) or 128 kB data blocks, and a fingerprint can be calculated for each data block. If the data to be backed up is small unstructured data, such as logs, documents, or images, the data to be backed up can be divided into 8 KB or 16 kB data blocks, and a fingerprint can be calculated for each data block. Thus, the fingerprint set of the backup task is formed based on the fingerprints of each data block, and a fingerprint feature vector is constructed based on the fingerprint set. The feature values ​​in the fingerprint feature vector can be fingerprints from the fingerprint set.

[0066] In one embodiment, when calculating a fingerprint for each data block, the calculated fingerprint can be truncated to reduce its size. For example, a 64KB data block can generate a 256-bit fingerprint value. However, to save storage space and improve processing efficiency, while ensuring the data collision rate, N bits (N < 256, N is a power of 2) can be used for representation. If the fingerprint value is a hash value, a compact hash sequence can be constructed based on the hash values ​​of each data block. The fingerprint feature vector generated based on this hash sequence represents the data characteristics of the data to be backed up, supporting data similarity comparison between backup nodes in cross-backup tasks and deduplication storage scenarios.

[0067] Step S204: Based on the feature distance between the fingerprint feature vectors corresponding to each backup task, merge the backup data of the backup tasks that meet the merging conditions in each backup task to generate multiple new backup tasks.

[0068] The feature distance between fingerprint feature vectors can be used to represent the similarity between backup tasks, and can be calculated based on the fingerprint feature vectors of two backup tasks. The smaller the feature distance, the more similar the data content of the data to be backed up in the two backup tasks.

[0069] In a specific implementation, the feature distance between fingerprint feature vectors can include, but is not limited to, Euclidean distance, Manhattan distance, Jaccard distance, etc.

[0070] In practice, if the feature distance between two or more backup tasks is within a certain threshold range (i.e., high similarity), the backup center server can merge the backup data of these tasks into a new backup task. After the merging operation, a new set of backup tasks can be generated. Each new backup task may contain a combination of multiple backup data, which reduces the complexity of the backup operation, simplifies the subsequent process of determining backup nodes, improves the efficiency of determining backup nodes, and allows highly similar backup data to be processed concurrently within the same backup task and assigned to the same backup node. Since the backup data being processed has a certain degree of similarity, it reduces the pressure on network transmission and improves the efficiency parameters of the backup node and the efficiency of deduplication.

[0071] Step S206: Determine the similarity between the new backup tasks and the backup nodes based on the fingerprint feature vectors corresponding to the new backup tasks and the fingerprint feature vectors of the cached data in the backup nodes.

[0072] In this context, a backup node refers to a server responsible for performing backup tasks in a global deduplication storage scenario. It can also use a deduplication pool to deduplicate the backup data, thereby deleting duplicate data and retaining a unique, non-duplicate data copy.

[0073] For the convenience of those skilled in the art, Figure 3An exemplary schematic diagram of a backup node is provided. The backup node includes a file storage and retrieval system with data classification capabilities. In this system, clustering algorithms group the most similar file sets together and iteratively expand the deduplication storage cluster to include more file sets until the expected capacity requirement fills the deduplication node. Simultaneously, data input and output are accelerated to minimize memory and computational overhead, forming a deduplication storage cluster with minimized storage requirements. In one embodiment, the file storage and retrieval system analyzes files using a storage processor, partitions and manages files and their metadata using storage areas, performs content-aware processing on files using a content-aware processor, promptly releases cache space in hot storage areas using a space release processor, and reads files to be read using a read processor. Therefore, on the one hand, by being aware of the content of files, the file storage and retrieval system can obtain metadata that includes the semantic features of the files, which facilitates the subsequent merging of various files based on the file metadata, thereby reducing the space resources occupied by file storage and the time resources spent on file retrieval. On the other hand, by partitioning and managing data through cold storage areas and hot storage areas, the characteristics of each cold storage area and hot storage area can be fully utilized, reducing the time resources spent on file retrieval. Based on the above file storage and retrieval system, both the space resources occupied by small file storage and the time resources spent on file retrieval can be reduced, thus improving the resource utilization rate during file storage and retrieval.

[0074] The cached data in the backup node is the data that has been stored in the backup node. The cached data in the backup node can also correspond to a fingerprint feature vector, which is used to characterize the data features of the cached data in the backup node.

[0075] In practice, the fingerprint feature vector of the cached data can be generated based on the fingerprints corresponding to each data block in the cached data. The cached data can be data randomly extracted from backup nodes, or it can be important data selected from backup nodes that meets the backup format requirements. Optionally, the fingerprints corresponding to each data block in the cached data can also be hash values ​​calculated using a hash function.

[0076] In practice, after identifying the new backup tasks, the fingerprint feature vectors corresponding to the new backup tasks can be recalculated for the data to be backed up in these tasks. Based on the fingerprint feature vectors of the new backup tasks and the fingerprint feature vectors of the cached data in the backup nodes, the similarity between the new backup tasks and the backup nodes can be determined. Optionally, Euclidean distance similarity, cosine similarity, Jaccard similarity, etc., can be calculated between the fingerprint feature vectors of the data to be backed up and the fingerprint feature vectors of the cached data to determine the similarity between the new backup tasks and the backup nodes.

[0077] The higher the similarity between the backup task and the backup node, the more the data to be backed up in the backup task overlaps with the existing cached data on the backup node. This allows similar data to be stored centrally on one node, improving deduplication efficiency. Furthermore, during the process of deduplication technology to remove duplicate data, there is no need to frequently transfer data between backup nodes, reducing network bandwidth consumption.

[0078] Step S208: Based on similarity and performance information of backup nodes, determine the target backup node for performing the backup task from multiple backup nodes in sequence.

[0079] The target backup node is used to store the backup data to be backed up for the backup task and to remove duplicate data from the backup data and cached data.

[0080] As an example, the performance information of backup nodes may include, but is not limited to, the storage speed, data storage density, network characteristics, and load status of backup nodes. Based on multiple metrics such as similarity and the performance information of backup nodes, a suitable backup node for data storage can be planned for each backup task.

[0081] In practice, for any backup task, the backup center server prioritizes backup nodes based on similarity and performance information. It selects a backup node with high similarity and excellent performance from among multiple nodes and uses this node as the target backup node for the task. Since the data to be backed up in this backup task has high similarity to the cached data already stored in the selected backup node, the complexity of backup task scheduling is simplified, the overall efficiency of the backup system is improved, and the goal of storage network load balancing is achieved. Then, for the next backup task, the similarity and performance information of the backup nodes are recalculated, and the most suitable target backup node for executing the next backup task is selected from among multiple nodes. This allows for the sequential determination of target backup nodes for multiple backup tasks.

[0082] If similar data is stored on different backup nodes, frequent data transfer between these nodes is necessary during deduplication, leading to significant network bandwidth consumption. Furthermore, dispersed storage of similar data can result in underutilization of backup node storage space, leading to low storage utilization. Therefore, by analyzing the data similarity between backup tasks and backup nodes, and combining this with performance information such as network characteristics, node storage density, and storage speed, the optimal backup node can be automatically selected. This reduces the complexity and error rate of manual node selection by users, improves the efficiency of backup node determination in global deduplication storage scenarios, and reduces the amount of redundant data transmitted over the network, thereby lowering network bandwidth usage and improving the efficiency of data backup and recovery.

[0083] As can be seen, this scheme describes a method for determining the backup node to perform backup tasks in a global deduplication storage scenario. By using the fingerprint feature vector and similarity analysis of the data, a suitable backup node is selected and the data deduplication storage is performed, which not only optimizes the allocation of backup tasks but also improves data storage and efficiency parameters.

[0084] The method, apparatus, computer equipment, computer-readable storage medium, and computer program product for determining backup nodes in the aforementioned global deduplication storage scenario determine multiple backup tasks. Each backup task includes data to be backed up, and each backup task corresponds to a fingerprint feature vector determined based on the data to be backed up. Based on the feature distance between the fingerprint feature vectors corresponding to each backup task, the data to be backed up in the backup tasks that meet the merging conditions are merged to generate multiple new backup tasks. The similarity between the new backup tasks and the backup nodes is determined based on the fingerprint feature vectors corresponding to the new backup tasks and the fingerprint feature vectors of the cached data in the multiple backup nodes. Based on the similarity and the performance information of the backup nodes, backup tasks are sequentially selected from the multiple backup nodes to execute the backup tasks. The target backup node is used to store the backup data to be backed up for backup tasks and to delete duplicate data in the backup data and cached data. By analyzing the backup data in the backup tasks and determining its fingerprint feature vector, highly similar backup data is merged, thereby efficiently classifying backup tasks. By comparing the fingerprint feature vectors of newly generated backup tasks with the cached data in the backup nodes, the similarity analysis between each backup task and each backup node is realized. Based on the similarity and the performance information of the backup nodes, the target backup nodes for executing backup tasks are determined in sequence. This ensures that the backup data is assigned to backup nodes with high data similarity and excellent performance, simplifies the complexity of backup task scheduling, and improves the efficiency of backup node determination in global deduplication storage scenarios.

[0085] In another embodiment, before merging the backup data of backup tasks that meet the merging conditions based on the feature distance between the fingerprint feature vectors corresponding to each backup task, the method further includes: determining multiple backup data blocks from the backup data in the backup task; determining multiple hash values ​​corresponding to any backup data block; using the minimum hash value among the multiple hash values ​​as the feature value of any backup data block; and constructing a fingerprint feature vector corresponding to the backup task based on the feature values ​​of each backup data block.

[0086] The data block to be backed up is a specific data block within the data to be backed up. Optionally, the data block to be backed up can be a data block randomly selected from the data to be backed up, and the size of the data block can be determined based on the file structure and specifications of the data to be backed up. For example, if the data to be backed up includes large files such as virtual machines or databases, the data block to be backed up can be a 64 kilobyte (KB) or 128 kB data block; if the data to be backed up is small unstructured data, such as logs, documents, or images, the data block to be backed up can be an 8 KB or 16 kB data block.

[0087] In practice, a single data block to be backed up can generate multiple hash values. The backup center server can calculate multiple hash values ​​corresponding to any data block to be backed up by constructing a sequence of hash functions. This sequence of hash functions can be defined using two unrelated hash functions.

[0088] As an example, suppose we define two unrelated hash functions as h1 and h2, and these two unrelated hash functions are a message digest (MD5) hash function and a secure hash (SHA1) function, then we can define the hash function sequence hi as:

[0089] hi = h1 + i * h2 mod p, 0 ≤ i <k。

[0090] Where i is an integer greater than or equal to 0 and less than k; p is a prime number greater than k; mod is the modulo operator; the larger k is, the more accurate the subsequent calculation of the similarity between feature vectors will be. Optionally, k can be 1024.

[0091] As another example, suppose we define two unrelated hash functions as h1 and h2, and these two unrelated hash functions are two multiplication and rotation hash (Murmur) functions with different initial seeds. Then we can define the hash function sequence hi as:

[0092] hi = h1 + i * h2 + i 2 , 0≤i <k。

[0093] Where i is an integer greater than or equal to 0 and less than k. The larger k is, the more accurate the subsequent calculation of the similarity between feature vectors will be. Optionally, k can be 1024.

[0094] As can be seen, by constructing the above hash function sequence, k hash values ​​corresponding to any data block to be backed up can be calculated. Compared with using k independent hash functions to calculate multiple hash values, the above method of constructing hash function sequence is simpler and more convenient.

[0095] After determining the multiple hash values ​​corresponding to each data block to be backed up, the minimum hash value of each data block can be taken and used as the feature value of that data block. After sequentially obtaining the feature values ​​of all data blocks to be backed up in the backup task, the fingerprint feature vector of the backup task can be obtained.

[0096] In one embodiment, assuming there are T backup tasks, the feature distance between the fingerprint feature vectors corresponding to each backup task is calculated, and a fingerprint feature vector is formed based on the calculated feature distance. The similarity matrix in the upper triangular shape can also be called a distance matrix. Taking the Jaccard distance as an example, this... The similarity upper triangular matrix (distance matrix) can be represented as shown in Table 1.

[0097] Table 1

[0098]

[0099] Among them, FP i A fingerprint feature vector that can be used to indicate backup tasks, J(FP) i FP j ) represents the Jaccard distance between the fingerprint feature vectors of the backup task.

[0100] As can be seen, by filtering and sampling the data to be backed up in the backup task, data blocks for fingerprint calculation are obtained. Furthermore, by calculating the feature distance between backup tasks, an upper triangular matrix of similarity between backup tasks is obtained. Subsequent similarity evaluation of backup tasks and backup nodes can quickly identify the optimal backup node and automatically perform load balancing of the storage network. During the processing, these feature vectors and feature matrices significantly reduce the system's resource requirements for memory and processors. Moreover, by using hash function sequences and continuously iterating and merging backup tasks, the accuracy of similarity evaluation between backup tasks and backup nodes can be further improved.

[0101] The technical solution of this embodiment uses hash values ​​to uniquely identify the data blocks to be backed up, and uses the smallest hash value among multiple hash values ​​as the feature value of the data blocks to be backed up, which improves the accuracy and efficiency of constructing the fingerprint feature vector corresponding to the backup task, thereby improving the accuracy and efficiency of merging the backup tasks in the future.

[0102] In another embodiment, the backup data of backup tasks that meet the merging conditions are merged according to the feature distance between the fingerprint feature vectors corresponding to each backup task to generate multiple new backup tasks. This includes: determining the two backup tasks with the smallest feature distance based on the feature distance between the fingerprint feature vectors corresponding to each backup task; if the total amount of data to be backed up by merging the data of the two backup tasks is less than the preset total amount of data, then the two backup tasks are merged to obtain multiple new backup tasks, and the step of determining the two backup tasks with the smallest feature distance based on the feature distance between the fingerprint feature vectors corresponding to each backup task is returned, until the number of new backup tasks is less than the preset number of categories.

[0103] To simplify the subsequent calculation of backup node selection, similar backup tasks can be merged. These merged backup tasks can use the same backup node, allowing backup tasks on each backup node to be processed concurrently. Since the backup data to be processed by the same backup node has a certain similarity, the pressure on network transmission is reduced and the storage efficiency is improved.

[0104] In another embodiment, the preset total data volume is determined based on the total data volume of multiple backup tasks and the preset number of categories; the preset number of categories is determined based on the performance information of multiple backup nodes and the number of nodes.

[0105] The preset number of categories indicates the number of new backup tasks generated after selectively merging multiple backup tasks. The preset number of categories can be determined based on performance information such as the number of backup nodes, network capabilities, and backup node storage density in the current backup system, and can be dynamically adjusted based on the performance information and number of backup nodes.

[0106] The preset total data volume can be used to indicate the upper limit of the total amount of data to be backed up in a backup task. If each backup task is treated as a separate class, to ensure the evenness of data storage, a range of data size for each class can be specified, that is, the preset total data volume of data to be backed up in a backup task can be defined. The preset total data volume can be determined based on the total data volume of multiple backup tasks and the preset number of categories.

[0107] In one embodiment, assuming the preset number of categories is cn, there are T backup tasks, i.e., T categories, and the preset total data volume can be expressed as: 1.2*total_size{T} / cn, where total_size{T} is the total data volume of all backup tasks. If the total data volume of a single backup task exceeds 1.2*total_size{T} / cn, then that backup task is classified as a separate category and is not merged.

[0108] For the convenience of those skilled in the art, Figure 4 An exemplary flowchart for merging backup tasks is provided. In a specific implementation, the backup center server can find the two backup tasks C with the smallest feature distance from the distance matrix shown in Table 1. i and C j If the total amount of data to be backed up in this backup task after merging is less than the preset total amount of data, then the two backup tasks C can be merged. i and C j Merge into a new class C ij Otherwise, return to the step of finding the two backup tasks with the smallest feature distance in the distance matrix; after merging, a new class C is obtained. ij Then, C can be calculated. ij The feature distance between backup tasks and other backup tasks is calculated, the distance matrix is ​​updated, and the steps to find the two backup tasks with the smallest feature distance from the distance matrix are returned. This merging continues until the number of new backup tasks is less than the preset number of categories, at which point the merging can stop, thus obtaining a new T. c One backup task, and T c Less than or equal to T.

[0109] For the convenience of those skilled in the art, Figure 5 An exemplary diagram illustrating the result of merging backup tasks is provided. Figure 5 In the diagram, T1 to T6 represent multiple backup tasks. T1 and T2 can be merged into a single class C1. Although T3 and C1 have high similarity (small feature distance), T3 has a large amount of data to be backed up (exceeding the preset total data volume). Therefore, T3 can be classified as a separate class C2. T4 to T6 meet the requirements for similarity (small feature distance) and backup task capacity (less than the preset total data volume), so they are merged into class C3. After merging, the classification result is obtained.

[0110] In another embodiment, before determining the similarity between the new backup tasks and the backup nodes based on the fingerprint feature vectors corresponding to the new backup tasks and the fingerprint feature vectors of the cached data in the backup nodes, the method further includes: determining cached data based on hot data in the backup nodes; determining multiple node data blocks from the cached data; and determining the fingerprint feature vectors of the cached data based on the fingerprints of the multiple node data blocks.

[0111] Hot data includes data in backup nodes that are accessed more frequently than a frequency threshold. In other words, hot data can refer to data that is frequently accessed or used in backup nodes.

[0112] Among them, the node data block is a specific data block in the cached data.

[0113] In practice, the backup center server can obtain hot data from the storage area of ​​the backup nodes and extract multiple node data blocks from the hot data. This extraction can be random or obtained through sampling and filtering. For multiple node data blocks of any backup node, the hash value of each node data block can be calculated using a hash function sequence. Based on the hash value of each node data block, a feature value is determined, thereby determining the fingerprint feature vector of the cached data.

[0114] In another embodiment, the backup center server can select F file sets from any backup node, and then extract M node data blocks from each file set. Based on the feature values ​​of the M node data blocks extracted from each of the F file sets, a fingerprint feature vector of the cached data is calculated. A larger value for F indicates a larger number of file sets selected, resulting in more accurate estimation of the similarity between the backup task and the backup node, but also increasing the computational load.

[0115] In practice, the backup center server determines the similarity between the new backup tasks and backup nodes based on the fingerprint feature vectors corresponding to the new backup tasks and the fingerprint feature vectors of the cached data in the backup nodes. This can be achieved by determining the fingerprint set for each backup task and the fingerprint set for the cached data in each backup node based on the feature values ​​in the fingerprint feature vectors. Then, for any backup task, the Jaccard similarity between the fingerprint set of any backup task and the fingerprint set of the cached data in each backup node can be calculated. The Jaccard similarity can be determined based on the ratio of the intersection to the union of the elements in the two sets. Finally, based on the Jaccard similarity between the fingerprint sets of each backup task and the fingerprint sets of the cached data in each backup node, a similarity matrix can be calculated. The values ​​in this similarity matrix include the Jaccard similarity between the fingerprint sets of each backup task and the fingerprint sets of the cached data in each backup node.

[0116] In the technical solution of this embodiment, since the fingerprint set stored by the backup node is massive, the backup node prioritizes comparing the fingerprint of the hot data in the hot storage area with the fingerprint of the data to be backed up in the backup task. This can maximize the guarantee that the intersection between the fingerprint set of the backup task and the fingerprint set of the backup node is not empty, thereby improving the accuracy of calculating the similarity between the backup task and the backup node, and thus improving the accuracy of determining the backup node in the global deduplication storage scenario.

[0117] In another embodiment, a target backup node for performing the backup task is sequentially determined from multiple backup nodes based on similarity and performance information of the backup nodes. This includes: obtaining the storage cost and efficiency cost of the backup node; the storage cost is determined based on the deduplication savings of the backup node, which includes the amount of data saved by the backup node after deleting duplicate data in the data to be backed up and the cached data; the efficiency cost is determined based on the network performance and backup efficiency of the backup node; the transmission cost of the backup node is determined based on similarity and data movement cost of the backup node; efficiency parameters of multiple backup nodes are determined based on storage cost, efficiency cost, and transmission cost; and the target backup node for performing the backup task is sequentially determined from multiple backup nodes based on the efficiency parameters of the multiple backup nodes.

[0118] Storage cost refers to the resource consumption or expense incurred by backup nodes when storing data, and is typically related to factors such as storage capacity, storage media, and deduplication efficiency. Nodes with lower storage costs mean greater savings in deduplication, i.e., requiring less physical space to store data. Storage cost is an important consideration when selecting backup nodes.

[0119] Deduplication savings refer to the amount of storage space saved by a backup node after using deduplication technology (removing duplicate data). This includes the savings after deleting duplicate data from the data to be backed up and the cached data. The greater the deduplication savings, the more efficiently the backup node can store data, reducing the physical storage usage.

[0120] Efficiency cost is the cost calculated based on the network performance and backup efficiency of the backup node. The network performance of the backup node determines the data transmission speed, while the backup efficiency determines the data processing and storage speed. The lower the efficiency cost, the higher the data transmission and storage processing capabilities of the backup node, and the faster it can complete the backup task.

[0121] Data movement cost refers to the cost incurred during the backup process of transferring data from the source node to the backup node. This cost may include network bandwidth consumption, data transmission time, and system resource usage. Lower data movement cost means higher efficiency in transferring data from the source node to the backup node.

[0122] Transmission cost is determined based on the similarity of backup nodes and the cost of data movement. Transmission cost represents the overall cost of transferring data from the source node to the backup node. The lower the transmission cost, the more efficient the data transfer to that node, and it is one of the important factors in selecting a target backup node.

[0123] Efficiency parameters are indicators derived by comprehensively considering storage costs, efficiency costs, and transmission costs, measuring the overall efficiency of a backup node when performing backup tasks. A higher efficiency parameter indicates that the backup node can complete backup tasks faster and more economically; therefore, among multiple backup nodes, the system will prioritize nodes with higher efficiency parameters.

[0124] The target backup node refers to the node selected from multiple backup nodes based on factors such as storage cost, efficiency cost, and transmission cost to perform backup tasks. This ensures that backup tasks are executed on the optimal node, maximizing the utilization efficiency of storage resources and optimizing backup performance.

[0125] In one embodiment, the efficiency parameter can be equal to the sum of storage cost and efficiency cost, divided by the transmission cost. The efficiency parameter can be expressed as:

[0126] ;

[0127] Where efficiency is the efficiency parameter, cost (saving) is the storage cost, cost (speed) is the efficiency cost, and cost (movment) is the transmission cost.

[0128] The technical solution of this embodiment automatically selects the optimal backup node by analyzing the data similarity between the backup task and the backup node, and by combining efficiency parameters determined by factors such as storage cost, efficiency cost and transmission cost. This improves the accuracy and efficiency of backup node selection, and the backup node can complete the backup task more quickly and economically, thereby improving the backup efficiency in the global deduplication storage scenario. At the same time, it reduces the complexity and error rate of users manually selecting nodes.

[0129] In another embodiment, determining a target backup node for performing a backup task from multiple backup nodes based on similarity and backup node performance information includes: for any backup task among the new multiple backup tasks, determining a target backup node for performing the backup task from multiple backup nodes based on the similarity between the backup task and the multiple backup nodes and the performance information of the backup nodes; if the target backup node has completed performing any backup task, taking other backup tasks besides the backup task as new multiple backup tasks, and taking other backup nodes besides the target backup node of the backup task as new multiple backup nodes, and returning to the step of determining the efficiency parameters of multiple backup nodes for any backup task among the new multiple backup tasks based on the similarity between the backup task and the multiple backup nodes and the performance information of the backup nodes.

[0130] For the convenience of those skilled in the art, Figure 6An exemplary flowchart of a method for selecting backup nodes is provided. In a specific implementation, the backup center server calculates the efficiency parameters of each backup node for any given backup task based on the similarity between any backup task and multiple backup nodes, as well as the performance information of the backup nodes. The backup nodes are then sorted from highest to lowest efficiency parameter. The selection of the backup node corresponding to the backup task can then begin, including the following steps: a) Determine the backup node with the highest efficiency parameter. b) Check whether the backup node meets all constraints, such as minimum savings, maximum migration, and maximum node capacity. Minimum savings refers to the minimum storage space that the backup node can save when storing data using data deduplication technology. Maximum migration refers to the maximum amount of data that the backup node needs to migrate during data deduplication to achieve data balance or system optimization goals. Maximum node capacity refers to the maximum amount of effective data that the backup node can store after applying data deduplication technology. c) If the constraints are met, select the backup node as the target backup node for executing the backup task and record the processing results. d. Remove the backup node so it is no longer used by other backup tasks. In other words, after the target backup node completes any backup task, other backup nodes besides the target backup node for that task are treated as new backup nodes. Each backup node can be excluded after completing one backup task, ensuring load balancing across all backup nodes, avoiding node overload or resource waste, and improving the reliability and stability of the entire backup system. e. Update the list of processed backup tasks. f. Update the status of the remaining backup nodes, recalculate the efficiency parameters of each backup node for the next backup task, and then return to step a until all constraints are met or all backup tasks have been processed. g. Output the final list of selected backup nodes and the processing results, including total deduplication savings and total migration costs.

[0131] In another embodiment, the backup center server sequentially determines the target backup node for performing backup tasks from multiple backup nodes based on similarity and performance information of the backup nodes. This includes: obtaining the storage cost and efficiency cost of the backup node; the storage cost is determined based on the deduplication savings of the backup node, which includes the amount of data saved by the backup node after deleting duplicate data in the data to be backed up and the cached data; the efficiency cost is determined based on the network performance and backup efficiency of the backup node; the transmission cost of the backup node is determined based on similarity and data movement cost of the backup node; the efficiency parameters of the multiple backup nodes are determined based on the storage cost, efficiency cost, and transmission cost; for any backup task in the new multiple backup tasks, the target backup node for performing any backup task is sequentially determined from the multiple backup nodes based on the efficiency parameters of the multiple backup nodes; if the target backup node has completed any backup task, the other backup tasks besides the target backup task are taken as new multiple backup tasks, and the other backup nodes besides the target backup node of the target backup task are taken as new multiple backup nodes, and the step of determining the efficiency parameters of the multiple backup nodes based on the similarity between the target backup task and the multiple backup nodes and the performance information of the backup nodes is returned for any backup task in the new multiple backup tasks.

[0132] The technical solution in this embodiment, through data similarity analysis and performance information, rationally selects backup nodes, reduces the amount of redundant data transmitted in the network, thereby reducing the overall redundancy of data, reducing network bandwidth usage, and improving the efficiency of data backup and recovery.

[0133] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0134] Based on the same inventive concept, this application also provides an apparatus for determining backup nodes in a global deduplication storage scenario, used to implement the method for determining backup nodes in the global deduplication storage scenario described above. The solution provided by this apparatus is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the apparatus for determining backup nodes in a global deduplication storage scenario provided below can be found in the limitations of the method for determining backup nodes in a global deduplication storage scenario described above, and will not be repeated here.

[0135] In one exemplary embodiment, such as Figure 7 As shown, a device for determining backup nodes in a global deduplication storage scenario is provided, comprising:

[0136] The task module 710 is used to determine multiple backup tasks; the backup tasks include data to be backed up; each backup task corresponds to a fingerprint feature vector determined based on the data to be backed up.

[0137] The merging module 720 is used to merge the backup data of backup tasks that meet the merging conditions in each backup task according to the feature distance between the fingerprint feature vectors corresponding to each backup task, and generate multiple new backup tasks.

[0138] The determination module 730 is used to determine the similarity between the new multiple backup tasks and the multiple backup nodes based on the fingerprint feature vectors corresponding to the new multiple backup tasks and the fingerprint feature vectors of the cached data in the multiple backup nodes; the backup nodes are used to store the data to be backed up and to delete duplicate data in the data to be backed up and the cached data.

[0139] The backup module 740 is used to sequentially determine the target backup node for performing the backup task from the plurality of backup nodes based on the similarity and performance information of the backup nodes.

[0140] In one embodiment, the merging module 720 is specifically used to determine the two backup tasks with the smallest feature distance based on the feature distance between the fingerprint feature vectors corresponding to each backup task; if the total amount of data to be backed up by merging the data included in the two backup tasks is less than a preset total amount of data, then the two backup tasks are merged to obtain multiple new backup tasks, and the step of determining the two backup tasks with the smallest feature distance based on the feature distance between the fingerprint feature vectors corresponding to each backup task is returned, until the number of tasks in the multiple new backup tasks is less than the preset number of categories.

[0141] In one embodiment, the preset total data volume is determined based on the total data volume of the plurality of backup tasks and the preset number of categories; the preset number of categories is determined based on the performance information and number of the plurality of backup nodes.

[0142] In one embodiment, the backup module 740 is configured to, for any one of the new plurality of backup tasks, determine a target backup node from the plurality of backup nodes for executing the any one backup task, based on the similarity between the any one backup task and the plurality of backup nodes and the performance information of the backup nodes.

[0143] If the target backup node completes any of the backup tasks, the other backup tasks besides the target backup node are taken as the new multiple backup tasks, and the other backup nodes besides the target backup node are taken as the new multiple backup nodes. Then, the step of determining the efficiency parameters of the multiple backup nodes based on the similarity between the target backup task and the multiple backup nodes and the performance information of the backup nodes is returned for any of the new multiple backup tasks.

[0144] In one embodiment, the backup module 740 is configured to obtain the storage cost and efficiency cost of the backup node; the storage cost is determined based on the deduplication savings of the backup node, the deduplication savings including the amount of data saved by the backup node after deleting duplicate data in the data to be backed up and the cached data; the efficiency cost is determined based on the network performance and backup efficiency of the backup node; the transmission cost of the backup node is determined based on the similarity and the data movement cost of the backup node; efficiency parameters of the plurality of backup nodes are determined based on the storage cost, the efficiency cost and the transmission cost; and a target backup node for performing the backup task is sequentially determined from the plurality of backup nodes based on the efficiency parameters of the plurality of backup nodes.

[0145] In one embodiment, the task module 710 is configured to determine a plurality of data blocks to be backed up from the data to be backed up in the backup task; determine a plurality of hash values ​​corresponding to any one of the data blocks to be backed up; take the smallest hash value among the plurality of hash values ​​as the feature value of any one of the data blocks to be backed up; and construct a fingerprint feature vector corresponding to the backup task based on the feature values ​​of each of the data blocks to be backed up.

[0146] In one embodiment, the determining module 730 is configured to, before determining the similarity between the new multiple backup tasks and the multiple backup nodes based on the fingerprint feature vectors corresponding to the new multiple backup tasks and the fingerprint feature vectors of cached data in the multiple backup nodes, further include: determining the cached data based on hot data in the backup nodes; the hot data includes data in the backup nodes whose access frequency is higher than a frequency threshold; determining multiple node data blocks from the cached data; and determining the fingerprint feature vector of the cached data based on the fingerprints of the multiple node data blocks.

[0147] The modules in the backup node determination device in the aforementioned global deduplication storage scenario can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0148] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores backup data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with an external backup center server via a network connection. When executed by the processor, the computer program implements a method for determining backup nodes in a global deduplication storage scenario.

[0149] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0150] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0151] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0152] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0153] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0154] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0155] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0156] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for determining backup nodes in a global deduplication storage scenario, characterized in that, The method includes: Multiple backup tasks are identified; each backup task includes data to be backed up; each backup task corresponds to a fingerprint feature vector determined based on the data to be backed up. Based on the feature distance between the fingerprint feature vectors corresponding to each backup task, the backup data to be backed up in each backup task that meets the merging conditions is merged to generate multiple new backup tasks. Based on the fingerprint feature vectors corresponding to the new multiple backup tasks and the fingerprint feature vectors of the cached data in the multiple backup nodes, the similarity between the new multiple backup tasks and the multiple backup nodes is determined; Based on the similarity and performance information of the backup nodes, a target backup node for executing the new backup tasks is sequentially determined from the plurality of backup nodes; the target backup node is used to store the data to be backed up for the new backup tasks and to delete duplicate data in the data to be backed up and the cached data.

2. The method according to claim 1, characterized in that, The step of merging the backup data of backup tasks that meet the merging conditions based on the feature distance between the fingerprint feature vectors corresponding to each backup task to generate multiple new backup tasks includes: Based on the feature distance between the fingerprint feature vectors corresponding to each of the backup tasks, determine the two backup tasks with the smallest feature distance; If the total amount of data to be backed up in the two backup tasks is less than the preset total amount of data, then the two backup tasks are merged to obtain a merged backup task, and the step of determining the two backup tasks with the smallest feature distance based on the feature distance between the fingerprint feature vectors corresponding to each backup task is returned, until the number of new backup tasks is less than the preset number of categories.

3. The method according to claim 2, characterized in that, The preset total data volume is determined based on the total data volume of the multiple backup tasks and the preset number of categories; the preset number of categories is determined based on the performance information and number of the multiple backup nodes.

4. The method according to claim 1, characterized in that, The step of sequentially determining the target backup node for executing the new multiple backup tasks from the plurality of backup nodes based on the similarity and performance information of the backup nodes includes: For any of the new multiple backup tasks, a target backup node for executing the backup task is determined from the multiple backup nodes based on the similarity between the backup task and the multiple backup nodes and the performance information of the backup nodes. If the target backup node has completed any of the backup tasks, the other backup tasks besides the any of the backup tasks are taken as the new plurality of backup tasks, and the other backup nodes besides the target backup node of the any of the backup tasks are taken as the new plurality of backup nodes. Then, the step of determining the target backup node for executing the any of the backup tasks is returned for any of the new plurality of backup tasks, based on the similarity between the any of the backup tasks and the plurality of backup nodes and the performance information of the backup nodes.

5. The method according to claim 1, characterized in that, The step of sequentially determining the target backup node for executing the new multiple backup tasks from the plurality of backup nodes based on the similarity and performance information of the backup nodes includes: The storage cost and efficiency cost of the backup node are obtained; the storage cost is determined based on the deduplication savings of the backup node, which includes the amount of data saved by the backup node after deleting duplicate data in the data to be backed up and the cached data; the efficiency cost is determined based on the network performance and backup efficiency of the backup node. Based on the similarity and the data movement cost of the backup node, the transmission cost of the backup node is determined; Based on the storage cost, the efficiency cost, and the transmission cost, determine the efficiency parameters of the plurality of backup nodes; Based on the efficiency parameters of the multiple backup nodes, the target backup node for executing the new multiple backup tasks is determined sequentially from the multiple backup nodes.

6. The method according to claim 1, characterized in that, Before merging the backup data of backup tasks that meet the merging conditions based on the feature distance between the fingerprint feature vectors corresponding to each backup task, the method further includes: Multiple data blocks to be backed up are determined from the data to be backed up in the backup task; Determine multiple hash values ​​corresponding to any of the data blocks to be backed up; The smallest hash value among the plurality of hash values ​​is used as the feature value of any of the data blocks to be backed up. Based on the feature values ​​of each of the data blocks to be backed up, a fingerprint feature vector corresponding to the backup task is constructed.

7. The method according to claim 1, characterized in that, Before determining the similarity between the new multiple backup tasks and the multiple backup nodes based on the fingerprint feature vectors corresponding to the new multiple backup tasks and the fingerprint feature vectors of cached data in the multiple backup nodes, the method further includes: The cached data is determined based on the hot data in the backup node; the hot data includes data in the backup node whose access frequency is higher than a frequency threshold. Multiple node data blocks are determined from the cached data; The fingerprint feature vector of the cached data is determined based on the fingerprints of the multiple node data blocks.

8. A device for determining backup nodes in a global deduplication storage scenario, characterized in that, The device includes: A task module is used to determine multiple backup tasks; each backup task includes data to be backed up; each backup task corresponds to a fingerprint feature vector determined based on the data to be backed up. The merging module is used to merge the backup data of backup tasks that meet the merging conditions in each backup task according to the feature distance between the fingerprint feature vectors corresponding to each backup task, and generate multiple new backup tasks. The determination module is used to determine the similarity between the new multiple backup tasks and the multiple backup nodes based on the fingerprint feature vectors corresponding to the new multiple backup tasks and the fingerprint feature vectors of cached data in the multiple backup nodes; The backup module is used to sequentially determine the target backup node for executing the new multiple backup tasks from the multiple backup nodes based on the similarity and performance information of the backup nodes; the target backup node is used to store the data to be backed up for the new multiple backup tasks, and to delete duplicate data in the data to be backed up and the cached data.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Business system backup method and device, equipment and storage medium

    CN113505027A

  • Method for quickly merging backup points in data deduplication system

    CN114077590A