Data deduplication method and device and electronic equipment

By distributing data between multiple processing devices for multiple rounds of local deduplication processing, the problem of low data deduplication efficiency caused by insufficient single-machine memory is solved, and efficient data deduplication and memory utilization are achieved.

CN119938655APending Publication Date: 2025-05-06DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411868324.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When processing large-scale data, existing data deduplication methods need to read all the data into memory, resulting in high memory pressure and low efficiency, especially in a stand-alone processing environment, which is difficult to effectively deduplicate.

Method used

The local deduplication strategy is adopted for multiple rounds of distribution of local deduplication, and the global data is distributed to multiple processing devices for local deduplication processing. The results of each round of processing are summarized into the input data of the next round until the stable or preset conditions are reached.

Benefits of technology

By distributing and processing data, the burden on stand-alone memory is reduced, the efficiency and processing capacity of data deduplication are improved, and the amount of data can be effectively reduced and the deduplication accuracy can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938655A_ABST
    Figure CN119938655A_ABST
Patent Text Reader

Abstract

The invention provides a data de-duplication method and device and electronic equipment, and the method comprises the steps: carrying out the local de-duplication processing of each text data set (global data) for training a large language model through employing a multi-round distribution local de-duplication strategy, and obtaining a first target de-duplication text set containing a plurality of pieces of target sample data; according to the data deduplication processing method and device, even under the actual limitation condition that the single-machine memory is not enough to support global data deduplication processing, the first processing device with the small memory is used for quickly carrying out multi-round local deduplication processing on the global data, the utilization rate of the memory is fully improved, the data volume of data deduplication is reduced, and the processing efficiency of data deduplication is improved. And then determining the plurality of pieces of target sample data as target training sample data for training the large language model. In this way, the data size of useless repeated data during large language model training can be effectively reduced, and the training effect and training efficiency of the large language model can be guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data deduplication method, device and electronic device. Background Art

[0002] With the rapid development of large language model technology, unsupervised training applications have become increasingly widespread, and the demand for unsupervised text data has increased dramatically. In the process of building training data, content duplication is a common phenomenon. Taking web page data as an example, a large number of reposts of the same page and regular small-scale updates of the same page can easily cause page content on the Internet to be repeated. Using such repeated training data for model training may cause the trained model to excessively "memorize" certain high-frequency content instead of learning real language patterns. In addition, repeated data will cause the training process to be unstable, thereby affecting model performance. Therefore, deduplication is a key step in the process of building data.

[0003] In existing data deduplication methods, it is usually necessary to read all data in TB (TeraByte) storage units into the memory of a specified machine device, which puts a large memory pressure on a single-machine processing environment and has a low data deduplication efficiency. Summary of the invention

[0004] In view of this, embodiments of the present application provide a data deduplication method, device, and electronic device to improve the data deduplication efficiency of existing data deduplication methods.

[0005] In a first aspect, an embodiment of the present application provides a data deduplication method, wherein the method comprises:

[0006] Acquire global data, wherein the global data is a set of text data for training a large language model;

[0007] Based on the global data, adopt multiple rounds of distribution of local deduplication strategies to perform local deduplication processing to obtain a target deduplication text data set, wherein the target deduplication text data set includes a number of target sample data;

[0008] The multi-round distribution local deduplication strategy includes: distributing global data to different first processing devices, and each of the first processing devices performs several rounds of local deduplication processing on the received local data; wherein the local data is a sub-data set of global data participating in data distribution in the current round, and the data set obtained by summarizing the local deduplication processing results of each of the first processing devices in the current round is the global data for the next round of local deduplication processing;

[0009] The plurality of target sample data are determined as target training sample data for training the large language model.

[0010] In some embodiments, after the step of distributing local deduplication strategies in multiple rounds based on the global data to perform local deduplication processing to obtain a target deduplication text data set, the method further includes:

[0011] Using a preset hash processing algorithm, split each piece of the target sample data into a plurality of hash segments;

[0012] Sending hash segments belonging to the same segment sequence number in each piece of the target sample data to the same second processing device, and each second processing device performing similarity calculation on the received hash segments;

[0013] Obtain the similarity calculation results output by each of the second processing devices. If the similarity calculation results meet the preset overlap condition, retain a first target sample data as the target reserved sample data, and delete the remaining first target sample data, wherein the first target sample data is the target sample data in which the similarity calculation results meet the preset overlap condition.

[0014] In some embodiments, based on the global data, adopting multiple rounds of distribution of local deduplication strategies to perform local deduplication processing to obtain a target deduplication text data set includes:

[0015] When performing local deduplication processing on the global data in the first round, the global data is randomly distributed to each of the first processing devices according to a preset random distribution strategy, and each of the first processing devices performs local deduplication processing on the received local data;

[0016] Obtain the local deduplication processing results output by the current round of each of the first processing devices, and distribute the summary data set of the local deduplication processing results output by the current round of each of the first processing devices to each of the first processing devices according to a preset redistribution strategy, and have each of the first processing devices perform the next round of local deduplication processing until the round of performing the local deduplication processing meets the preset termination round, or until the data volume of the summary data set output by each of the first processing devices converges.

[0017] In some embodiments, the use of a preset hash processing algorithm to split each piece of the target sample data into a plurality of segments includes:

[0018] Using a minimization hash processing algorithm, according to a preset number of stripes, the hash signature of each piece of the target sample data is divided into hash segments of the preset number of stripes;

[0019] Mapping hash segments belonging to the same stripe sequence number to a target hash bucket in the same second processing device, and performing similarity calculation on the hash segments in the same target hash bucket using a minimization hash processing algorithm or a locality sensitive hash processing algorithm;

[0020] If the hash segments in the same target hash bucket collide, the target sample data corresponding to the two colliding hash segments are output.

[0021] In some embodiments, after obtaining the similarity calculation results output by each of the second processing devices, if the similarity calculation results meet the preset coincidence condition, retaining a first target sample data as the target reserved sample data, and deleting the remaining first target sample data, the method further includes:

[0022] The remaining plurality of target reserved sample data are determined as target training sample data for training the large language model.

[0023] In a second aspect, the present application provides a data deduplication device, the device comprising:

[0024] A data acquisition module, used to acquire global data, wherein the global data is a set of text data for training a large language model;

[0025] A local deduplication processing module is used to adopt a multi-round distribution local deduplication strategy to perform local deduplication processing based on the global data to obtain a target deduplication text data set, wherein the target deduplication text data set includes a plurality of target sample data;

[0026] The multi-round distribution local deduplication strategy includes: distributing global data to different first processing devices, and each of the first processing devices performs several rounds of local deduplication processing on the received local data; wherein the local data is a sub-data set of global data participating in data distribution in the current round, and the data set obtained by summarizing the local deduplication processing results of each of the first processing devices in the current round is the global data for the next round of local deduplication processing;

[0027] A data output module is used to determine the plurality of target sample data as target training sample data for training the large language model.

[0028] In some embodiments, the local deduplication processing module is specifically used to:

[0029] When performing local deduplication processing on the global data in the first round, the global data is randomly distributed to each of the first processing devices according to a preset random distribution strategy, and each of the first processing devices performs local deduplication processing on the received local data;

[0030] Obtain the local deduplication processing results output by the current round of each of the first processing devices, and distribute the summary data set of the local deduplication processing results output by the current round of each of the first processing devices to each of the first processing devices according to a preset redistribution strategy, and have each of the first processing devices perform the next round of local deduplication processing until the round of performing the local deduplication processing meets the preset termination round, or until the data volume of the summary data set output by each of the first processing devices converges.

[0031] In some embodiments, the apparatus further comprises:

[0032] The global deduplication processing module is used to perform the following processing on each piece of the target sample data output by the local deduplication processing module:

[0033] Using a preset hash processing algorithm, split each piece of the target sample data into a plurality of hash segments;

[0034] Sending hash segments belonging to the same segment sequence number in each piece of the target sample data to the same second processing device, and each second processing device performing similarity calculation on the received hash segments;

[0035] Obtain the similarity calculation results output by each of the second processing devices. If the similarity calculation results meet the preset overlap condition, retain a first target sample data as the target reserved sample data, and delete the remaining first target sample data, wherein the first target sample data is the target sample data in which the similarity calculation results meet the preset overlap condition.

[0036] In some embodiments, the use of a preset hash processing algorithm to split each piece of the target sample data into a plurality of segments includes:

[0037] Using a minimization hash processing algorithm, according to a preset number of stripes, the hash signature of each piece of the target sample data is divided into hash segments of the preset number of stripes;

[0038] Mapping hash segments belonging to the same stripe sequence number to a target hash bucket in the same second processing device, and performing similarity calculation on the hash segments in the same target hash bucket using a minimization hash processing algorithm or a locality sensitive hash processing algorithm;

[0039] If the hash segments in the same target hash bucket collide, the target sample data corresponding to the two colliding hash segments are output.

[0040] In some embodiments, after obtaining the similarity calculation results output by each of the second processing devices, if the similarity calculation results meet a preset overlap condition, retaining one first target sample data as target reserved sample data and deleting the remaining first target sample data, the data output module is specifically used to determine the remaining several target reserved sample data as target training sample data for training the large language model.

[0041] In a third aspect, an embodiment of the present application provides an electronic device, wherein the electronic device includes:

[0042] Processor; and

[0043] Memory for storing programs,

[0044] Wherein, the program includes instructions, and when the instructions are executed by the processor, the processor executes the data deduplication method described in the first aspect.

[0045] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium storing computer instructions, characterized in that the computer instructions are used to enable a computer to execute the data deduplication method described in the first aspect.

[0046] Beneficial effects of this application:

[0047] The present application provides a data deduplication method, device and electronic device, wherein the method uses a multi-round distribution local deduplication strategy to perform local deduplication processing on each text data set (i.e., global data) for training a large language model, obtains a first target deduplication text set containing several target sample data, and then determines the several target sample data as the target training sample data for training the large language model. In this way, the amount of useless duplicate data during the training of the large language model can be effectively reduced, which helps to ensure the training effect and training efficiency of the large language model.

[0048] And because the embodiment of the present application adopts the method of distributing the global data to different first processing devices for performing several rounds of local deduplication processing, each round of local deduplication processing can remove part of the duplicate data, and each first processing device shares the amount of data required for deduplication processing. In this way, even in the realistic limitation that the memory of a single machine is insufficient to support the global data deduplication processing, the first processing device with smaller memory can quickly perform multiple rounds of local deduplication processing on the global data, which fully improves the memory utilization rate, reduces the amount of data to be deduplicated, and helps to improve the processing efficiency of data deduplication. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Further details, features and advantages of the present application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0050] Figure 1 A schematic diagram of a process flow of a data deduplication method provided in an embodiment of the present application is shown;

[0051] Figure 2 Another schematic diagram of a data deduplication method provided in an embodiment of the present application is shown;

[0052] Figure 3 Another schematic diagram of a data deduplication method provided in an embodiment of the present application is shown;

[0053] Figure 4 Another schematic diagram of a data deduplication method provided in an embodiment of the present application is shown;

[0054] Figure 5 A schematic diagram of a process flow of a data deduplication device provided in an embodiment of the present application is shown;

[0055] Figure 6 A schematic diagram of a logical structure of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0056] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are only for exemplary purposes and are not intended to limit the scope of protection of the present application.

[0057] It should be understood that the various steps described in the method implementation of the present application can be performed in different orders and / or performed in parallel. In addition, the method implementation may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0058] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0059] It should be noted that the modifications of "one" and "plurality" mentioned in the present application are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0060] Before describing the data deduplication method, device, and electronic device provided by the present application, the technical terms involved in this application are described here:

[0061] LSH: Locality Sensitive Hashing, a locality sensitive hashing algorithm, is a fast approximate, nearest neighbor search technology for processing massive high-dimensional data. Specifically, LSH maps data items to data buckets (a common name for data containers) through hash functions, so that similar data items can be mapped to the same data bucket with a high probability, which can reduce the search scale and speed up the search.

[0062] MinhashLSH: A fast similarity search algorithm that combines Minhash and locality sensitive hashing (LSH) techniques, mainly used to handle the approximate nearest neighbor search problem of large-scale data. Specifically, by using multiple hash functions to map high-dimensional data to low-dimensional space, the amount of distance calculation between data is reduced. The algorithm first uses Minhash technology to hash the data set, and then uses the LSH strategy to group these hash values. Similar data items will be mapped to the same data bucket, which helps to improve data search efficiency.

[0063] LLM: Large Language Model is an artificial intelligence model that is trained with massive text data, has a large number of model parameters and can perform language tasks.

[0064] Jaccard similarity: An indicator used to measure the similarity between two data sets, defined as the ratio between the size of the intersection and the size of the union of the two data sets.

[0065] As described in the background technology, a large amount of duplicate data will cause instability in the LLM model training process and affect the performance of the LLM model. Therefore, data deduplication becomes a key step in the process of building training data. When processing large-scale data, the prior art usually uses an approximate hashing algorithm to improve computational efficiency. For example, MinhashLSH is used to focus on the overall similarity between multiple documents, rather than a complete match. Compared with the complete match method, the use of MinhashLSH to focus on the overall similarity between multiple documents can improve computational efficiency. Among them, MinhashLSH combines the idea of ​​the Minhash algorithm and the technology of LSH local sensitive hashing, and uses a hash function to convert the original text features into a set of signatures, which are used to estimate the Jaccard similarity between two data sets.

[0066] For each data set, the Minhash algorithm selects multiple hash functions, calculates the hash value of each element in the data set, and then takes the minimum value for each hash function. The minimum value is the Minhash value of the data set. In this way, for two data sets, if their Minhash values ​​under multiple hash functions are similar or the same, the two data sets are considered to be similar. In order to improve processing efficiency, the Minhash algorithm is often used in conjunction with local sensitive hashing LSH. The LSH is an approximate nearest neighbor search algorithm that hashes data objects into multiple "data buckets" so that similar data objects are hashed into the same data bucket as much as possible, thereby improving processing efficiency by reducing the number of object pairs that need to be directly compared.

[0067] The specific processing flow is: use the Minhash algorithm to generate a signature for each text in the data set (this signature is a hash signature), and then split each Minhash signature into multiple smaller data fragments, each of which contains a portion of the hash value in the Minhash signature. For each fragment, use the LSH algorithm to hash it into multiple data buckets, and similar content is more easily mapped to the same data bucket. Among them, the data bucket specifically refers to the hash bucket. During the hash processing process, the hash values ​​of the same address are attributed to the same subset, which is the hash bucket.

[0068] However, during the execution of the MinhashLSH algorithm, the hash signature needs to be divided into bands first, and then the hash fragments on the same band need to be bucketed by LSH mapping. This is a global data operation, which requires reading the hash fragments at the corresponding positions of the full amount of data for comparison in order to find similar content. In the context of large language models, the data used for deduplication is often counted in TB units. All of this data needs to be written into the memory of the same processing device and then read from the memory of the processing device, which places high memory requirements on a single machine device. For scenarios with multiple independent devices, it is difficult to achieve the goal of global deduplication, which affects the deduplication effect.

[0069] In view of this, the present application provides a data deduplication method, device and electronic device, wherein, in the first aspect, the present application provides a data deduplication method, which is applied to any electronic device with data deduplication function, including but not limited to personal mobile terminals, computers or servers, etc. As an embodiment, the data deduplication method can be applied to a training device for training a large language model, so that when training a large language model, the data deduplication method is called to perform deduplication processing on the global data participating in the model training. As another embodiment, the data deduplication method can be applied to a management device in a database storing global data, and the management device performs deduplication processing on the global data in the database regularly or irregularly, so that the deduplication training sample data can be directly obtained from the database when the large language model is subsequently trained.

[0070] In some embodiments, Figure 1 As shown, the data deduplication method includes the following steps:

[0071] S11, obtaining global data, wherein the global data is a set of text data for training a large language model;

[0072] S12, based on the global data, adopting multiple rounds of distribution of local deduplication strategies to perform local deduplication processing to obtain a target deduplication text data set, wherein the target deduplication text data set includes a plurality of target sample data;

[0073] The multi-round distribution local deduplication strategy includes: distributing global data to different first processing devices, and each of the first processing devices performs several rounds of local deduplication processing on the received local data; wherein the local data is a sub-data set of global data participating in data distribution in the current round, and the data set obtained by summarizing the local deduplication processing results of each of the first processing devices in the current round is the global data for the next round of local deduplication processing;

[0074] S13: Determine the plurality of target sample data as target training sample data for training the large language model.

[0075] The method uses a multi-round distribution local deduplication strategy to perform local deduplication processing on each text data set (i.e., global data) for training a large language model, obtains a first target deduplication text set containing a number of target sample data, and then determines the several target sample data as the target training sample data for training the large language model. In this way, the amount of useless duplicate data during large language model training can be effectively reduced, which helps to ensure the training effect and training efficiency of the large language model.

[0076] And because the embodiment of the present application adopts the method of distributing the global data to different first processing devices for performing several rounds of local deduplication processing, each round of local deduplication processing can remove part of the duplicate data, and each first processing device shares the amount of data required for deduplication processing. In this way, even in the realistic limitation that the memory of a single machine is insufficient to support the global data deduplication processing, the first processing device with smaller memory can quickly perform multiple rounds of local deduplication processing on the global data, which fully improves the memory utilization rate, reduces the amount of data to be deduplicated, and helps to improve the processing efficiency of data deduplication.

[0077] In some embodiments, the local deduplication processing of global data in the above steps S11 to S13 may be a first stage of data deduplication processing. On the basis of the first stage of data deduplication processing, the method further includes a second stage of data deduplication processing, specifically including the following steps S14 to S16:

[0078] S14, using a preset hash processing algorithm to split each piece of the target sample data into a plurality of hash segments;

[0079] S15, sending the hash segments belonging to the same segment sequence number in each piece of the target sample data to the same second processing device, and each second processing device performs similarity calculation on the received hash segments;

[0080] S16. Obtain the similarity calculation results output by each of the second processing devices. If the similarity calculation results meet the preset overlap condition, retain a first target sample data as the target reserved sample data, and delete the remaining first target sample data, wherein the first target sample data is the target sample data in which the similarity calculation results meet the preset overlap condition.

[0081] As an implementation method, the target sample data obtained by the local deduplication processing in the first stage can be determined as the training sample data for training the large language model. Compared with the initial global data, the degree of similarity between the training sample data and the training sample data at this time is greatly reduced, and the amount of data is relatively small. However, there may still be a situation where the amount of training sample data obtained is still higher than the memory of a single machine and cannot be read in completely. Therefore, in an embodiment of the present application, the target sample data obtained by the local deduplication processing in the first stage can be subjected to the second stage of global deduplication processing by executing steps S14 to S16. The data deduplication method adopted in the present application is a two-stage deduplication scheme, in which local deduplication is performed in the first stage and global deduplication is adopted in the second stage.

[0082] The first stage of local deduplication is mainly carried out through multiple rounds of deduplication to minimize the amount of data that needs to be processed, thereby alleviating the memory pressure of the second stage of global deduplication. The second stage adopts global deduplication, and the core is to use a multi-machine bucketing strategy to reduce the memory pressure of a single machine. In this way, the embodiment of the present application adopts a two-stage technical concept of deduplication processing: the local deduplication processing in the first stage can reduce the pressure of the global data volume on the memory of a single machine, and then use the multi-machine bucketing strategy to distribute the hash fragments to multiple machines and devices, so that a single machine is only responsible for the hash mapping of one hash fragment, which effectively improves the utilization of memory, so that the entire system of multiple machines can process more data, thereby achieving the effect of improving the efficiency of global deduplication.

[0083] The following is a detailed description of the above steps S11 to S16 in conjunction with the accompanying drawings:

[0084] In the embodiment of the present application, global data refers to a set of text data used to train a large language model. The text data set contains a large amount of text data, and the types of text data include: web page text data, book text data, conversation text data, etc. Among them, the global data can be the stock data stored in the preset database, or it can be the data obtained from the Internet in real time through the web crawler technology. Based on this, when executing step S11, the global data can be obtained by reading the preset database, or it can be obtained from the Internet through the web crawler technology. The specific way of obtaining the global data is not strictly limited in this article.

[0085] In the embodiment of the present application, due to the wide range of sources and types of global data, data duplication and data similarity are unavoidable. The reason for data duplication may be that the original data is reproduced. The reason for data similarity may be that the original data has been added or deleted, or the version of the book text data has been changed. In the embodiment of the present application, the amount of global data is massive, usually counted in TB.

[0086] In some embodiments, the multi-round distribution local deduplication strategy refers to: distributing and deduplicating global data according to a preset distribution round. The preset distribution round refers to the number of times the global data is distributed, and distribution refers to the process of splitting the global data and sending it to different first processing devices respectively.

[0087] Based on this, when executing step S12, a multi-round distribution local deduplication strategy is adopted based on global data to perform the first stage of local deduplication processing. In some embodiments, the above step S12 can be implemented by the following steps:

[0088] S12-1. When performing local deduplication processing on the global data in the first round, the global data is randomly distributed to each of the first processing devices according to a preset random distribution strategy, and each of the first processing devices performs local deduplication processing on the received local data.

[0089] S12-2. Obtain the local deduplication processing results output by each of the first processing devices in the current round, and distribute the summary data set of the local deduplication processing results output by each of the first processing devices in the current round to each of the first processing devices according to the preset redistribution strategy. Each of the first processing devices will perform the next round of local deduplication processing until the round of performing local deduplication processing meets the preset termination round, or until the data volume of the summary data set output by each of the first processing devices converges.

[0090] That is to say, when the global data is locally deduplicated in the first round, the global data is distributed to each first processing device according to the preset random distribution strategy. The local deduplication of global data in subsequent rounds is to distribute the global data to each first processing device according to the preset redistribution strategy, wherein the preset redistribution strategy can continue to use the data distribution strategy adopted in the first round of local deduplication, that is, all rounds of local deduplication use the preset random distribution strategy. The preset redistribution strategy can also be distinguished from the data distribution strategy adopted in the first round of local deduplication, that is, whether it is a preset random distribution strategy or a preset redistribution strategy, both are to split large data into small data and then send them to different first processing devices for service. The specific data distribution strategy to be used can be flexibly selected according to actual needs, and this application does not make strict restrictions.

[0091] The preset termination round can be set according to historical experience or according to the data volume of the global data input in the first round, wherein the larger the data volume of the global data is, the larger the corresponding preset termination round is. In the embodiment of the present application, the data volume of the summary data set output by each first processing device converges to the data volume of the summary data set output by the current time and the data volume of the summary data set output by the previous time gradually decreases and stabilizes at a certain value.

[0092] In the embodiment of the present application, since a single first processing device only performs deduplication processing on the received local data when performing local deduplication processing, there may be a situation where the local data in one first processing device is the same as the local data in other first processing devices. By using the embodiment of the present application, local deduplication processing is performed through multiple rounds, and the local data in different first processing devices can be mixed and then distributed in multiple rounds. After multiple rounds of processing, the situation where the local data in each first processing device has the same data can be effectively reduced, thereby reducing the data volume of the global data.

[0093] In the embodiment of the present application, the preset random distribution strategy refers to splitting the incoming global data in a random size manner, and sending each sub-data set of the split global data to different first processing devices respectively. Figure 2 In the manner shown, different local data are respectively subjected to local deduplication processing by different first processing devices.

[0094] As an example, there are 5 first processing devices pre-allocated for data deduplication, and the available memory of each machine device is: free memory of machine device A: 150GB, free memory of machine device B: 300GB, free memory of machine device C: 280GB, free memory of machine device D: 140GB, free memory of machine device E: 260GB. At this time, 1TB of global data can be randomly divided into 5 local data of 128GB, 256GB, 256GB, 128GB, and 256GB, respectively, and sent to machine devices A to E, and machine devices A to E perform deduplication processing on the received local data respectively. It can be seen that at this time, there is no need to request that the memory space of each processing device must be greater than the total amount of global data 1TB.

[0095] As an implementation mode, when the memory space of the first processing device is sufficient, the global data can be distributed equally, and the global data can be evenly divided into several local data, and then the local data can be sent to different first processing devices respectively, and the first processing devices can perform deduplication processing.

[0096] In the embodiment of the present application, the global data not only refers to the training data of the large language model obtained for the first time, but may include the summary of the processing results obtained after the local deduplication processing of each first processing device, that is, the summary data set of the local deduplication processing results output by each first processing device in the current round may also be referred to as global data. In other words, the input global data of the local deduplication processing of the subsequent round is the summary data set obtained after the summary of the output results of the local deduplication processing of the previous round, that is, Figure 3As shown, the summary data set is regarded as new global data, and is distributed to each first processing device in another round according to a preset redistribution strategy, and each first processing device performs local deduplication processing.

[0097] In an embodiment of the present application, the first processing device performs local deduplication processing on the received local data, specifically including: comparing the various pieces of data in the received local data to see if there is completely identical data. If so, only one piece of data is retained and the rest are deleted. Among them, when comparing the various pieces of data in the received local data to see if there is completely identical data, it can be done by calculating whether the values ​​of the various pieces of data are completely identical. If they are completely identical, it can be determined that two or more pieces of data with completely identical values ​​are identical. In this case, only one piece of data needs to be retained, and the rest is useless data and can be directly deleted. In an embodiment of the present application, deleting a piece of data can specifically refer to deleting the piece of data from the data set of the global data, or the first processing device can directly destroy the piece of data and no longer output the piece of data.

[0098] In some embodiments, the preset hash processing algorithm can be a minimization hash processing algorithm, i.e., a MinhashLSH algorithm, or a local sensitive hash algorithm. The specific algorithm type can be flexibly selected according to actual needs. In the global processing process of the second stage, when executing steps S14 and S15, it can be implemented by the following steps:

[0099] Using a minimization hash processing algorithm, according to a preset number of stripes, the hash signature of each piece of the target sample data is divided into hash segments of the preset number of stripes;

[0100] Mapping hash segments belonging to the same stripe sequence number to a target hash bucket in the same second processing device, and performing similarity calculation on the hash segments in the same target hash bucket using a minimization hash processing algorithm or a locality sensitive hash processing algorithm;

[0101] If the hash segments in the same target hash bucket collide, the target sample data corresponding to the two colliding hash segments are output.

[0102] In the embodiment of the present application, the stripe band technology refers to dividing continuous data into data blocks of the same size and distributing them to different disks. The preset number of stripes is the basis for data division when executing the stripe technology, that is, the preset number of stripes represents the data segmentation score. When executing step S14, the target sample data can be regarded as a continuous data, or the hash signature of the target sample data can be regarded as a continuous data, and then the hash signature of the target sample data is divided into n hash segments of the same size according to the number n, and each hash segment corresponds to the data stored in the data block. As an implementation method, it can be as follows Figure 4 As shown, the hash signature of sample data 1 is divided into band1, band2, band3...bandn.

[0103] Since the hash signatures of each target sample data obtained after being processed by the hash algorithm are of equal length, the hash signatures of each target sample data can be divided into n parts of equal size through striping technology, that is, Figure 4 As shown, each piece of sample data can be divided into band1, band2, band3, ... bandn. At this time, the hash fragments belonging to the same stripe number are mapped to the same second processing device, that is, Figure 4 As shown, the band1 hash fragment of each sample data is mapped to hash bucket 1 in machine 1, the band2 hash fragment of each sample data is mapped to hash bucket 2 in machine 2, and so on.

[0104] Among them, using the minimization hash processing algorithm or the locality sensitive hash processing algorithm, similarity calculation is performed on hash segments in the same target hash bucket, which can be achieved by calculating the Jaccard similarity between hash segments in the same target hash bucket. Then, if the hash segments in the same target hash bucket have a similarity greater than a preset similarity threshold, it can be considered that the hash segments in the same target hash bucket collide, and the target sample data corresponding to the two colliding hash segments are output. For example, it can be as follows Figure 4 As shown, if the calculated similarity between band1 of sample 1 and band1 of sample 4 in bucket 1 is greater than the preset similarity threshold, it can be determined that band1 of sample 1 and band1 of sample 4 collide, and sample 1 and sample 4 are output.

[0105] In the embodiment of the present application, the preset coincidence condition may refer to a similarity greater than a preset similarity threshold. If there are multiple similarity calculation results that meet the preset coincidence condition, only one of the first target sample data is retained as the target reserved sample data, and the remaining first target sample data is deleted. The first target sample data is the target sample data whose similarity calculation results meet the preset coincidence condition. For example, if band1 of sample 1, band1 of sample 2, and band1 of sample 4 meet the preset similarity threshold, only the first target sample data: sample 1 may be retained, and the other samples 2 and sample 4 are deleted.

[0106] In some embodiments, after executing step S16, the method further includes:

[0107] The remaining plurality of target reserved sample data are determined as target training sample data for training the large language model.

[0108] This step is the same as S13, and both are for transferring the deduplicated data to the training device of the large language model so that the training device can train the large language model based on the received deduplicated data. As an implementation method, the remaining several target reserved sample data can be directly transferred to the training device for training the large language model, and the training device can train the large language model based on the received target reserved sample data, which helps to ensure the training accuracy of the large language model.

[0109] Compared with the traditional method of setting up multiple hash buckets in the same machine and then mapping each data to different hash buckets, and having the machine perform global deduplication processing, in the embodiment of the present application, by setting up a hash bucket in a machine, a single machine and a single bucket are implemented, so that the memory pressure of the single machine is reduced by n times compared with the traditional processing method. The actual verification results of the inventor show that by using 10 1TB machines, global deduplication of 10TB of data can be achieved at the same time. In this way, the memory utilization rate is effectively improved, so that more data can be processed, which is helpful to achieve global deduplication of massive data and also helps to improve the accuracy of deduplication.

[0110] In a second aspect, the present application provides a data deduplication device, wherein, Figure 5 As shown, the device 50 comprises:

[0111] A data acquisition module 501 is used to acquire global data, wherein the global data is a set of text data for training a large language model;

[0112] A local deduplication processing module 502 is used to perform local deduplication processing based on the global data by distributing local deduplication strategies in multiple rounds to obtain a target deduplication text data set, wherein the target deduplication text data set includes a plurality of target sample data;

[0113] The multi-round distribution local deduplication strategy includes: distributing global data to different first processing devices, and each of the first processing devices performs several rounds of local deduplication processing on the received local data; wherein the local data is a sub-data set of global data participating in data distribution in the current round, and the data set obtained by summarizing the local deduplication processing results of each of the first processing devices in the current round is the global data for the next round of local deduplication processing;

[0114] The data output module 503 is used to determine the plurality of target sample data as target training sample data for training the large language model.

[0115] In some embodiments, the local deduplication processing module is specifically used to:

[0116] When performing local deduplication processing on the global data in the first round, the global data is randomly distributed to each of the first processing devices according to a preset random distribution strategy, and each of the first processing devices performs local deduplication processing on the received local data;

[0117] Obtain the local deduplication processing results output by the current round of each of the first processing devices, and distribute the summary data set of the local deduplication processing results output by the current round of each of the first processing devices to each of the first processing devices according to a preset redistribution strategy, and have each of the first processing devices perform the next round of local deduplication processing until the round of performing the local deduplication processing meets the preset termination round, or until the data volume of the summary data set output by each of the first processing devices converges.

[0118] In some embodiments, the apparatus further comprises:

[0119] The global deduplication processing module is used to perform the following processing on each piece of the target sample data output by the local deduplication processing module:

[0120] Using a preset hash processing algorithm, split each piece of the target sample data into a plurality of hash segments;

[0121] Sending hash segments belonging to the same segment sequence number in each piece of the target sample data to the same second processing device, and each second processing device performing similarity calculation on the received hash segments;

[0122] Obtain the similarity calculation results output by each of the second processing devices. If the similarity calculation results meet the preset overlap condition, retain a first target sample data as the target reserved sample data, and delete the remaining first target sample data, wherein the first target sample data is the target sample data in which the similarity calculation results meet the preset overlap condition.

[0123] In some embodiments, the use of a preset hash processing algorithm to split each piece of the target sample data into a plurality of segments includes:

[0124] Using a minimization hash processing algorithm, according to a preset number of stripes, the hash signature of each piece of the target sample data is divided into hash segments of the preset number of stripes;

[0125] Mapping hash segments belonging to the same stripe sequence number to a target hash bucket in the same second processing device, and performing similarity calculation on the hash segments in the same target hash bucket using a minimization hash processing algorithm or a locality sensitive hash processing algorithm;

[0126] If the hash segments in the same target hash bucket collide, the target sample data corresponding to the two colliding hash segments are output.

[0127] In some embodiments, after obtaining the similarity calculation results output by each of the second processing devices, if the similarity calculation results meet a preset overlap condition, retaining one first target sample data as target reserved sample data and deleting the remaining first target sample data, the data output module is specifically used to determine the remaining several target reserved sample data as target training sample data for training the large language model.

[0128] Among them, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with the relevant laws and regulations and do not violate public order and good morals.

[0129] The names of the messages or information exchanged between multiple devices in the embodiments of the present application are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0130] In a third aspect, the exemplary embodiments of the present application further provide an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to enable the electronic device to perform a method according to an embodiment of the present application when executed by the at least one processor.

[0131] The exemplary embodiment of the present application further provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform a method according to an embodiment of the present application.

[0132] The exemplary embodiments of the present application further provide a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is used to enable the computer to execute the method according to the embodiment of the present application.

[0133] refer to Figure 6, the structural block diagram of the electronic device 600 that can be used as the server or client of the present application will now be described, which is an example of a hardware device that can be applied to various aspects of the present application. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples, and are not intended to limit the implementation of the present application described and / or required herein.

[0134] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0135] A plurality of components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. The input unit 606 may be any type of device capable of inputting information to the electronic device 600, and the input unit 606 may receive input digital or character information, and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 607 may be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 608 may include, but is not limited to, a disk, an optical disk. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0136] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above. For example, in some embodiments, the aforementioned data deduplication method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via the ROM 602 and / or the communication unit 609. In some embodiments, the computing unit 601 may be configured to perform the aforementioned data deduplication method in any other appropriate manner (e.g., by means of firmware).

[0137] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0138] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0139] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0141] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0142] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.

Claims

1. A data deduplication method, characterized in that: The method comprises: Acquire global data, wherein the global data is a set of text data for training a large language model; Based on the global data, adopt multiple rounds of distribution of local deduplication strategies to perform local deduplication processing to obtain a target deduplication text data set, wherein the target deduplication text data set includes a number of target sample data; The multi-round distribution local deduplication strategy includes: distributing global data to different first processing devices, and each of the first processing devices performs several rounds of local deduplication processing on the received local data; wherein the local data is a sub-data set of global data participating in data distribution in the current round, and the data set obtained by summarizing the local deduplication processing results of each of the first processing devices in the current round is the global data for the next round of local deduplication processing; The plurality of target sample data are determined as target training sample data for training the large language model.

2. The method according to claim 1, characterized in that After the step of distributing a local deduplication strategy in multiple rounds based on the global data to perform local deduplication processing to obtain a target deduplication text data set, the method further includes: Using a preset hash processing algorithm, split each piece of the target sample data into a plurality of hash segments; Sending hash segments belonging to the same segment sequence number in each piece of the target sample data to the same second processing device, and each second processing device performing similarity calculation on the received hash segments; Obtain the similarity calculation results output by each of the second processing devices. If the similarity calculation results meet the preset overlap condition, retain a first target sample data as the target reserved sample data, and delete the remaining first target sample data, wherein the first target sample data is the target sample data in which the similarity calculation results meet the preset overlap condition.

3. The method according to claim 1, characterized in that Based on the global data, a local deduplication strategy is distributed in multiple rounds to perform local deduplication processing to obtain a target deduplication text data set, including: When performing local deduplication processing on the global data in the first round, the global data is randomly distributed to each of the first processing devices according to a preset random distribution strategy, and each of the first processing devices performs local deduplication processing on the received local data; Obtain the local deduplication processing results output by the current round of each of the first processing devices, and distribute the summary data set of the local deduplication processing results output by the current round of each of the first processing devices to each of the first processing devices according to a preset redistribution strategy, and have each of the first processing devices perform the next round of local deduplication processing until the round of performing the local deduplication processing meets the preset termination round, or until the data volume of the summary data set output by each of the first processing devices converges.

4. The method according to claim 2, characterized in that: The preset hash processing algorithm is used to split each piece of the target sample data into a plurality of segments, including: Using a minimization hash processing algorithm, according to a preset number of stripes, the hash signature of each piece of the target sample data is divided into hash segments of the preset number of stripes; Mapping hash segments belonging to the same stripe sequence number to a target hash bucket in the same second processing device, and performing similarity calculation on the hash segments in the same target hash bucket using a minimization hash processing algorithm or a locality sensitive hash processing algorithm; If the hash segments in the same target hash bucket collide, the target sample data corresponding to the two colliding hash segments are output.

5. The method according to claim 2, characterized in that: After obtaining the similarity calculation results output by each of the second processing devices, if the similarity calculation results meet the preset coincidence condition, retaining a first target sample data as target reserved sample data, and deleting the remaining first target sample data, the method further includes: The remaining plurality of target reserved sample data are determined as target training sample data for training the large language model.

6. A data deduplication device, characterized in that: The device comprises: A data acquisition module, used to acquire global data, wherein the global data is a set of text data for training a large language model; A local deduplication processing module is used to adopt a multi-round distribution local deduplication strategy to perform local deduplication processing based on the global data to obtain a target deduplication text data set, wherein the target deduplication text data set includes a plurality of target sample data; The multi-round distribution local deduplication strategy includes: distributing global data to different first processing devices, and each of the first processing devices performs several rounds of local deduplication processing on the received local data; wherein the local data is a sub-data set of global data participating in data distribution in the current round, and the data set obtained by summarizing the local deduplication processing results of each of the first processing devices in the current round is the global data for the next round of local deduplication processing; A data output module is used to determine the plurality of target sample data as target training sample data for training the large language model.

7. The device according to claim 6, characterized in that The local deduplication processing module is specifically used for: When performing local deduplication processing on the global data in the first round, the global data is randomly distributed to each of the first processing devices according to a preset random distribution strategy, and each of the first processing devices performs local deduplication processing on the received local data; Obtain the local deduplication processing results output by the current round of each of the first processing devices, and distribute the summary data set of the local deduplication processing results output by the current round of each of the first processing devices to each of the first processing devices according to a preset redistribution strategy, and have each of the first processing devices perform the next round of local deduplication processing until the round of performing the local deduplication processing meets the preset termination round, or until the data volume of the summary data set output by each of the first processing devices converges.

8. The device according to claim 6, characterized in that The device also includes: The global deduplication processing module is used to perform the following processing on each of the target sample data output by the local deduplication processing module: Using a preset hash processing algorithm, split each piece of the target sample data into a plurality of hash segments; Sending hash segments belonging to the same segment sequence number in each piece of the target sample data to the same second processing device, and each second processing device performing similarity calculation on the received hash segments; Obtaining similarity calculation results output by each of the second processing devices, and if the similarity calculation results meet a preset coincidence condition, retaining a first target sample data as target reserved sample data, and deleting the remaining first target sample data, wherein the first target sample data is the target sample data in which the similarity calculation results meet the preset coincidence condition; The preset hash processing algorithm is used to split each piece of the target sample data into a plurality of segments, including: Using a minimization hash processing algorithm, according to a preset number of stripes, the hash signature of each piece of the target sample data is divided into hash segments of the preset number of stripes; Mapping hash segments belonging to the same stripe sequence number to a target hash bucket in the same second processing device, and performing similarity calculation on the hash segments in the same target hash bucket using a minimization hash processing algorithm or a locality sensitive hash processing algorithm; If the hash segments in the same target hash bucket collide, output the target sample data corresponding to the two colliding hash segments; After obtaining the similarity calculation results output by each of the second processing devices, if the similarity calculation results meet the preset overlap condition, retaining one first target sample data as the target reserved sample data, and deleting the remaining first target sample data, the data output module is specifically used to determine the remaining several target reserved sample data as the target training sample data for training the large language model.

9. An electronic device, characterized in that: The electronic device comprises: Processor; and Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 5.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1-5.

Citation Information

Cited By

  • Data deduplication method and device, electronic equipment, storage medium and product

    CN121233577A

  • Two-stage data deduplication method and system without damaging data timeliness

    CN121807819A