Data processing method, device and equipment

By cleaning, classifying, checking and deduplication of the data to be processed, the problem of low data processing efficiency in existing data governance methods is solved, and efficient identification of duplicate content and improving data processing efficiency is achieved.

CN119939308APending Publication Date: 2025-05-06CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411999295.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing data governance methods have the problem of low data processing efficiency.

Method used

By cleaning, sorting, checking and deduplication of the to be processed data, the target data for model training is generated. The specific steps include detecting and formatting the data to be processed, classifying and processing to generate data of multiple categories, checking dilution based on the similarity of the category data, and finally labeling the deduplication data to generate target data.

Benefits of technology

It improves data quality and processing efficiency, can efficiently identify duplicate content when data is expanded on a large scale, and improves the accuracy of duplicate detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939308A_ABST
    Figure CN119939308A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer data processing, provides a data processing method, device and equipment, and is used for solving the problem of relatively low data processing efficiency of a data treatment method in related technologies. According to the embodiment of the invention, firstly, data cleaning processing is carried out on to-be-processed data to obtain to-be-deduplicated data, then, classification processing is carried out on the to-be-deduplicated data to obtain multiple categories of data, duplicate checking and duplicate removing processing is carried out on the multiple categories of data, and finally, target data for model training is obtained based on the deduplicated data. According to the embodiment of the invention, the data quality is improved by cleaning the data, then the cleaned data is subjected to duplicate checking and duplicate removal processing, it is ensured that duplicate contents can be efficiently identified during large-scale expansion of the data, the accuracy of duplicate detection is improved, the efficiency of data processing is improved, and the technical scheme provided by the invention is simple in method and easy to implement. And the universality is good.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer data processing technology, and in particular to data processing methods, devices and equipment. Background Art

[0002] Fine-tuning data governance refers to the systematic management of data used for training and optimization during the fine-tuning phase of large models to ensure the quality and applicability of the data. With the widespread application of large models, the number and complexity of their parameters continue to increase, and the demand for training data also grows accordingly. High-quality data can not only improve the model's learning ability on specific tasks, but also enhance its generalization ability in other tasks.

[0003] In related technologies, data governance includes data cleaning and data deduplication. However, existing data governance methods have the problem of low data processing efficiency.

[0004] Therefore, it is particularly necessary to establish an effective data governance solution. Summary of the invention

[0005] The purpose of this application is to provide a data processing method, device and equipment to solve the problem of low data processing efficiency in data governance methods in related technologies.

[0006] In a first aspect, the present application provides a data processing method, the method comprising:

[0007] Perform data cleaning on the data to be processed to obtain the data to be deduplicated;

[0008] Classifying the data to be deduplicated to obtain multiple categories of data; each category of the multiple categories of data includes multiple data sets;

[0009] For the multiple categories of data, determining a duplicate checking result for each category of data based on the similarities between multiple data sets included in each category of data;

[0010] Based on the duplicate checking results of the data of the multiple categories, the data to be deduplicated is deduplicated to obtain deduplicated data;

[0011] Based on the deduplicated data, target data for model training is obtained.

[0012] In a possible implementation manner, performing data cleaning on the data to be processed to obtain the data to be deduplicated includes:

[0013] Detecting the data to be processed and filtering out unnecessary characters in the data to be processed;

[0014] The data to be processed after filtering out unnecessary characters is adjusted to a set format to obtain the data to be deduplicated.

[0015] In a possible implementation, before determining the duplicate checking result of each category of data based on the similarity between the multiple data sets included in each category of data, the method further includes:

[0016] For each of the multiple categories of data, perform the following operations respectively:

[0017] Determine the minimum hash value of each data set in a plurality of data sets included in the data of the category;

[0018] Generate a hash matrix corresponding to the data of the category based on the minimum hash value of each data set included in the data of the category;

[0019] Based on the hash matrix corresponding to the data of the category, the similarities between multiple data sets included in the data of the category are determined.

[0020] In a possible implementation manner, each of the multiple data sets has multiple minimum hash values; and the step of respectively determining the minimum hash value of each of the multiple data sets included in the data of the category includes:

[0021] For each data set in the multiple data sets, multiple hash functions are respectively used to determine the minimum hash value of the data set, and the minimum hash value corresponding to each hash function is obtained.

[0022] In a possible implementation manner, the minimum hash value corresponding to each hash function refers to a minimum value among multiple hash values ​​obtained by performing hash processing on each of the multiple data included in the data set using the hash function.

[0023] In a possible implementation, the multiple categories of data include instruction data, input data, and output data; and the deduplication processing of the to-be-deduplicated data based on the duplicate checking results of the multiple categories of data includes:

[0024] If the instruction data comes from the first instruction database, and the duplicate checking result of any one category of the input data and the output data is duplicate data, then the data to be deduplicated is deduplicated;

[0025] If the instruction data comes from the second instruction database, and the duplicate checking result of any one of the input data, the input data, and the output data is duplicate data, the data to be deduplicated is deduplicated.

[0026] In a possible implementation, obtaining target data for model training based on the deduplicated data includes:

[0027] The large model is used to add annotations to the deduplicated data to obtain the target data.

[0028] In a second aspect, the present application provides a data processing device, the device comprising:

[0029] A data cleaning module is configured to perform data cleaning on the data to be processed to obtain data to be deduplicated;

[0030] A data classification module is configured to classify the to-be-deduplicated data to obtain data of multiple categories; each category of the data of the multiple categories includes multiple data sets;

[0031] A data duplication checking module is configured to determine a duplication checking result of each category of data based on the similarities between multiple data sets included in each category of data for the multiple categories of data;

[0032] A data deduplication module is configured to perform deduplication processing on the data to be deduplicated based on the duplicate checking results of the multiple categories of data to obtain deduplicated data;

[0033] The target data determination module is configured to obtain target data for model training based on the deduplicated data.

[0034] In a possible implementation manner, the data to be processed is cleaned to obtain the data to be deduplicated, and the data cleaning module is configured as follows:

[0035] Detecting the data to be processed and filtering out unnecessary characters in the data to be processed;

[0036] The data to be processed after filtering out unnecessary characters is adjusted to a set format to obtain the data to be deduplicated.

[0037] In a possible implementation, before determining the duplicate checking result of each category of data based on the similarity between the multiple data sets included in each category of data for the multiple categories of data, the data duplicate checking module is further configured to:

[0038] For each of the multiple categories of data, perform the following operations respectively:

[0039] Determine the minimum hash value of each data set in a plurality of data sets included in the data of the category;

[0040] Generate a hash matrix corresponding to the data of the category based on the minimum hash value of each data set included in the data of the category;

[0041] Based on the hash matrix corresponding to the data of the category, the similarities between multiple data sets included in the data of the category are determined.

[0042] In a possible implementation, each of the multiple data sets has multiple minimum hash values; the minimum hash value of each data set in the multiple data sets included in the data of the category is determined respectively, and the data duplication checking module is configured as follows:

[0043] For each data set in the multiple data sets, multiple hash functions are respectively used to determine the minimum hash value of the data set, and the minimum hash value corresponding to each hash function is obtained.

[0044] In a possible implementation manner, the minimum hash value corresponding to each hash function refers to a minimum value among multiple hash values ​​obtained by performing hash processing on each of the multiple data included in the data set using the hash function.

[0045] In a possible implementation, the multiple categories of data include instruction data, input data, and output data; the data to be deduplicated is deduplicated based on the duplicate checking results of the multiple categories of data, and the data deduplication module is configured as follows:

[0046] If the instruction data comes from the first instruction database, and the duplicate checking result of any one category of the input data and the output data is duplicate data, then the data to be deduplicated is deduplicated;

[0047] If the instruction data comes from the second instruction database, and the duplicate checking result of any one category of data among the instruction data, the input data, and the output data is duplicate data, then the data to be deduplicated is deduplicated.

[0048] In a possible implementation, the target data for model training is obtained based on the deduplicated data, and the target data determination module is configured as follows:

[0049] The large model is used to add annotations to the deduplicated data to obtain the target data.

[0050] In a third aspect, the present application provides an electronic device, including:

[0051] Processor and memory;

[0052] The memory is used to store the processor executable instructions;

[0053] The processor is configured to execute the instructions to implement the data processing method as described in any one of the first aspects of the present application.

[0054] In a fourth aspect, the present application provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the data processing method as described in any one of the first aspects of the present application.

[0055] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the data processing method as described in any one of the first aspects of the present application.

[0056] The technical solution provided by the embodiments of the present application brings at least the following beneficial effects:

[0057] The embodiment of the present application provides a data processing method, which improves data quality by cleaning data, and then checks and removes duplicates from the cleaned data, thereby ensuring that duplicate content can be efficiently identified even when the data is expanded on a large scale, improving the accuracy of duplicate detection, and improving the efficiency of data processing. The technical solution provided by the present application is simple in method and has good versatility.

[0058] It should be understood that the above general description and the following detailed description are only exemplary and explanatory and cannot limit the present application. Based on the common sense in the art, the above preferred conditions can be combined arbitrarily to obtain the preferred embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. Obviously, the drawings introduced below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0060] Figure 1 A schematic diagram of an application scenario of a data processing method provided in an embodiment of the present application;

[0061] Figure 2 A schematic diagram of the overall process of a data processing method provided in an embodiment of the present application;

[0062] Figure 3 A schematic diagram of a process for performing data cleaning on data to be processed provided in an embodiment of the present application;

[0063] Figure 4A schematic diagram of a process for determining similarity between data sets provided in an embodiment of the present application;

[0064] Figure 5 A schematic diagram of a process for performing deduplication processing on data to be deduplicated based on duplicate checking results of multiple categories of data provided in an embodiment of the present application;

[0065] Figure 6 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;

[0066] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Among them, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0068] Furthermore, in the description of the embodiments of the present application, unless otherwise specified, “ / ” means or. For example, A / B can mean A or B. The “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.

[0069] In the following, the terms "first", "second", and "first" are used for descriptive purposes only and should not be understood as suggesting or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first", "second", or "first" may explicitly or implicitly include one or more of the features.

[0070] The following is an explanation of the professional terms and technologies involved in this application:

[0071] LSH (Locality-Sensitive Hashing): Efficiently perform similarity search by mapping similar data points to the same or similar hash buckets. It is suitable for processing high-dimensional data such as text, images, or feature vectors, and can quickly find approximately similar objects in large-scale data.

[0072] Fine-tuning data governance refers to the systematic management of data used for training and optimization during the fine-tuning phase of large models to ensure the quality and applicability of the data. With the widespread application of large models, the number and complexity of their parameters continue to increase, and the demand for training data also grows accordingly. High-quality data can not only improve the model's learning ability on specific tasks, but also enhance its generalization ability in other tasks.

[0073] In related technologies, data governance mainly includes two parts: traditional data cleaning and data deduplication using existing tools. The cleaning work is mainly rule-driven outlier cleaning. This cleaning method lacks certain flexibility and is difficult to deal with problems such as data imbalance and noise interference. Deduplication work is mainly static batch processing. The deduplicated data set no longer accepts new data. This method is suitable for initial one-time deduplication. It is efficient at a single time but easily causes a sudden increase in memory usage, which is not conducive to processing real-time data and has low efficiency.

[0074] Therefore, it is particularly necessary to establish an effective data governance solution.

[0075] In view of this, the present application provides a data processing method, device and equipment to solve the problem of low data processing efficiency in data governance methods in related technologies.

[0076] The inventive concept of the present application can be summarized as follows: first, data cleaning is performed on the data to be processed to obtain data to be deduplicated; then, the data to be deduplicated is classified to obtain data of multiple categories; duplicate checking and deduplication processing are performed on the data of multiple categories; finally, target data for model training is obtained based on the deduplicated data.

[0077] In summary, the data processing method provided in the embodiment of the present application improves data quality by cleaning the data, and then checks and removes duplicates on the cleaned data, thereby ensuring that duplicate content can be efficiently identified even when the data is expanded on a large scale, improving the accuracy of duplicate detection, and improving the efficiency of data processing. In addition, the technical solution provided by the present application is simple in method and has good versatility.

[0078] After introducing the main inventive ideas of the embodiments of the present application, the following briefly introduces the application scenarios to which the technical solutions of the embodiments of the present application can be applied. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not limited. In specific implementation, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.

[0079] For ease of understanding, a data processing method provided in an embodiment of the present application is described in detail below with reference to the accompanying drawings:

[0080] like Figure 1As shown, it is a schematic diagram of an application scenario of a data processing method provided by an embodiment of the present application. The figure includes: a network 10, a server 20, and a memory 30. The server 20 obtains the data to be processed through the network. Through the method provided by the embodiment of the present application, the data processing method can be cleaned, duplicated, and deduplicated, thereby improving the efficiency of data processing and obtaining high-quality data for training models.

[0081] The description in this application only details a single server, but those skilled in the art should understand that the network 10, server 20 and memory 30 shown are intended to represent the operation of the electronic device, server and memory involved in the technical solution of this application. The single server and memory are described in detail at least for the convenience of explanation, and do not imply any restrictions on the number, type or location of the server. It should be noted that if additional modules are added to the illustrated environment or individual modules are removed from it, it will not change the underlying concepts of the example embodiments of this application. In addition, although for the convenience of explanation, Figure 1 A bidirectional arrow from the storage 30 to the server 20 is shown in the figure, but those skilled in the art can understand that the sending and receiving of the above data also needs to be implemented through the network 10.

[0082] It should be noted that the memory in the embodiment of the present application may be, for example, a cache system, a hard disk storage, a memory storage, etc. In addition, the data processing method proposed in the present application is not only applicable to Figure 1 The application scenario shown can also be used in other possible application scenarios, and the embodiments of the present application are not limited thereto.

[0083] Based on the above description, the present application provides a data processing method, the overall process of which is as follows: Figure 2 As shown, including the following:

[0084] In step 201, data cleaning is performed on the data to be processed to obtain data to be deduplicated.

[0085] In a possible implementation, the data to be processed is cleaned, such as Figure 3 As shown, it can be implemented as:

[0086] In step 301, the data to be processed is detected, and unnecessary characters in the data to be processed are filtered out.

[0087] In one possible implementation, the task of data cleaning is to identify and process various anomalies in the data used for fine-tuning the large model, including empty annotations, long-tail generation, repetition, truncation, deletion, special characters, and format anomalies. The anomaly categories involved in this application and the corresponding recognition algorithms are shown in the following table:

[0088]

[0089]

[0090] In step 302, the data to be processed after filtering out unnecessary characters is adjusted to a set format to obtain data to be deduplicated.

[0091] The above data cleaning steps can automatically detect and filter unnecessary characters, correct format anomalies and content anomalies, and achieve standardized processing of content. Compared with traditional manual cleaning methods, by connecting to large models, the degree of automation is high, which effectively reduces human errors, ensures the clarity and standardization of data, and reduces resource consumption for subsequent deduplication work.

[0092] In step 202, the data to be deduplicated is classified to obtain multiple categories of data, wherein each category of the multiple categories of data includes multiple data sets, for example, the multiple categories are Instruction data, Input data, and Output data.

[0093] It should be added that, before determining the duplicate checking results of each category of data based on the similarities between the multiple data sets contained in the data of each category, the embodiment of the present application will also perform word segmentation on the data of each category. For example, Instruction and Output are Chinese texts, and the string sets A and B corresponding to Instruction and Output are obtained through word segmentation. Input is a log text, which is not suitable for Chinese word segmentation. Feature division can be performed according to the format of different logs, and the features that contribute most to the risk assessment task are selected for segmentation to obtain the string set C corresponding to Input. The string sets A, B, and C are the data of each category, and are composed of multiple data sets.

[0094] In step 203, for multiple categories of data, based on the similarities between multiple data sets included in each category of data, a duplicate checking result of each category of data is determined.

[0095] In a possible implementation, for multiple categories of data, based on the similarities between multiple data sets contained in each category of data, before determining the duplicate check result of each category of data, the present application will also determine the similarities between the data sets, and the process is as follows: Figure 4 As shown, the following steps are included:

[0096] For each of the multiple categories of data, perform the following operations:

[0097] In step 401, the minimum hash value of each data set in a plurality of data sets included in the data of a category is determined respectively.

[0098] In a possible implementation, each of the multiple data sets has multiple minimum hash values, and the minimum hash value of each of the multiple data sets included in the data of the category is determined respectively, which can be implemented as follows:

[0099] For each data set in the multiple data sets, multiple hash functions are respectively used to determine the minimum hash value of the data set, and the minimum hash value corresponding to each hash function is obtained. The minimum hash value corresponding to each hash function refers to the minimum value among the multiple hash values ​​obtained by using the hash function to perform hash processing on each data in the multiple data contained in the data set.

[0100] In step 402, a hash matrix corresponding to the data of the category is generated based on the minimum hash value of each data set included in the data of the category.

[0101] For example, the data of instruction type contains multiple data sets, namely A1, A2, and A3, each of which contains 4 data. The data of data set A1 are A11, A12, A13, and A14, the data of data set A2 are A21, A22, A23, and A24, and the data of data set A3 are A31, A32, A33, and A34. The hash functions include X and Y. Hash function X is used to hash each data in data set A1 to obtain 4 hash values. The minimum value of the 4 hash values ​​is the minimum hash value XA1 corresponding to the hash function. Similarly, the minimum hash values ​​of hash function X and hash function Y corresponding to data set A1 are XA1 and YA1 respectively. Similarly, the minimum hash values ​​of each data set of multiple data sets A1, A2, and A3 are XA1, YA1, XA2, YA2, XA3, and YA3 respectively. Therefore, the hash matrix corresponding to the data of instruction type is as follows Each column corresponds to a data set.

[0102] In step 403, based on the hash matrix corresponding to the data of the category, the similarities between the multiple data sets included in the data of the category are determined.

[0103] In a possible implementation, the present application estimates the Jaccard similarity between sets based on a hash matrix, and if the similarity of the hash values ​​of two sets is higher than a preset similarity threshold, it is determined that the two sets are duplicates.

[0104] In step 204, based on the duplicate checking results of multiple categories of data, the data to be deduplicated is deduplicated to obtain deduplicated data.

[0105] In a possible implementation, the multiple categories of data include instruction data, input data, and output data; based on the duplicate checking results of the multiple categories of data, the duplicate data to be deduplicated is deduplicated, such as Figure 5 As shown, it can be implemented as:

[0106] In step 501, if the instruction data comes from the first instruction database, and the duplicate checking result of any category of data in the input data and the output data is duplicate data, deduplication processing is performed on the data to be deduplicated.

[0107] In step 502, if the instruction data comes from the second instruction database, and the duplicate checking result of any category of data among the instruction data, input data, and output data is duplicate data, deduplication processing is performed on the data to be deduplicated.

[0108] For example, depending on the source of the instruction data and the specific duplication of the input data, output data, and instruction data, the judgment logic of whether to perform deduplication processing is as shown in the following table:

[0109]

[0110] Among them, the preset instruction library is the first instruction database, the large model generation instruction library is the second instruction database, and “-” indicates that this type of data is not considered.

[0111] It should be supplemented that the duplicate checking results of multiple categories of data are determined based on the similarity between multiple data sets contained in each category of data. For example, the Instruction data contains 10 data sets, and a similarity is determined for every two data sets. If the similarity is determined to be greater than the preset similarity threshold, the two data sets are determined to be duplicates. Among the 10 data sets, the number of duplicate data sets is greater than the preset number threshold, then the duplicate checking result of the Instruction data is determined to be duplicate data.

[0112] For example, the instruction data contains 10 data sets, a similarity is determined for every two data sets, and it is determined that there are 7 duplicate data sets. The preset number threshold is 6. Then the duplicate check result of the Instruction instruction data shows that the instruction data is duplicate data, and the instruction data is deduplicated. The specific deduplication processing can be to remove 6 duplicate data sets and retain 1 duplicate data set. Finally, the deduplicated data is obtained.

[0113] In step 205, target data for model training is obtained based on the deduplicated data.

[0114] In a possible implementation, the present application will also use a large model to add annotations to the deduplicated data to obtain the target data, wherein the large model can be a language model LLM. A multi-round annotation mechanism is adopted in the implementation process. Each data in the deduplicated data will be independently annotated multiple times to integrate the opinions of multiple models to improve the accuracy of the annotation; when the annotation results are inconsistent, adjustments are made through majority voting or other correction rules to ensure that the final annotation results are credible.

[0115] For example, in the first round of labeling, model No. 1 labels the log data as "normal log", model No. 2 labels the log data as "Web attack", model No. 3 labels the log data as "potential attack risk", etc. According to the correction rules, the labeling content of model No. 1 is discarded, and the second round of labeling is carried out until the labeling is consistent.

[0116] In a possible implementation, the embodiment of the present application will also regularly detect the quality changes of the data to be processed and the data after deduplication through statistical analysis and data verification tools, including analysis in two dimensions, the sample distribution dimension and the model training dimension. In the sample distribution dimension, the ratio of black samples to white samples in the data to be processed and the data after deduplication is determined at each preset time interval. If the ratio of black samples to white samples is within the preset ratio range, it is determined that the ratio of black samples to white samples is reasonably distributed.

[0117] For example, the distribution of the ratio of black samples to white samples is shown in the following table:

[0118] Before treatment After treatment Black Sample 453770 256017 White Sample 170983 96106 Black and white ratio 2.65:1 2.66:1

[0119] Before and after data processing, the distribution ratio of black samples to white samples did not change much. Therefore, this data processing met the requirements in terms of sample distribution dimension.

[0120] In the model training dimension, evaluation is performed by comparing indicators such as accuracy, precision, recall, F1 score, and training time.

[0121] The above-mentioned data evaluation mechanism has real-time monitoring and dynamic feedback capabilities, which ensures the reliability of data during data processing and facilitates the timely discovery and correction of potential problems.

[0122] In another possible implementation, the data processing method specifically includes the following contents:

[0123] Each sample contains three fields: input, output, and instruction; for the input field, perform the following operations:

[0124] First, the text of the field is divided into multiple sub-texts, and several sub-texts are combined into a set, which is called the "input set of a certain sample";

[0125] For each subtext in the input set, a series of hash functions are applied to generate the corresponding hash value, and the minimum hash value generated by each hash function is selected as the hash signature of the subtext;

[0126] Based on the minimum hash value, a locality-sensitive hash bucket is constructed. The goal of this hash bucket is to map similar text sets into the same hash bucket, and dissimilar sets into different buckets; the minimum hash value is divided into multiple "blocks", for example, every b minimum hash values ​​are a group, and each group is a sub-signature; each sub-signature is used as a hash value, and the input field text is mapped to the hash bucket using the hash value. For samples with the same sub-signature, they will be mapped to the same bucket.

[0127] The above hash function design ensures that text sets with higher Jaccard similarity will be mapped to the same hash bucket, while text sets with lower similarity are more likely to be mapped to different hash buckets;

[0128] For example, assuming that the minimum hash value is divided into multiple blocks, if the first two blocks (subsignatures) of two field texts are exactly the same, the two field texts will be mapped to the same hash bucket.

[0129] When calculating sample similarity and removing duplicates, only the field texts in the same hash bucket need to be compared;

[0130] For the field texts in each bucket, their Jaccard similarity is calculated. Jaccard similarity evaluates the similarity of field texts by comparing the overlap between subtext sets of field texts. If the similarity between two field texts is higher than the set threshold, the two field texts are considered to be duplicate field texts and need to be deduplicated.

[0131] It should be added that the processing of the output field and the instruction field is the same as that of the input field, including:

[0132] Divide the field text into a sub-text set;

[0133] Apply a hash function to each subtext to generate a minimum hash value;

[0134] A hash bucket is constructed based on the minimum hash value and sample similarity is calculated, and finally duplicate field texts are deduplicated.

[0135] In summary, the data processing method provided in the embodiment of the present application improves data quality by cleaning the data, and then checks and removes duplicates on the cleaned data, thereby ensuring that duplicate content can be efficiently identified even when the data is expanded on a large scale, improving the accuracy of duplicate detection, and improving the efficiency of data processing. In addition, the technical solution provided by the present application is simple in method and has good versatility.

[0136] Based on the same inventive concept, the embodiment of the present application also provides a data processing device, such as Figure 6 As shown, the device 600 includes:

[0137] The data cleaning module 601 is configured to perform data cleaning on the data to be processed to obtain the data to be deduplicated;

[0138] The data classification module 602 is configured to classify the to-be-deduplicated data to obtain data of multiple categories; each category of the data of the multiple categories includes multiple data sets;

[0139] The data duplication checking module 603 is configured to determine the duplication checking result of each category of data based on the similarity between the multiple data sets included in each category of data for the multiple categories of data;

[0140] The data deduplication module 604 is configured to perform deduplication processing on the data to be deduplicated based on the duplicate checking results of the multiple categories of data to obtain deduplicated data;

[0141] The target data determination module 605 is configured to obtain target data for model training based on the deduplicated data.

[0142] In a possible implementation manner, the data to be processed is cleaned to obtain the data to be deduplicated, and the data cleaning module is configured as follows:

[0143] Detecting the data to be processed and filtering out unnecessary characters in the data to be processed;

[0144] The data to be processed after filtering out unnecessary characters is adjusted to a set format to obtain the data to be deduplicated.

[0145] In a possible implementation, before determining the duplicate checking result of each category of data based on the similarity between the multiple data sets included in each category of data for the multiple categories of data, the data duplicate checking module is further configured to:

[0146] For each of the multiple categories of data, perform the following operations respectively:

[0147] Determine the minimum hash value of each data set in a plurality of data sets included in the data of the category;

[0148] Generate a hash matrix corresponding to the data of the category based on the minimum hash value of each data set included in the data of the category;

[0149] Based on the hash matrix corresponding to the data of the category, the similarities between multiple data sets included in the data of the category are determined.

[0150] In a possible implementation, each of the multiple data sets has multiple minimum hash values; the minimum hash value of each data set in the multiple data sets included in the data of the category is determined respectively, and the data duplication checking module is configured as follows:

[0151] For each data set in the multiple data sets, multiple hash functions are respectively used to determine the minimum hash value of the data set, and the minimum hash value corresponding to each hash function is obtained.

[0152] In a possible implementation manner, the minimum hash value corresponding to each hash function refers to a minimum value among multiple hash values ​​obtained by performing hash processing on each of the multiple data included in the data set using the hash function.

[0153] In a possible implementation, the multiple categories of data include instruction data, input data, and output data; the data to be deduplicated is deduplicated based on the duplicate checking results of the multiple categories of data, and the data deduplication module is configured as follows:

[0154] If the instruction data comes from the first instruction database, and the duplicate checking result of any one category of the input data and the output data is duplicate data, then the data to be deduplicated is deduplicated;

[0155] If the instruction data comes from the second instruction database, and the duplicate checking result of any one category of data among the instruction data, the input data, and the output data is duplicate data, then the data to be deduplicated is deduplicated.

[0156] In a possible implementation, the target data for model training is obtained based on the deduplicated data, and the target data determination module is configured as follows:

[0157] The large model is used to add annotations to the deduplicated data to obtain the target data.

[0158] Refer to the following Figure 7 The electronic device 130 according to this embodiment of the present application is described. Figure 7The electronic device 130 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0159] like Figure 7 As shown, the electronic device 130 is in the form of a general electronic device. The components of the electronic device 130 may include but are not limited to: the at least one processor 131, the at least one memory 132, and a bus 133 connecting different system components (including the memory 132 and the processor 131).

[0160] Bus 133 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a processor, or a local bus using any of a variety of bus architectures.

[0161] The memory 132 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 1321 and / or a cache memory 1322 , and may further include a read-only memory (ROM) 1323 .

[0162] The memory 132 may also include a program / utility 1325 having a set (at least one) of program modules 1324, such program modules 1324 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0163] The electronic device 130 may also communicate with one or more external devices 134 (e.g., keyboards, pointing devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 130, and / or communicate with any device that enables the electronic device 130 to communicate with one or more other electronic devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 135. Furthermore, the electronic device 130 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 136. As shown, the network adapter 136 communicates with other modules for the electronic device 130 via a bus 133. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 130, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0164] In an exemplary embodiment, the present application further provides a computer-readable storage medium including instructions, such as a memory 132 including instructions, and the above instructions can be executed by a processor 131 of an electronic device 130 to complete the above data processing method. Optionally, the computer-readable storage medium can be a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0165] In an exemplary embodiment, a computer program product is also provided, including a computer program, and when the computer program is executed by the processor 131, the data processing method provided in the present application is implemented.

[0166] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0167] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0168] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0170] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A data processing method, characterized in that: The method comprises: Perform data cleaning on the data to be processed to obtain the data to be deduplicated; Classifying the data to be deduplicated to obtain multiple categories of data; each category of the multiple categories of data includes multiple data sets; For the multiple categories of data, determining a duplicate checking result for each category of data based on the similarities between multiple data sets included in each category of data; Based on the duplicate checking results of the data of the multiple categories, the data to be deduplicated is deduplicated to obtain deduplicated data; Based on the deduplicated data, target data for model training is obtained.

2. The method according to claim 1, characterized in that: The data to be processed is cleaned to obtain the data to be deduplicated, including: Detecting the data to be processed and filtering out unnecessary characters in the data to be processed; The data to be processed after filtering out unnecessary characters is adjusted to a set format to obtain the data to be deduplicated.

3. The method according to claim 1, characterized in that Before determining the duplicate checking result of each category of data based on the similarity between the multiple data sets included in each category of data, the method further includes: For each of the multiple categories of data, perform the following operations respectively: Determine the minimum hash value of each data set in a plurality of data sets included in the data of the category; Generate a hash matrix corresponding to the data of the category based on the minimum hash value of each data set included in the data of the category; Based on the hash matrix corresponding to the data of the category, the similarities between multiple data sets included in the data of the category are determined.

4. The method according to claim 3, characterized in that Each of the multiple data sets has multiple minimum hash values; The step of respectively determining the minimum hash value of each data set in a plurality of data sets included in the data of the category includes: For each data set in the multiple data sets, multiple hash functions are respectively used to determine the minimum hash value of the data set, and the minimum hash value corresponding to each hash function is obtained.

5. The method according to claim 4, characterized in that The minimum hash value corresponding to each hash function refers to the minimum value among multiple hash values ​​obtained by using the hash function to perform hash processing on each of the multiple data included in the data set.

6. The method according to claim 1, characterized in that The multiple categories of data include instruction data, input data, and output data; and the deduplication processing of the to-be-deduplicated data based on the duplicate checking results of the multiple categories of data includes: If the instruction data comes from the first instruction database, and the duplicate checking result of any one category of the input data and the output data is duplicate data, then the data to be deduplicated is deduplicated; If the instruction data comes from the second instruction database, and the duplicate checking result of any one category of data among the instruction data, the input data, and the output data is duplicate data, then the data to be deduplicated is deduplicated.

7. The method according to claim 1, characterized in that The step of obtaining target data for model training based on the deduplicated data includes: The large model is used to add annotations to the deduplicated data to obtain the target data.

8. A data processing device, characterized in that: The device comprises: A data cleaning module is configured to perform data cleaning on the data to be processed to obtain data to be deduplicated; A data classification module is configured to classify the to-be-deduplicated data to obtain data of multiple categories; each category of the data of the multiple categories includes multiple data sets; The data duplication checking module is configured to determine the duplication checking result of each category of data based on the similarity between the multiple data sets included in each category of data for the multiple categories of data; A data deduplication module is configured to perform deduplication processing on the data to be deduplicated based on the duplicate checking results of the multiple categories of data to obtain deduplicated data; The target data determination module is configured to obtain target data for model training based on the deduplicated data.

9. An electronic device, characterized in that: include: A memory for storing program instructions; A processor is used to call the program instructions stored in the memory, and execute the data processing method according to any one of claims 1 to 7 according to the obtained program instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the data processing method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The computer program product comprises: a computer program code, and when the computer program code is run on a computer, the computer is enabled to execute the data processing method according to any one of claims 1 to 7.