Data deduplication processing method and device, electronic equipment and vehicle
By reordering the text data word segmentation and position, and generating feature values to identify duplicate text data, the situation in which the data similar but partial changes cannot be effectively removed in the prior art, and the effect of data deduplication processing is improved.
Patent Information
- Application Number
- CN202311652086.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art can only remove exactly the same text data, and cannot effectively remove the situation in which most of the data are similar but only a small part of the data changes, resulting in poor data deduplication processing.
By performing word segmentation processing on multiple text data to be processed, the initial arrangement position of word segmentation elements is determined, and the positional rearrangement of word segmentation elements is performed based on this position multiple times. The first word segmentation element after each rearrangement is obtained as the feature value of the text data, and the duplicate text data is identified and removed based on the duplicate information of the feature value.
When most of the data are similar but only a small part of the data changes in the text data, similar text data can be effectively removed and the effect of data deduplication processing can be improved.
Smart Images

Figure CN120104597A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data deduplication processing method, device, electronic equipment and vehicle. Background Art
[0002] During the training process of a large language model (LLM), it is necessary to prepare massive text data at the TB level. In order to avoid overfitting caused by multiple uses of the same data in the pre-trained model, the text data to be trained needs to be deduplicated.
[0003] Currently, the value of the fifth version of the message digest algorithm (MD5) corresponding to each text data can be calculated. At least two text data must have the same MD5 value to determine whether the text data is repeated, and the text data with the same MD5 value can be deduplicated.
[0004] However, only the MD5 values corresponding to exactly the same text data will be the same, which will result in only being able to remove exactly the same text data during the deduplication process of the text data. In the case where most of the text data is similar but only a small part of the data has changed, it is impossible to deduplicate similar text data, which results in poor data deduplication effect. Summary of the invention
[0005] In view of this, the present application provides a data deduplication processing method, device, electronic device and vehicle, the main purpose of which is to improve the current existing technology that can only remove completely identical text data. For situations where most of the data in the text data is similar but only a small part of the data has changed, similar text data cannot be deduplicated, which leads to a technical problem that the data deduplication effect is poor.
[0006] In a first aspect, the present application provides a data deduplication processing method, comprising:
[0007] Perform word segmentation processing on the multiple text data to be processed respectively to obtain each word segmentation element corresponding to each text data;
[0008] Determine the initial arrangement position of each word segmentation element in each text data;
[0009] Rearranging the word segmentation elements in each text data multiple times based on the initial arrangement position, and obtaining the first word segmentation element in the text data after each position rearrangement as the feature value of the text data;
[0010] Identifying repeated text data in the plurality of text data according to repeated information of the characteristic value of each text data;
[0011] The repeated text data is deduplicated.
[0012] In a second aspect, the present application provides a data deduplication processing device, comprising:
[0013] A processing module is configured to perform word segmentation processing on the multiple text data to be processed respectively to obtain each word segmentation element corresponding to each text data;
[0014] A determination module is configured to determine the initial arrangement position of each word segmentation element in each text data;
[0015] An acquisition module is configured to rearrange the positions of the word segmentation elements in each text data based on the initial arrangement positions for multiple times, and obtain the first word segmentation element in the text data after each position rearrangement as a feature value of the text data;
[0016] an identification module configured to identify repeated text data in the plurality of text data according to repeated information of a characteristic value of each text data;
[0017] The processing module is also configured to perform deduplication processing on the repeated text data.
[0018] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0019] In a fourth aspect, the present application provides an electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the method described in the first aspect when executing the computer program.
[0020] In a fifth aspect, the present application provides a vehicle, comprising: the device as described in the second aspect, or the electronic device as described in the fourth aspect.
[0021] By means of the above technical scheme, the present application provides a data deduplication processing method, device, electronic device and vehicle, which first perform word segmentation processing on the multiple text data to be processed respectively to obtain the word segmentation elements corresponding to each text data; determine the initial arrangement position of each word segmentation element in each text data; then rearrange the word segmentation elements in each text data based on the initial arrangement position multiple times, and obtain the first word segmentation element in the text data after each position rearrangement as the feature value of the text data; then identify the repeated text data existing in the multiple text data based on the feature value of each text data; and deduplicate the repeated text data. Compared with the current existing technology, the present application, in the process of deduplicating text data, segments each text data, and uniformly shuffles the segmentation elements of each text data for multiple times, and obtains the characteristic value of the text data by the first segmentation element in the text data after each position rearrangement, and screens the repeated text data by comparing the repeated information of the characteristic value. In the case that most of the data in the text data is similar but only a small part of the data has changed, the similar text data can correspond to the characteristic value whose repeated information meets the conditions, so that the similar text data in which most of the data is similar but only a small part of the data has changed can be effectively removed, thereby improving the effect of data deduplication processing.
[0022] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0025] Figure 1 A schematic diagram of a process flow of a data deduplication processing method provided in an embodiment of the present application is shown;
[0026] Figure 2 A schematic diagram of a process flow of a data deduplication processing method provided in an embodiment of the present application is shown;
[0027] Figure 3A structural schematic diagram of a data deduplication processing device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0028] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the scheme of the present application will be further described below. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.
[0029] In order to improve the technical problem that only identical text data can be removed in the existing technology, when most of the text data is similar but only a small part of the data is changed, it is impossible to deduplicate similar text data, which leads to poor deduplication effect. This embodiment provides a data deduplication method, such as Figure 1 As shown, the method includes:
[0030] Step 101: perform word segmentation processing on multiple text data to be processed respectively to obtain each word segmentation element corresponding to each text data.
[0031] In the embodiment of the present application, word segmentation may be analyzing and processing text data with words as the smallest unit, and word segmentation elements may be words obtained through word segmentation processing.
[0032] For example, a complete sentence can be converted into multiple words after word segmentation. For example, "Today is a sunny day". After word segmentation, three words "today", "weather" and "sunny" can be obtained. These three words "today", "weather" and "sunny" are the word segmentation elements corresponding to "Today is a sunny day".
[0033] Step 102: Determine the initial arrangement position of each word segmentation element in each text data.
[0034] In an embodiment of the present application, the initial arrangement position of the word segmentation elements may be the initial arrangement position of each word segmentation element in a set of word segmentation elements of a plurality of text data, and the initial arrangement position may be randomly generated.
[0035] Step 103 , rearrange the word segmentation elements in each text data based on the initial arrangement position for multiple times, and obtain the first word segmentation element in the text data after each position rearrangement as the feature value of the text data.
[0036] In the embodiment of the present application, the multiple rearrangements of the arrangement positions of the word segmentation elements of each text data are all performed based on the initial arrangement positions in step 102 .
[0037] Optionally, the feature value of the text data is the first word segmentation element after the text data is rearranged based on the initial arrangement position each time.
[0038] For example, for text data A, the first word segmentation element after rearrangement is a 1 , a 1 That is, the feature value corresponding to the sorted text data A. The first word segmentation element after the next rearrangement is a 2 , a 2 is the feature value corresponding to the next sorted text data A; for text data B, the first word segmentation element after rearrangement is b, and b is the feature value corresponding to the current sorted text data B. The first word segmentation element after the next rearrangement is b 2 , b 2 That is, the feature value corresponding to the next sorting of the text data B will not be given one by one here.
[0039] Step 104: Identify repeated text data in the plurality of text data according to the repeated information of the characteristic value of each text data.
[0040] The repetition information of the feature value may include complete repetition of the feature value, or partial repetition of the feature value, such as repetition within the same segment area, and so on.
[0041] In the embodiment of the present application, each time the word segmentation elements of the text data are rearranged in position, a feature value corresponding to the text data can be determined. After the text data is rearranged multiple times, multiple feature values that can replace the text data can be obtained.
[0042] Optionally, if there are at least two replacement text data whose feature values are completely the same, it can be identified that there are duplicate text data in the at least two text data.
[0043] Step 105: De-duplicate the duplicate text data.
[0044] In the embodiment of the present application, the deduplication processing of the duplicate text data may be to keep only one of at least two duplicate text data and remove all the others.
[0045] Compared with the current existing technology, in the process of deduplicating text data, this embodiment segments each text data, and uniformly shuffles the segmentation elements of each text data for multiple times. The feature value of the text data is obtained by the first segmentation element in the text data after each position rearrangement, and the repeated text data is screened by comparing the repeated information of the feature value. In the case that most of the data in the text data is similar but only a small part of the data has changed, the similar text data can correspond to the feature value whose repeated information meets the conditions, so that the similar text data in which most of the data is similar but only a small part of the data has changed can be effectively removed, thereby improving the effect of data deduplication processing.
[0046] Further, as a refinement and extension of the above embodiment, in order to fully illustrate the specific implementation process of the method of the embodiment of the present disclosure, the embodiment of the present application provides the following Figure 2 The specific method shown includes:
[0047] Step 201: perform word segmentation processing on multiple text data to be processed respectively to obtain each word segmentation element corresponding to each text data.
[0048] For example, taking the text data to be processed as C1=I came to XX company, C2=XX company press conference as an example, the text data to be processed is segmented, and each segmentation element corresponding to each text data can be obtained as follows:
[0049] C1=I came to XX company and got: {I, came to, XX company};
[0050] C2 = XX company press conference to obtain: {XX company, press conference}.
[0051] Step 202: Determine the arrangement position of each word segmentation element in each text data.
[0052] Optionally, step 202 may specifically include: performing a union operation on each word segmentation element corresponding to each text data to obtain a word segmentation element set; determining the sorting number of each word segmentation element in the word segmentation element set; and determining the initial sorting number corresponding to each word segmentation element in each text data with reference to the sorting number of each word segmentation element in the word segmentation element set.
[0053] Among them, the initial sort number is used to determine the initial arrangement position.
[0054] Exemplarily, based on step 201, the word segmentation elements corresponding to the text data C1 and C2 are subjected to a union operation to obtain {I, came to, XX company, press conference}, and the sorting number of each word segmentation element can be determined by Table 1, which can be shown as follows:
[0055] Table 1
[0056]
[0057]
[0058] In Table 1, the row number may represent the sorting number of each word segmentation element in the word segmentation element set, and the element may be each word segmentation element in the word segmentation element set.
[0059] Exemplarily, in the word segmentation element set, the sorting number corresponding to the word segmentation element "I" is "1", and in the text data C1, the initial sorting number corresponding to the word segmentation element "I" is "1", and the initial arrangement position is the first row; correspondingly, in the word segmentation element set, the sorting number corresponding to the word segmentation element "come" is "3", and in the text data C1, the initial sorting number corresponding to the word segmentation element "come" is "3", and the initial arrangement position is the third row; in the word segmentation element set, the sorting number corresponding to the word segmentation element "XX Company" is "4", and in the text data C1, the initial sorting number corresponding to the word segmentation element "XX Company" is "4", and the initial arrangement position is the fourth row. In the text data C1, the initial arrangement positions of the word segmentation elements can be shown in Table 2 below:
[0060] Table 2
[0061] Line Number element 1 I 3 come 4 XX Company
[0062] For example, in the word segmentation element set, the sorting number corresponding to the word segmentation element "press conference" is "2", and in the text data C2, the initial sorting number corresponding to the word segmentation element "press conference" is "2", and the initial arrangement position is the second row; correspondingly, in the word segmentation element set, the sorting number corresponding to the word segmentation element "XX company" is "4", and in the text data C2, the initial sorting number corresponding to the word segmentation element "XX company" is "4", and the initial arrangement position is the fourth row. In the text data C2, the initial arrangement positions of the word segmentation elements can be shown in the following Table 3:
[0063] Table 3
[0064] Line Number element 2 Press Conference 4 XX Company
[0065] Step 203: Use multiple hash functions to calculate the arrangement positions of the word segmentation elements after each position rearrangement.
[0066] Optionally, step 203 may specifically include: determining the number of word segmentation elements in the word segmentation element set; using the number of word segmentation elements in the word segmentation element set and the initial sorting number corresponding to each word segmentation element in each text data as parameters, using multiple hash functions to calculate the re-sorting number of the word segmentation element each time; and determining the arrangement position of the word segmentation element after each re-arrangement according to the re-sorting number each time.
[0067] Optionally, step 203 specifically further includes: calculating a reordering number of the word segmentation elements in the target text data according to a formula of the first Hash function.
[0068] The formula of the first hash function is shown in Formula 1 below:
[0069] h1(i)=(i+1)%n (Formula 1)
[0070] In Formula 1, i represents the initial sorting number of the word segmentation element, n represents the number of word segmentation elements in the word segmentation element set, h1(i) represents the re-sorting number of the word segmentation element with initial sorting number i calculated by the first hash function, and % represents the remainder.
[0071] Exemplarily, based on step 202, the hash function of formula 1 is used to perform a first rearrangement on table 1 to obtain table 4 as shown below:
[0072] Table 4
[0073] Line Number element C1 C2 1 XX Company 1 1 2 I 1 0 3 Press Conference 0 1 4 come 1 0
[0074] Exemplarily, based on Table 4, in the word segmentation element set, the sorting number corresponding to the word segmentation element "XX Company" is "1", and in the text data C1, the first rearrangement sorting number corresponding to the word segmentation element "XX Company" is "1", and the first rearrangement arrangement position is the first row; correspondingly, in the word segmentation element set, the sorting number corresponding to the word segmentation element "I" is "2", and in the text data C1, the first rearrangement sorting number corresponding to the word segmentation element "I" is "2", and the first rearrangement arrangement position is the second row; in the word segmentation element set, the sorting number corresponding to the word segmentation element "come" is "4", and in the text data C1, the first rearrangement sorting number corresponding to the word segmentation element "come" is "4", and the first rearrangement arrangement position is the fourth row. In the text data C1, the first rearrangement arrangement position of the word segmentation element can be shown in Table 5 below:
[0075] Table 5
[0076] Line Number element 1 XX Company 2 I 4 come
[0077] Exemplarily, based on Table 4, in the word segmentation element set, the sorting number corresponding to the word segmentation element "XX Company" is "1", and in the text data C2, the first rearrangement sorting number corresponding to the word segmentation element "XX Company" is "1", and the first rearrangement arrangement position is the first row; correspondingly, in the word segmentation element set, the sorting number corresponding to the word segmentation element "press conference" is "3", and in the text data C2, the first rearrangement sorting number corresponding to the word segmentation element "press conference" is "3", and the first rearrangement arrangement position is the third row. In the text data C2, the first rearrangement arrangement position of the word segmentation element can be shown in the following Table 6:
[0078] Table 6
[0079] Line Number element 1 XX Company 3 Press Conference
[0080] Correspondingly, based on Table 5 and Table 6, after the first rearrangement, the first element of each set is used as the characteristic value of the set. It can be obtained that the characteristic value of set C1 Feture1 = "XX Company" and the characteristic value of set C2 Feture2 = "XX Company".
[0081] Optionally, step 203 specifically further includes: calculating another reordering number of the word segmentation elements in the target text data according to a formula of a second Hash function.
[0082] The formula of the second hash function is shown in Formula 2 below:
[0083] h2(i)=(i-1)%n (Formula 2)
[0084] In Formula 2, i represents the initial sorting number of the word segmentation element, n represents the number of word segmentation elements in the word segmentation element set, and h2(i) represents another re-sorting number of the word segmentation element with initial sorting number i calculated by the second hash function.
[0085] Exemplarily, using the hash function of Formula 2 to rearrange Table 1 for the second time can obtain Table 7 as shown below:
[0086] Table 7
[0087] Line Number element C1 C2 1 Press Conference 0 1 2 come 1 0 3 XX Company 1 1 4 I 1 0
[0088] Exemplarily, based on Table 7, in the word segmentation element set, the sorting number corresponding to the word segmentation element "come" is "2", and in the text data C1, the second rearrangement sorting number corresponding to the word segmentation element "come" is "2", and the second rearrangement arrangement position is the first row; correspondingly, in the word segmentation element set, the sorting number corresponding to the word segmentation element "XX company" is "3", and in the text data C1, the second rearrangement sorting number corresponding to the word segmentation element "XX company" is "3", and the second rearrangement arrangement position is the third row; in the word segmentation element set, the sorting number corresponding to the word segmentation element "I" is "4", and in the text data C1, the second rearrangement sorting number corresponding to the word segmentation element "I" is "4", and the second rearrangement arrangement position is the fourth row. In the text data C1, the second rearrangement arrangement position of the word segmentation element can be shown in the following Table 8:
[0089] Table 8
[0090] Line Number element 1 come 2 XX Company 4 I
[0091] Exemplarily, based on Table 7, in the word segmentation element set, the sorting number corresponding to the word segmentation element "press conference" is "1", and in the text data C2, the second rearrangement sorting number corresponding to the word segmentation element "press conference" is "1", and the second rearrangement arrangement position is the first row; correspondingly, in the word segmentation element set, the sorting number corresponding to the word segmentation element "XX company" is "3", and in the text data C2, the first rearrangement sorting number corresponding to the word segmentation element "XX company" is "3", and the first rearrangement arrangement position is the third row. In the text data C2, the first rearrangement arrangement position of the word segmentation element can be shown in the following Table 9:
[0092] Table 9
[0093] Line Number element 1 Press Conference 3 XX Company
[0094] Correspondingly, based on Table 8 and Table 9, after the second rearrangement, the first element of each set is used as the feature value of the set. It can be obtained that the feature value of set C1 Feture1 = "come", and the feature value of set C2 Feture2 = "press conference".
[0095] Step 204: Rearrange the initial arrangement positions of the word segmentation elements in each text data according to the arrangement positions calculated each time.
[0096] Exemplarily, based on step 203, the characteristic value sets of sets C1 and C2 after the two hash functions shown in formula 1 and formula 2 can be obtained as shown in Table 10 below:
[0097] Table 10
[0098] Hash Functions C1 C2 h1 XX Company XX Company h2 come Press Conference
[0099] Optionally, based on Table 10, the size of the text set can be reduced to 2 by selecting two hash functions. For long texts, the number of hash functions can be appropriately expanded to reduce the size of the text set to a suitable value.
[0100] For example, if the number of hash functions selected for text data A1 and text data A2 is k, and k is much smaller than the number of word segmentation elements of text data A1 and text data A2, the text data can be reduced in dimension, and the time complexity of the calculation after dimension reduction is O(k*k). This greatly improves the efficiency of text similarity calculation and saves memory space.
[0101] For TB-level massive pre-training data, the amount of data is often as much as billions. If the Jaccard distance between texts is compared globally and calculated, it will bring huge memory and computing overhead, resulting in very low data deduplication efficiency. The embodiment of the present application improves the deduplication efficiency of massive data by reducing the time complexity and space complexity of the algorithm through improvements on the current deduplication strategy.
[0102] Correspondingly, most of the pre-trained corpus documents come from article data on the Internet, in which a large part of the data will have a few keyword changes, but the text semantics are similar. For example: some articles have only some differences in the title, and the content is basically the same; and the articles quote a lot of content from other articles. The embodiment of the present application can also ensure the deduplication effect while improving the deduplication efficiency. It can identify the above-mentioned text similarities to a certain extent and remove similar texts.
[0103] Step 205: Identify repeated text data in the plurality of text data according to the repeated information of the characteristic value of each text data.
[0104] Optionally, step 205 may specifically include: determining at least two text data having the same characteristic value as repeated text data.
[0105] Optionally, the characteristic value of each text data includes the same number of word segmentation elements; accordingly, step 205 specifically also includes: dividing the word segmentation elements included in the characteristic value of each text data into the same number of parts according to the arrangement order of the word segmentation elements in the characteristic value, and marking the order of the number of parts; comparing whether the word segmentation elements corresponding to each text data in the same order of number of parts are the same, such as for each text data, comparing whether the word segmentation elements in each part are the same; if the word segmentation element in the Nth part of the first text data is the same as the word segmentation element in the Nth part of at least one second text data, then determining that the first text data and the at least one second text data are repeated text data.
[0106] Wherein, N is a positive integer.
[0107] In the embodiment of the present application, although the computing power of calculating the similarity between two texts can be reduced by data dimensionality reduction, for n documents, it is still necessary to calculate the similarity between two documents to remove duplicates of similar texts, and the time complexity is O(n*n). In order to solve this problem, the data is divided into the same number of copies and the order of the copies is marked; for each text data, the word segmentation elements in each copy are compared in turn to see if they are the same.
[0108] Exemplarily, taking dividing into k parts into k buckets as an example, in the embodiment of the present application, similar documents are divided into the same bucket, so that similar text deduplication can be completed by only comparing the amount of data in each bucket.
[0109] Accordingly, a specific example is used to show the method of bucketing similar data. For example, there are 5 texts now. After the data dimension is reduced by 12 hash functions, the corresponding feature value set of each text is obtained as shown in Table 11. The feature value is represented by the index corresponding to the word in Table 11 for convenience. Table 11 is as follows:
[0110] Table 11
[0111]
[0112]
[0113] Table 11 is the eigenvalue table corresponding to C1-C5. The eigenvalue table can be called a signature matrix, an eigenvalue matrix, etc. The name of the eigenvalue table is not specifically limited in this embodiment.
[0114] The signature matrix is divided into b bands, each of which consists of 3 rows. For each band, the MD5 value of each column of feature data in the band can be calculated by a hash function.
[0115] In the embodiment of the present application, the same hash function can be used for all intervals to calculate the MD5 value of each column in each interval, or different hash functions can be used for different intervals to calculate the MD5 value of each column. However, for the same column vector in different intervals, even if the MD5 values corresponding to the feature data are the same, the corresponding text data will not be considered as similar data.
[0116] Optionally, as long as the feature values corresponding to two text data have the same two columns in a certain interval, the two text data are considered to have a relatively high similarity and are regarded as repeated text data; and for multiple text data that do not fall into the same bucket as other text data in all intervals, they are considered to have low similarity and are directly ignored.
[0117] For example, in Table 11, the eigenvalues of the 1st and 4th columns in interval b1 are both [2, 3, 4], so these two columns will fall into the same bucket under interval 1. Therefore, regardless of whether these two columns fall into the same bucket in the remaining three intervals, these two text data will become duplicate text data. The two columns that are not equal in interval 1 have another three chances to become duplicate text, because they only need to be equal once in the remaining three intervals.
[0118] Exemplarily, based on Table 11, the feature values are divided into four parts, and the similarity of the text data considered to be repeated text data can be determined to be 25% through the Jaccard distance.
[0119] Optionally, the Jaccard distance can determine whether two sets are equal. It is currently widely used to determine whether two texts are similar. It is called the Jaccard similarity algorithm. Its basic principle is shown in the following formula 3:
[0120] Jac(X,Y)=|X∩Y| / |X∪Y| (Formula 3)
[0121] For example, set X = {a, b, c}, Y = {b, c, d}. Then Jac(X, Y) = 2 / 4 = 0.50. That is, set X and set Y have 50% of the same elements. That is, the number of intersections of two sets is divided by the number of unions of the two sets. The range is between [0, 1].
[0122] When used for text similarity judgment, the text needs to be segmented first and converted into a word set. For example, the following two texts are segmented as follows: C1 = I came to XX company and the segmentation is {I, came, XX company}; C2 = XX company press conference and the segmentation is {XX company, press conference}.
[0123] The length of set C1 is len(C1)=3, and the length of set C2 is len(C2)=2. Then the Jaccard similarity between the two texts is: 1 / 4=0.25. The greater the Jaccard similarity, the more similar the two texts are. The time complexity of calculating the Jaccard distance between two texts is O(len(C1)*len(C2)). It takes a lot of time to calculate the similarity of long texts, so data dimensionality reduction is required.
[0124] Step 206: De-duplicate the duplicate text data.
[0125] In the embodiment of the present application, based on step 205, duplicate text data can be obtained and deduplication processing can be performed on the duplicate text data. After deduplication processing for at least two duplicate text data, only one text data is left. Text data that is not duplicate text data can be directly ignored.
[0126] In the embodiment of the present application, in terms of data deduplication time optimization, the original 48 hours for deduplication of 40G documents is optimized to only 2 hours; in terms of data deduplication memory optimization, the memory required for deduplication of 40G documents in the database (hudi) format used to store massive data is optimized from 800G to 80G; in terms of data deduplication effect, the problem that MD5 deduplication cannot determine a few texts with different keywords but the same semantics as similar texts is solved. At the same time, both long and short text data can have relatively accurate identification of similar texts. Of the 270G documents from the Internet, 210G remains after deduplication using the method of the embodiment of the present application, and the deduplication ratio is 22.2%.
[0127] Compared with the current prior art, in the process of deduplication of text data, this embodiment achieves the purpose of compressing the text set and realizing data dimension reduction by designing a specific hash function and limiting the number of functions, thereby improving the computing efficiency and reducing the memory space required to be used; using the characteristic value of the text data to replace the text data, when there is a large part of the data similar but only a small part of the data changes in the text data, the similar text data can correspond to the same characteristic value, so that the similar text data with most of the data similar but only a small part of the data changes can be effectively removed, thereby improving the effect of data deduplication processing. By dividing the text data into the same number of parts, only a part of the characteristic values of the text data can be compared, and similar texts can be identified, achieving the purpose of deduplication of similar texts, greatly reducing the time used for deduplication of massive data, and improving the efficiency of deduplication processing.
[0128] Further, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides a data deduplication processing device, such as Figure 3 As shown, the device includes: a processing module 31, a determination module 32, an acquisition module 33, and an identification module 34.
[0129] The processing module 31 is configured to perform word segmentation processing on the multiple text data to be processed respectively, and obtain each word segmentation element corresponding to each text data;
[0130] A determination module 32 is configured to determine the initial arrangement position of each word segmentation element in each text data;
[0131] The acquisition module 33 is configured to rearrange the positions of the word segmentation elements in each text data based on the initial arrangement positions for multiple times, and obtain the first word segmentation element in the text data after each position rearrangement as the feature value of the text data;
[0132] The identification module 34 is configured to identify repeated text data in the plurality of text data according to the repeated information of the characteristic value of each text data;
[0133] The processing module 31 is further configured to perform deduplication processing on the repeated text data.
[0134] In a specific application scenario, the determination module 32 is specifically configured to perform a union operation on each word segmentation element corresponding to each text data to obtain a word segmentation element set; determine the sorting number of each word segmentation element in the word segmentation element set; and determine the initial sorting number corresponding to each word segmentation element in each text data with reference to the sorting number of each word segmentation element in the word segmentation element set, wherein the initial sorting number is used to determine the initial arrangement position.
[0135] In a specific application scenario, the acquisition module 33 is specifically configured to use multiple hash functions to calculate the arrangement positions of the word segmentation elements after each position rearrangement; according to the arrangement positions calculated each time, the initial arrangement positions of the word segmentation elements in each text data are rearranged.
[0136] In a specific application scenario, the acquisition module 33 is also specifically configured to determine the number of segmentation elements in the segmentation element set; using the number of segmentation elements in the segmentation element set and the initial sorting number corresponding to each segmentation element in each text data as parameters, multiple hash functions are used to calculate the re-sorting number of the segmentation element each time; according to the re-sorting number each time, the arrangement position of the segmentation element after each position re-arrangement is determined.
[0137] In a specific application scenario, the acquisition module 33 is further configured to calculate a reordering number of the word segmentation element in the target text data according to the formula of the first hash function, the formula of the first hash function including: h1(i)=(i+1)%n, wherein i represents the initial sorting number of the word segmentation element, n represents the number of word segmentation elements in the word segmentation element set, and h1(i) represents a reordering number of the word segmentation element with initial sorting number i calculated by the first hash function.
[0138] In a specific application scenario, the acquisition module 33 is further configured to calculate another reordering number of the word segmentation element in the target text data according to the formula of the second Hash function, and the formula of the second Hash function includes: h2(i) = (i-1)% n, where i represents the initial sorting number of the word segmentation element, n represents the number of word segmentation elements in the word segmentation element set, and h2(i) represents another reordering number of the word segmentation element with the initial sorting number i calculated by the second Hash function. In a specific application scenario, the identification module 34 is further configured to determine at least two text data with the same feature value as duplicate text data.
[0139] In a specific application scenario, the feature value of each text data includes the same number of word segmentation elements; accordingly, the identification module 34 is further configured to divide the word segmentation elements included in the feature value of each text data into the same number of parts according to the arrangement order of the word segmentation elements in the feature value, and mark the order of the number of parts; compare whether the word segmentation elements corresponding to each text data in the same order of number of parts are the same; if the word segmentation element in the Nth part of the first text data is the same as the word segmentation element in the Nth part of at least one second text data, then determine that the first text data and the at least one second text data are repeated text data, where N is a positive integer.
[0140] It should be noted that for other corresponding descriptions of the functional units involved in the data deduplication processing device provided in this embodiment, reference can be made to Figure 1 and Figure 2 The corresponding description in will not be repeated here.
[0141] Based on the above Figure 1 and Figure 2 The method shown in the figure, accordingly, the embodiment of the present disclosure also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned Figure 1 and Figure 2 The method shown.
[0142] Based on this understanding, the technical solution of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of the present disclosure.
[0143] Based on the above Figure 1 and Figure 2 The method shown, and Figure 3 In order to achieve the above-mentioned purpose, the embodiment of the present disclosure also provides an electronic device, which can be configured on a computer terminal, etc. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to achieve the above-mentioned Figure 1 and Figure 2 The method shown.
[0144] In some embodiments, the above-mentioned physical device may also include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface may include a display, an input unit such as a keyboard, etc., and the optional user interface may also include a USB interface, a card reader interface, etc. The network interface may include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.
[0145] Those skilled in the art will appreciate that the above-mentioned physical device structure provided in the embodiments of the present disclosure does not constitute a limitation on the physical device, and may include more or fewer components, or a combination of certain components, or different arrangements of components.
[0146] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the above-mentioned physical device, and supports the operation of the information processing program and other software and / or programs. The network communication module is used to realize the communication between the components inside the storage medium, and the communication with other hardware and software in the information processing physical device.
[0147] Based on the above electronic device, the embodiment of the present disclosure further provides a vehicle, which may specifically include: Figure 3 The device shown in the figure or the electronic device as described above. The vehicle can be a new energy vehicle or a traditional vehicle.
[0148] Through the description of the above disclosed implementation methods, the technical personnel in the field can clearly understand that the present disclosure can be implemented by means of software plus the necessary general hardware platform, or by hardware. By applying the scheme of the embodiment of the present disclosure, compared with the current prior art, in the process of deduplication of text data, the present embodiment achieves the purpose of compressing the text set by designing a specific hash function and limiting the number of functions, thereby achieving data dimension reduction, thereby improving the computational efficiency and reducing the memory space required to be used; using the eigenvalue of the text data to replace the text data, in the case where most of the data in the text data is similar but only a small part of the data changes, the similar text data can correspond to the same eigenvalue, thereby effectively removing the similar text data with most of the data similar but only a small part of the data changes, thereby improving the effect of data deduplication processing. By dividing the text data into the same number of copies, only a part of the eigenvalues of the text data can be compared, and similar texts can be identified, achieving the purpose of deduplication of similar texts, greatly reducing the time used for deduplication of massive data, and improving the efficiency of deduplication processing.
[0149] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprises" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0150] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features applied herein.
Claims
1. A data deduplication processing method, It is characterized in that include: Perform word segmentation processing on the multiple text data to be processed respectively to obtain each word segmentation element corresponding to each text data; Determine the initial arrangement position of each word segmentation element in each text data; Rearranging the word segmentation elements in each text data multiple times based on the initial arrangement position, and obtaining the first word segmentation element in the text data after each position rearrangement as the feature value of the text data; Identifying repeated text data in the plurality of text data according to repeated information of the characteristic value of each text data; The repeated text data is deduplicated.
2. The method according to claim 1, It is characterized in that The step of determining the initial arrangement position of each word segmentation element in each text data includes: Perform a union operation on each word segmentation element corresponding to each text data to obtain a word segmentation element set; Determine the sorting number of each word segmentation element in the word segmentation element set; Referring to the sorting number of each word segmentation element in the word segmentation element set, an initial sorting number corresponding to each word segmentation element in each text data is determined, and the initial sorting number is used to determine the initial arrangement position.
3. The method according to claim 2, It is characterized in that The multiple rearrangement of the initial arrangement positions of the word segmentation elements in each text data includes: Use multiple hash functions to calculate the arrangement position of word segmentation elements after each position re-arrangement; According to the arrangement position calculated each time, the initial arrangement position of the word segmentation elements in each text data is rearranged.
4. The method according to claim 3, It is characterized in that The method of using multiple hash functions to calculate the arrangement positions of word segmentation elements after each position rearrangement includes: Determining the number of segmentation elements in the segmentation element set; Using the number of word segmentation elements in the word segmentation element set and the initial sorting number corresponding to each word segmentation element in each text data as parameters, multiple hash functions are used to calculate the re-sorting number of the word segmentation element each time; According to the re-ranking number each time, the arrangement position of the word segmentation element after each position re-arrangement is determined.
5. The method according to claim 4, It is characterized in that Taking the number of word segmentation elements in the word segmentation element set and the initial sorting number corresponding to each word segmentation element in each text data as parameters, multiple hash functions are used to calculate the re-sorting number of the word segmentation element each time, including: According to the formula of the first Hash function, a reordering number of the word segmentation element in the target text data is calculated, and the formula of the first Hash function includes: h1(i)=(i+1)%n Among them, i represents the initial sorting number of the word segmentation element, n represents the number of word segmentation elements in the word segmentation element set, and h1(i) represents the re-sorting number of the word segmentation element with initial sorting number i calculated by the first hash function.
6. The method according to claim 5, It is characterized in that After calculating and obtaining the reordering number of the word segmentation elements in the target text data, it also includes: According to the formula of the second Hash function, another reordering number of the word segmentation elements in the target text data is calculated, and the formula of the second Hash function includes: h2(i)=(i-1)%n Among them, i represents the initial sorting number of the word segmentation element, n represents the number of word segmentation elements in the word segmentation element set, and h2(i) represents another re-sorting number of the word segmentation element with initial sorting number i calculated by the second hash function.
7. The method according to claim 1, It is characterized in that The step of identifying repeated text data in the plurality of text data according to the repeated information of the characteristic value of each text data comprises: At least two text data having the same feature value are determined as repeated text data.
8. The method according to claim 1, It is characterized in that The feature value of each text data includes the same number of word segmentation elements; The step of identifying repeated text data in the plurality of text data according to the repeated information of the characteristic value of each text data comprises: Divide the word segmentation elements included in the feature value of each text data into the same number of copies according to the arrangement order of the word segmentation elements in the feature value, and mark the order of the number of copies; Compare whether the corresponding word segmentation elements of each text data in the same order are the same; If the word segmentation element in the Nth portion of the first text data is the same as the word segmentation element in the Nth portion of at least one second text data, then the first text data and the at least one second text data are determined to be repeated text data, where N is a positive integer.
9. A data deduplication processing device, It is characterized in that The device comprises: A processing module is configured to perform word segmentation processing on the multiple text data to be processed respectively to obtain each word segmentation element corresponding to each text data; A determination module is configured to determine the initial arrangement position of each word segmentation element in each text data; An acquisition module is configured to rearrange the positions of the word segmentation elements in each text data based on the initial arrangement positions for multiple times, and obtain the first word segmentation element in the text data after each position rearrangement as a feature value of the text data; An identification module configured to identify repeated text data in the plurality of text data according to repeated information of a characteristic value of each text data; The processing module is also configured to perform deduplication processing on the repeated text data.
10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
11. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, It is characterized in that When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
12. A vehicle, It is characterized in that include: The device as claimed in claim 9, or the electronic device as claimed in claim 11.