Text data parallel semantic deduplication method, system and equipment and medium
By combining the BERT model, simhash algorithm and global deduplication graph methods, the problems of low efficiency of large-scale text deduplication, neglect of semantic similarity and insufficient multi-core processing strategies in the existing technology are solved, and efficient and accurate text semantic deduplication is achieved, improving data processing efficiency and quality.
Patent Information
- Application Number
- CN202510235681.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
Existing text deduplication technologies require a large amount of memory and computing resources when processing large-scale data, and mainly focus on structural similarity and ignore semantic similarity. The lack of effective strategies in multi-core parallel processing environments leads to inefficiency.
Combining the BERT model, simhash algorithm and global deduplication graph, parallel semantic deduplication of large-scale text data is achieved through preprocessing, semantic feature extraction, similarity calculation and global deduplication graph management.
It improves the efficiency and accuracy of text deduplication, makes full use of the computing power of multi-core clusters, reduces the use of memory resources, and enhances the consideration of text semantic similarity, ensuring the data quality after deduplication.
Smart Images

Figure CN120179756A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text processing, and more specifically, relates to a method, system, device, and medium for parallel semantic deduplication of text data. Background Art
[0002] In the field of text processing, deduplication technology aims to identify and remove text entries with the same or highly similar content. During the pre-training phase of large language models, ensuring data diversity is crucial for improving model performance, where data diversity includes structural and semantic diversity. However, duplicate items in text data may have a negative impact on this diversity, thereby reducing the effect of model training. Therefore, implementing a text deduplication strategy is essential for optimizing the performance of large models.
[0003] When processing large-scale text data, hash algorithms are often used to map text into fixed-length signatures for quickly comparing the similarity between texts. The Simhash algorithm, as an efficient text deduplication technology, generates a hash signature of the text through a series of steps such as feature extraction, hash value calculation, weighted summation, and dimensionality reduction, and uses the Hamming distance to evaluate the similarity between texts. Similarly, the Minhash algorithm is also a commonly used method in the field of text deduplication.
[0004] Although the above algorithms can achieve text data deduplication to a certain extent, when processing large-scale data sets, they usually need to load all data into memory at once for comparison, which poses a high requirement for the memory capacity of computing devices. At the same time, to determine whether a data item is duplicate, it must be compared with all other items in the data set, which will result in significant performance overhead. In addition, the above methods mainly focus on the structural similarity of texts while ignoring semantic similarity, which also has an important impact on the training of large models. Moreover, the above methods lack effective strategies to coordinate the work of multiple processing cores in a multi-core parallel processing environment to avoid duplicate labor and conflicts during the deduplication process, thus failing to fully utilize the advantages of parallel computing. Summary of the Invention
[0005] Aiming at the above problems, the purpose of the present invention is to provide a method, system, device, and medium for parallel semantic deduplication of text data, which effectively realizes parallel semantic deduplication of large-scale text data and improves the efficiency and accuracy of data processing by combining the BERT model, the simhash algorithm, and the global deduplication graph.
[0006] To achieve the above object, the present invention is realized through the following technical solutions: In a first aspect, an embodiment of the present application provides a method for parallel semantic deduplication of text data, including: Collect large-scale text data to be deduplicated, generate sample data after preprocessing, perform word segmentation and word frequency statistics using the BERT vocabulary, and calculate the word frequency vector of each sample data; Use the pre-trained BERT model to extract the semantic features of the sample data, calculate the word semantic weights of the words in the sample data according to the word frequency vector, and calculate the similarity of the sample data based on the semantic features and the simhash algorithm; According to the semantic similarity of the sample data, maintain the pre-deletion dictionary and the global deduplication graph, and perform the deletion operation of the sample data based on the global deduplication graph to form a deduplicated data set.
[0007] In an optional embodiment, the collecting large-scale text data to be deduplicated, generating sample data after preprocessing, performing word segmentation and word frequency statistics using the BERT vocabulary, and calculating the word frequency vector of each sample data includes: Collect large-scale text data to be deduplicated, and generate a unified encoding format through standardized text format processing; Use regular expressions to clean the text data, removing the preset characters, privacy information, and noise information therein; Use the BERT vocabulary to perform word segmentation on the text data, generate multiple sample data, and calculate the word frequency of the preset words in the sample data to generate a word frequency vector.
[0008] In an optional embodiment, the calculating the word frequency of the preset words in the sample data and generating a word frequency vector includes: Calculate the word frequency of word w in sample data i through the following formula :
[0009] where, represents the number of times word w appears in sample data i, represents the number of words in sample data i.
[0010] In an optional embodiment, the using the pre-trained BERT model to extract the semantic features of the sample data, calculating the word semantic weights of the words in the sample data according to the word frequency vector, and calculating the similarity of the sample data based on the semantic features and the simhash algorithm includes: Input sample data i into the pre-trained BERT model, and take the output of the last hidden layer as the semantic feature of this sample data i , take the attention weights of the first layer and the last layer and calculate the mean value as the word semantic weight of sample data i ; The index, semantic feature of sample data i , word semantic weight and the word frequency vector are stored in the memory of each processor core; For each sample data i, according to the k-nearest neighbor principle, k sample data with the Euclidean distance of semantic features closest to the semantic features of sample data i are selected and placed in the memory for comparison; Use the TF-IDF algorithm and word semantic weights to calculate the word weights of sample data i; Using the said word weights, obtain the hash signature of sample data i based on the simhash algorithm to measure sample similarity.
[0011] In an alternative embodiment, the using the TF-IDF algorithm and word semantic weights to calculate the word weights of sample data i includes: According to the k sample data with the Euclidean distance of semantic features closest to the semantic features of sample data i, determine (k + 1) sample data, and calculate the inverse document frequency of word w in sample data i through the following formula :
[0012] where, is the number of samples with word w in sample i data and its k nearest neighbors; According to the word frequency , the inverse document frequency and the word semantic weight , calculate the word weight of word w in sample data i through the following formula :
[0013] where, represents the word semantic weight of word w in sample data i.
[0014] In an alternative embodiment, the maintaining a pre-deletion dictionary and a global deduplication graph according to the semantic similarity of sample data, and performing the deletion operation of sample data based on the global deduplication graph to form a deduplicated data set includes: For sample data i, compare it with its k nearest neighbor samples; If the Hamming distance between the hash signature of a certain nearest neighbor sample and sample data i is less than the preset threshold, it is considered that the nearest neighbor sample is repeated with sample data i, and add the key as the index of sample data i and the value as the index of the repeated nearest neighbor sample to the pre-deletion dictionary; Merge the pre-deletion dictionaries corresponding to each sample data to construct a global deduplication graph, use the data samples as nodes, and if two sample data are repeated, construct an undirected edge for the two nodes; In the global deduplication graph, determine the sample data to be deleted; Perform an actual deletion operation on the sample data to be deleted; Merge the dataset after the deletion operation to form the final deduplicated dataset.
[0015] In an optional implementation, the determining the sample data to be deleted in the global deduplication graph includes: In the global deduplication graph, each connected component represents a set of duplicate sample data; In each connected component, according to the length of the sample data corresponding to the nodes therein, retain the sample data with the longest length, and store the indexes of the remaining sample data into the list to be deleted, so as to perform the deletion operation according to the list to be deleted.
[0016] In a second aspect, an embodiment of the present application further provides a text data parallel semantic deduplication system, including: A text preprocessing module, configured to collect large-scale text data to be deduplicated, generate sample data after preprocessing, perform word segmentation and word frequency statistics using a BERT vocabulary, and calculate the word frequency vector of each sample data; A text similarity calculation module, configured to use a pre-trained BERT model to extract the semantic features of the sample data, calculate the word semantic weights of the words in the sample data according to the word frequency vector, and calculate the similarity of the sample data based on the semantic features and the simhash algorithm; A text deduplication module, configured to maintain a pre-deletion dictionary and a global deduplication graph according to the semantic similarity of the sample data, and perform a deletion operation on the sample data based on the global deduplication graph to form a deduplicated dataset.
[0017] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the text data parallel semantic deduplication method described in any one of the above are implemented.
[0018] In a fourth aspect, an embodiment of the present application further provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the text data parallel semantic deduplication method described in any one of the above are implemented.
[0019] From the above technical solutions, it can be seen that the present invention has the following advantages: In the text data parallel semantic deduplication method provided by this application, through three stages of text preprocessing, similarity calculation, and text deduplication, efficient deduplication of large-scale text data is achieved. In the preprocessing stage, the text is subjected to format standardization, simple cleaning, and word frequency statistics. In the similarity calculation stage, the BERT model is used to extract semantic features, and the similarity of the text is calculated in combination with the simhash algorithm. In the deduplication stage, the repeatability of the text is judged through the k-nearest neighbor principle and Hamming distance, a pre-deletion dictionary is maintained, and a global deduplication graph is constructed based on this. High-quality samples are retained through connected component analysis, and duplicates are deleted. This invention adopts a parallel processing mechanism, significantly improving the deduplication efficiency. Finally, a dataset that is more refined in structure and semantics is formed, providing a high-quality data foundation for the pre-training of large models.
[0020] Through a parallel processing mechanism, this application makes full use of the computing power of the multi-core cluster, reduces the total time required for deduplication, and improves the efficiency of text deduplication.
[0021] By only storing necessary numerical data rather than the complete text content in memory, this application adapts to the processing requirements of large-scale datasets and effectively reduces the occupation of memory resources.
[0022] This application enhances the consideration of text semantic similarity. By using the BERT model to extract the semantic features of the text and combining the simhash algorithm to calculate the similarity of the text, it can more accurately identify and retain semantically valuable text.
[0023] This application optimizes the parallel deduplication strategy. By constructing a global deduplication graph and connected component analysis, it effectively solves the conflict problem in parallel processing and ensures the quality of the data after deduplication. Brief Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 It is a schematic flow chart of the text data parallel semantic deduplication method provided by this application.
[0026] Figure 2 It is a schematic flow chart of the text preprocessing method provided by this application.
[0027] Figure 3 It is a schematic flow chart of the text similarity calculation method provided by this application.
[0028] Figure 4 It is a schematic flow chart of the text deduplication method provided by this application.
[0029] Figure 5 This is a schematic structural diagram of the text data parallel semantic deduplication system provided for this application.
[0030] Figure 6 This is a schematic structural diagram of the electronic device provided for this application. Detailed implementation manners
[0031] In the following, in the specific steps of the text data parallel semantic deduplication method to be described in detail below, various embodiments of the present disclosure will be described more fully. The present disclosure can have various embodiments, and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents, and / or alternative solutions falling within the spirit and scope of the various embodiments of the present disclosure.
[0032] Hereinafter, the term "comprising" or "may comprise" that can be used in various embodiments of the present disclosure indicates the presence of the disclosed functions, operations, or elements, and does not limit the addition of one or more functions, operations, or elements. In addition, as used in various embodiments of the present disclosure, the terms "comprising", "having" and their cognates are only intended to indicate the presence of specific features, numbers, steps, operations, elements, components, or combinations of the foregoing items, and should not be construed as precluding the existence or addition of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items first.
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] Please refer to Figure 1 Shown is a method flow chart of a text data parallel semantic deduplication method in a specific embodiment. The method includes: S1: Collect large-scale text data to be deduplicated, generate sample data after preprocessing, perform word segmentation and word frequency statistics using the BERT vocabulary, and calculate the word frequency vector of each sample data.
[0035] Exemplarily, first collect large-scale text data to be deduplicated, generate a unified encoding format through standardized text format processing; then use regular expressions to clean the text data; finally, perform word segmentation and word frequency statistics on the text data using the BERT vocabulary to generate a word frequency vector.
[0036] S2: Extract the semantic features of the sample data using a pre-trained BERT model, calculate the word semantic weights of the words in the sample data based on the word frequency vectors, and calculate the similarity of the sample data based on the semantic features and the simhash algorithm.
[0037] Exemplarily, first use the BERT model to extract the semantic features of the text; then use the word semantic weights calculated by the attention weights of BERT.
[0038] At this time, according to the k-nearest neighbor principle, select the k data items closest to the target text data; and use the TF-IDF algorithm and the word semantic weights to calculate the word weights of the samples.
[0039] Finally, use the above word weights to obtain the hash signature of the sample based on the simhash algorithm to measure the sample similarity.
[0040] It should be noted that this step will be executed in parallel on multiple cores of the processor to improve the deduplication efficiency when facing large-scale data.
[0041] S3: According to the semantic similarity of the sample data, maintain a pre-deletion dictionary and a global deduplication graph, and perform the deletion operation of the sample data based on the global deduplication graph to form a deduplicated data set.
[0042] Exemplarily, first compare the Hamming distance of the hash signatures between the text data to determine whether the text data is repeated; for the repeated data, do not immediately perform the deletion operation, but maintain a pre-deletion dictionary.
[0043] Then, construct a global deduplication graph at the central node to merge the pre-deletion dictionary; further, in the global deduplication graph, retain one sample with the longest length in each connected domain, and mark the rest for deletion; At this time, perform the actual deletion operation on each processor core.
[0044] Finally, merge the data sets to form the final deduplicated data set.
[0045] In this embodiment, by cleverly combining the deep semantic understanding ability of the BERT model, the fast similarity evaluation of the simhash algorithm, and the efficient management strategy of the global deduplication graph, an innovative and efficient solution is provided for the parallel semantic deduplication of large-scale text data. When processing massive text data, this method first uses the BERT model to extract deep semantic features of the text, which can not only capture the subtle semantic differences between texts, but also greatly enhance the accuracy and robustness of deduplication. Subsequently, through the simhash algorithm, the high-dimensional semantic feature space is mapped to a low-dimensional hash space, thereby realizing the fast evaluation of text similarity and greatly improving the deduplication efficiency.
[0046] In addition, this method also introduces the concept of a global deduplication graph. By finely managing the sample data and its similarity, an intuitive and easy-to-operate deduplication framework is constructed. In this framework, each data sample is regarded as a node, and the relationship between similar samples is represented as an edge between nodes. This graphical representation not only makes the identification of duplicate data intuitive and easy to implement, but also provides the possibility to implement more complex deduplication strategies (such as making deduplication decisions based on factors such as data length and timestamp).
[0047] In summary, this method realizes the efficient and accurate parallel semantic deduplication of large-scale text data by integrating advanced semantic understanding models, efficient similarity evaluation algorithms, and intuitive deduplication management strategies, which not only improves the efficiency of data processing, but also significantly enhances the quality and usability of data. This has important practical significance and application value for fields such as data mining, information retrieval, and text analysis.
[0048] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.
[0049] Refer to Figure 2 As shown, this embodiment discloses a text preprocessing method, including: S101: Collect large-scale text data to be deduplicated, and generate a unified encoding format through standardized text format processing.
[0050] S102: Use regular expressions to clean the text data, removing the preset characters, privacy information, and noise information therein.
[0051] S103: Use the BERT vocabulary to tokenize the text data, generate multiple sample data, and calculate the word frequency of the preset words in the sample data to generate a word frequency vector.
[0052] Among them, the word frequency of word w in sample data i is calculated by the following formula :
[0053] In the above formula, represents the number of times the word w appears in the sample data i, represents the number of words in the sample data i.
[0054] In this embodiment, the preprocessing of the large-scale text data set is realized by standardizing the text format, cleaning the text using regular expressions, and performing word segmentation and word frequency statistics using the BERT vocabulary in sequence.
[0055] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.
[0056] Referring to Figure 3 shown, this embodiment discloses a text similarity calculation method, which specifically includes the following steps: S201: Input the sample data i into the pre-trained BERT model, and take the output of the last hidden layer as the semantic feature of the sample data i , take the attention weights of the first layer and the last layer and calculate the mean value as the word semantic weight of the sample data i .
[0057] S202: Store the index, semantic feature and word frequency vector of the sample data i into the memory of each processing core.
[0058] S203: For each piece of sample data i, according to the k-nearest neighbor principle, select k pieces of sample data with the Euclidean distance of the semantic feature closest to the semantic feature of the sample data i and put them into the memory for comparison.
[0059] S204: Determine (k + 1) pieces of sample data according to the k pieces of sample data with the Euclidean distance of the semantic feature closest to the semantic feature of the sample data i, and calculate the inverse document frequency of the word w in the sample data i through the following formula :
[0060] where is the number of samples with the word w in the sample i data and its k nearest neighbors.
[0061] S205: According to the word frequency , inverse document frequency and word semantic weight , calculate the word weight of the word w in the sample data i through the following formula :
[0062] Among them, represents the word semantic weight of word w in sample data i.
[0063] Through the above two steps, the word weight of sample data i is calculated by using the TF-IDF algorithm and the word semantic weight.
[0064] S206: Use the word weight to obtain the hash signature of sample data i based on the simhash algorithm to measure the sample similarity.
[0065] Specifically, after normalizing the word weights of all words in sample data i the hash signature of sample data i is obtained by using the simhash algorithm, and this signature can be used to calculate the similarity between different samples.
[0066] It can be seen that in the actual application process, by traversing all samples and cyclically executing the steps of the text similarity calculation method provided in this embodiment, the hash signature of each sample can be obtained.
[0067] In this embodiment, the BERT model is used to extract the semantic features of the text, and the simhash algorithm is combined to calculate the similarity of the text, so that the deduplication process not only considers the structural similarity of the text, but also fully considers the semantic similarity of the text. This helps to more accurately identify and retain unique texts that are semantically valuable and improves the quality of the data set.
[0068] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.
[0069] Refer to Figure 4 As shown, this embodiment discloses a text deduplication method, which specifically includes the following steps: S301: Compare sample data i with its k nearest neighbor samples; If the Hamming distance between the hash signature of a certain nearest neighbor sample and sample data i is less than a preset threshold, it is considered that this nearest neighbor sample is repeated with sample data i, and the index with the key of sample data i and the value of the index of the repeated nearest neighbor sample are added to the pre-deletion dictionary; if there is no nearest neighbor sample whose Hamming distance from the hash signature of sample data i is less than the preset threshold, no processing is performed; this step is executed to traverse all samples.
[0070] Among them, the pre-deletion dictionary is a deletion dictionary initialized and generated in each processing core. It is constructed before this step is executed.
[0071] S302: Merge the pre - deletion dictionaries corresponding to each sample data to construct a global deduplication graph.
[0072] Specifically, first randomly select a processing core as the central node, and merge the pre - deletion dictionaries of all cores with this as the center. Then, construct a global deduplication graph at the central node. In the global deduplication graph, the nodes are sample data. If two sample data are duplicates, an undirected edge is constructed.
[0073] S303: Determine the sample data to be deleted in the global deduplication graph.
[0074] Since in the global deduplication graph, each connected component represents a group of duplicate sample data. Therefore, in each connected component, according to the length of the sample data corresponding to the nodes therein, retain the sample data with the longest length, and store the indexes of the remaining sample data into the list of data to be deleted, so as to perform the deletion operation according to the list of data to be deleted.
[0075] S304: Perform the actual deletion operation on the sample data to be deleted.
[0076] Exemplarily, each core of the processor performs the actual deletion operation according to the list of data to be deleted.
[0077] S305: Merge the datasets after the deletion operation to form the final deduplicated dataset.
[0078] Specifically, after all cores of the processors complete the deletion operation, merge the datasets to form the final deduplicated dataset.
[0079] It can be seen from this embodiment that the present invention adopts a parallel processing mechanism to distribute large - scale text data to different processing cores for pre - processing, duplicate checking, and deduplication, significantly improving the efficiency of text deduplication. Compared with the traditional serial deduplication method, the present invention can make full use of the computing power of multi - core clusters such as spark, greatly reducing the total time required for deduplication.
[0080] In the process of deduplication in this embodiment, only numerical data such as the index and its hash signature of each sample need to be stored in the memory, rather than the complete text content, thus greatly reducing the occupation of memory resources. This method is particularly suitable for processing large - scale datasets and can process more data with limited hardware resources.
[0081] In this embodiment, by constructing a global deduplication graph and connected - component analysis, the conflict problem that may occur in parallel processing by multiple cores is effectively solved. The strategy of retaining the sample with the longest length in each connected component ensures the data quality after deduplication.
[0082] It can be seen that the text data set after deduplication by the method of the present invention provides a more refined and high-quality data basis for the pre-training of large models. This helps to improve the learning effect of the model in the pre-training stage, and further improve the performance of the model in downstream tasks.
[0083] Based on the above embodiments, it can be known that the parallel semantic deduplication method for text data disclosed by the present invention has certain improvements and enhancements in terms of deduplication efficiency, memory occupancy, semantic similarity consideration, conflict resolution strategy, and model pre-training performance compared with the prior art.
[0084] As Figure 5 shown, the following are embodiments of the parallel semantic deduplication system for text data provided by the present disclosure. This system and the parallel semantic deduplication method for text data in the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the parallel semantic deduplication system for text data can refer to the embodiments of the parallel semantic deduplication method for text data.
[0085] A parallel semantic deduplication system for text data, comprising: A text preprocessing module, configured to collect large-scale text data to be deduplicated, generate sample data after preprocessing, perform word segmentation and word frequency statistics using a BERT vocabulary, and calculate the word frequency vector of each sample data.
[0086] A text similarity calculation module, configured to use a pre-trained BERT model to extract semantic features of sample data, calculate the word semantic weights of words in the sample data according to the word frequency vector, and calculate the similarity of the sample data based on the semantic features and the simhash algorithm.
[0087] A text deduplication module, configured to maintain a pre-deletion dictionary and a global deduplication graph according to the semantic similarity of the sample data, and perform a deletion operation on the sample data based on the global deduplication graph to form a deduplicated data set.
[0088] The parallel semantic deduplication system for text data provided in this embodiment realizes efficient and accurate parallel semantic deduplication of large-scale text data by integrating the deep semantic understanding of the BERT model, the fast similarity evaluation of the simhash algorithm, and the efficient management of the global deduplication graph. First, the BERT model is used to extract the deep semantic features of the text, enhancing the accuracy of deduplication. Secondly, the simhash algorithm maps the semantic features to the hash space, quickly evaluating the text similarity and improving the deduplication efficiency. Finally, the global deduplication graph manages the sample data, and through graphical representation, an intuitive and easy-to-operate deduplication decision is realized. This comprehensive technical means not only improves the efficiency of data processing, but also ensures the quality and usability of the data, and has important practical significance and application value in multiple fields.
[0089] Figure 6 Schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.
[0090] The text data parallel semantic deduplication method provided by the embodiments of the present application can be applied to an electronic device. Those skilled in the art can understand that the electronic device structure involved in the embodiments of the present invention does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.
[0091] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, a key, a camera, a display screen, and a SIM card interface, etc.
[0092] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), etc., an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0093] Among them, the processor may be the nerve center and command center of the electronic device. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.
[0094] A memory can also be set in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can save the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the system efficiency.
[0095] The external memory interface can be used to connect to an external memory card, such as a MicroSD card, to implement the storage capacity expansion of the electronic device. The external memory card communicates with the processor through the external memory interface to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.
[0096] The internal memory can be used to store computer-executable program codes, and the computer-executable program codes include instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory can include a program storage area and a data storage area. The internal memory can include a high-speed random access memory and can also include non-volatile memories, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0097] The wireless communication function of the electronic device can be implemented through an antenna, a wireless communication module, a modulation and demodulation processor, a baseband processor, etc.
[0098] The wireless communication module can provide wireless communication solutions applied to the electronic device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSSs), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0099] The electronic device can implement audio functions, etc., through an audio module, a speaker, a receiver, a microphone, a headphone interface, an application processor, etc.
[0100] The electronic device can implement a shooting function through an ISP, a camera, a video codec, a GPU, a display screen, an application processor, etc.
[0101] An electronic device can implement a display function through a GPU, a display screen, an application processor, etc.
[0102] The GPU is a microprocessor for image processing, connecting the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs, which execute program instructions to generate or change display information.
[0103] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0104] The above-mentioned electronic device realizes the parallel semantic deduplication method of text data in the present application. By integrating the deep semantic analysis of the BERT model, the fast similarity calculation of the simhash algorithm, and the management strategy of the global deduplication graph, it provides an efficient solution for the parallel semantic deduplication of large-scale text data. First, use the BERT model to extract deep semantic features of the text to ensure the accuracy of deduplication. Secondly, convert the high-dimensional semantic features into low-dimensional hash values through the simhash algorithm to quickly evaluate the text similarity and improve the deduplication efficiency. Finally, construct a global deduplication graph to intuitively manage the sample data and its similarity relationship, and realize the accurate identification and deletion of duplicate data. The integration of these technical means achieves the beneficial effect of improving the data processing efficiency while ensuring the high quality and usability of data deduplication.
[0105] In the storage medium provided by the present application, there is a program product capable of implementing the parallel semantic deduplication method of text data.
[0106] The parallel semantic deduplication method of text data includes: collecting large-scale text data to be deduplicated, generating sample data after preprocessing, performing word segmentation and word frequency statistics using the BERT vocabulary, and calculating the word frequency vector of each sample data; Using a pre-trained BERT model to extract the semantic features of sample data, calculating the word semantic weights of the words in the sample data according to the word frequency vector, and calculating the similarity of the sample data based on the semantic features and the simhash algorithm; According to the semantic similarity of the sample data, maintaining a pre-deletion dictionary and a global deduplication graph, and performing the deletion operation of the sample data based on the global deduplication graph to form a deduplicated data set.
[0107] In some possible implementation manners, the parallel semantic deduplication method of text data in the present disclosure may be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of this specification.
[0108] The storage medium of the present disclosure may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0109] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A text data parallel semantic deduplication method, characterized in that: include: Collect large-scale text data to be deduplicated, generate sample data after preprocessing, and use the BERT vocabulary to perform word segmentation and word frequency statistics to calculate the word frequency vector of each sample data; Use the pre-trained BERT model to extract the semantic features of the sample data, calculate the semantic weights of the words in the sample data based on the word frequency vector, and calculate the similarity of the sample data based on the semantic features and the simhash algorithm; According to the semantic similarity of the sample data, the pre-deletion dictionary and the global deduplication map are maintained, and the sample data deletion operation is performed based on the global deduplication map to form a deduplicated data set.
2. The text data parallel semantic deduplication method according to claim 1, characterized in that: The method collects large-scale text data to be deduplicated, generates sample data after preprocessing, and uses the BERT vocabulary to perform word segmentation and word frequency statistics to calculate the word frequency vector of each sample data, including: Collect large amounts of text data to be deduplicated and generate a unified encoding format through standardized text format processing; Use regular expressions to clean text data and remove preset characters, privacy information, and noise information; Use the BERT vocabulary to segment text data, generate multiple sample data, calculate the frequency of preset words in the sample data, and generate a frequency vector.
3. The text data parallel semantic deduplication method according to claim 2, characterized in that: The step of calculating the word frequency of a preset word in the sample data and generating a word frequency vector includes: The frequency of word w in sample data i is calculated by the following formula: : in, represents the number of times word w appears in sample data i, Represents the number of words in sample data i.
4. The text data parallel semantic deduplication method according to claim 3, characterized in that: The method uses the pre-trained BERT model to extract the semantic features of the sample data, calculates the word semantic weights of the words in the sample data according to the word frequency vector, and calculates the similarity of the sample data based on the semantic features and the simhash algorithm, including: Input sample data i into the pre-trained BERT model and take the output of the last hidden layer as the semantic feature of sample data i , take the attention weights of the first and last layers and calculate the mean as the semantic weight of the word of sample data i ; The index and semantic features of sample data i , word semantic weight and word frequency vector Stored in the memory of each processor core; For each sample data i, according to the k-nearest neighbor principle, select k sample data whose semantic features are closest to the semantic features of sample data i and put them into memory for comparison; Calculate the word weight of sample data i using TF-IDF algorithm and word semantic weight; Using the word weights, a hash signature of sample data i is obtained based on the simhash algorithm to measure sample similarity.
5. The text data parallel semantic deduplication method according to claim 4, characterized in that: The method of calculating the word weight of sample data i by using the TF-IDF algorithm and the word semantic weight includes: According to the sample data whose Euclidean distance between k semantic features and the semantic features of sample data i is the closest, (k+1) sample data are determined, and the inverse document frequency of word w in sample data i is calculated by the following formula: : in, is the number of samples with word w in sample i and its k nearest neighbors; According to word frequency , inverse document frequency and word semantic weight , the word weight of word w in sample data i is calculated by the following formula : in, Represents the word semantic weight of word w in sample data i.
6. The text data parallel semantic deduplication method according to claim 5, characterized in that: The method includes maintaining a pre-deletion dictionary and a global deduplication map according to the semantic similarity of the sample data, and performing a deletion operation on the sample data based on the global deduplication map to form a deduplication data set, including: For sample data i, compare it with its k nearest neighbor samples; If the hash signature Hamming distance between a neighbor sample and sample data i is less than the preset threshold, the neighbor sample is considered to be a duplicate of sample data i, and a key is added to the pre-deletion dictionary as the index of sample data i, and the value is the index of the duplicate neighbor sample; Merge the pre-deletion dictionary corresponding to each sample data to build a global deduplication graph, with data samples as nodes. If two sample data are repeated, an undirected edge is built between the two nodes. In the global deduplication graph, determine the sample data to be deleted; Perform actual deletion operations on the sample data to be deleted; After merging and deleting the data set, the final deduplicated data set is formed.
7. The text data parallel semantic deduplication method according to claim 6, characterized in that: Determining the sample data to be deleted in the global deduplication graph includes: In the global deduplication graph, each connected domain represents a set of repeated sample data; In each connected domain, according to the length of the sample data corresponding to the nodes therein, the sample data with the longest length is retained, and the indexes of the remaining sample data are stored in the list to be deleted, so as to perform the deletion operation according to the list to be deleted.
8. A text data parallel semantic deduplication system, characterized in that: The system adopts the text data parallel semantic deduplication method as claimed in any one of claims 1 to 7; The system comprises: The text preprocessing module is used to collect large-scale text data to be deduplicated, generate sample data after preprocessing, and use the BERT vocabulary to perform word segmentation and word frequency statistics to calculate the word frequency vector of each sample data; The text similarity calculation module is used to extract the semantic features of the sample data using the pre-trained BERT model, calculate the semantic weights of the words in the sample data according to the word frequency vector, and calculate the similarity of the sample data based on the semantic features and the simhash algorithm; The text deduplication module is used to maintain a pre-deletion dictionary and a global deduplication map according to the semantic similarity of the sample data, and to perform a sample data deletion operation based on the global deduplication map to form a deduplicated data set.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the text data parallel semantic deduplication method as described in any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for parallel semantic deduplication of text data as claimed in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Text recognition method and device, computer equipment and readable storage medium
CN120373303A