Document duplicate checking method and device based on semantics
By generating and quantifying the semantic vector set of documents, similar documents are retrieved, and sentence-level comparison is performed, the problem of low document diversion checking efficiency in the existing technology is solved, and efficient document diversion checking is achieved.
Patent Information
- Application Number
- CN202210182346.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-02-25
AI Technical Summary
The prior art is less efficient when searching semantic related documents from massive historical documents and performing document comparisons, making it difficult to efficiently realize document pounding.
By generating a semantic vector set of documents, quantization of vectors, searching out the historical documents closest to the document to be checked, and sentence segmentation and sentence pair combinations are performed on the document to filter out similar sentence pairs.
It greatly shortens the calculation time, improves the efficiency of document plagiarism checking, and enables faster search and comparison of documents.
Smart Images

Figure CN114564935B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a semantic-based document duplication checking method and device. Background Art
[0002] This section is intended to provide a background or context to the embodiments of the invention recited in the claims. No description herein is admitted to be prior art by inclusion in this section.
[0003] When electronic media are writing documents, or when certain documents need to be queried for duplication, etc., it is necessary to find the subject from a large number of historical documents, especially semantically related documents, and then compare the found documents with the documents to be checked to check for duplicates. Therefore, an efficient document duplication checking method is currently needed. Summary of the invention
[0004] The embodiment of the present invention provides a semantic-based document duplication checking method for checking duplicate documents with high efficiency. The method comprises:
[0005] Generate a semantic vector set of a document set, wherein the document set includes a document to be checked for duplicates and a plurality of historical documents;
[0006] Perform vector quantization on the semantic vector set to obtain a compressed vector set;
[0007] Based on the compressed vector set, the historical document closest to the document to be checked for duplicates is retrieved, and the historical document closest to the document to be checked for duplicates is determined as a similar document;
[0008] Segment the sentences of the document to be checked for duplicates to obtain a first sentence set, and segment the sentences of similar documents to obtain a second sentence set;
[0009] Combining sentences in the first sentence set and the second sentence set in pairs to obtain multiple sets of sentence pairs;
[0010] Filter out similar sentence pairs from multiple groups of sentence pairs.
[0011] The embodiment of the present invention also provides a semantic-based document duplication checking device, which is used to check the duplication of documents with high efficiency. The device includes:
[0012] A semantic vector set generation module, used to generate a semantic vector set of a document set, wherein the document set includes a document to be checked for duplicates and a plurality of historical documents;
[0013] A vector quantization module, used for performing vector quantization on a semantic vector set to obtain a compressed vector set;
[0014] A similar document determination module is used to retrieve the historical document closest to the document to be checked for duplicates based on the compressed vector set, and determine the historical document closest to the document to be checked for duplicates as a similar document;
[0015] A sentence segmentation module is used to segment the sentences of the document to be checked for duplicates to obtain a first sentence set, and to segment the sentences of similar documents to obtain a second sentence set;
[0016] A sentence pair obtaining module, used for combining sentences in the first sentence set and the second sentence set in pairs to obtain multiple groups of sentence pairs;
[0017] The similar sentence pair screening module is used to screen out similar sentence pairs from multiple groups of sentence pairs.
[0018] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned semantic-based document duplication checking method when executing the computer program.
[0019] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned semantic-based document duplication checking method.
[0020] An embodiment of the present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned semantic-based document duplication checking method.
[0021] In an embodiment of the present invention, a semantic vector set of a document set is generated, and the document set includes a document to be checked for duplicates and multiple historical documents; the semantic vector set is vectorized to obtain a compressed vector set; based on the compressed vector set, the historical document closest to the document to be checked for duplicates is retrieved, and the historical document closest to the document to be checked for duplicates is determined as a similar document; the sentences of the document to be checked for duplicates are segmented to obtain a first sentence set, and the similar documents are segmented to obtain a second sentence set; the sentences in the first sentence set and the second sentence set are combined in pairs to obtain multiple groups of sentence pairs; similar sentence pairs are screened out from the multiple groups of sentence pairs. Compared with the technical solution in the prior art that directly checks for duplicates of similar documents and similar sentences through cosine similarity, the semantic description vector is used, and then the semantic vector set is vectorized to retrieve similar documents, which greatly shortens the calculation time and improves the efficiency of checking for duplicates. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:
[0023] Figure 1 Flow chart of a semantic-based document duplication checking method in an embodiment of the present invention;
[0024] Figure 2 A flowchart of generating a semantic vector set in an embodiment of the present invention;
[0025] Figure 3 is a schematic diagram of vector quantization in an embodiment of the present invention;
[0026] Figure 4 A flowchart of vector quantization in an embodiment of the present invention;
[0027] Figure 5 A schematic diagram of the distance calculation between a historical document and a document to be checked for duplicates in an embodiment of the present invention;
[0028] Figure 6 Schematic diagram of a document duplication checking device based on semantics in an embodiment of the present invention;
[0029] Figure 7 Schematic diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] To make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0031] Figure 1 Flow chart of the semantic-based document duplication checking method in an embodiment of the present invention, Figure 1 As shown, including:
[0032] Step 101, generating a semantic vector set of a document set, wherein the document set includes a document to be checked for duplicates and a plurality of historical documents;
[0033] Step 102, performing vector quantization on the semantic vector set to obtain a compressed vector set;
[0034] Step 103, based on the compressed vector set, retrieve the historical document closest to the document to be checked for duplicates, and determine the historical document closest to the document to be checked for duplicates as a similar document;
[0035] Step 104, segment the sentences of the document to be checked for duplicates to obtain a first sentence set, and segment the sentences of similar documents to obtain a second sentence set;
[0036] Step 105, combining sentences in the first sentence set and the second sentence set in pairs to obtain multiple groups of sentence pairs;
[0037] Step 106, screening out similar sentence pairs from the multiple groups of sentence pairs.
[0038] In an embodiment of the present invention, compared with the technical solution in the prior art of directly checking for duplicates of similar documents and similar sentences through cosine similarity, a semantic description vector is used, and then the semantic vector set is vectorized and quantized to retrieve similar documents, which greatly shortens the calculation time and improves the efficiency of checking for duplicates.
[0039] In step 101, a semantic vector set of a document set is generated, where the document set includes a document to be checked for duplicates and a plurality of historical documents.
[0040] Among them, historical documents can be said to be a large number of documents that need to be compared, which can be obtained from the database.
[0041] Figure 2 This is a flow chart of generating a semantic vector set in an embodiment of the present invention. In one embodiment, generating a semantic vector set of a document set includes:
[0042] Step 201: for each document in the document set, input the document into a semantic training model to obtain a semantic two-dimensional matrix, wherein the first dimension of the two-dimensional matrix is the length information of the sentence, and the second dimension is the vector of the semantic information;
[0043] Step 202, along the first dimension of the two-dimensional matrix, add the vectors of the second dimension and take the average, obtain the semantic vector of the document, and add the semantic vector of the document to the semantic vector set.
[0044] In one embodiment, before adding the vectors of the second dimension and taking the average along the first dimension of the two-dimensional matrix, the method further includes:
[0045] Step 203, when the length of the sentence in the first dimension of the two-dimensional matrix is greater than a preset length, the portion exceeding the preset length is truncated;
[0046] Step 204: When the length of the sentence in the first dimension of the two-dimensional matrix is less than a preset length, a preset symbol is used to supplement the first dimension.
[0047] There can be multiple semantic training models as long as a semantic two-dimensional matrix can be obtained, wherein the BERT pre-training model is one of the semantic training models. In the embodiment of the present invention, the document obtained by the BERT pre-training model inference will obtain a matrix of size (512, 768), wherein the first dimension describes the sentence length information, and along the first dimension of the two-dimensional matrix, the vectors of the second dimension are added and the average is taken to obtain a one-dimensional semantic vector of length 768, which represents the vector representation of the entire document. Steps 203 and 204 provide a processing method when the length of the sentence in the first dimension does not meet the preset length, that is, the part exceeding the preset length (e.g., 512) is truncated, and when it is less than the preset length (e.g., 512), the first dimension is supplemented with a preset symbol, and the preset symbol can be 0, *, etc.
[0048] Generally, the semantic vectors obtained in step 101 have the ability to express semantics. The similarity between two documents can be obtained by calculating the similarity between two vectors. The cosine similarity can be used to calculate the similarity between two vectors. However, in a scenario where the historical documents are very large, the complexity of the pairwise comparison method is N squared, which is basically not feasible in practical applications. Therefore, the embodiment of the present invention adopts the idea of the pq vector algorithm.
[0049] In step 102, vector quantization is performed on the semantic vector set to obtain a compressed vector set.
[0050] Figure 3 is a schematic diagram of vector quantization in an embodiment of the present invention, Figure 4 is a flowchart of vector quantization in an embodiment of the present invention, Figure 3 and Figure 4 Correspondingly, in one embodiment, vector quantization is performed on the semantic vector set to obtain a compressed vector set, including:
[0051] Step 401, dividing the vector dimension of the semantic vector set to obtain multiple groups of semantic sub-vectors, the number of semantic sub-vectors in each group of semantic sub-vectors is the number of semantic vectors in the semantic vector set, and the dimension of the semantic sub-vector is smaller than the dimension of the semantic vector;
[0052] Step 402, clustering each group of semantic sub-vectors to obtain a plurality of cluster centers corresponding to each group of semantic sub-vectors, wherein the number of the plurality of cluster centers corresponding to each group of semantic sub-vectors is less than the number of semantic sub-vectors;
[0053] Step 403, for each semantic sub-vector in each group of semantic sub-vectors, find the class center that is closest to the semantic sub-vector among the multiple class centers corresponding to the group of semantic sub-vectors, and mark it as the label of the semantic sub-vector;
[0054] Among them, the labels of all semantic sub-vectors constitute the compressed vector set.
[0055] In one embodiment, the K-means clustering method is used to cluster each subset of semantic vectors.
[0056] In Figure 3 the number of semantic vectors in the semantic vector set is N, and the dimension of the semantic vectors is D; the vector dimension of the semantic vector set is sliced to obtain m groups of semantic sub-vectors. The number of semantic sub-vectors in each group of semantic vectors is N, and the dimension of the semantic vectors is D / m. The K-means clustering method is used to cluster each group of semantic sub-vectors. Each group has K cluster centers, K < N, and the dimension of the cluster centers is D / m. For each semantic sub-vector in each group of semantic sub-vectors, find the cluster center closest to the semantic sub-vector among the multiple cluster centers corresponding to the group of semantic sub-vectors, and mark it as the label of the semantic sub-vector to complete vector quantization and obtain a compressed vector set. The data volume size before compression is N×D×4×8 bits (float), and the data volume size after compression is N×log2K×m, which greatly reduces the data volume.
[0057] In step 103, based on the compressed vector set, retrieve the historical document closest to the document to be checked for duplication, and determine the historical document closest to the document to be checked for duplication as the similar document.
[0058] In one embodiment, retrieving the historical document closest to the document to be checked for duplication based on the compressed vector set includes:
[0059] Construct multiple distance tables. Among them, each distance table corresponds to a group of semantic sub-vectors. Each distance table stores the labels of any two cluster centers corresponding to the multiple cluster centers of each group of semantic sub-vectors as index values and the distances between the any two cluster centers as distance values;
[0060] For each semantic sub-vector of the document to be checked for duplication, based on the label of the semantic sub-vector, query the distance value between the semantic sub-vector and the semantic sub-vectors of each historical document from the distance table corresponding to the group where the semantic sub-vector is located; perform a summation calculation on the obtained multiple distance values to obtain the distance summation value for each historical document;
[0061] Determine the historical document with the smallest distance summation value as the historical document closest to the document to be checked for duplication.
[0062] Figure 5This is a schematic diagram of the distance calculation between the historical document and the document to be checked for duplicates in an embodiment of the present invention. The distance between each semantic subvector x and y corresponding to the two documents is transformed into the distance between the two class centers q(x) and q(y). The distance between the class centers q(x) and q(y) can be calculated in advance and stored in a distance table. When a formal query is made, the distance table can be directly checked to obtain the distance, which greatly shortens the calculation time.
[0063] In step 104, the sentences of the document to be checked for duplicates are segmented to obtain a first sentence set, and the sentences of similar documents are segmented to obtain a second sentence set.
[0064] For example, there is a document A, which can be transformed into the first sentence set {A1, A2, A3, …, An} according to sentence segmentation, and there is a document B, which can be transformed into the second sentence set {B1, B2, B3, …, Bm} according to sentence segmentation.
[0065] In step 105, sentences in the first sentence set and the second sentence set are combined in pairs to obtain multiple groups of sentence pairs.
[0066] The sentences in the first sentence set and the second sentence set are combined in pairs to form combinations of AiBj. There are a total of m×n such combinations.
[0067] In step 106, similar sentence pairs are screened out from the multiple groups of sentence pairs.
[0068] In one embodiment, similar sentence pairs are screened out from multiple groups of sentence pairs, including:
[0069] The edit distance between each group of sentence pairs is calculated, and when the edit distance is less than a preset threshold, the sentence pairs are determined to be similar sentence pairs.
[0070] Specifically, the edit distance between each set of sentence pairs can be performed in parallel.
[0071] In one embodiment, the following formula is used to calculate the edit distance between each group of sentence pairs:
[0072]
[0073] Among them, lev a,b (i,j) is sentence a i and sentence b j The edit distance between .
[0074] In summary, in the method proposed in the embodiment of the present invention, a semantic vector set of a document set is generated, and the document set includes a document to be checked for duplicates and multiple historical documents; the semantic vector set is vectorized to obtain a compressed vector set; based on the compressed vector set, the historical document closest to the document to be checked for duplicates is retrieved, and the historical document closest to the document to be checked for duplicates is determined as a similar document; the sentences of the document to be checked for duplicates are segmented to obtain a first sentence set, and the similar documents are segmented to obtain a second sentence set; the sentences in the first sentence set and the second sentence set are combined in pairs to obtain multiple groups of sentence pairs; similar sentence pairs are screened out from the multiple groups of sentence pairs. Compared with the technical solution in the prior art that directly checks for duplicates of similar documents and similar sentences through cosine similarity, the semantic description vector is used, and then the semantic vector set is vectorized to retrieve similar documents, which greatly shortens the calculation time and improves the efficiency of checking for duplicates.
[0075] The present invention also provides a semantic-based document duplicate checking device, as described in the following embodiments. Since the principle of solving the problem by the device is similar to that of the semantic-based document duplicate checking method, the implementation of the device can refer to the implementation of the semantic-based document duplicate checking method, and the repeated parts will not be repeated.
[0076] Figure 6 The schematic diagram of the semantic-based document duplicate checking device in an embodiment of the present invention comprises:
[0077] A semantic vector set generation module 601 is used to generate a semantic vector set of a document set, wherein the document set includes a document to be checked for duplicates and a plurality of historical documents;
[0078] A vector quantization module 602, used to perform vector quantization on the semantic vector set to obtain a compressed vector set;
[0079] A similar document determination module 603 is used to retrieve the historical document closest to the document to be checked for duplicates based on the compressed vector set, and determine the historical document closest to the document to be checked for duplicates as a similar document;
[0080] Sentence segmentation module 604, used for segmenting sentences of the document to be checked for duplicates to obtain a first sentence set, and for segmenting sentences of similar documents to obtain a second sentence set;
[0081] A sentence pair obtaining module 605 is used to combine sentences in the first sentence set and the second sentence set in pairs to obtain multiple groups of sentence pairs;
[0082] The similar sentence pair screening module 606 is used to screen similar sentence pairs from multiple groups of sentence pairs.
[0083] In one embodiment, the semantic vector set generation module is specifically used to:
[0084] For each document in the document set, the document is input into the semantic training model to obtain a two-dimensional matrix of semantics, wherein the first dimension of the two-dimensional matrix is the length information of the sentence, and the second dimension is the vector of semantic information;
[0085] Along the first dimension of the two-dimensional matrix, the vectors of the second dimension are added and the average is taken to obtain the semantic vector of the document, and the semantic vector of the document is added to the semantic vector set.
[0086] In one embodiment, the semantic vector set generation module is specifically used to:
[0087] Along the first dimension of the two-dimensional matrix, before taking the average after adding the vectors of the second dimension, when the length of the sentence in the first dimension of the two-dimensional matrix is greater than the preset length, the part exceeding the preset length is truncated; when the length of the sentence in the first dimension of the two-dimensional matrix is less than the preset length, the first dimension is supplemented with a preset symbol.
[0088] In one embodiment, the vector quantization module is specifically used for:
[0089] The vector dimension of the semantic vector set is divided to obtain multiple groups of semantic sub-vectors, the number of semantic sub-vectors in each group of semantic sub-vectors is the number of semantic vectors in the semantic vector set, and the dimension of the semantic sub-vector is smaller than the dimension of the semantic vector;
[0090] Clustering each group of semantic sub-vectors to obtain a plurality of class centers corresponding to each group of semantic sub-vectors, wherein the number of the plurality of class centers corresponding to each group of semantic sub-vectors is less than the number of semantic sub-vectors;
[0091] For each semantic sub-vector in each group of semantic sub-vectors, find the class center that is closest to the semantic sub-vector among the multiple class centers corresponding to the group of semantic sub-vectors, and mark it as the label of the semantic sub-vector;
[0092] Among them, the labels of all semantic sub-vectors constitute the compressed vector set.
[0093] In one embodiment, the vector quantization module is specifically used for:
[0094] K-means clustering method is used to cluster each semantic vector subset.
[0095] In one embodiment, the similar document determination module is specifically used to:
[0096] Constructing multiple distance tables, wherein each distance table corresponds to a group of semantic sub-vectors, each distance table uses the labels of any two class centers of the multiple class centers corresponding to each group of semantic sub-vectors as index values, and uses the distance between the any two class centers as the distance value for storage;
[0097] For each semantic sub-vector of the document to be checked for duplicates, based on the label of the semantic sub-vector, query the distance value between the semantic sub-vector and the semantic sub-vector of each historical document from the distance table corresponding to the group to which the semantic sub-vector belongs; sum up the obtained multiple distance values to obtain the summed distance value with each historical document;
[0098] The historical document with the smallest distance sum is determined to be the historical document closest to the document to be checked for duplicates.
[0099] In one embodiment, the similar sentence pair screening module is specifically used for:
[0100] The edit distance between each group of sentence pairs is calculated, and when the edit distance is less than a preset threshold, the sentence pairs are determined to be similar sentence pairs.
[0101] In one embodiment, the similar sentence pair screening module is specifically used for:
[0102] The following formula is used to calculate the edit distance between each pair of sentences:
[0103]
[0104] Among them, lev a,b (i,j) is sentence a i and sentence b j The edit distance between .
[0105] In summary, in the device proposed in the embodiment of the present invention, a semantic vector set of a document set is generated, and the document set includes a document to be checked for duplicates and multiple historical documents; the semantic vector set is vectorized to obtain a compressed vector set; based on the compressed vector set, the historical document closest to the document to be checked for duplicates is retrieved, and the historical document closest to the document to be checked for duplicates is determined as a similar document; the sentences of the document to be checked for duplicates are segmented to obtain a first sentence set, and the similar documents are segmented to obtain a second sentence set; the sentences in the first sentence set and the second sentence set are combined in pairs to obtain multiple groups of sentence pairs; similar sentence pairs are screened out from the multiple groups of sentence pairs. Compared with the technical solution in the prior art that directly checks for duplicates of similar documents and similar sentences through cosine similarity, the semantic description vector is used, and then the semantic vector set is vectorized to retrieve similar documents, which greatly shortens the calculation time and improves the efficiency of checking for duplicates.
[0106] An embodiment of the present invention further provides a computer device, Figure 7It is a schematic diagram of a computer device in an embodiment of the present invention. The computer device 700 includes a memory 710, a processor 720, and a computer program 730 stored in the memory 710 and executable on the processor 720. When the processor 720 executes the computer program 730, the above-mentioned semantic-based document duplication checking method is implemented.
[0107] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned semantic-based document duplication checking method.
[0108] An embodiment of the present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned semantic-based document duplication checking method.
[0109] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0110] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0111] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0112] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0113] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A semantic-based document duplication checking method, characterized in that: include: Generate a semantic vector set of a document set, wherein the document set includes a document to be checked for duplicates and a plurality of historical documents; Perform vector quantization on the semantic vector set to obtain a compressed vector set; Based on the compressed vector set, the historical document closest to the document to be checked for duplicates is retrieved, and the historical document closest to the document to be checked for duplicates is determined as a similar document; Segment the sentences of the document to be checked for duplicates to obtain a first sentence set, and segment the sentences of similar documents to obtain a second sentence set; Combining sentences in the first sentence set and the second sentence set in pairs to obtain multiple sets of sentence pairs; Filter out similar sentence pairs from multiple groups of sentence pairs; Perform vector quantization on the semantic vector set to obtain a compressed vector set, including: The vector dimension of the semantic vector set is divided to obtain multiple groups of semantic sub-vectors, the number of semantic sub-vectors in each group of semantic sub-vectors is the number of semantic vectors in the semantic vector set, and the dimension of the semantic sub-vector is smaller than the dimension of the semantic vector; Clustering each group of semantic sub-vectors to obtain a plurality of class centers corresponding to each group of semantic sub-vectors, wherein the number of the plurality of class centers corresponding to each group of semantic sub-vectors is less than the number of semantic sub-vectors; For each semantic sub-vector in each group of semantic sub-vectors, find the class center that is closest to the semantic sub-vector among the multiple class centers corresponding to the group of semantic sub-vectors, and mark it as the label of the semantic sub-vector; Among them, the labels of all semantic sub-vectors constitute the compressed vector set.
2. The method according to claim 1, characterized in that Generate a semantic vector set for the document set, including: For each document in the document set, the document is input into the semantic training model to obtain a two-dimensional matrix of semantics, wherein the first dimension of the two-dimensional matrix is the length information of the sentence, and the second dimension is the vector of semantic information; Along the first dimension of the two-dimensional matrix, the vectors of the second dimension are added and the average is taken to obtain the semantic vector of the document, and the semantic vector of the document is added to the semantic vector set.
3. The method according to claim 2, characterized in that Along the first dimension of the two-dimensional matrix, the vectors of the second dimension are added and averaged, and also include: When the length of the sentence in the first dimension of the two-dimensional matrix is greater than a preset length, the portion exceeding the preset length is truncated; When the length of the sentence in the first dimension of the two-dimensional matrix is less than a preset length, the first dimension is supplemented with a preset symbol.
4. The method according to claim 1, characterized in that K-means clustering method is used to cluster each semantic vector subset.
5. The method according to claim 1, characterized in that Based on the compressed vector set, retrieve the historical documents that are closest to the document to be checked for duplicates, including: Constructing multiple distance tables, wherein each distance table corresponds to a group of semantic sub-vectors, each distance table uses the labels of any two class centers of the multiple class centers corresponding to each group of semantic sub-vectors as index values, and uses the distance between the any two class centers as the distance value for storage; For each semantic sub-vector of the document to be checked for duplicates, based on the label of the semantic sub-vector, query the distance value between the semantic sub-vector and the semantic sub-vector of each historical document from the distance table corresponding to the group to which the semantic sub-vector belongs; sum up the obtained multiple distance values to obtain the summed distance value with each historical document; The historical document with the smallest distance sum is determined to be the historical document closest to the document to be checked for duplicates.
6. The method according to claim 1, characterized in that From multiple groups of sentence pairs, similar sentence pairs are selected, including: The edit distance between each group of sentence pairs is calculated, and when the edit distance is less than a preset threshold, the sentence pairs are determined to be similar sentence pairs.
7. The method according to claim 6, characterized in that The following formula is used to calculate the edit distance between each pair of sentences: Among them, lev a,b (i,j) is sentence a i and sentence b j The edit distance between .
8. A semantic-based document duplication checking device, characterized in that: include: A semantic vector set generation module, used to generate a semantic vector set of a document set, wherein the document set includes a document to be checked for duplicates and a plurality of historical documents; A vector quantization module, used for performing vector quantization on a semantic vector set to obtain a compressed vector set; A similar document determination module is used to retrieve the historical document closest to the document to be checked for duplicates based on the compressed vector set, and determine the historical document closest to the document to be checked for duplicates as a similar document; A sentence segmentation module is used to segment the sentences of the document to be checked for duplicates to obtain a first sentence set, and to segment the sentences of similar documents to obtain a second sentence set; A sentence pair obtaining module, used for combining sentences in the first sentence set and the second sentence set in pairs to obtain multiple groups of sentence pairs; A similar sentence pair screening module is used to screen similar sentence pairs from multiple groups of sentence pairs; The vector quantization module is specifically used for: The vector dimension of the semantic vector set is divided to obtain multiple groups of semantic sub-vectors, the number of semantic sub-vectors in each group of semantic sub-vectors is the number of semantic vectors in the semantic vector set, and the dimension of the semantic sub-vector is smaller than the dimension of the semantic vector; Clustering each group of semantic sub-vectors to obtain a plurality of class centers corresponding to each group of semantic sub-vectors, wherein the number of the plurality of class centers corresponding to each group of semantic sub-vectors is less than the number of semantic sub-vectors; For each semantic sub-vector in each group of semantic sub-vectors, find the class center that is closest to the semantic sub-vector among the multiple class centers corresponding to the group of semantic sub-vectors, and mark it as the label of the semantic sub-vector; Among them, the labels of all semantic sub-vectors constitute the compressed vector set.
9. The device according to claim 8, characterized in that The semantic vector set generation module is specifically used for: For each document in the document set, the document is input into the semantic training model to obtain a two-dimensional matrix of semantics, wherein the first dimension of the two-dimensional matrix is the length information of the sentence, and the second dimension is the vector of semantic information; Along the first dimension of the two-dimensional matrix, the vectors of the second dimension are added and the average is taken to obtain the semantic vector of the document, and the semantic vector of the document is added to the semantic vector set.
10. The device according to claim 9, characterized in that The semantic vector set generation module is specifically used for: Along the first dimension of the two-dimensional matrix, before taking the average after adding the vectors of the second dimension, when the length of the sentence in the first dimension of the two-dimensional matrix is greater than the preset length, the part exceeding the preset length is truncated; when the length of the sentence in the first dimension of the two-dimensional matrix is less than the preset length, the first dimension is supplemented with a preset symbol.
11. The device according to claim 8, characterized in that The vector quantization module is specifically used for: K-means clustering method is used to cluster each semantic vector subset.
12. The device according to claim 8, characterized in that The similar document determination module is specifically used for: Constructing multiple distance tables, wherein each distance table corresponds to a group of semantic sub-vectors, each distance table uses the labels of any two class centers of the multiple class centers corresponding to each group of semantic sub-vectors as index values, and uses the distance between the any two class centers as the distance value for storage; For each semantic sub-vector of the document to be checked for duplicates, based on the label of the semantic sub-vector, query the distance value between the semantic sub-vector and the semantic sub-vector of each historical document from the distance table corresponding to the group to which the semantic sub-vector belongs; sum up the obtained multiple distance values to obtain the summed distance value with each historical document; The historical document with the smallest distance sum is determined to be the historical document closest to the document to be checked for duplicates.
13. The device according to claim 8, characterized in that The similar sentence pair screening module is specifically used for: The edit distance between each group of sentence pairs is calculated, and when the edit distance is less than a preset threshold, the sentence pairs are determined to be similar sentence pairs.
14. The device according to claim 13, characterized in that The similar sentence pair screening module is specifically used for: The following formula is used to calculate the edit distance between each pair of sentences: Among them, lev a,b (i,j) is sentence a i and sentence b j The edit distance between .
15. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
17. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Document duplicate checking method and device based on semantic analysis
CN108804418A
Document duplicate checking method and system based on semantic analysis
CN111325015A